How a Talend Job for Apache Spark works | Talend Studio Help
Skip to main content

How a Talend Job for Apache Spark works

Last updated: 9/30/2026
A Talend Spark Job generates code for either the Spark Batch or Spark Structured Streaming engine, depending on the Job type you select.

A Talend Spark Batch Job can be run in the following runtime modes and environments:

  • Amazon EMR
  • Amazon EMR Serverless
  • Azure HDInsight
  • Azure Synapse
  • Cloudera Data Engineering
  • Cloudera on Cloud
  • Cloudera on Premises
  • Databricks
  • Availability-noteBeta
    Databricks Serverless
  • Google Managed Service for Apache Spark (formerly Dataproc)
  • Kubernetes
  • Local
  • Livy Knox
  • Spark-submit scripts
  • Standalone
  • Yarn Cluster

A Talend Spark Structured Streaming Job can be run in the following runtime modes and environments:

For details on configuring each distribution, see Defining connection details in Spark configuration.

Information noteNote: A Talend Spark Job is not equivalent to a Spark Job as defined in the Apache Spark documentation; a single Job can generate one or more Spark Jobs depending on how you design it. For more information, see Apache Spark glossary.

When you run the Job, statistics are displayed in the design workspace to indicate the progress of the Spark computations.

The execution information of a Talend Spark Job is logged by the HistoryServer service of the cluster being used. To view this information, open the service web console. The name of the Job in the console is automatically constructed to be ProjectName_JobName_JobVersion, for example, LOCALPROJECT_wordcount_0.1.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!