Skip to main content

How a Talend Job for Apache Spark works

Last updated: 8/4/2026
A Talend Spark Job generates code for either the Spark Batch or Spark Streaming engine, depending on the job type you select.

A Talend Spark Job can be run in the following modes and distributions:

  • Cloudera Data Engineering
  • Cloudera Private Cloud
  • Cloudera Public Cloud
  • Databricks
  • Availability-noteBeta
    Databricks Serverless
  • Dataproc
  • EMR
  • EMR Serverless
  • HDInsight
  • Kubernetes
  • Livy Knox
  • Local
  • Spark-submit scripts
  • Standalone
  • Synapse
  • Yarn cluster

For details on configuring each distribution, see Defining connection details in Spark configuration.

Information noteNote: A Talend Spark Job is not equivalent to a Spark Job as defined in the Apache Spark documentation; a single Job can generate one or more Spark Jobs depending on how you design it. For more information, see Apache Spark glossary.

When you run the Job, statistics are displayed in the design workspace to indicate the progress of the Spark computations.

The execution information of a Talend Spark Job is logged by the HistoryServer service of the cluster being used. To view this information, open the service web console. The name of the Job in the console is automatically constructed to be ProjectName_JobName_JobVersion, for example, LOCALPROJECT_wordcount_0.1.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!