How a Talend Job for Apache Spark works
Last updated: 8/4/2026
A Talend
Spark Job generates code for either the Spark Batch or Spark Streaming engine, depending
on the job type you select.
A Talend Spark Job can be run in the following modes and distributions:
- Cloudera Data Engineering
- Cloudera Private Cloud
- Cloudera Public Cloud
- Databricks
- Databricks Serverless
- Dataproc
- EMR
- EMR Serverless
- HDInsight
- Kubernetes
- Livy Knox
- Local
- Spark-submit scripts
- Standalone
- Synapse
- Yarn cluster
For details on configuring each distribution, see Defining connection details in Spark configuration.
Information noteNote: A Talend Spark Job is not equivalent to a Spark Job as defined in the Apache Spark
documentation; a single Job can generate one or more Spark Jobs depending on how you
design it. For more information, see Apache Spark glossary.
When you run the Job, statistics are displayed in the design workspace to indicate the progress of the Spark computations.
The execution information of a Talend Spark Job is logged by the HistoryServer service of the cluster being used. To view this information, open the service web console. The name of the Job in the console is automatically constructed to be ProjectName_JobName_JobVersion, for example, LOCALPROJECT_wordcount_0.1.