Defining Spark Batch connection details
Last updated: 9/30/2026Complete the Spark Universal connection configuration in the Spark configuration tab of the Run view of your Job. This configuration is effective on a per-Spark Batch Job basis.
Talend Studio allows you to run Spark Jobs on a Spark Universal distribution.
For Spark Batch Jobs, the Runtime mode/environment list is ordered as shown in the following table, which also provides connection details for each mode or environment:
| Mode or environment | Description |
|---|---|
| Amazon EMR |
Talend Studio submits Jobs and collects the execution information of your Job from Amazon EMR
service. For more information, see Defining Amazon EMR connection parameters with Spark Universal. |
| Amazon EMR Serverless |
Talend Studio submits Jobs and collects the execution information of your Job from Amazon EMR
Serverless service. For more information, see Defining Amazon EMR Serverless connection parameters with Spark Universal. |
| Azure HDInsight |
Talend Studio submits Jobs and collects the execution information of your Job from
Azure HDInsight. For more information, see Defining Azure HDInsight connection parameters with Spark Universal. |
| Azure Synapse |
Talend Studio submits Jobs and collects the execution information of your Job
from Azure Synapse Analytics. For more information, see Defining Azure Synapse Analytics connection parameters with Spark Universal. |
| Cloudera Data Engineering |
Talend Studio submits Jobs and collects the execution information of your Job from
Cloudera Data Engineering service. For more information, see Defining Cloudera Data Engineering connection parameters with Spark Universal. |
| Cloudera on Cloud |
Talend Studio submits Jobs and collects the execution information of your Job from
Cloudera on Cloud. For more information, see Defining Cloudera on Cloud connection parameters with Spark Universal. |
| Cloudera on Premises |
Talend Studio submits Jobs and collects the execution information of your Job from
Cloudera on Premises. For more information, see Defining Cloudera on Premises connection parameters with Spark Universal. |
| Databricks |
Talend Studio submits Jobs and collects the execution information of your Job from
Databricks. The Spark driver runs either on a Job Databricks cluster or on an
all-purpose Databricks cluster on GCP, AWS, or Azure. For more information, see Defining Databricks connection parameters with Spark Universal. |
| Databricks Serverless |
Talend Studio submits Jobs and collects the execution information of your Job
from Databricks Serverless compute. Serverless compute is fully managed by
Databricks and requires no cluster configuration. Available on AWS and Azure
only. For more information, see Defining Databricks Serverless connection parameters with Spark Universal. |
| Google Managed Service for Apache Spark (formerly Dataproc) |
Talend Studio submits Jobs and collects the execution information of your Job
from Google Managed Service for Apache Spark. For more information, see Defining Google Managed Service for Apache Spark (formerly Dataproc) connection parameters with Spark Universal. |
| Kubernetes |
Talend Studio submits Jobs and collects the execution information of your Job
from Kubernetes. The Spark driver runs on the cluster managed by
Kubernetes and can run independently from Talend Studio. For more information, see Defining Kubernetes connection parameters with Spark Universal. |
| Local |
Talend Studio builds the Spark environment in itself at runtime to run the Job
locally in Talend Studio. With this mode, each processor of the local machine is used as a
Spark worker to perform the computations. For more information, see Defining Local connection parameters with Spark Universal. |
| Livy Knox |
Talend Studio submits Jobs and collects the execution information of your Job from Livy
Knox. For more information, see Defining Livy Knox connection parameters with Spark Universal. |
| Spark-submit scripts |
Talend Studio submits Jobs and collects the execution information of your Job
from Yarn and ApplicationMaster of your cluster, typically an HPE
Data Fabric cluster. The Spark driver runs on the cluster and can
run independently from Talend Studio. For more information, see Defining Spark-submit scripts connection parameters with Spark Universal |
| Standalone |
Talend Studio connects to a Spark-enabled cluster to run the Job from this
cluster. For more information, see Defining Standalone connection parameters with Spark Universal. |
| Yarn cluster |
Talend Studio submits Jobs and collects the execution information of your Job
from Yarn and ApplicationMaster. The Spark driver runs on the
cluster and can run independently from Talend Studio. For more information, see Defining Yarn cluster connection parameters with Spark Universal. |