Defining Databricks connection parameters with Spark Structured Streaming | Talend Studio Help
Skip to main content Skip to complementary content

Defining Databricks connection parameters with Spark Structured Streaming

Last updated: 9/30/2026

Configure a Databricks connection for a Spark Structured Streaming Job.

Procedure

  1. Click the Run view beneath the design workspace, then click the Structured Streaming configuration view.
  2. Select Built-in from the Property type list.
  3. Select Universal from the Distribution list, select the Spark version, then select Databricks from the Runtime mode/environment list.
  4. Select the Define a streaming timeout (ms) check box and, in the field that is displayed, enter the time frame at the end of which the Spark Structured Streaming Job automatically stops running.
  5. Complete the Databricks configuration parameters.
    Parameter Usage
    Cloud provider Select the cloud provider to be used between AWS, Azure and GCP.
    Run mode Select the mode you want to use to run your Job on Databricks cluster when you execute your Job in Talend Studio. With Create and run now, a new Job is created and run immediately on Databricks and with Runs submit, a one-time run is submitted without creating a Job on Databricks.
    Enable Unity Catalog Select this check box to leverage Unity Catalog. Then, you need to specify the Unity Catalog related information in the Catalog, Schema, and Volume parameters.
    Information noteImportant: All the parameters have to be created on Databricks with granted permissions for all authorized users before using them in Talend Studio.
    Use pool You can select this check box to leverage a Databricks pool. If you do, you must indicate the Pool ID instead of the Cluster ID. You must also select Job clusters from the Cluster type drop-down list.
    Endpoint Enter the URL address of your workspace.
    Cluster ID Enter the ID of the Databricks cluster to be used. This ID is the value of the spark.databricks.clusterUsageTags.clusterId property of your Spark cluster. You can find this property on the properties list in the Environment tab in the Spark UI view of your cluster.
    Authentication mode Select the method to authenticate to Databricks in the drop-down list:
    • Personal access token: authenticate with a personal access token (PAT). For more information, see Databricks personal access token authentication from the Databricks documentation.

      In Authentication token, enter the authentication token generated for your Databricks user account.

    • OAuth2: authenticate with OAuth2. For more information, see Authenticate access to Databricks using OAuth token federation from the Databricks documentation.

      In Client ID and Secret ID, enter the OAuth client credentials generated for your Databricks service principal.

    Dependencies folder Enter the directory that is used to store your Job related dependencies on Databricks Filesystem at runtime, putting a slash (/) at the end of this directory. For example, enter /jars/ to store the dependencies in a folder named jars. This folder is created on the fly if it does not exist then.

    From Databricks 15.4 LTS, the default library location is moved to WORKSPACE, instead of DBFS.

    Project ID Enter the ID of your Google Platform project where the Databricks project is located.

    This field is only available when you select GCP from the Cloud provider drop-down list.

    Bucket Enter the name of the bucket you use for Databricks from Google Platform.

    This field is only available when you select GCP from the Cloud provider drop-down list.

    Workspace ID Enter the ID of your Google Platform workspace respecting the following format: databricks-workspaceid.

    This field is only available when you select GCP from the Cloud provider drop-down list.

    Google credentials Enter the directory in which the JSON file containing your service account key is stored in the Talend JobServer machine.

    This field is only available when you select GCP from the Cloud provider drop-down list.

    Poll interval when retrieving Job status (in ms) Enter the time interval (in milliseconds) at the end of which you want Talend Studio to ask Spark for the status of your Job.
    Cluster type From the drop-down list, select the type of cluster you want to use. For more information, see About Databricks clusters.
    Information noteNote: When you run the Job using Talend Studio with Java 17, you need to set the JNAME=zulu17-ca-amd64 environment variable:
    • on Databricks side for job clusters
    • in Init scripts using the set_java17_dbr.sh script on S3 for all-purpose clusters

    DBFS is no longer supported as Init scripts location. For all versions of Databricks, it is replaced to WORKSPACE.

    Use policy Select this check box to enter the name of the policy to be used by your Job cluster. You can use a policy to limit the ability to configure clusters based on a set of rules.

    For more information about cluster policies, see Manage cluster policies from the official Databricks documentation.

    Use DBFS to upload dependencies Select this check box to upload dependencies to DBFS.
    Enable ACL

    Select this check box to use access control lists (ACLs) to configure permission to access workspace or account level objects.

    In ACL permission, you can configure permission to access workspace objects with CAN_MANAGE, CAN_MANAGE_RUN, IS_OWNER, or CAN_VIEW.

    In ACL type, you can configure permission to use account-level objects with User, Group, or Service Principal.

    In Name, enter the name you were given by the administrator.

    This option is available when Cluster type is set to Job clusters. For more information, see the Databricks documentation.

    Autoscale Select or clear this check box to define the number of workers to be used by your Job cluster. If you select this check box, autoscaling is enabled. Then define the minimum number of workers in Min workers and the maximum number of workers in Max workers. Your Job cluster is scaled up and down in this scope based on its workload.
    • If you select this check box, autoscaling is enabled. Then define the minimum number of workers in Min workers and the maximum number of workers in Max workers. Your Job cluster is scaled up and down in this scope based on its workload.

      According to the Databricks documentation, autoscaling works best with Databricks runtime versions 3.0 or onwards.

    • If you clear this check box, autoscaling is deactivated. Then define the number of workers a Job cluster is expected to have. This number does not include the Spark driver node.
    Node type and Driver node type Select the node types for the workers and the Spark driver node. These types determine the capacity of your nodes and their pricing by Databricks.

    For more information about these node types and the Databricks Units they use, see Supported Instance Types from the Databricks documentation.

    Enable credentials passthrough Select this check box to disable user credential passthrough when connecting to Databricks Universal. When this option is selected, users' individual credentials are not used for authentication to data sources.
    Configure cluster logs Select this check box to define where to store your Spark logs for a long term.
    Custom tags Select this check box to add custom tags as key-value pairs to your Databricks resources.
    Init scripts DBFS is no longer supported as Init scripts location. For all versions of Databricks, it was replaced to WORKSPACE.
    Do not restart the cluster when submitting Select this check box to prevent Talend Studio restarting the cluster when Talend Studio is submitting your Jobs. However, if you make changes in your Jobs, clear this check box so that Talend Studio restarts your cluster to take these changes into account.
  6. Select the Set tuning properties check box to define the tuning parameters for your Spark Structured Streaming Job.
    Information noteImportant: You must define the tuning parameters otherwise you can get an error (400 - Bad request).
  7. In the Advanced properties table, add any Spark properties you want to override the defaults set by Talend Studio.

Results

The connection details are complete.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!