Defining Yarn cluster connection parameters with Spark Structured Streaming | Talend Studio Help
Skip to main content Skip to complementary content

Defining Yarn cluster connection parameters with Spark Structured Streaming

Last updated: 9/30/2026

Configure a Yarn cluster connection for a Spark Structured Streaming Job.

Procedure

  1. Click the Run view beneath the design workspace, then click the Structured Streaming configuration view.
  2. Select Built-in from the Property type list.
  3. Select Universal from the Distribution list, select the Spark version, then select Yarn cluster from the Runtime mode/environment list.
  4. In Path to custom Hadoop configuration jar, enter the path to the Hadoop configuration JAR file for your cluster.
  5. Select the Define a streaming timeout (ms) check box and, in the field that is displayed, enter the time frame at the end of which the Spark Structured Streaming Job automatically stops running.
  6. If you run your Spark Job on Windows, specify the location of the winutils.exe program:
    • If you want to use your own winutils.exe file, select the Define the Hadoop home directory check box and enter its folder path.
    • Otherwise, leave the Define the Hadoop home directory check box clear. Talend Studio will generate and use a directory automatically for this Job.
  7. Select Use custom classpath to specify additional classpath entries for the Job.
  8. Enter the authentication information by specifying your username. You can also use Kerberos to authenticate by selecting the Use Kerberos authentication checkbox.
  9. In Username, enter the username required by the cluster.
  10. Select the Set tuning properties check box to define the tuning parameters for your Spark Structured Streaming Job.
    Information noteImportant: You must define the tuning parameters otherwise you can get an error (400 - Bad request).
  11. Select the Enable spark event logging check box to enable the Spark application logs of the Job to be persistent in the file system.
    The parameters relevant to Spark logs are displayed:
    • Compress Spark event logs: Select this check box to compress the logs.
    • Spark event logs directory: Enter the directory in which Spark events are logged.
    • Spark history server address: Enter the location of the history server.

    Your cluster administrator may have set these properties in the configuration files. Contact the administrator to get the exact values.

  12. In the Advanced properties table, add any Spark properties you want to override the defaults set by Talend Studio.

Results

The connection details are complete.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!