tGSConfiguration properties for Apache Spark Structured Streaming | Talend Components for Jobs Help
Skip to main content Skip to complementary content

tGSConfiguration properties for Apache Spark Structured Streaming

Last updated: 9/30/2026

Use these properties to configure tGSConfiguration running in the Spark Structured Streaming Job framework.

The Spark Structured Streaming tGSConfiguration component belongs to the Storage family.

The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.

Basic settings

Properties Description

Property Type

Select the way the connection details will be set.

  • Built-In: The connection details will be set locally for this component. You need to specify the values for all related connection properties manually.

  • Repository: The connection details stored centrally in Repository > Metadata will be reused by this component.

    You need to click the [...] button next to it and in the pop-up Repository Content dialog box, select the connection details to be reused, and all related connection properties will be automatically filled in.

Set the following connection properties:

Properties Description

Project ID

Enter the ID of your Google Cloud Platform project.

If you are not certain about your project ID, confirm it in the Manage Resources page of your Google Cloud Platform services.

Google Storage bucket

Enter the name of the bucket to be used by the whole Job. Then the File components such as tFileInputDelimited or tFileOutputDelimited use the directories in this bucket.

For example, if you enter my_bucket in this field and enter /user/ychen in the Folder field of tFileInputDelimited, tFileInputDelimited reads data from gs://my_bucket/user/ychen.

Temp Folder

Enter the path to the temporary folder that the component uses for intermediate data processing. The default value is /tmp.

Set the following authentication properties:

Properties Description

Path to Google Credentials file

Enter the path to the credentials file associated to the user account to be used. This file must be stored in the machine in which your Talend Job is actually launched and executed.

If you use Talend JobServer to run your Job, store the credentials file not only in the machine of the Talend JobServer, in which the Job is launched, but also in the worker machines of the Spark cluster, in which the Job is executed; if you do not use the Talend JobServer, store the credentials file in your local machine from which you launch the Job and in the worker machines of the Spark cluster.

Use P12 credentials file format

When the Google credentials file to be used is in P12 format, select this check box and then in the Service account Id field that is displayed, enter the ID of the service account for which this P12 credentials file has been created.

Usage

Usage guidance Description
Usage rule

You can use multiple tGSConfiguration components in one Job to provide connection configuration to Google Storage.

From R2022-12 onwards of Talend Studio 8.0, tGSConfiguration supports Spark Universal in the Local mode and Databricks on Google Cloud Platform.

This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job.

Spark Connection

You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job.

This connection is effective on a per-Job basis.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!