tPubSubOutput properties for Apache Spark Structured Streaming | Talend Components for Jobs Help
Skip to main content Skip to complementary content

tPubSubOutput properties for Apache Spark Structured Streaming

Last updated: 9/30/2026

These properties are used to configure tPubSubOutput running in the Spark Structured Streaming Job framework.

The Spark Structured Streaming tPubSubOutput component belongs to the Messaging family.

The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.

Basic settings

Properties Description
Schema and Edit schema

A schema is a row description. It defines the fields (columns) processed by the component. When you create a Spark Job, avoid the reserved word line when naming fields.

  • Built-In: create and store the schema locally for this component only.

  • Repository: reuse a schema created and stored in the Repository across projects and Job designs.

Click Edit schema to modify the schema. If you modify a Repository schema, the available options include:

  • View schema: view the schema without modifying it.

  • Change to built-in property: change the schema to Built-in for local changes.

  • Update repository connection: modify the Repository schema and choose whether to propagate the changes to other Jobs.

Google Cloud configuration Enter the name of the subscription to use or create when the selected Topic operation manages a subscription. Pub/Sub publishes messages to the topic; the subscription receives messages from that topic.
Pub/Sub topic Enter the name of the Pub/Sub topic to which you want to publish messages.
Pub/Sub subscription Enter the name of the Pub/Sub subscription associated with the topic to which you want to publish messages.
Topic operation Select how to handle the topic and its subscription:
  • Use existing topic and subscription: publishes messages to an existing topic and uses the existing subscription for that topic.
  • Create topic and subscription if they don't exist: creates the topic and subscription if they do not already exist.
Set trigger Select this check box to configure how often the streaming query processes data. When selected, choose one of the following trigger types:
  • Available now: processes all available data in a single batch, then stops the query.
  • Fixed interval: runs the query at a fixed time interval. Enter the interval duration in the field that appears, for example 10 seconds.
Output mode Select the output mode from the drop-down list:
  • Append: adds new rows without modifying existing rows.
  • Update: writes only the rows that were updated since the last trigger.
Use ordering key Select this check box to enable Pub/Sub message ordering. When enabled, the partitionKey column from the input data is used as the message ordering key.
Checkpoint location Enter the path to the directory where Spark stores checkpoint data for the streaming query.

Usage

Usage guidance Description
Usage rule

This component is used as an end component and requires an input link.

This component publishes serialized messages to a Google Cloud Pub/Sub topic.

This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job.

Spark Connection

You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job.

This connection is effective on a per-Job basis.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!