tJava properties for Apache Spark Structured Streaming | Talend Components for Jobs Help
Skip to main content Skip to complementary content

tJava properties for Apache Spark Structured Streaming

Last updated: 9/30/2026

Use these properties to configure tJava running in the Spark Structured Streaming Job framework.

The Spark Structured Streaming tJava component belongs to the Custom Code family.

The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.

Basic settings

Properties Description
Schema and Edit schema

A schema is a row description. It defines the number of fields (columns) to be processed by the component. Depending on your Java code, tJava can pass the input dataset as-is or produce a transformed output dataset.

  • Built-In: You create and store the schema locally for this component only.

  • Repository: You have already created the schema and stored it in the Repository. You can reuse it in various projects and Job designs.

Click Edit schema to make changes to the schema. If you make changes, the schema automatically becomes built-in.

  • View schema: choose this option to view the schema only.

  • Change to built-in property: choose this option to change the schema to Built-in for local changes.

  • Update repository connection: choose this option to change the schema stored in the repository and decide whether to propagate the changes to all the Jobs upon completion.

    If you just want to propagate the changes to the current Job, you can select No upon completion and choose this schema metadata again in the Repository Content window.

Code Enter the Java code to process the incoming DataFrame. The DataFrame is available using the variable name composed of the component label and the connection label.

For information about the Spark Java API, see the Apache Spark Java API documentation.

Advanced settings

Properties Description
Classes code

Define the classes that you need to use in the code written in the Code field in the Basic settings view.

It is recommended to define new classes in this field, instead of in the Code field, so as to avoid eventual exceptions in serialization.

Import

Enter the Java code to import, if necessary, external libraries used in the Code field of the Basic settings view.

Usage

Usage guidance Description
Usage rule

This component is used as an end component and requires an input link.

This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job.

Code example

The input dataset is available as ds_input_tJava_1, and the output dataset is available as ds_output_tJava_1. You can modify the input data and pass the modified output dataset to downstream components.

// Print the input stream to the console.
      // ds_input_tJava_1.writeStream()
      //     .outputMode("append")
      //     .format("console")
      //     .option("truncate", "false")
      //     .option("numRows", 50)
      //     .start().awaitTermination();

      // Add a column containing the uppercase value of newColumn1 and pass the result downstream.
      // Add import static org.apache.spark.sql.functions.*; to the IMPORT section.
      ds_output_tJava_1 = ds_input_tJava_1.withColumn(
         "newColumn1",
         upper(col("newColumn1"))
      );
Spark Connection

You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job.

This connection is effective on a per-Job basis.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!