tJava properties for Apache Spark Structured Streaming
Last updated: 9/30/2026Use these properties to configure tJava running in the Spark Structured Streaming Job framework.
The Spark Structured Streaming tJava component belongs to the Custom Code family.
The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.
Basic settings
| Properties | Description |
|---|---|
| Schema and Edit schema |
A schema is a row description. It defines the number of fields (columns) to be processed by the component. Depending on your Java code, tJava can pass the input dataset as-is or produce a transformed output dataset.
Click Edit schema to make changes to the schema. If you make changes, the schema automatically becomes built-in.
|
| Code | Enter the Java code to process the incoming DataFrame. The DataFrame is
available using the variable name composed of the component label and the
connection label. For information about the Spark Java API, see the Apache Spark Java API documentation. |
Advanced settings
| Properties | Description |
|---|---|
| Classes code |
Define the classes that you need to use in the code written in the Code field in the Basic settings view. It is recommended to define new classes in this field, instead of in the Code field, so as to avoid eventual exceptions in serialization. |
| Import |
Enter the Java code to import, if necessary, external libraries used in the Code field of the Basic settings view. |
Usage
| Usage guidance | Description |
|---|---|
| Usage rule |
This component is used as an end component and requires an input link. This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job. |
| Code example |
The input dataset is available as ds_input_tJava_1, and the output dataset is available as ds_output_tJava_1. You can modify the input data and pass the modified output dataset to downstream components. |
| Spark Connection |
You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job. This connection is effective on a per-Job basis. |