tMap properties for Apache Spark Structured Streaming | Talend Components for Jobs Help
Skip to main content Skip to complementary content

tMap properties for Apache Spark Structured Streaming

Last updated: 9/30/2026

Use these properties to configure tMap running in the Spark Structured Streaming Job framework.

The Spark Structured Streaming tMap component belongs to the Processing family.

This component is available in Talend Real-Time Big Data Platform and Talend Data Fabric.

Basic settings

Properties Description
Map editor Use the map editor to define the tMap routing and transformation properties.

When you click the Property Settings button at the top of the input area, a Property Settings dialog box is displayed in which you can set the following parameters:

  • If you do not want to handle execution errors, select the Die on error check box (selected by default). It will kill the Job if there is an error.

  • To maximize the data transformation performance in a Job that handles multiple lookup input flows with large amounts of data, you can select the Lookup in parallel check box.

  • Temp data directory path: enter the path where you want to store the temporary data generated for lookup loading. For more information on this folder, see Solving memory limitation issues in tMap use.

  • Max buffer size (nb of rows): enter the size of physical memory, in number of rows, you want to allocate to processed data.

Mapping links display as Select how mapping links are displayed:
  • Auto: the default setting displays links as curves.
  • Curves: the mapping displays as curves.
  • Lines: the mapping displays as straight lines. The Lines option can slightly improve performance.
Use replicated join (when all lookup tables can fit in memory) Select this check box to perform a replicated join between the input flows. By replicating each lookup table into memory, this type of join does not require an additional shuffle-and-sort step, thus speeding up the whole process.

Ensure that all lookup tables fit in memory.

Usage

Usage guidance Description
Usage rule

This component is used as an intermediate step.

This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job.

Note that in this documentation, unless otherwise explicitly stated, a scenario presents only Standard Jobs, that is to say traditional Talend data integration Jobs.

Spark Connection

You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job.

This connection is effective on a per-Job basis.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!