tAvroInput properties for Apache Spark Structured Streaming | Talend Components for Jobs Help
Skip to main content Skip to complementary content

tAvroInput properties for Apache Spark Structured Streaming

Last updated: 9/30/2026

Use these properties to configure tAvroInput running in the Spark Structured Streaming Job framework.

The Spark Structured Streaming tAvroInput component belongs to the File family.

The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.

Basic settings

Properties Description
Define a storage configuration component

Select the configuration component to be used to provide the configuration information for the connection to the target file system such as HDFS.

If you leave this check box clear, the target file system is the local system.

The configuration component to be used must be present in the same Job. For example, if you have dropped a tHDFSConfiguration component in the Job, you can select it to write the result in a given HDFS system.

Property Type

Select the way the file path and the schema will be set.

  • Built-In: The file path and the schema will be set locally for this component.

  • Repository: The file details stored centrally in Repository > Metadata will be reused by this component.

    You need to click the [...] button next to it and in the pop-up Repository Content dialog box, select the file to be reused, and all related properties will be automatically filled in.

Schema and Edit schema

A schema is a row description. It defines the number of fields (columns) to be processed and passed on to the next component. When you create a Spark Job, avoid the reserved word line when naming the fields.

  • Built-In: You create and store the schema locally for this component only.

  • Repository: You have already created the schema and stored it in the Repository. You can reuse it in various projects and Job designs.

Click Edit schema to make changes to the schema. If you make changes, the schema automatically becomes built-in.

  • View schema: choose this option to view the schema only.

  • Change to built-in property: choose this option to change the schema to Built-in for local changes.

  • Update repository connection: choose this option to change the schema stored in the repository and decide whether to propagate the changes to all the Jobs upon completion.

    If you just want to propagate the changes to the current Job, you can select No upon completion and choose this schema metadata again in the Repository Content window.

Folder/File Browse to or enter the path to the Avro data to read.

If the path points to a folder, the component reads all Avro files in that folder. Spark ignores subfolders unless you configure recursive reading in the Spark settings.

To specify multiple paths, separate each path with a comma (,).

To read from cloud storage, add the corresponding configuration component to the Job. For example, use tS3Configuration for Amazon S3, tGSConfiguration for Google Cloud Storage, or tAzureFSConfiguration for Azure Data Lake Storage.

Die on error

Select the checkbox to stop the execution of the Job when an error occurs.

Clear the checkbox to skip any rows on error and complete the process for error-free rows. When errors are skipped, you can collect the rows on error using a Row > Reject link.

Advanced settings

Properties Description
Set minimum partitions

Select this check box to control the number of partitions to be created from the input data over the default partitioning behavior of Spark.

In the displayed field, enter, without quotation marks, the minimum number of partitions you want to obtain.

When you want to control the partition number, you can generally set at least as many partitions as the number of executors for parallelism, while bearing in mind the available memory and the data transfer pressure on your network.

Use hierarchical mode

Select this check box to map the binary (including hierarchical) Avro schema to the flat schema defined in the schema editor of the current component. If the Avro message to be processed is flat, leave this check box clear.

Once selecting it, you need set the following parameter(s):

  • Mapping: create the map between the schema columns of the current component and the data stored in the hierarchical Avro message to be handled. In the Node column, you need to enter the JSON path pointing to the data to be read from the Avro message.

Usage

Usage guidance Description
Usage rule

This component is used as a start component and requires an output link.

This component provides a static dataset and is intended to be used as a lookup input for tMap only.

This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job.

Spark Connection

You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job.

This connection is effective on a per-Job basis.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!