tHiveOutput properties for Apache Spark Structured Streaming | Talend Components for Jobs Help
Skip to main content Skip to complementary content

tHiveOutput properties for Apache Spark Structured Streaming

Last updated: 9/30/2026

These properties are used to configure tHiveOutput running in the Spark Structured Streaming Job framework.

The Spark Structured Streaming tHiveOutput component belongs to the Databases family.

The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.

Information noteImportant: Talend does not support the import of schema for complex data types such as array, struct, and map.

Basic settings

Properties Description

Property Type

Select the way the file path and the schema will be set.

  • Built-In: The file path and the schema will be set locally for this component.

  • Repository: The file details stored centrally in Repository > Metadata will be reused by this component.

    You need to click the [...] button next to it and in the pop-up Repository Content dialog box, select the file to be reused, and all related properties will be automatically filled in.

External table Select this check box to write data to an external Hive table. When selected, the Table property changes to Path, and Spark writes data directly in the configured format without using Hive's table management.
Table / Path When External table is cleared, enter the name of the Hive table to write to. When External table is selected, enter the path to the output directory.
Format Select the file format for the output data. The default format is parquet.
Output mode Append: adds rows to the output table without modifying existing rows.
Information noteRestriction: In Spark Structured Streaming Jobs, tHiveOutput supports only the Append output mode. The Complete and Update modes are not supported for file-based sinks in Structured Streaming.
Checkpoint location Enter the path to the directory where Spark stores the checkpoint data for this streaming query. Checkpointing enables fault tolerance and allows a failed query to resume from where it stopped.
Set trigger Select this check box to configure how often the streaming query processes data. When selected, choose one of the following trigger types:
  • Available now: processes all available data in a single batch, then stops the query.
  • Fixed interval: runs the query at a fixed time interval. Enter the interval duration in the field that appears, for example 10 seconds.
Set query name Select this check box to assign a name to the streaming query. In the Query name field, enter the name.

Usage

Usage guidance Description
Usage rule

This component is used as an end component and requires an input link.

This component should use a tHiveConfiguration component present in the same Job to connect to Hive.

This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job.

Spark Connection

You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job.

This connection is effective on a per-Job basis.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!