tHiveStreamInput | Talend Components for Jobs Help
Skip to main content Skip to complementary content

tHiveStreamInput

Last updated: 9/30/2026
Reads streaming data from Hive.

tHiveStreamInput properties for Apache Spark Structured Streaming

These properties are used to configure tHiveStreamInput running in the Spark Structured Streaming Job framework.

The Spark Structured Streaming tHiveStreamInput component belongs to the Databases family.

The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.

Basic settings

Properties Description
External table Select this check box to read data from an external Hive table. Clear it to read data from an internal Hive table.
Hive Storage Configuration Select the tHiveConfiguration component that provides the Hive connection settings.
Schema and Edit schema

A schema is a row description. It defines the fields (columns) processed by the component. When you create a Spark Job, avoid the reserved word line when naming fields.

  • Built-In: create and store the schema locally for this component only.

  • Repository: reuse a schema created and stored in the Repository across projects and Job designs.

Click Edit schema to modify the schema. If you modify a Repository schema, the available options include:

  • View schema: view the schema without modifying it.

  • Change to built-in property: change the schema to Built-in for local changes.

  • Update repository connection: modify the Repository schema and choose whether to propagate the changes to other Jobs.

Input source Select the source of the data from the drop-down list:
  • Hive table: reads data from a Hive table.
  • Hive query: reads data returned by a Hive query.
Database Enter the name of the database in the Hive service. This property is available when Hive table or Hive query is selected.
Table name Enter the name of the Hive table from which to read data. This property is available when Hive table or Hive query is selected.
Hive Stream Temp View name Enter the name of the temporary view used to execute a SQL query against a Hive table. This property is available when Hive query is selected.
Hive query Enter the Hive query to use to select the data. The query must reference only one table and must use the Hive Stream Temp View name instead of the original table name. This property is available when Hive query is selected.
Path Enter the path to the files of the external Hive table. This property is available when External table is selected.
Format Select the format of the external table files. This property is available when External table is selected.
Enable watermarking Select this check box to activate watermarking and select the time mode to be used:
  • Event time: Select to use a timestamp column representing when each event occurred, then select the column in the Watermark column (timestamp) drop-down list.
  • Processing time: Select to use the time when Spark reads each record.

In the Watermark delay field, specify how long Spark waits for late data before finalizing each window.

Advanced settings

Properties Description
Register Hive UDF jars

Add the Hive user-defined function (UDF) jars you want tHiveInput to use. Note that you must define a function alias for each UDF to be used in the Temporary UDF functions table.

Once you add one row to this table, click it to display the [...] button and then click this button to display the jar import wizard. Through this wizard, import the UDF jar files you want to use.

A registered function is often used in a Hive query that you edit in the Hive Query field in the Basic settings view. Note that this Hive Query field is displayed only when you select Hive query from the Input source list.

Temporary UDF functions

Complete this table to give each imported UDF class a temporary function name to be used in the Hive query in the current tHiveInput component.

Usage

Usage guidance Description
Usage rule

This component is used as a start component and requires an output link.

This component should use a tHiveConfiguration component present in the same Job to connect to Hive.

This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job.

Spark Connection

You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job.

This connection is effective on a per-Job basis.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!