tHiveStreamInput
Last updated: 9/30/2026tHiveStreamInput properties for Apache Spark Structured Streaming
These properties are used to configure tHiveStreamInput running in the Spark Structured Streaming Job framework.
The Spark Structured Streaming tHiveStreamInput component belongs to the Databases family.
The streaming version of this component is available in Talend Real-Time Big Data Platform and in Talend Data Fabric.
Basic settings
| Properties | Description |
|---|---|
| External table | Select this check box to read data from an external Hive table. Clear it to read data from an internal Hive table. |
| Hive Storage Configuration | Select the tHiveConfiguration component that provides the Hive connection settings. |
| Schema and Edit schema |
A schema is a row description. It defines the fields (columns) processed by the component. When you create a Spark Job, avoid the reserved word line when naming fields.
Click Edit schema to modify the schema. If you modify a Repository schema, the available options include:
|
| Input source | Select the source of the data from the drop-down list:
|
| Database | Enter the name of the database in the Hive service. This property is available when Hive table or Hive query is selected. |
| Table name | Enter the name of the Hive table from which to read data. This property is available when Hive table or Hive query is selected. |
| Hive Stream Temp View name | Enter the name of the temporary view used to execute a SQL query against a Hive table. This property is available when Hive query is selected. |
| Hive query | Enter the Hive query to use to select the data. The query must reference only one table and must use the Hive Stream Temp View name instead of the original table name. This property is available when Hive query is selected. |
| Path | Enter the path to the files of the external Hive table. This property is available when External table is selected. |
| Format | Select the format of the external table files. This property is available when External table is selected. |
| Enable watermarking | Select this check box to activate watermarking and select the time mode to
be used:
In the Watermark delay field, specify how long Spark waits for late data before finalizing each window. |
Advanced settings
| Properties | Description |
|---|---|
| Register Hive UDF jars |
Add the Hive user-defined function (UDF) jars you want tHiveInput to use. Note that you must define a function alias for each UDF to be used in the Temporary UDF functions table. Once you add one row to this table, click it to display the [...] button and then click this button to display the jar import wizard. Through this wizard, import the UDF jar files you want to use. A registered function is often used in a Hive query that you edit in the Hive Query field in the Basic settings view. Note that this Hive Query field is displayed only when you select Hive query from the Input source list. |
| Temporary UDF functions |
Complete this table to give each imported UDF class a temporary function name to be used in the Hive query in the current tHiveInput component. |
Usage
| Usage guidance | Description |
|---|---|
| Usage rule |
This component is used as a start component and requires an output link. This component should use a tHiveConfiguration component present in the same Job to connect to Hive. This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job. |
| Spark Connection |
You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job. This connection is effective on a per-Job basis. |