tJapaneseTokenize properties for Apache Spark Structured Streaming
Last updated: 9/30/2026These properties are used to configure tJapaneseTokenize running in the Spark Structured Streaming Job framework.
The Spark Structured Streaming tJapaneseTokenize component belongs to the Data Quality family.
This component is supported on Local Spark 3.5.x and Databricks/EMR with Spark 3.x.
The component in this framework is available in Talend Data Management Platform, Talend Big Data Platform, Talend Real-Time Big Data Platform, Talend Data Services Platform, and in Talend Data Fabric.
Basic settings
| Properties | Description |
|---|---|
|
Schema and Edit Schema |
|
|
Tokenization |
The columns from the output schema are added to the Column column in the Tokenization table. For each of the schema columns containing Japanese text to be tokenized, select the corresponding check box in the Tokenize column. You can select the check box in the header row to select all schema columns. |
Usage
| Usage guidance | Description |
|---|---|
| Usage rule |
This component is used as an intermediate step. This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job. |
| Spark Connection |
You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job. This connection is effective on a per-Job basis. |