tSynonymSearch properties for Apache Spark Structured Streaming
Last updated: 9/30/2026These properties are used to configure tSynonymSearch running in the Spark Structured Streaming Job framework.
The Spark Structured Streaming tSynonymSearch component belongs to the Data Quality family.
This component is supported on Local Spark 3.5.x and Databricks/EMR with Spark 3.x.
This component is available in Talend Real-Time Big Data Platform and Talend Data Fabric.
Basic settings
| Properties | Description |
|---|---|
|
Schema and Edit schema |
|
|
Limit of each group |
Type in a number to indicate the maximum display of the reference entries matched to each group of the input data. Each row of input data is recognized as one group by this component. If the entries count exceeds the indicated limit, this component displays the ones scored the highest. For further information about the scores used on the matched entries, see Default schema columns. |
|
Columns to search |
Complete this table to provide parameters used to match the input data and the reference entries in a given index. The columns to be completed are: - Input column: select the column(s) of interest from the input data schema. - Reference output column: select the column(s) from the output data schema to present the matched reference entries found in the given synonym index. Index path: enter the path to the index you need to search in the cluster. The value must be enclosed in double quotation marks. When using Spark Local mode, use a path to
a local folder:
Otherwise, use a path to the folder where the index is stored in HDFS. The path must start with hdfs://. You cannot use a path to a local folder. - Search mode: select the search mode you want to use to match input strings against index strings. For further information about available search modes, see Search modes for Index rules. - Score threshold (available for all modes): set a numerical value above 0.0 by which you want to filter the results. Set the threshold to 0.0 to disable the filter. The score value is returned by the Lucene engine and can be anything above 0. The higher the score is, the higher is the similarity of the match. Use the threshold to remove low scoring matches from the output results. There is no easy way to decide about a good threshold value. It will depend on the input data and the indexed data. - Max edits (based on the Levenshtein algorithm and available for the Match all fuzzy and Match any fuzzy modes): select an edit distance, 1 or 2, from the list. Any terms within the edit distance from the input data are matched. With a max edit distance 2, for example, you can have up to 2 insertions, deletions, or substitutions Fuzzy match gains much in performance with Max edits for fuzzy match. Information noteNote:
Jobs migrated in Talend Studio from older releases run correctly, but results might be slightly different because Max edits for fuzzy match is now used in place of Minimum similarity for fuzzy match. - Word distance (available for the Match partial mode): select from the list the maximum number of words allowed to come inside a sequence of words that may be found in the index, default value is 1. - Limit: type in a number to indicate the maximum reference entries to be matched to each record of the corresponding input column you have selected. |
Usage
| Usage guidance | Description |
|---|---|
| Usage rule |
This component is used as an intermediate step. This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job. |
| Spark Connection |
You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job. This connection is effective on a per-Job basis. |