tPatternMasking properties for Apache Spark Structured Streaming
Last updated: 9/30/2026These properties are used to configure tPatternMasking running in the Spark Structured Streaming Job framework.
The Spark Structured Streaming tPatternMasking component belongs to the Data Quality family.
This component is supported on Local Spark 3.5.x and Databricks/EMR with Spark 3.x.
This component is available in Talend Real-Time Big Data Platform and Talend Data Fabric.
Basic settings
| Properties | Description |
|---|---|
|
Schema and Edit Schema |
|
|
Modifications |
Define in the table what fields to change and how to change them: Column to mask: Select the column from the input flow for which you want to generate similar data by modifying its values. You can mask data from different columns but you need to follow the order of the fields you want to mask. Each column is processed sequentially, meaning that data masking operations will be performed on the data from the first column, the second column, and so on. In a colum, each data field is a fixed length field, except the last data field. For fixed length fields, each value must contain the same number of characters, for example: "30001,30002,30003" or "FR,EN". In a column, the last Enumeration or Enumeration from file data field is a variable length field. For variable length fields, each value might not always contain the same number of characters, for example: "30001,300023,30003" or "FR,ENG".
Field type: Select the field type the data belongs to.
In the Values, Path, Range and Date Range, values must be enclosed in double quotes. When the input data is invalid, meaning that a value does not match the pattern defined in the component, the generated value is null. |
Usage
| Usage guidance | Description |
|---|---|
| Usage rule |
This component is used as an intermediate step. This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job. |
| Spark Connection |
You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job. This connection is effective on a per-Job basis. |