tDataMasking properties for Apache Spark Structured Streaming
Last updated: 9/30/2026These properties are used to configure tDataMasking running in the Spark Structured Streaming Job framework.
The Spark Structured Streaming tDataMasking component belongs to the Data Quality family.
This component is supported on Local Spark 3.5.x and Databricks/EMR with Spark 3.x.
This component is available in Talend Real-Time Big Data Platform and Talend Data Fabric.
Basic settings
| Properties | Description |
|---|---|
|
Schema and Edit Schema |
|
|
Modifications |
Define in the table what fields to change and how to change them: Input Column: Select the column from the input flow that contains the data to be masked. The supported data types are: Date, Double, Float, Integer, Long and String. These modifications are based on the function you select in the Function column.
Category: select a category of
masking functions from the list.
Function: Select the function that will hide or obfuscate the original data with substitutes. For example, you can replace digits or letters with the substitute of your choice, replace values with synonyms from an index file or nullify values. The functions you can select from the Function list depend on the data type of the input column. For example, if the column type is Long, you can use the Numeric variance function. If the column type is String, the Numeric variance function will not be available. Also, the Function list for a Date column is date-specific, it allows you to decide the type of modification you want to do on date values. Method: Select the Basic method or one FF1 algorithm (Format-Preserving Encryption (FPE)), FF1 with AES or FF1 with SHA-2: The Basic method is the default algorithm. Information noteNote: As the masking methods are stronger, it is recommended to use the FF1
algorithms rather than the Basic method.
The FF1 with AES method is based on the Advanced Encryption Standard in CBC mode. The FF1 with SHA-2 method depends on the secure hash function HMAC-256. Information noteNote: Java 8u161 is the minimum
required version to use the FF1 with AES method.
To be able to use this FPE method with Java versions earlier than 8u161, download the
Java Cryptography Extension (JCE) unlimited strength jurisdiction policy files from
Oracle website.
The FF1 with AES and FF1 with SHA-2 methods require a password to be specified in the Password or 256-bit key for FF1 methods field of the Advanced settings to generate unique masked values. The Alphabet list is only available for functions that use Format-Preserving Encryption algorithms. When using the Character handling functions, such as Replace all, Replace characters between two positions, Replace all digits with FPE methods, you must select an alphabet. Characters that belong to the selected alphabets are masked with characters from the same character type within the selected alphabet. When selecting the Best guess alphabet, masked values contain characters from all alphabets represented in the input values. Best guess is the default alphabet. Any unrecognized character is copied to the output as is. Extra Parameter: This field is used by some of the functions, it will be disabled when not applicable. When applicable, enter a number or a letter to decide the behavior of the function you have selected. When you set Function to
Generate from file/list, define the file path in Extra
Parameter. Set the file path as follows:
Keep format: this function is only used on Strings. Select this check box to keep the input format when using the Generate account number and keep original country, Generate credit card number and keep original bank, Bank Account Masking, Credit Card Masking, Phone Masking and SSN Masking functions or categories. That is to say, if there are spaces, dots ('.'), hyphens ('-') or slashes ('/') in the input, those characters are kept in the output. If you select this check box when using Phone Masking functions, the characters that are not numbers from the input are copied to the output as is. |
Usage
| Usage guidance | Description |
|---|---|
| Usage rule |
This component is used as an intermediate step. This component, along with the Spark Structured Streaming component Palette it belongs to, appears only when you are creating a Spark Structured Streaming Job. |
| Spark Connection |
You need to use the Structured Streaming Configuration tab in the Run view to define the connection to a Spark cluster for the whole Job. This connection is effective on a per-Job basis. |