tMatchGroup

Creates groups of similar data records in any source data including large volumes of data by using one or several match rules.

tMatchGroup compares columns in both standard input data flows and in Spark input data flows by using matching methods and groups similar encountered duplicates together.

Several tMatchGroup components can be used sequentially to match data against different blocking keys. This will refine the groups received by each of the tMatchGroup components through creating different data partitions that overlap previous data blocks and so on.

In defining a group, the first processed record of each group is the master record of the group. The other records are computed as to their distances from the master records and then are distributed to the due master record accordingly.

This component is not shipped with your Talend Studio by default. You need to install it using the Feature Manager. For more information, see Installing features using the Feature Manager.

Did this page help you?

If you find any issues with this page or its content – a typo, a missing step, or a technical error – please let us know!

Leave your feedback here

tMatchGroup

In this section

Did this page help you?