Configure quality pipelines - Precisely Data Integrity Suite

Data Integrity Suite

Product
Spatial_Analytics
Data_Integration
Data_Enrichment
Data_Governance
Precisely_Data_Integrity_Suite
geo_addressing_1
Data_Observability
Data_Quality
dis_core_foundation
Services
Spatial Analytics
Data Integration
Data Enrichment
Data Governance
Geo Addressing
Data Observability
Data Quality
Core Foundation
ft:title
Data Integrity Suite
ft:locale
en-US
PublicationType
pt_product_guide
copyrightfirst
2000
copyrightlast
2026

Quality pipelines detect and correct duplicates, nonstandard formats, missing values, typos, misplaced characters, and inconsistent case. Pipelines chain steps such as standardization, de-duplication, cleaning, validation, and reformatting. Each step uses output from the previous step. A pipeline transforms source data into clean, standardized data and outputs it to your source dataset, a new dataset, or a file. This helps ensure your data is accurate, consistent, unique, and valid before it reaches its destination.

View all pipelines in a table. Create, edit, delete, rename, or duplicate pipelines.

Go to Quality > Pipelines to view this page.

  • Search: Type part of a pipeline name to filter results.
  • Create Pipeline: Select to create a new pipeline.
  • Refresh: Updates the pipeline list.
  • Name: Select the context menu next to a pipeline name to Edit, Delete, Rename, or Duplicate it. Select Run configurations to create, edit, or run a configuration.
    What's changed: In the new user experience, Edit appears as Open Pipeline. For more information on the new experience, see About the new user experience.
  • Dataset: The dataset your pipeline processes. Select a dataset name to edit it.
  • Status: Indicates whether the pipeline is error-free or invalid.
  • Modified By: Shows who last updated the pipeline.
  • Last Modified: Shows when the pipeline was last updated.
Tip: When you create or edit a pipeline, expand the Suggestions panel to view recommendations. Column suggestions are based on semantic types. Entity suggestions (delimited by the entity bar above column headings) are based on the semantic types of columns in that entity. The panel shows up to 10 suggestions by default. Each suggestion identifies a recommended step and the column it applies to. If no suggestions exist for a column or entity, they are not grouped or categorized.

Examples:

  • Address entity: Add the Verify Address and Geocoding step.
  • Full Name or Company Name columns: Add the Parse Name step.
  • Email or Mobile Phone columns: Add the Parse Email or Parse Phone Number step.
  • Semantic First Name column: Add first name standardization in the Standardize Field step.

Pipeline editor interface: When you edit a pipeline, the pipeline editor initially displays two panels. The upper panel shows steps in the pipeline. You can add, remove, or edit steps in this panel. The lower panel displays sample data in table format. Each column represents a data field while each row represents a sample record. The top row shows name, semantic type, and field type for each data field.

When you select a step in the upper panel, it highlights columns affected by that step. Highlighted columns show data as it appears after the step is complete. While a step is selected, you can select the Transformation Preview checkbox to display sample values in the highlighted columns as they appear before and after the transformation.

Filter rows that match data in a pipeline:

  1. Above the sample data table, in the Column/Data box, choose Data.
  2. In the adjacent Search Field Name box, enter a string that matches the data that you are looking for in the table.

As you type, rows are hidden that do not contain matching data.

When you finish typing the search string, only those rows with data that matches the search string are visible in the table.

You can add a new step anywhere in a pipeline; at the beginning or end, before or after an existing step, or between steps. When you add or edit a step, the pipeline editor expands a settings panel from the right side of the window.

After you configure transformation options, use the Preview button before you save the settings to view the result of the transformation.

When you are satisfied with the results of a transformation, save and include it in the pipeline. Semantic type detection is applied to new columns as they are added by a step. The same semantic type as the original column is normally applied to a copied column. If the analysis detects the semantic type for a new column, it is displayed beneath the column name.

Create quality pipeline

To create a quality pipeline:
  1. Go to Quality > Pipelines.
  2. Select + Create Pipeline.
  3. Select the dataset that you want to use to create the pipeline.
  4. Select Generate Sample or upload sample data.
    • The sample retrieves records sequentially, specified here, starting with the first record in the dataset. You can specify between 1 and 2000 rows. The default value is 100. You may choose to increase this value to capture additional variability from the source data.
    • The Sample Preview and Pipeline Grid page shows 500 records if the sample was generated with 500 or more sample rows.
    What's changed: In the new user experience, Generate Sample appears as Upload Sample. For more information on the new experience, see About the new user experience.
  5. Select the fields that you want to include in your sample and select Generate Sample.
  6. Select Create Pipeline.
A quality pipeline is created and shown with your selected dataset fields. On this page you can view sample data as you add transformation steps. The pipeline page consists of three panes. The upper pane shows transformation steps in the pipeline. The table under that shows the state of the data before or after a transformation step. A third pane expands from the side when you add or edit a transformation step. It exposes options that you can edit for a selected transform.

You can edit the pipeline name by selecting the default name. Add and configure transformation steps as needed.

After the pipeline is created, you can add more fields by using the Add Input option.

Warning:
  • You can only choose the dataset that belongs to the same Agent as the primary input.
  • If an incompatible input is added to the pipeline, you may not be able to configure it in the Run configuration, and an error will be thrown stating that incompatible inputs have been added to the pipeline. This will also be reflected in the data validation errors.

The Add Input interface provides dataset view and filtering options, including curated datasets from multiple sources. You can filter datasets by specific properties to quickly access datasets required for analysis and reporting. For more information, see View and filter catalog assets.

Manage pipeline operations

Go to Quality > Pipelines.
  1. To duplicate a pipeline:
    1. In the Name column, find the pipeline that you want to duplicate. In the Search keyword box, you can enter a keyword to filter pipelines shown in the table.
    2. Select the context menu in the Name column, then select Duplicate Pipeline. This opens the Duplicate Pipeline dialog.
    3. In the New pipeline name box, enter a name for the duplicated pipeline, then select Duplicate.
    4. A duplicate pipeline initially consists of the same dataset and transformation steps as the original pipeline. Configurations are not copied.
    5. After you complete this procedure, the duplicated pipeline appears on the Pipelines page. You can now edit settings for the duplicated pipeline.
  2. To edit a pipeline:
    1. In the Name column, find the pipeline that you want to edit. In the Search box, you can match all or part of a pipeline title to filter pipelines shown in the table.
    2. Select the context menu in the Name column, then select Edit. Completing this procedure opens the pipeline page.
    3. On this page, you can view sample data as you add transformation steps to the pipeline.
    4. The pipeline page consists of three panes. The upper pane shows transformation steps in the pipeline. The table under that shows the state of the data before or after a transformation step. A third pane expands from the side when you add or edit a transformation step. It exposes options that you can edit for a selected transform.
    5. You can now view sample data from the associated dataset and add transformation steps to the pipeline.
  3. To rename a pipeline:
    1. In the Name column, find the pipeline that you want to rename. In the Search keyword box, you can enter a keyword to filter pipelines shown in the table.
    2. Select the context menu in the Name column, then select Rename.
    3. In the edit box that appears in the Name column, type a new name, then press Enter.
    4. A name must start with a letter. It can contain letters, numbers, dashes, and underscores.
  4. To delete a pipeline:
    1. Find the pipeline that you want to delete in the pipelines table. In the Search keyword box, you can enter a keyword to filter pipelines shown in the table.
    2. Select the context menu in the Name column, then select Delete.
    3. Select Yes to confirm the deletion.
Tip: Alternatively, you can perform each pipeline operation as a separate task: duplicate a pipeline, edit a pipeline, rename a pipeline, or delete a pipeline from the Pipelines page using individual step-by-step procedures.

Quality pipeline settings

Pages in the Data Quality Pipeline Settings dialog let you view and edit entities and run configurations for a pipeline. This page is displayed when you select the Settings button on the Pipeline page.
The Dataset entities pane lists entities that are defined in the pipeline. You can select an entity in this pane to view and edit settings.
  • Entity type: Shows the type of entity.
  • Entity name: Specifies the name for the entity that shows above columns on the pipeline page. You can edit the name shown in this field.
  • Map fields for this entity: You can change fields mapped to the entity field types.

The Select a Run Configuration pane shows run configurations that have been defined for this pipeline. Depending on permissions assigned to your role, you can select a run configuration in this pane to view or edit settings for the run configuration.
  • Name: Specifies the name for the run configuration.
  • Connection: Specifies the connection to the source data.
  • Source dataset: This specifies the dataset on which to run the pipeline. By default, the dataset used to build the pipeline shows here. Select the Browse button to choose a dataset in a connected database. Input fields for the pipeline must match input fields in the source dataset schema.
  • Target options:
    • Append to target dataset: Output is appended to columns in the target dataset. You can use this option when the schema of the pipeline output matches the target dataset schema. The pipeline engine verifies that the pipeline output schema and the target dataset schema match before it commences output to the target dataset.
    • Overwrite to target dataset: Output overwrites the target dataset. Choose one of two options from the dropdown list.
      • Truncate Dataset: Empties the dataset, then writes data to the same dataset. This choice deletes all data but preserves the dataset definition. You can use this option when the schema of the pipeline output matches the target dataset schema. The pipeline engine verifies that the pipeline output schema and the target dataset schema match before it commences output to the target dataset.
      • Drop Dataset: Drop the target dataset, then create and write data to a new dataset. This choice recreates the dataset definition on the host server. The pipeline output does not have to match the target dataset schema and no verification is performed by the pipeline engine.
  • Target dataset: This specifies the dataset in which to store the output results from the pipeline. Select the Browse button to choose a dataset in a connected database.
  • Pipeline engine: Specifies the pipeline engine to run pipeline jobs. Select one of the pipeline engines from the list or select the Add button to create a new pipeline engine.

Profiling guidelines for pipeline performance for agent based connections

When you use an agent based connection to profile large datasets in pipelines, follow these best practices and system requirements to ensure performance and stability.

Category Details
Storage requirements
  • Profiling large datasets requires temporary disk usage by the Spark engine.
  • Allocate at least 1 TB of local storage on the Agent virtual machine to manage temporary data spills during profiling and to prevent failure during job runs.
Database table optimization
  • To improve profiling performance, source database tables should be well-optimized.
  • Recommended table optimizations include:
    • Primary Keys (PK)
    • Unique Keys (UK)
    • Appropriate Indexes
Memory requirements
  • The profiling engine demands significant memory, particularly with large and wide datasets.
  • Memory allocation should be adjusted based on the volume of input tables to prevent out-of-memory errors.
Parallel profiling jobs
  • The capacity for running parallel profiling jobs is based on the host machine's configuration (CPU and RAM) and the memory allocated per profiling pipeline.
  • For example, a machine with 64 GB RAM and 16 CPU cores can run a maximum of 2 parallel profiling jobs concurrently, with each job allocated 24 GB of memory.
Performance tuning with CPU
  • Increasing the CPU allocation for the profiling pipeline engine can improve profiling performance.
  • Note: This improvement is most effective when the source tables are equipped with Primary Keys (PK), Unique Keys (UK), and indexes.
  • For example, assigning one CPU core allows the system to run a single task at a time, while two cores enable parallel task runs, reducing the processing time. However, allocated CPU cores should match the host machine's actual cores. Over provisioning or assigning more cores than available can cause resource contention, leading to unexpected behavior or performance issues during runs.