The Data Quality service within the Data Integrity Suite plays a vital role in ensuring the success of your data upload and integration processes. It provides you with the tools necessary to identify, analyze, and rectify any issues present in the data you upload. By doing so, this service enables you to guarantee that data originating from various sources is reliable, consistent, and accurate.
The Data Quality service helps maintain the integrity of your data by:
- Validating data against defined rules and standards.
- Correcting inconsistencies such as duplicate records, missing fields, or incorrect formats.
- Enriching data by filling gaps or appending additional information from external sources.
- Standardizing formats to ensure uniformity across datasets.
Data Quality configuration guide
The service provides a robust library of features that transform raw business data into actionable insights, uncovering the "who," "where," and "why" behind operations. Users can effortlessly connect to auto-cataloged data sources, create or import sample data to enhance data quality, and design processes in the cloud with an intuitive interface. Experience real-time data dynamics during the design phase and implement rules seamlessly across diverse environments for streamlined execution.
- Select the datasource type: Choose a datasource that will be the target of your data quality services. Select the one where your essential data is stored.
- Establish and catalog a connection: After setting up the datasource, establish a connection and catalog it to guarantee that data quality can access it as required. This step is crucial for maintaining a reliable link to your data assets.
- Create quality pipeline: Design and develop a data quality pipeline that orchestrates the flow of data through various quality checks and transformation processes.
- Configure transformation steps: Implement specific transformation rules and operations that clean, standardize, and transform the incoming data. This might include tasks such as removing duplicates, correcting erroneous entries, converting data formats, and so on.
- Apply run configuration and validate pipeline: Run the data quality pipeline using a run configuration that specifies parameters such as batch size, execution schedule, and error handling strategies. This configuration ensures that the pipeline runs efficiently and effectively, producing reliable and consistent outputs.
- Manage quality jobs: Supervise quality jobs to ensure optimal performance and results. This includes overseeing the job execution, resource allocation, and troubleshooting any issues that arise during the data quality processes.