Subcapability 04 of 05 · Data & Analytics

Data Quality & Master Data

Six data quality dimensions and MDM Hub Architecture that eliminate the three versions of the same customer across systems and transform data from a source of debate into a verifiable base for every executive decision.

What is at stake

The CRM system lists 45,000 customers. The ERP lists 38,000. The BI platform lists 42,000. Which number is real? That question consumes management capacity every week and has no answer while three versions of the same customer coexist without a unification process with a formal owner.

What it is, in practice

A company where the same business question receives different answers depending on which system or area was consulted has a data problem, not a communication problem. Monthly revenue calculated by sales, finance and the ERP system rarely coincides when there is no active Master Data Management (MDM). The alignment meeting to discover which number is correct is a recurring, invisible operational cost that also delays the decision itself.

How we work

Measurable gains

What changes in the result when this subcapability matures.

Frequently asked questions

What is MDM and why do data unification projects usually fail?

MDM, Master Data Management, is the discipline that establishes the authoritative version of critical business entities with a governance process to maintain that version over time. Unification projects fail because they treat MDM as an IT project with a delivery date, not as a continuous discipline with a business owner. When the project ends without a formal owner responsible for maintaining the golden record, data diverges again within six months.

What are the six data quality dimensions?

Accuracy verifies that data reflects business reality. Completeness verifies that required fields are present. Consistency verifies that the same data has the same value across different systems. Timeliness verifies that data is updated within the timeframe needed for the decision. Validity verifies that data is within the accepted domain of values. Uniqueness verifies that there are no duplicates of the same entity. Each dimension has a measurable KPI and an operational threshold.

What is data observability and how does it differ from data quality?

Data quality measures the six dimensions of data at a point in time. Data observability monitors data behavior continuously over time, detecting anomalous changes in distribution, volume and schema before they reach the consumer. The two are complementary: quality defines the expected standard, observability detects when the standard changes.

Does low-quality data really impact AI results?

Yes. An AI model learns patterns from training data. Data with untreated null fields, entity duplicates, inconsistency across systems and out-of-domain values produces patterns that do not reflect business reality. The model learns noise along with signal. In production, the incorrect prediction or classification appears with the same confidence as the correct one, making the error difficult to detect.

How to prioritize which datasets to address first?

Priority follows the financial impact of the decision the dataset feeds. Datasets that feed revenue decisions, such as pricing, demand forecasting and credit scoring, have the highest direct impact. Within those datasets, priority goes to the quality dimension with the largest gap relative to the minimum threshold for the use case. Data profiling makes this mapping before any treatment initiative.

Want clarity on where to invest first?

A complete technology capability assessment with an evolution roadmap connected to financial result.