Customer Data Integration: How It Works and How to Start
Customer data integration (CDI) combines customer information from separate systems, standardizes its meaning, and makes it available for analysis and operational use. It applies the broader data integration process described by IBM to customer records. Building customer profiles also requires deduplication, matching, and field selection, so that information from different sources can be associated with the same customer.

What customer data is integrated?
The AWS customer data platform reference architecture combines information from CRM systems, contact centers, email, websites, mobile applications, and point-of-sale transactions through batch inputs and event streams.
SAP’s customer profile documentation describes customer identifiers, profile attributes, consents, and service events. The roles of these data types are:
- Identifiers associate records with customer or source accounts.
- Attributes describe the customer, including contact information.
- Activities record service interactions; transactions capture purchases and order details.
- Processing permissions govern how customer information may be stored and shared.
In Microsoft’s unification workflow, profile tables contain customer information, while purchases and other activities with a one-to-many relationship are kept outside the profile-unification input. The practical modeling lesson is to preserve individual transactions and relate them to profiles, rather than treating repeated purchases as duplicate customers.
What is customer data integration used for?
- Customer service: a consolidated CRM view can present purchase history, order status, and outstanding service issues together.
- Customer analysis: data warehouses consolidate information from multiple sources for historical analysis, reports, and dashboards.
- Segmentation and personalization: the AWS reference architecture uses unified profiles for segmentation, campaign systems, and website or mobile personalization.
Define the required output for each use. For service, specify the records and relationships that must appear in the customer view. For analysis, specify the customer definition and reporting period. For activation, specify the audience fields, destination, and processing permissions. These specifications provide concrete criteria for accepting the integration.
Customer data integration vs. CDP, CRM, and data warehouse
Customer data integration applies the process of combining and harmonizing data to customer information. The systems below serve different roles in that process.
| Term | Main role |
|---|---|
| Customer data integration | The integration process that harmonizes information from multiple sources for analytical and operational use. |
| Customer data platform (CDP) | Under the CDP Institute’s definition, software responsible for maintaining persistent unified customer records and making customer data services available to other systems. |
| Customer relationship management (CRM) | CRM software manages customer relationships and interactions, including contact information, sales opportunities, and service issues. |
| Data warehouse | A central analytical repository that receives data from transactional and other systems for reporting, queries, and analysis. |
The CDP Institute’s description of deployment models allows customer data services to run inside the platform or use external components, including warehouses. When specifying the integration, assign ownership for customer identity, profile structure, governance, and downstream access, alongside the choice of storage location.
Common customer data integration methods
AWS distinguishes ETL from ELT: extract, transform, load prepares data before loading it into the destination; extract, load, transform performs the transformation after loading. Specify the update schedule and required output separately from the transformation order.
| Method | How it works | What to establish |
|---|---|---|
| Batch loading | Collects and loads records periodically, using the batch processing approach described by AWS. | Refresh schedule, reporting cutoff, and handling of missed loads. |
| API and event ingestion | Receives data through application APIs or pushed events, both supported in the AWS reference design. | Source capabilities and the delay acceptable at the destination. |
| Change data capture (CDC) | Captures ongoing database changes; AWS DMS reads database logs for supported sources. | Initial loading, change-stream configuration, and replication lag. |
| Data virtualization | Presents an integrated view without physically extracting, transforming, or loading data. | Query access and where mapping, matching, and access controls will run. |
CDC does not mean instantaneous synchronization. AWS DMS explicitly states that its CDC replication is not real time: latency depends on source workload, network conditions, replication resources, target capacity, and data characteristics. For customer data integration, assess the delay through the full route, including profile processing and delivery to the consuming application.
AWS describes warehouse data as available through BI tools and SQL clients. Specify whether the integration must support historical reporting, current application data, or both, then validate the required outputs and refresh schedules.
How customer data integration works
Map source fields to a common model
Mapping establishes what source fields mean and how they relate. In Salesforce’s mapping workflow, incoming fields are mapped to standardized data model objects before identity resolution creates unified profiles. This separates the structure of incoming records from the structure needed for downstream processing.
Salesforce uses fully qualified keys consisting of a source key and a source qualifier to prevent identical key values from different systems being misinterpreted. Preserve both the identifier’s value and its source context in the mapping.
Standardize and deduplicate records
Microsoft’s unification guidance recommends normalization, deduplication within source tables, and progressive testing of matching rules. Its normalization and deduplication steps address variations in data entry and multiple source rows representing the same customer.
The same guidance warns against matching on names alone or heavily repeated values and recommends testing whether fuzzy matching improves results enough to justify its processing time. Treat normalized values as matching inputs, and validate the resulting record relationships separately.
Resolve customer identity
Identity resolution connects records through methods such as Salesforce’s exact, normalized, and fuzzy matching: equal values, equivalent values after formatting normalization, or specified similarities, respectively.
Define the customer entity before applying those rules. Microsoft supports household grouping separately from individual profiles, assigning a shared cluster identifier to related people. Keep the entity being matched explicit in the data model and in the output definition.
Reconcile conflicting field values
For selecting values in the assembled profile, Microsoft’s field-unification settings support source priority, most-recent, and least-recent selection. Its recency-based selection requires a suitable date or numeric field from every participating table.
Microsoft documents grouping address fields under one merge policy to avoid combining incompatible pieces from different source records. Record the selection policy alongside the mapping, including how missing values and conflicting sources are handled.
Deliver profiles to their destinations
The CDP Institute lists APIs, database queries, and file extracts as access methods for customer data. Specify the destination interface, customer entity, included fields, update requirements, and access rules. Validate the delivered output against that specification, including the data available in the consuming application.
Data quality, permissions, and correction
Microsoft’s troubleshooting documentation recommends checking source accuracy and completeness, primary keys, normalization, and matched records. Its primary-key warning explains that using a demographic field as a table’s key can remove records sharing that value during deduplication. Keep the source record identifier distinct from a field used as matching evidence.
SAP’s processing-purpose mechanism governs both inbound storage and outbound sharing. Its governance enforcement can discard incoming events lacking a required purpose and restrict outgoing attributes, activity indicators, and segment membership according to active purposes.
For implementation, identify the authoritative source of preferences and processing-purpose information, the fields permitted at each destination, and the route for updating restrictions. Treat those as explicit parts of the integration specification. SAP’s documentation describes purpose status and timestamps stored with profiles, providing a concrete mechanism for carrying that context through processing.
The AWS reference design uses least-privilege access and encryption in transit and at rest. Establish equivalent controls for the actual systems holding and receiving customer information.
Microsoft documents profile merges and splits when source data or rules change. Its troubleshooting process calls for correcting source problems, rerunning unification, and validating the result. Preserve the trace from output to source when checking a repair.
How to measure integration quality
Evaluate identity matching against checked record pairs. Applying Google’s precision and recall definitions to customer identity resolution gives two distinct measures:
- Match precision: correct proposed links divided by all proposed links.
- Match recall: correctly found links divided by all true links in the checked set.
For this application, check the reference labels establishing which records belong together. A false positive or false negative represents an incorrect proposed link or a missed true relationship, respectively. Google’s guidance ties the choice of metric to the costs of different errors, so select acceptance criteria for the intended customer-data use.
Include source completeness, rejected records, required fields, and output age in delivery checks. To investigate missing or unexpected records, Microsoft recommends examining source and unification output tables. Specify the expected population, time period, and denominator for each check.
Establish the first integration before expanding it
Use the documented workflows above to define a bounded first release. Write down the customer entity, required sources, field meanings, identity rules, reconciliation policies, processing permissions, and destination output. Assign ownership for source corrections and for the rules that produce the profile.
Then verify the complete route: source arrival, mapping, deduplication, matching, field selection, and delivery. Check traceability and correction as part of that verification. Microsoft recommends adding unification rules progressively and tracking their results; apply the same approach when extending the initial customer view to more sources or uses.