Master Data Management Strategy: A Practical Guide
When researchers asked 75 executives to check 100 recently created records from their own departments for obvious errors, 47% of those records turned out to contain at least one critical mistake. Only 3% of the departments scored acceptably, and that was using the loosest possible standard. The study, published in Harvard Business Review by Tadhg Nagle, Thomas Redman and Dave Sammon, has not been meaningfully contradicted since. Master data management exists to stop exactly that.
Read that alongside a second finding and the picture gets uncomfortable. Redman estimates in MIT Sloan Management Review that the cost of bad data runs to 15% to 25% of revenue for most companies, and that roughly two thirds of that cost can be removed permanently by fixing root causes rather than symptoms.
Two thirds of a quarter of revenue is not a rounding error. It is the size of most companies' entire technology budget. And the largest single root cause, in every organization I have worked with, is the same: the business has no agreed answer to the question of which record is the real one.
That question is what master data management exists to answer. Which customer record is authoritative when sales, finance and support each hold a different version. Which product code the warehouse and the storefront both mean. Which supplier the accounts payable system is actually paying. Master data is the small set of entities that show up everywhere and belong to no single department, and it is precisely because they belong to no one that they rot.
This guide is about building a master data management strategy that survives contact with an actual company. It is not a product comparison. Product comparisons age out in six months and most of the ones you will find were paid for by the vendors listed in them. What follows is the sequence I use: how to pick the first domain, how to build a business case that finance will approve, what the architecture choices really cost, where governance and MDM stop overlapping, and a 90 day plan that produces a measurable result rather than another stalled data initiative.
Why master data management programs stall
Most MDM programs do not fail loudly. They stall quietly, usually somewhere in month nine, and the pattern repeats with enough regularity that it is worth naming the four causes directly.
They start as technology projects. A platform gets selected, an environment gets provisioned, and a team starts loading data into it. Nine months later there is a working system with clean records that no operational process actually reads from. The system becomes a very expensive reporting side effect, because nobody changed the order entry screen, the customer onboarding flow or the supplier setup process to consume it.
They try to master everything at once. Customer, product, supplier, location, employee, chart of accounts, all in the first phase, because the business case was written to justify a large budget and a large budget needs a large scope. Every domain added multiplies the number of source systems, stakeholders and matching rules involved. Programs that start with one domain finish; programs that start with five negotiate.
Nobody owns the definition. The technical team can build the matching engine, but it cannot decide whether a customer is a legal entity, a billing account or a person. That decision belongs to the business, it is genuinely contested between departments, and until someone with authority makes it, every matching rule is a guess dressed up as configuration.
The business case is written in the language of data, not money. "Single source of truth" is not a benefit. It is a mechanism. Finance approves projects that reduce a cost or protect a revenue stream, and if you cannot express what duplicate customer records cost you in specific terms, such as wasted marketing spend, credit exposure that was misjudged, or invoices sent to the wrong entity, the program will be approved reluctantly and defunded at the first budget squeeze.
There is a broader version of this problem worth keeping in view. DalleMule and Davenport observed in Harvard Business Review that less than half of an organization's structured data is actively used in decisions, and less than 1% of its unstructured data is analyzed at all. MDM is the defensive half of a data strategy, the half concerned with control, consistency and a single version of the truth. It only pays off when the offensive half, the part that actually uses data to make money, has something specific it is waiting on.
What master data actually is, and what it is not
Precision here saves months of argument later, because scope creep in MDM almost always enters through a definitional gap.
Master data is the set of core business entities that are referenced across multiple processes and systems, and that persist over time. Customers, products, suppliers, locations, assets, employees, accounts. The defining test is not importance but reuse: if the same entity is created and maintained independently in three systems, it is master data whether or not anyone has called it that.
Transactional data is what happens to master data. Orders, invoices, shipments, payments, tickets. Transactions are high volume, they are immutable once posted, and they are not what MDM manages. They are what MDM protects, because a transaction attached to the wrong customer record is a transaction you will spend money to correct.
Reference data is the controlled vocabulary that master and transactional data both use. Country codes, currency codes, units of measure, industry classifications, internal status codes. Small, slow moving, and disproportionately damaging when inconsistent, because reference data mismatches break joins silently.
Metadata describes all of the above: where a field comes from, what it means, who owns it, how it is calculated. Related discipline, different problem, and confusing the two is a common way to buy the wrong tool.
The practical distinction that matters most: master data is the data you would have to reconcile if you acquired a company tomorrow. If a data set would not appear in that reconciliation exercise, it is probably not master data, and putting it in scope will cost you time you do not have.
One more boundary worth drawing. Master data management and data governance are adjacent, not identical. Governance sets the policies, assigns the ownership and defines the rules. MDM is the operational machinery that enforces those rules on a specific set of entities, day after day, at the moment records are created. You can have governance without MDM, and many organizations do, which is why they have excellent policy documents and terrible customer records. You cannot usefully have MDM without governance, because the matching engine will simply automate whatever inconsistency the policy failed to resolve. The policy layer is covered in depth in the practical guide to building a data governance framework.
Master data management strategy: where to start
This is the decision that determines whether the program delivers anything in year one, and it is usually made badly, by committee, on the basis of which department complained loudest.
The correct method is to pick one domain, and to pick it using three criteria applied in order.
First criterion: where the pain converts to money fastest. Not where the data is worst. Where the bad data is costing something you can name. Duplicate customer records that cause credit limits to be assessed on a fragment of the real exposure. Product data inconsistencies that cause pricing errors at the point of sale. Supplier duplicates that let the same vendor be paid twice. If you cannot name a number, you have not found the right domain yet, you have found the noisiest complaint.
Second criterion: where the source systems are fewest. Every additional system that creates records in a domain adds a matching problem, a stakeholder and an integration. A domain with two source systems and a clear business owner will deliver in a quarter. The same domain with nine source systems and no owner will consume a year. Start where the topology is simplest, even if the pain is slightly lower, because the first delivery is what buys you the second phase.
Third criterion: where a downstream process is genuinely waiting. MDM only creates value when something consumes the mastered record. Before committing, name the specific process that will change: the onboarding screen that will call the service, the pricing engine that will read the golden record, the credit check that will use the consolidated exposure. If no process changes, you are building a cleaner copy of a problem.
In practice, for mid market companies, this analysis points to customer or supplier data far more often than to product data, for a reason worth stating plainly. Product data is usually the domain with the loudest complaints and the most contested ownership, because it sits between merchandising, operations and finance. Customer and supplier data tend to have a clearer natural owner and a more direct line to money.
There is a fourth consideration that overrides all three when it applies. If the company is mid way through a system replacement, an acquisition integration or an ERP migration, the domain those projects are already touching is the one to master, because the reconciliation work is happening anyway and you can attach the permanent fix to a budget that already exists. The same logic that governs sequencing in legacy system modernization applies here: attach the structural fix to the change that is already funded.
The four architecture patterns, and what each one really costs
Once the domain is chosen, the architecture decision follows. There are four established patterns, and vendors will tell you their platform supports all four, which is technically true and practically misleading, because the cost is in the operating model, not the software.
Registry. The MDM system holds only identifiers and the matching logic. Source systems keep their own records; the hub knows that customer 4471 in the CRM and customer A-9920 in the billing system are the same entity, and can assemble a consolidated view on request. Cheapest and fastest to stand up, least invasive, and it never fixes the underlying data. Right choice when the immediate need is reporting and analytics, wrong choice when the need is to stop bad records being created.
Consolidation. Records are copied into a central hub, matched and merged into a golden record, which then feeds reporting and downstream analytical systems. The hub is authoritative for reading but not for writing. Moderate effort, delivers a clean view quickly, and leaves the source systems free to keep generating the same duplicates tomorrow. The most common landing point in practice, and often an honest one for phase one.
Coexistence. The hub masters the records and writes cleaned data back to the source systems, which continue to be used for entry. Harder, because it requires each source system to accept updates it did not originate, which is a technical problem in some systems and a political one in all of them. This is the point at which MDM starts changing operational reality rather than describing it.
Centralized, or transactional. The hub is the only place master records are created and maintained; source systems consume them. Cleanest end state, highest change management cost, and only realistic where the organization has the authority to change how people work in the source systems. Attempting this as a first phase is the single most common way to spend two years and deliver nothing.
The honest recommendation, which very few vendors will make, is to plan for consolidation in phase one and coexistence for the specific fields that matter most in phase two. Full centralization is a destination, not a starting position, and treating it as a starting position is how programs lose their sponsor.
The build versus buy question sits alongside this and deserves its own analysis rather than a default answer, because a small registry for two systems is a genuinely reasonable internal build while a multi domain hub almost never is. The framework for that decision is set out in the guide to build versus buy decisions for AI and data software.
Matching, survivorship and the golden record
This is the technical core, and it is where non technical sponsors lose the thread and start approving decisions they do not understand. It is simpler than it is usually made to sound.
Matching decides whether two records refer to the same real world entity. Deterministic matching uses exact rules: same tax identifier, same registration number, same email. Probabilistic matching scores similarity across multiple fields and calls a match above a threshold. Real implementations use both, deterministic first for the cases where an authoritative identifier exists, probabilistic for the rest.
The decision that matters is not the algorithm. It is the threshold, and specifically which of the two errors you would rather make. A false merge combines two real customers into one record, which corrupts credit exposure, billing and sometimes privacy obligations, and is expensive to unwind. A false split leaves duplicates in place, which costs marketing spend and reporting accuracy. In regulated and financial contexts, set the threshold conservatively and route the ambiguous middle band to a human queue. In marketing contexts, the tolerance is different. Make this an explicit business decision, minuted, with the tradeoff written out, because it will be revisited under pressure the first time something goes wrong.
Survivorship decides which value wins when matched records disagree. The naive rule is most recent wins, and it is wrong often enough to cause real damage: a stale but verified address beats a recent one typed into a field with no validation. Better survivorship rules are per attribute and source aware. The tax identifier survives from the finance system, the shipping address from the logistics system, the contact preference from the customer portal, because each of those systems is the one where that attribute is actually maintained.
The golden record is the output. Two properties determine whether it is trusted. It must carry lineage, meaning that for every field you can see which source contributed it and when. And it must be reversible, meaning a bad merge can be undone without a database restore. A hub that cannot unmerge is a hub whose operators will stop merging.
Add one operational rule that is more important than any of the above: keep the manual review queue small enough that it is actually worked. A queue with four thousand items is not a queue, it is an archive. If the ambiguous band is generating more items than your stewards can clear in a week, the thresholds are wrong or the scope is too wide. The correct response is to tighten scope, not to hire more stewards.
Data quality: the work that decides the outcome
Matching engines do not fix data quality. They expose it, and they expose it at volume, which is why the third month of an MDM program is usually the month the sponsor gets nervous.
Six dimensions are worth measuring, and they should be measured before the platform is selected rather than after.
Completeness. What percentage of records have the fields that downstream processes actually require. Not all fields: the required ones. Measuring completeness across every attribute produces a frightening number that nobody can act on.
Uniqueness. The estimated duplicate rate within each source system, measured on a sample rather than assumed. In my experience across mid market implementations, customer duplicate rates between 8% and 20% are normal, and organizations consistently estimate their own rate at roughly half the true figure.
Validity. Whether values conform to their expected format and domain. Tax identifiers that pass a checksum, country codes that exist, dates that are not in 1900.
Consistency. Whether the same entity carries the same value across systems. This is the dimension that MDM directly addresses and the one most likely to be catastrophically bad without anyone knowing.
Timeliness. How long it takes a real world change to appear in the system. A customer that moved eight months ago and is still billed at the old address is a timeliness failure, not a completeness failure, and the fixes are entirely different.
Accuracy. Whether the value matches reality. The hardest to measure, because it requires an external reference, and the one most often quietly skipped.
The critical discipline is to measure these on a sample before buying anything. A two week exercise on a thousand randomly selected records, scored by people who know the business, produces a baseline you can defend, a business case you can price, and a defence against the first vendor who tells you their platform will handle it. It also produces the single most useful artifact of the whole program: a defensible number for how bad the situation actually is.
Cleaning is where the effort lands, and it is worth being blunt about the split. In the programs I have seen deliver, the platform is a minority of the total cost. The majority is the human work of deciding what the right value is, in cases where two systems disagree and both have a plausible claim. No vendor demo shows this work, because it is unglamorous and it cannot be automated away. The organizations that budget for it finish. The ones that treat it as an implementation detail run out of money in month seven.
Governance, roles and who actually decides
MDM without a decision making structure produces an expensive system that no one is allowed to correct. Four roles are needed, and in a mid market company they are four hats, not four full time people.
The executive sponsor. Someone senior enough to settle a dispute between two department heads about what a customer is. This is the role that cannot be delegated and cannot be filled by a committee. If the program does not have it, do not start.
The data owner, per domain. A named business leader accountable for the definition, the quality standard and the acceptable thresholds for that domain. The head of commercial operations for customer data, the head of procurement for supplier data. Business roles, not technical ones.
Data stewards. The people who work the exception queue, resolve ambiguous matches and maintain reference values. Usually part time, drawn from the operational teams who already know the data. The most common staffing mistake is to make stewardship a technical function, at which point every ambiguous record gets escalated because the steward has no basis to decide.
The technical lead. Owns the platform, the integrations, the matching configuration and the monitoring. Reports into the program, not the other way around.
The forum that makes this work is a short, regular decision meeting, not a governance board that meets quarterly to review a dashboard. Thirty minutes, weekly during the build phase, with a standing agenda of exactly two items: what decisions are blocked, and what the exception queue looks like. Programs with this meeting make decisions in days. Programs without it accumulate a backlog of unresolved definitional questions, and every one of those questions is a place where the implementation team will guess.
A useful test of whether governance is real: ask who can reject a new source system's request to create customer records directly. If the answer is nobody, you do not have master data management, you have a reporting layer with ambitions.
What master data management costs
Anyone who quotes a number without asking how many source systems you have, how many records are in scope and what your duplicate rate looks like is selling something. Here is the cost structure, which stays valid even as list prices move.
Platform licensing or subscription. Usually priced on record volume, domains, or both. Two traps: the price steps up sharply when you add the second domain, and the entry tier often excludes the capabilities you will discover you need, typically real time services, workflow for stewardship, or write back to source systems.
Implementation services. Configuration, matching rule development, integration build. Typically the largest single line in year one, and typically underestimated in the initial proposal because the proposal was written before anyone profiled the data.
Data profiling and remediation. The human work described above. Budget it explicitly, in person days, or it will be taken from the implementation budget mid project and the implementation will be cut short.
Integration. Each source system connection is its own project with its own maintenance. Price them individually, and ask specifically what happens when a synchronization fails silently, because a failing integration that reports success is worse than no integration at all.
Stewardship, ongoing. The steady state cost of working the exception queue. In mid market implementations this rarely runs below a quarter of a full time role per active domain, and it does not go away after go live.
Change management and training. Every process that consumes mastered data changes, and the people doing that process need to know why the screen behaves differently. This is the line that budgets cut first and the one that determines whether anyone uses the result.
Ongoing platform maintenance. Plan for roughly 20% of annual subscription as a recurring evolution budget for rule tuning, new attributes and source system changes.
The honest sizing heuristic: take the year one platform cost and multiply by two and a half to three for the fully loaded first year including services, remediation and internal time. If a proposal comes in dramatically below that, the difference has not been saved, it has been deferred into change requests that will be negotiated when you have no leverage.
The metrics that matter, and two that mislead
The standard vendor dashboard counts records under management. That number tells you almost nothing. Six measures are worth reporting, and two are worth actively avoiding.
Duplicate rate per domain, measured over time. The headline operational measure. Report it per source system as well as consolidated, because an improving average can conceal one system that is generating duplicates faster than they can be merged.
Percentage of new records created through the governed path. If records can still be created directly in source systems without validation, this number tells you how much of the problem you are still manufacturing. It is the single best leading indicator of whether the program will hold.
Exception queue age, not size. The oldest unresolved item in the queue. Size fluctuates with volume; age tells you whether the stewardship model is staffed correctly.
Downstream consumption. How many operational processes read from the hub, not from their local copy. This is the measure that separates a real MDM program from a clean warehouse, and it is the one most often left off the dashboard because it is uncomfortable.
Time to onboard a new entity. How long from a new customer or supplier being requested to being usable across systems. It converts directly into commercial speed and it is the number operational leaders care about.
Match precision on a sampled audit. Periodically pull a sample of automatic merges and have a human check them. The rate of incorrect merges is the number that tells you whether your thresholds are still right as data volumes change.
The two misleading measures: total records under management, which grows regardless of whether anything improved, and data quality score as a single composite number, which averages away the one dimension that is actually broken. Composite scores are excellent for steering committee slides and useless for deciding what to do on Monday.
For the broader question of choosing measures before choosing tools rather than after, the reasoning is developed in the guide to data driven decision making.
Readiness scorecard: should you start now
One point for each statement that is true. This does not measure how sophisticated you are. It measures the probability that the program reaches production without stalling.
Business case
- We can name a specific process that is failing because of duplicate or inconsistent records.
- We have estimated, on sampled data rather than intuition, what that failure costs per year.
- We have chosen exactly one domain for phase one.
- We can name the downstream process that will consume the mastered record.
Ownership
- There is a named executive sponsor with authority to settle definitional disputes.
- There is a named business owner for the chosen domain, not a technical owner.
- Stewards have been identified from operational teams and their time has been agreed.
- Someone has the authority to refuse a new source system permission to create records directly.
Data
- We know how many systems create records in the chosen domain.
- We have measured the duplicate rate on a sample rather than estimating it.
- We know which system is authoritative for each key attribute.
- We have agreed what the entity actually is, in writing, across departments.
Delivery
- Budget covers three years and includes remediation and stewardship, not just licensing.
- We have baseline measurements for the metrics we intend to move.
Reading the result
12 to 14 points: you are ready to go to vendor selection.
8 to 11: start, but close the gaps in the ownership section first. Those are the ones that surface in month six, when a definitional dispute stops the build and there is nobody empowered to settle it.
4 to 7: spend four weeks profiling the data and agreeing the definition before selecting anything. A program started without those two artifacts costs substantially more and delivers substantially later.
Below 4: the problem is not the platform. What is needed first is a decision about who owns the core entities of the business, and it is worth making that decision with someone who has seen enough of these programs to recognize within hours where yours will break. An initial assessment costs a fraction of what a stalled implementation costs, and its main output is a clear answer to whether to proceed now or in six months.
The 30, 60, 90 day plan
Calibrated for a mid market company, single domain, with or without an existing hub to replace.
Days 1 to 30: measure and define
Week 1. Confirm the executive sponsor and the domain owner in writing, including the time commitment. Choose the single domain for phase one using the three criteria above, and write down the specific downstream process that will consume the result.
Week 2. Inventory every system that creates or maintains records in the chosen domain. Not the systems you think of first: all of them, including the departmental database and the spreadsheet that a regional team maintains. For each, record who creates records, at what volume, and with what validation.
Week 3. Profile a sample. A thousand records, randomly selected, scored by people who know the business against the six quality dimensions. This is the artifact that will carry the business case, so do it properly rather than delegating it to a script.
Week 4. Write the entity definition and get it signed. One page: what a customer is, what a customer is not, what the authoritative identifier is, which attributes are mandatory. Circulate it to every department that creates records and resolve the objections now, because every objection you do not resolve now becomes a configuration decision made by someone with less context.
Days 31 to 60: design and select
Week 5. Decide the architecture pattern for phase one and write down the reasoning. Registry, consolidation or coexistence, with the phase two target stated separately so the choice is understood as a stage rather than a compromise.
Week 6. Draft the matching and survivorship rules on paper, per attribute, before any vendor conversation. This document turns vendor demos from a sales exercise into a test, because you can ask each vendor to configure your rules rather than show you theirs.
Week 7. Run vendor evaluations against your own sampled data. Refuse the standard demo. Give each vendor the same extract, including the hard cases you already know about, and compare what comes back. The hard cases are the whole point; the standard demo works by construction.
Week 8. Price integrations individually, negotiate exit terms and data return format while you still have leverage, and take two reference calls with organizations of comparable size. Ask one question: what would you do differently.
Days 61 to 90: build small and prove it
Weeks 9 and 10. Configure the hub for the single domain, load the source data, and run matching. Expect the first results to be wrong and plan for two tuning cycles rather than one. Work the exception queue with real stewards, not with the implementation team, because the queue is where you find out whether the definition was actually agreed.
Week 11. Connect one downstream consumer. One. The onboarding screen, the credit check, the pricing service. This is the step that converts the program from a data exercise into an operational change, and it is the step that most phase ones postpone.
Week 12. Measure the same metrics from week 3 and compare. Report the duplicate rate, the exception queue age and the consumption count. Decide whether to extend to a second domain, tune the current one, or stop, and put the decision in writing with the numbers that drove it.
At day 90 you will not have a single view of the enterprise. You will have something more useful: a mastered domain that one real process depends on, and a defensible number to take to whoever approves the next phase.
Four real cases
Different sectors and sizes, to show that the method holds rather than the product. All anonymized by industry.
A sports distribution company. The presenting problem was sales performance. The actual problem was that commercial data and customer relationship data lived in two systems that did not reconcile, so nobody could see a complete picture of any account. We unified the underlying data and built a segmentation on top of it that had not previously been possible, which changed who got contacted and with what message. Sales grew 30%. No new tool was introduced. What changed was which data ran through the existing ones and who was looking at it.
A hotel. Revenue was flat at around 9 million. The systems worked, but nobody was cross referencing occupancy against booking channels, and pricing decisions were being made on instinct. We restructured the reporting and the dynamic pricing policy. Revenue moved to 10 million without adding a single room.
A medical center. The bottleneck was scheduling and cancellation handling. The staff had built genuinely effective informal procedures for filling released slots, but those procedures lived in the heads of three people and in a paper notebook. We formalized them and rebuilt the booking and recovery flow. Delivered capacity rose 20% with no new equipment and no additional clinicians.
An agritourism business. Small operation, no structured systems, everything dependent on individual initiative. Here the sequence mattered more than anywhere else: we defined the guest acquisition and management process first, then chose the minimum tool that supported it. Guest numbers doubled.
The common thread applies directly to master data. In none of the four cases did the result come from a product feature. It came from deciding in advance which number had to move, and from making explicit what had previously lived in a few people's heads. That is what master data management is, stripped of the vocabulary: the point at which the definition of a customer, a product or a supplier stops depending on which system you happened to open.
If you recognized your own situation in more than one of those cases, the useful question is not which platform to buy. It is which core business entity currently has three different definitions inside your company, and what that has cost you in the last twelve months.
What to do on Monday
Three concrete actions, none of which requires budget or approval.
Pull the same entity from three systems and compare it. Take your twenty largest customers, or your twenty highest spend suppliers, and pull their records from every system that holds them. Line them up in a spreadsheet. Count the fields that disagree. That count is your real starting point, and it is almost always worse than the estimate you would have given beforehand.
Ask four people in different departments to define a customer. Do it separately, in writing, in two sentences each. If the four answers do not match, you have found the actual blocker, and no platform selection will resolve it.
Write the three numbers that have to move. With a threshold and a date. Duplicate rate in the chosen domain, time to onboard a new entity, and one downstream cost you can name in money. If you cannot express them numerically, it is not yet time to talk to a vendor, and that is the most valuable thing you can learn this week.
Everything else follows from those three. And if the conclusion at this point is that the problem is larger than the tooling, that is usually the correct reading: it means master data is one part of a wider decision about how the organization defines and distributes the entities its whole operation depends on, and that is exactly the kind of work worth structuring with someone who has done it before, while the contract is still unsigned and the options are still open.
FAQ
What is a master data management strategy and where should you start?
A master data management strategy is the plan for deciding which core business entities will be centrally defined and maintained, in what order, using which architecture, and enforced by whom. Where to start is the decision that determines whether it delivers. Pick one domain, chosen on three criteria applied in order: where bad data converts into a cost you can name in money, where the number of source systems creating records is smallest, and where a specific downstream process is genuinely waiting to consume the result. Programs that begin with one domain finish. Programs that begin with five spend their first year negotiating scope rather than delivering anything.
What is the difference between master data management and data governance?
Governance sets policy: what the rules are, who owns which data, what quality thresholds apply, and who can decide when there is a dispute. Master data management is the operational machinery that enforces those rules on a specific set of entities at the moment records are created and maintained. You can have governance without MDM, which is why many organizations have excellent policy documents and poor customer records. You cannot usefully run MDM without governance, because the matching engine will simply automate whichever inconsistency the policy never resolved. In practice governance decides what a customer is; MDM makes sure there is only one of each.
How much does master data management cost?
Platform licensing is the visible line and rarely the largest. The full first year cost includes implementation services, integration for each source system, the human work of profiling and remediating data, change management for every process that will consume the mastered record, and internal stewardship time. A workable sizing heuristic is to take the year one platform cost and multiply by two and a half to three for the fully loaded first year. From year two, budget roughly 20% of the annual subscription for rule tuning and evolution, plus ongoing stewardship, which rarely runs below a quarter of a full time role per active domain. A quote well below that has deferred the difference into change requests.
How long does a master data management implementation take?
Twelve to twenty four months for a multi domain program, driven mostly by the number of source systems and the state of the data rather than by the platform. But full rollout is the wrong milestone to plan against. The right one is ninety days, by which point a single domain should be mastered and one real downstream process should be consuming the result, with the same metrics measured before and after. The time is not consumed by configuration. It is consumed by profiling the data and by resolving the definitional disagreements between departments about what the entity actually is.
What is a golden record, and how is it created?
A golden record is the single consolidated version of an entity, assembled from multiple source records that have been identified as referring to the same real world thing. It is produced in two steps. Matching decides which records belong together, using exact rules where an authoritative identifier exists and probabilistic scoring elsewhere. Survivorship decides which value wins when matched records disagree, ideally per attribute and source aware, so the tax identifier survives from finance and the shipping address from logistics. A golden record is only trustworthy if it carries lineage, meaning you can see which source contributed each field, and if merges can be reversed without a database restore.
Which master data domain should be tackled first?
For mid market companies the analysis usually points to customer or supplier data rather than product data. Product data typically generates the loudest complaints but has the most contested ownership, sitting between merchandising, operations and finance, which means the definitional argument alone can consume a quarter. Customer and supplier domains tend to have a clearer natural owner and a more direct line to a monetary consequence: credit exposure assessed on a fragment of the real relationship, or the same vendor paid twice. The exception overrides everything else: if an ERP migration or acquisition integration is already underway, master the domain that project is touching, because the reconciliation work is funded already.
Do we need an MDM platform, or can we solve this with data quality tools?
They solve different problems and the distinction matters commercially. Data quality tools profile, standardize, validate and deduplicate data, usually in batch, usually for analytics. An MDM platform adds persistent identity resolution, a maintained golden record, stewardship workflow for ambiguous cases, and services that operational systems can call in real time. If the requirement is a clean data set for reporting, quality tooling may be sufficient and considerably cheaper. If the requirement is that the order entry screen stops creating the fourth copy of an existing customer, quality tooling will not get you there, because the problem is not the state of the data but the moment of creation.
How do you measure whether master data management is working?
Six measures, tracked from a baseline taken before implementation. Duplicate rate per domain, reported per source system as well as consolidated. The percentage of new records created through the governed path, which is the best leading indicator of whether the program will hold. Exception queue age rather than size, because size fluctuates with volume while age reveals whether stewardship is staffed correctly. The number of operational processes actually consuming from the hub rather than from a local copy. Time to onboard a new customer or supplier. And match precision, sampled and checked by a human periodically. Avoid total records under management and single composite quality scores: both move without anything improving.