Data Governance Framework: A Practical Guide
Gartner predicts that eighty percent of data and analytics governance initiatives will fail by 2027, and the stated reason is not budget or tooling. It is the absence of a real or manufactured crisis to force the issue. Read that again, because it is the most useful sentence written about this subject in the last five years: governance programs do not die from lack of money, they die from lack of consequence.
That prediction sits next to a second one from the same firm. Through 2026, organizations will abandon sixty percent of AI projects unsupported by AI-ready data, a forecast drawn from a survey of 248 data management leaders in which sixty-three percent said they either do not have the right data management practices for AI or are not sure whether they do.
So the picture is this. Companies are spending record amounts on AI while sitting on data they cannot vouch for, and the standard response is to launch a data governance program that statistically will not survive to its third birthday.
I have watched this from close range in a dozen companies, and the failure is almost never technical. It is that the program is designed as a policy exercise when it should have been designed as a delivery exercise attached to a decision someone is already trying to make. This guide is about building the second kind. It contains the structure of a working framework, the roles that actually need to exist, a ninety day build sequence, the cost lines nobody quotes, and the specific points where AI changes the requirements rather than just adding a slide to the deck.
What a Data Governance Framework Actually Is
A data governance framework is the written answer to four questions, applied consistently across an organization. Who owns this data. What are we allowed to do with it. How do we know it is correct. What happens when it is not.
Everything else in the discipline is elaboration. Catalogs, lineage tools, stewardship councils, maturity models: all of them are machinery built to answer those four questions at scale. When a program collapses, it is usually because the machinery got purchased before the four questions had a single concrete answer for a single real dataset.
Three things a framework is not, and each misunderstanding costs a quarter.
It is not a document. Companies produce a forty page data governance policy, circulate it, get sign off, and then discover that nothing changed because nothing in anyone's daily work referenced it. A framework that is not embedded in an approval, a release gate, or a report someone actually reads is a document, not a framework.
It is not a tool purchase. The catalog does not create ownership. It records ownership that a human being agreed to. Buying the catalog first produces an expensive inventory of orphaned tables, which is roughly what you already had, plus a subscription.
It is not centralization. The instinct to route every data decision through a central team creates a queue, and the queue creates workarounds. The teams you were trying to govern will simply build their own copy of the data outside your perimeter, and now you have less visibility than when you started.
The Only Metric That Matters at the Start
Before designing anything, measure one number: the percentage of your most-used business metrics that two different teams calculate identically.
Pick your ten most-cited numbers. Revenue, active customers, churn, pipeline, headcount, whatever your leadership actually argues about. Ask two independent teams to produce each one for the same period. Count how many match.
In most mid-sized companies the answer is between three and six out of ten. That gap is your business case, and it is far more persuasive to a CFO than any maturity model, because it translates directly into meetings spent reconciling numbers instead of deciding things.
Why Most Data Governance Programs Fail
The Gartner prediction names the mechanism, but it is worth unpacking the four specific ways I see programs die.
They start with an inventory instead of a decision. A team spends five months cataloging four thousand tables. At the end they have a catalog and no change in behavior, because nobody was blocked on not having a catalog. The programs that survive start with one decision that is currently being made badly, and fix the data underneath it.
They assign stewardship to people with no authority. A data steward who cannot say no to a schema change, cannot hold a release, and does not control any budget is a person with a title and a spreadsheet. The role only works when it carries a specific, small, real power.
They measure activity instead of outcome. Number of tables cataloged, percentage of fields with descriptions, policies published. All of these can go up while the ten metrics still disagree. Governance metrics must be outcome metrics or they will be gamed within two quarters.
They arrive as an imposition from a function nobody asked. This is the one Gartner is pointing at. Without a crisis, real or engineered, governance reads as overhead. The engineered version is legitimate and I recommend it: pick a real regulatory deadline, a real failed audit, a real board question that could not be answered, and hang the program on it.
Deloitte's Chief Data Officer Survey 2025 shows the tension precisely. Data governance was the top priority for CDOs overall at fifty-one percent, and rose to sixty-three percent among lower maturity organizations, while higher maturity organizations had already moved on to AI and data products. Meanwhile eighty-seven percent of CDOs report into the C-suite, but fifty-four percent believe they are currently less influential than their executive peers. Reporting line without influence is exactly the condition in which a governance program becomes a document.
The Six Components of a Working Framework
Strip away the vendor diagrams and a functioning framework has six parts. Each one has a test for whether it exists in reality or only on a slide.
1. Scope and Domains
Data is divided into domains that map to how the business is organized, not to how the database is organized. Customer, product, finance, employee, supplier. Each domain has a single accountable owner at the business level.
The test: can you name, out loud, the person accountable for customer data, and would that person recognize the title if you asked them?
2. Roles and Decision Rights
Three roles minimum. The domain owner is a business leader who is accountable for the data in their domain and who has the authority to approve or refuse changes to its definitions. The data steward does the daily work: definitions, quality rules, issue triage. The data custodian is technical, usually in engineering, and implements controls.
The test: when a definition changes, does a specific named person have to approve it, and can that person say no.
3. Definitions and the Business Glossary
One agreed definition per business term, with the calculation logic, the source of truth, and the owner. This is the least glamorous component and the one with the highest return.
The test: search for your definition of "active customer" and count how many exist. If more than one, you do not have a glossary, you have documents.
4. Quality Rules and Measurement
Quality is not an abstract virtue. It is a set of specific rules attached to specific fields, measured continuously, with a threshold and an owner. Completeness, validity, uniqueness, timeliness, consistency across systems.
The test: can you produce, today, a number for the quality of your top five datasets, and does anyone receive it on a schedule.
5. Access, Classification and Protection
Every dataset carries a classification, and the classification drives access. Public, internal, confidential, restricted. Access is granted by role, reviewed on a cycle, and revoked when the role changes.
The test: pick a person who left the company four months ago and check whether their access is gone from every system, not just the identity provider.
6. Lifecycle and Retention
How long each class of data is kept, where it goes at end of life, and who signs off on deletion. This is the component most often skipped and the one regulators ask about first.
The test: is there a dataset in your warehouse older than your stated retention period. There almost certainly is.
MIT Sloan Management Review made the underlying argument well in its work on data governance in the modern organization: governance has to be designed around how decisions get made, not around the shape of the technology stack. That framing is the difference between a program that changes behavior and one that produces artifacts.
Roles: Who Owns What, and Who Can Say No
The role design is where most frameworks quietly fail, because they are copied from a large enterprise into a company one twentieth the size.
In a company under 200 people, you do not need a council. You need one executive sponsor, two or three part-time domain owners who are the business leaders already accountable for those areas, and one person, often in analytics engineering, who does stewardship as forty percent of their job. Anything more elaborate will not be staffed and will decay into a calendar invite that gets declined.
In a company between 200 and 2,000, you need the same structure plus a small forum that meets monthly with a real agenda: definition disputes, quality exceptions, access requests that do not fit the standard pattern. The forum must have decision authority. A forum that only escalates is a bottleneck with catering.
Above 2,000, the domain model needs formal federation, with domain teams operating semi-autonomously against central standards, and a central function that owns the standards, the platform and the audit.
Three design rules hold at every size.
Accountability sits with the business, not with IT. If the head of engineering owns the definition of revenue, the definition will be technically correct and commercially useless, and the commercial team will build a shadow version.
Every role needs one real power. For the steward, it is usually the authority to hold a release when a quality gate fails. Small, specific, and enough to make the role matter.
Nobody gets a governance role as a second job with no time allocated. Write the percentage into the role description and defend it, or accept that the work will not happen.
The organizational half of this problem is the same one that determines whether any technology program lands, and I have covered the mechanics of it separately in the guide on AI change management.
How to Build a Data Governance Framework in 90 Days
This is the sequence I use. It is deliberately narrow, and the narrowness is the point.
Days 1 to 30: Anchor and Measure
Week 1. Choose the anchor. One decision, currently made badly, that leadership cares about. A board metric that cannot be reconciled. An audit finding. A regulatory deadline. A model that failed because its training data was wrong. Write it down in one sentence and get the sponsor to agree that this is what the program is for.
Week 2. Run the ten metric test described earlier. Two teams, ten numbers, same period, count the matches. This is your baseline and your business case in a single artifact.
Week 3. Map the lineage of the three metrics that matter most to the anchor decision. Not the whole warehouse. Three metrics, from source system to the report where leadership sees them. You will find at least one surprise per metric, and typically the surprise is a manual step in a spreadsheet.
Week 4. Name the domain owners for the domains those three metrics touch. Get verbal agreement from each. Write the three role descriptions with time allocations. Publish nothing yet.
Days 31 to 60: Define and Instrument
Week 5. Write the definitions for the three metrics, with calculation logic, source of truth and owner. Circulate for dispute. The disputes are the valuable part: they surface the disagreements that have been costing you reconciliation meetings for years.
Week 6. Resolve the disputes in a single working session with the domain owners in the room. Somebody has to lose an argument. That is the moment the framework becomes real, and if it does not happen, you do not yet have governance.
Week 7. Implement quality rules on the fields those three metrics depend on. Start with completeness and validity, which are cheap and catch most of the damage. Set thresholds from observed data, not from aspiration.
Week 8. Wire the quality results into a channel a human being reads, and give a named person the authority to hold a release when a gate fails. Test that authority once, deliberately, on something low risk.
Days 61 to 90: Prove and Extend
Week 9 and 10. Run the loop. Quality checks fire, issues get triaged, definitions get referenced in disputes. Log every issue and its resolution time, because that log is your evidence.
Week 11. Re-run the ten metric test. The three you worked on should now match. The other seven will not, and that contrast is your argument for the next phase.
Week 12. Write a two page report: baseline, what changed, cost, what the next three domains would take. Present it to the sponsor with a specific ask. Then decide whether to extend, adjust or stop, and put the decision in writing.
At day ninety you will not have a governed enterprise. You will have something more useful: documented proof that the model works on your data with your people, and a number to show whoever approves the next phase.
Policies That People Actually Follow
Most data policies fail the same test: they describe an ideal state without describing what a person should do differently on Tuesday.
A usable policy has four properties.
It is short enough to read. Two pages per topic. If a policy needs a summary, the summary is the policy and the rest is reference material.
It names the moment it applies. Not "data should be classified appropriately" but "when a new table is created in the warehouse, the creator sets the classification field before the first scheduled run, and the pipeline fails without it." Policy that is enforced by the system beats policy that is enforced by memory, every time.
It has an exception path. A policy with no legitimate way to get an exception will be violated silently instead of openly. Give people a documented way to request one, with an owner and a time limit, and you will find out where the policy does not fit reality.
It has an owner and a review date. An unreviewed policy becomes wrong within eighteen months and then actively harmful, because it gives people confident bad guidance under an official banner.
One practical note on tone. Policies written in legal register get ignored by the engineers who have to implement them. Write them in the language of the people who will execute them, and keep the legal register for the annexes.
Data Quality: The Metrics That Matter
Six dimensions get taught. Three of them earn their keep in the first year.
Completeness. Percentage of required fields populated. Cheap to measure, immediately actionable, and the source of most downstream breakage.
Validity. Percentage of values that conform to the expected format or the allowed set. Catches the classic failures: dates in the wrong century, country codes invented by a form, a status field with fourteen values when the process has four.
Consistency across systems. The same entity, counted in two systems, producing the same number. This is the expensive one to implement and the one that maps most directly to the trust problem, because it is exactly the mismatch that leadership notices.
The other three, uniqueness, timeliness and accuracy, matter, but they are harder to instrument and easier to defer. Accuracy in particular is often unmeasurable without an external reference, and programs that start there stall.
Two rules on measurement.
Set thresholds from observed distributions, not from round numbers. A completeness target of ninety-nine percent set on day one, against a field that is currently at seventy-two percent, produces a permanently red dashboard that everyone learns to ignore within a month. Set the threshold slightly above current performance and ratchet it.
Report the trend, not the absolute. Nobody acts on "completeness is 91%." People act on "completeness fell from 94% to 91% after the release on the 14th."
The broader discipline of building measurement that changes decisions, rather than measurement that decorates a dashboard, is the subject of the guide on data driven decision making, and the two topics are effectively the same problem seen from opposite ends.
Governance for AI: What Actually Changes
AI does not replace data governance. It raises the cost of not having it, and it adds four requirements that traditional frameworks do not cover.
Provenance becomes mandatory, not nice to have. With a report, a wrong number gets caught by a human who knows the business. With a model, a wrong training input gets amplified silently across thousands of outputs. You need to be able to say which data trained or grounded a given output, and most warehouses cannot answer that today.
Rights and permissions travel with the data into the model. If a dataset was collected under a consent that does not cover model training, using it for training is a problem that no amount of anonymization talk will fix afterwards. This has to be checked before ingestion, and the check has to be a field in your catalog, not a conversation.
Freshness requirements change shape. A monthly refresh is fine for a report and dangerous for a retrieval system that will confidently cite last quarter's pricing. Grounding data needs its own freshness SLA, separate from the analytics one.
Evaluation data becomes a governed asset. The test set that decides whether a model ships is now one of the most consequential datasets in the company. It needs an owner, version control and access restrictions, and almost nobody treats it that way in year one.
This is the point where data governance and AI governance meet without being the same thing. Data governance answers whether the inputs are trustworthy. AI governance answers whether the system built on them is safe, fair and accountable, which is a separate discipline with its own controls, covered in the guide on AI governance for business. Running the second without the first is the specific failure mode behind the sixty percent abandonment forecast.
If you are still deciding whether to build these capabilities internally or buy them, the trade-off is structural rather than financial, and I worked through it in the build versus buy framework.
Tooling: What to Buy, and When
The market sells four categories, and the right sequence is not the order the vendors present them.
A catalog records what data exists, what it means and who owns it. Buy it when manual documentation has become the bottleneck, which is usually somewhere past a few hundred actively used tables. Before that, a well-maintained spreadsheet genuinely outperforms a badly-maintained catalog.
A quality platform runs rules continuously and alerts on breaches. This is the first purchase I recommend for most mid-sized companies, because it produces evidence immediately and evidence is what keeps the program funded.
A lineage tool traces data from source to consumption. Valuable, and often available inside the warehouse or transformation layer you already pay for. Check before buying standalone.
An access governance layer manages who can see what, with review cycles. Buy when the number of systems makes manual review unreliable, or when a regulator or auditor asks a question you cannot answer.
Two rules that save a lot of money.
Nothing gets bought before one domain is governed manually. If you cannot govern one domain with a spreadsheet and a monthly meeting, buying a platform will produce an expensive version of the same failure.
Check what your existing stack already does. Modern warehouses, transformation tools and identity providers now cover a meaningful share of catalog, lineage and access review. A surprising number of governance purchases duplicate capability the company is already paying for.
The Cost Structure Nobody Quotes
Vendors quote licences. Licences are rarely the largest line.
Platform licences. Priced per user, per data source, or per volume scanned. The volume-based models are the ones that surprise people at the second renewal.
Implementation and integration. Connecting the tool to every system that matters. In the programs I have seen, this runs one to two times the first-year licence.
Steward time. The largest real cost and the one that never appears in a business case. Two to three days a week of a capable person's time, for at least the first two quarters. If nobody is funding that, the program is not funded.
Definition work. The hours of business people arguing productively about what a term means. Expensive people, real hours, and the highest-return spend in the entire programme.
Remediation. Fixing the data the quality rules find. This is the line that catches everyone. The rules are easy; the backlog they expose is the actual work.
Ongoing operation. Budget roughly a fifth of the licence annually for evolution, plus the permanent steward allocation.
The honest first-year estimate: take the annual licence, multiply by two to two and a half, and add the fully-loaded cost of the steward time. If a vendor's proposal comes in dramatically below that, the difference will reappear as change requests during implementation.
Readiness Scorecard
One point per statement that is true. This measures the probability that a program survives, not how modern you are.
Anchor and sponsorship
- There is one named executive who will lose something if this fails.
- We can state in one sentence which decision the program is meant to improve.
- There is a real deadline, audit finding or board question driving urgency.
Baseline
- We have run the ten metric test and know the score.
- We know the lineage of our three most important metrics, end to end.
- We know which of our critical datasets have no named owner.
Roles
- Domain owners are business leaders, not technologists.
- At least one person has a written time allocation for stewardship.
- Somebody has the authority to hold a release for a data quality failure.
Definitions and quality
- There is one agreed definition for our top five business terms.
- Quality rules exist on at least one critical dataset and are measured on a schedule.
- Quality results reach a human being who is expected to act.
Sustainability
- Budget covers three years, including steward time and remediation.
- There is a written review cycle for policies and definitions.
Reading the score
Twelve to fourteen: you are ready to extend across domains.
Eight to eleven: proceed, but close the gaps in the anchor and roles sections first. Those are the ones that fail at month six.
Four to seven: stop and spend four weeks on the baseline. A program launched without the ten metric test has no evidence, and a program without evidence loses its budget in the first cost review.
Below four: the problem is not governance tooling. It is that no decision currently depends on this data in a way anyone can name, and that is worth resolving before committing a three-year budget. A short diagnostic engagement costs a fraction of a misdirected program and is the fastest way to establish whether now is the right moment or whether six months from now is.
Four Real Cases
Different sectors and sizes, to show that the method holds rather than the product. All anonymized by sector.
A sports distribution company. The problem was not a missing platform. Sales data and inventory data lived in separate worlds with different definitions of the same product, and nobody could see the whole picture. We unified the data foundation and built a segmentation on top of it that had not existed before, with immediate consequences for what to stock and whom to contact. Sales grew thirty percent. The tools barely changed; what changed was the definitions underneath and who was looking at the numbers.
A hotel. Revenue was flat around nine million. The systems worked, but occupancy data and booking channel data were never joined, and pricing decisions were made on instinct. We restructured the reporting and the dynamic pricing policy on top of a single agreed set of definitions. Revenue moved to ten million with no additional rooms.
A medical center. The bottleneck was scheduling and cancellation handling, with staff running informal procedures to recover freed slots. We formalized and rebuilt the booking and recovery flow, which required first agreeing what counted as an available slot, a definitional problem before it was a process problem. Delivered capacity rose twenty percent with the same facility, no new equipment and no new doctors.
An agriturismo. Small operation, no structured systems, everything driven by personal initiative. Here sequence mattered more than anywhere else: we defined the guest acquisition and management process first, then chose the minimum tooling that supported it. Guests doubled.
The common thread is the same every time. In none of the four cases did the result come from a software feature. It came from deciding first which number had to move, and then building the data structure around that decision instead of around a vendor's catalog.
If you recognized your own situation in more than one of those, the useful question is not which governance platform to buy. It is which decision you are currently making without trustworthy data, and that is the kind of knot worth untangling before a multi-year budget is committed.
Regulation: GDPR, the EU AI Act, and Sector Rules
Three regulatory realities shape framework design, and all three are cheaper to handle at design time than retrofitted.
GDPR and equivalents require you to know what personal data you hold, why you hold it, on what legal basis, how long you keep it, and who has access. Every one of those is a field in a governance framework. Companies that built a catalog for analytics reasons find that the compliance answer falls out of it almost free; companies that built compliance documentation separately end up maintaining two inconsistent inventories.
The EU AI Act phases in obligations that depend on risk classification, with data governance requirements attached to high-risk systems, including provisions on training data quality, representativeness and documentation. If any system you build could land in that category, the data lineage and provenance work stops being optional hygiene and becomes evidence you will need to produce.
Sector rules add their own layer. Financial services carry model risk and record-keeping obligations. Healthcare carries special-category data rules and consent complexity. Public sector work carries transparency and archiving duties. Identify yours before designing retention and classification, because retrofitting retention rules onto a live warehouse is one of the more expensive corrections available.
One practical instruction that saves months. Get whoever handles data protection into the framework design at week one, not at the review before launch. Their requirements are structural, and structural requirements discovered late get implemented badly.
What to Do Monday
Three concrete actions, none of which requires a budget.
Run the ten metric test. Pick your ten most-cited numbers, ask two teams to produce each independently for the same period, count the matches. It takes a week of part-time effort and produces the single most persuasive artifact you will have.
Name an owner for your three most important datasets. Out loud, to their face, with agreement. If you cannot get agreement, you have learned something more valuable than any tool evaluation would have told you.
Write down the anchor decision. One sentence describing the decision this program exists to improve. If you cannot write it, you are not ready to start, and discovering that this week is worth more than discovering it in month seven.
Everything else follows from those three. The relationship between trustworthy data, adoption and return on investment is where the economics of the whole thing are decided, and I have treated it separately in the guides on enterprise AI adoption and AI for knowledge management.
And if the feeling at this point is that the problem is bigger than the framework, that is usually the correct reading. It means data governance is one piece of a wider decision about how the organization makes decisions at all, which is exactly the kind of work worth structuring with someone who has done it before, while there is still room to shape the program rather than defend it.
FAQ
How do you build a data governance framework from scratch?
Start with an anchor decision rather than an inventory. Pick one decision leadership currently makes badly because the data cannot be trusted, then measure the baseline by asking two independent teams to calculate your ten most-cited business metrics for the same period and counting how many match. Map the lineage of the three metrics that matter most to the anchor, name business-level owners for those domains, agree the definitions in a session where somebody has to lose an argument, and implement completeness and validity rules on the fields those metrics depend on. Ninety days later, re-run the ten metric test and use the contrast between the fixed metrics and the rest as your case for the next phase.
What are the components of a data governance framework?
Six, and each has a test for whether it exists in reality. Scope and domains, mapped to how the business is organized with one accountable owner each. Roles and decision rights, meaning domain owner, steward and custodian, where at least one role carries real authority. A business glossary with one agreed definition, calculation logic, source of truth and owner per term. Quality rules attached to specific fields with thresholds and continuous measurement. Access, classification and protection with periodic review. Lifecycle and retention with a named signer for deletion. If a component cannot pass its test on a real dataset today, it is on a slide rather than in operation.
How long does data governance take to implement?
Full coverage across an enterprise takes years and never finishes, which is why that is the wrong target to set. The right target is ninety days to a governed slice: three metrics, one or two domains, definitions agreed, quality rules running, and one person with the authority to hold a release. Extending to additional domains then runs in six to twelve week increments, each with its own before-and-after measurement. Programs that attempt enterprise-wide coverage in the first phase are the ones that spend five months producing a catalog and no behavior change.
Why do most data governance initiatives fail?
Gartner predicts eighty percent will fail by 2027, attributing it to the absence of a real or manufactured crisis. In practice the failure shows up four ways: the program starts with an inventory instead of a decision, so nothing was ever blocked on its absence; stewardship is assigned to people with no authority to say no; success is measured by activity such as tables cataloged rather than by outcome; and the whole thing arrives as overhead from a function nobody asked. The corrective is to attach the program to a real deadline, audit finding or board question, and to give at least one role a small but genuine power such as holding a release when a quality gate fails.
What is the difference between data governance and AI governance?
Data governance answers whether the inputs are trustworthy: who owns this data, what may we do with it, is it correct, what happens when it is not. AI governance answers whether the system built on those inputs is safe, fair, documented and accountable, which brings in model risk, evaluation, monitoring for drift and human oversight. They are separate disciplines with separate controls, and the sequence matters. Gartner forecasts that through 2026 organizations will abandon sixty percent of AI projects unsupported by AI-ready data, which is precisely what running AI governance on top of ungoverned data produces.
How much does data governance cost?
Licences are rarely the biggest line. First-year cost includes platform licences, implementation and integration at typically one to two times the first-year licence, steward time at two to three days a week of a capable person for at least two quarters, the hours of business people agreeing definitions, and remediation of the problems the quality rules expose, which is the line that surprises everyone. A workable estimate is the annual licence multiplied by two to two and a half, plus the fully loaded cost of steward time, with roughly a fifth of the licence budgeted annually thereafter for evolution. A proposal far below that will resurface as change requests.
Who should own data governance in an organization?
Accountability belongs to the business, not to IT. Domain owners should be the business leaders already accountable for customer, finance, product or employee outcomes, because a definition owned by engineering will be technically correct and commercially useless, and the commercial team will build a shadow version. IT and data engineering hold the custodian role, implementing controls. Under two hundred people you need one executive sponsor, two or three part-time domain owners and one person doing stewardship as a defined fraction of their job. Councils and formal federation only make sense at larger scale, and imposing them early guarantees they will not be staffed.
Do we need a data catalog to start?
No, and buying one first is a common expensive mistake. A catalog records ownership that humans already agreed to; it does not create that agreement. Below a few hundred actively used tables, a well-maintained spreadsheet outperforms a badly-maintained catalog. The purchase that usually earns its keep first is a quality platform, because it produces evidence immediately and evidence is what keeps a program funded. Before buying anything, govern one domain manually for a quarter, and check what your existing warehouse, transformation layer and identity provider already cover, since a surprising share of governance purchases duplicate capability you are already paying for.