Publication

The Federated Data Governance Program

Making data trustworthy enough to put AI on top of it.

Version 1.0 September 28, 2026 Jeremy Taylor About 22 minutes

Download the PDF (37 pages)

The Federated Data Governance Program: People, an AI agent, and Applications draw context from The Catalog, a flat-file cabinet, which in turn is fed by five Data Domains.

Introduction

There's no shortage of discussion that AI needs context, and that context needs governance. Both are true. People want to make it work.

The value is in tying three things together: an architecture where agents are required to use approved data, a process to manage that data, and tooling to implement it. That's what this program is. It's the minimum structure that makes data trustworthy enough to put AI on top of it.

Governance fails when it isn't attached to value and a project requirement. People want to be recognized for delivery and impact, and governance for the sake of governance stalls. So this program is limited, focused, and built directly into the technology implementation. The desire to be part of the deployed agent meets the reality of governing the data the agent uses.

I wrote The Federated Data Governance Program from two directions: the Carnegie Mellon University Chief Data and AI Officer executive program, and my own responsibility for standing up enterprise data governance. This is the companion to Pillar 3, Data Management & Intelligence, of The Enterprise AI Operating Model. It turns that pillar's catalog, ownership, and grading into a program you can run.

The program builds on work by Robert S. Seiner, Zhamak Dehghani, and DAMA. Sources and Influences credits each one and lists what is original here.

How to use this

This is for a mid-size business, described by its shape rather than its revenue: a central data team, data owned by the business, and data created across many teams, systems, and locations. Every role and cadence is sized for that shape. A larger organization can add to it, and a smaller one can take that first step in Where to Start.

Three readers will get the most from it:

  • a data leader asked to get the company's data ready for AI, with no governance program yet
  • a leader whose existing program has stalled and needs resetting around AI
  • a practitioner (an architect or catalog owner) asked for trusted data, with no mechanism to achieve it or declare it

Part 1 explains why the program exists and what it rests on. Part 2 sets out who decides what. Part 3 is how data earns certification. The appendices are reference for the people running the program, and Where to Start is a first step that needs no platform.

The model names no products, so everything specific is a placeholder. The PDF's Appendix A says what to swap and what to keep; Appendix C maps each capability to whatever you run.

Part 1 - Program Overview

Why governance, and what it covers

Data underpins nearly every meaningful decision and product the organization ships, and the volume, complexity, and regulatory expectations around that data keep growing. Without clear ownership, shared standards, and a shared language, teams duplicate work, lose confidence in reports, create compliance exposure, and slow down the analytics and AI initiatives that should be accelerating.

Data is managed the way any company asset is managed: to protect its value and to grow it. For a business putting AI to work, that comes down to one test. Is the data trustworthy enough to put AI on top of it? The program exists to pass that test with the least structure that will do it, and it breaks down into seven goals.

  1. Clear accountability with a named Owner for every domain and every shared asset
  2. Shared language of one definition and name per business concept
  3. Trusted data includes a certification that tells consumers how far to rely on it
  4. Secure by default means classification, privacy, and access are built in from the start
  5. Explorable so people and AI tools can find what exists and what it means
  6. Lifecycle status on every asset so consumers know what is active and what is going away
  7. Scale by automating the metadata capture but not the judgment

The catalog

The catalog is the curated inventory of the enterprise data assets, and the heart of the program. Anything registered in it is an asset. Three audiences use it: people browse it to see what exists, or query it through AI tools connected over the Model Context Protocol (MCP) to test an idea for a new product or capability; AI agents take their context from it, finding the right data by concept name, Tier, and certification; and specialized applications scale from it through its integration options: cached data, exported optimizations, and published contracts.

Everything else in this model exists to make what the catalog says true. The catalog is the context layer, not the data layer: it describes the data and points to it, but it does not hold or serve it. This model assumes a separate data layer that people and agents can query, a data lake or warehouse, or the operational databases themselves, exposed to agents through MCP or an API under the requesting person's access. Building that layer is outside the scope of this model.

A data domain is a logical grouping of data owned by the part of the business that knows it best. Domains follow how the business is organized (customer, product, finance, orders) instead of systems or technologies, and each domain has a named Owner.

A data product is a curated, documented, and supported dataset that other parts of the organization can find, access, and trust. It comes with a data contract: a published agreement describing the schema, refresh cadence and SLA, quality expectations, access rules, and example queries. A dataset shared beyond its origin team without these things is not yet a data product.

A dataset table plus a data contract listing owner, schema and definitions, refresh and SLA, quality expectations, access rules, and example queries.
A data product is a dataset plus its contract. Field names and SLA values are illustrative; the Owner defines the real contract.

Federated ownership

The teams that create data own it. A technology team that builds an application owns the datasets that application produces. A business sponsor who buys an application owns the datasets it produces, even when a vendor runs it. The central data team provides the platform, the standards, and the catalog, and builds shared and aggregate products until the domains can.

The idea comes from data mesh, a sociotechnical approach: an organizational agreement about who owns data, with architectural consequences. This model takes the ownership half on day one and the engineering half as capacity allows. Workloads move to the domains as soon as they can carry them.

Some data originates outside the organization's control: franchisees, partners, vendors, customers. Ownership still lands on whoever brings it in. That Owner chooses between system-enforced rules at intake and cleanup after the fact, and records the choice in the contract.

Seven source domains feed a central data team, which builds shared products for an aggregate domain and for consumers: people, AI agents, and applications.
Seven source domains plus one aggregate domain. Account Health Insights composes data from Account, Subscription, and Usage instead of originating any of its own.

Substitute. The eight domains in the source document are illustrative. Use your own set, and keep the distinction between source domains, which align to the systems that create data, and aggregate domains, which compose from others. Aggregate domains consume definitions instead of creating them, and they are where overlapping concepts get reconciled (see Certifying a product or a domain, below).

Minimally invasive governance

Governance should feel like a tailwind, not a tax. Minimally invasive means four things. The first two adapt principles from Robert S. Seiner's Non-Invasive Data Governance™; see Sources and Influences.

  • Inherent responsibility. Stewardship is a relationship to data. If you can see, update, or define data, you already carry responsibility for it. That is where the role model in Part 2 starts, and it means most of the program describes people who are already doing the work.
  • Apply, don't reinvent. Most people already do the right things informally. The program codifies that, fills gaps, and reduces friction. It adds no approval layers for their own sake.
  • Federation over central control. The center defines what good looks like; domains decide how to achieve it within those guardrails.
  • Automation over policy enforcement. Wherever possible, the catalog, CI/CD, and quality tooling make the right thing the easy thing.

Guiding principles

These principles apply across all data governance activities. Each is a rule a reasonable person could argue against, which is what makes it useful when a decision is close.

  • Minimally invasive. Codify what people already do; add friction only where it pays back.
  • A person decides. Named Owners may use automation to support decisions.
  • Protected by default. In new data and AI applications, unclassified means Restricted.
  • Visible. Definitions, lineage, quality, certification, Tier, and lifecycle status live in the catalog, where data consumers (humans or machines) look.
  • Govern in proportion to value. Effort follows importance. Tier sets the target; not every asset needs Gold.

Part 2 - Roles and Decision Rights

Who is accountable for what, mapped to the five program layers. This part reconciles enterprise governance roles with data mesh execution roles into a single taxonomy. Every role lives at exactly one layer, assigned by the work it does instead of its seniority.

Everyone is already a steward

Before any titles, the role model rests on one idea:

  • If you can see the data, you are responsible for how you use it.
  • If you can update or enter the data, you are responsible for the quality of inputs.
  • If you define data used by your part of the organization, you are responsible for keeping that definition consistent with the enterprise standard.

Being a data steward describes a relationship to data. It is not a job title or a box on an org chart. Most people in the organization are already implicit stewards; the program makes that explicit and gives them the tools to do it well.

AI agents use data too, but an agent is not a steward. Responsibility for what an agent does with data rests with the person it acts for, which is why agents act under that person's access.

The five program layers

The program runs in five layers. The executive layer provides authority and air cover, rarely involved day to day. The strategic layer is the Council: it sets standards and priorities, owns the certification rubric, and resolves cross-domain conflicts on a periodic cadence. The tactical layer is where most of the recurring work happens: owning the domains and certifying their data. The operational layer is everyone who creates, updates, or uses data, the implicit stewards, most of whom carry no governance title. The supporting layer is central services (program management, compliance, security, the data platform).

Five bands stacked from Executive at the top to Operational at the bottom, each paired with its role: Sponsor, Council, Data Owner and Data Steward, Implicit Steward. A Supporting column runs alongside, holding the Program Manager, Regulatory and Compliance, Information Security, and the Data Platform.
Each role sits in the layer where its work happens. The darker the band, the rarer the decision. The supporting layer runs alongside the other four and enables them without deciding.

None of the roles in these layers is a full-time position. Each makes explicit a decision someone already makes, and at a mid-size business one person often holds several: one Owner may own several domains, and the Program Manager is often the head of the central data team.

Unified role taxonomy

The taxonomy merges enterprise governance roles with data mesh execution roles. Every role lives at one program layer.

LayerRoleNamed resourcePrimary accountability
ExecutiveSponsorChief data or technology executiveAdvocate for the program; provide resources and authority; align with company strategy; resolve escalations.
StrategicEnterprise Data CouncilThe Owners of the most-used domains, the head of the central data team, the regulatory and compliance leadSet standards and priorities; own the certification rubric; resolve cross-domain conflicts and competing definitions; review program measures.
TacticalData OwnerManager level: the lead whose team builds the application, or the business sponsor who bought it. One person may own several domains.Value of data within the domain; approve domain policies and access; set Tier and certify the domain's assets (the work may be delegated to a trusted person; the accountability may not). Publish and maintain data contracts for the domain's products; see quality issues through to resolution.
TacticalData StewardAssigned by the Owner, typically an analyst or data engineer who already knows the dataDefine schemas, definitions, lineage, quality rules, and example queries for the domain; keep catalog entries current; triage quality issues to the Owner.
OperationalImplicit StewardEveryone who sees, updates, or defines dataUse data within policy; judge fitness for use from the catalog before relying on it; raise gaps with the Owner or Steward; take responsibility for what they enter or define.
SupportingProgram ManagerOften the head of the central data team, part timeCoordinate roles and teams; run the Council's cadence; track the measures; develop training and data literacy.
SupportingRegulatory and ComplianceLegalMonitor regulatory obligations; identify risk; audit governance processes.
SupportingInformation SecuritySecurityPartner on classification, personal data, and access policy.
SupportingData PlatformThe central data team, with DevOps and cloud operationsRun the catalog, warehouse, and quality tooling the program depends on: uptime, configuration, access controls, backup and recovery.

The PDF's Appendix B also carries a full RACI for common governance activities, mapped against these same roles.

Decision rights and escalation

  • Inside a domain, the Owner decides: definitions, quality thresholds, access lists, Tier, and certification.
  • Across domains, the Council decides: enterprise standards, the certification rubric, and any concept two domains name or define differently. Domains decide whether their assets meet the bar; they do not set the bar.
  • Regulatory exposure or material risk escalates through the Program Manager to the Sponsor, with legal and compliance consulted.

Escalation should be the exception. The program is designed so that most decisions are made at the lowest reasonable layer.

Part 3 - Certifying Data

How data earns Bronze, Silver, and Gold certification, whether it's a single table or an entire domain.

Purpose

Certification is what makes the catalog trustworthy enough to put AI on top of it. An agent choosing between two sources reads the Tier to find which one is authoritative, then reads the certification to decide how far to rely on it. A person deciding whether to build a report on a table does the same.

The unit of certification is an asset: whatever you register in the catalog. The level you register it at is its grain, a term borrowed from dimensional modeling. Four grains are useful in practice: a column, a table or dataset, a data product, and a whole domain.

This is a data maturity model instead of a domain maturity model, deliberately. The same rubric applies at every grain. Maturity belongs to the data at whatever grain you register it. Certification is assigned by the Owner, recorded in the catalog, and shown to consumers there.

The three certification levels

Not medallion layers. Bronze, silver, and gold are also the layer names in the widely used medallion architecture, where they describe how far through a pipeline data has been processed: raw, cleaned, curated. That is a different axis. Here the levels describe how well an asset is described and how much its Owner stands behind it. A raw landing table can be certified Gold if it is fully described, owned, monitored, and classified, and a curated serving table can sit at Bronze if nobody has documented it. If your platform already uses medallion names for pipeline stages, say so when you publish the rubric, or rename these levels.

Each level is a promise, and each includes everything in the level below it.

Before Bronze: registered, uncertified. Anything the catalog discovers, or anyone registers, enters the inventory at once and without a certification. Inventory does not wait on classification: you cannot classify what you have not found. In new data and AI applications, an uncertified asset is treated as Restricted until its Owner confirms a classification. Limited access is the cost of being unclassified, and classifying is how an Owner lifts it. Registering existing data does not revoke anyone's current access.

Bronze: known. Someone answers for it. The asset has a named Owner, a confirmed sensitivity classification, its source system, and a line on what it is for. People and agents can find it and know whom to ask; nobody has yet vouched for its meaning or its quality.

Silver: described. It is explained and measured. A Steward handles day-to-day questions, the fields that matter are defined in plain language, upstream lineage is traced, core quality measures are checked regularly, and the refresh cadence is stated. Consumers can use it with care and know where to go when something looks wrong.

Gold: committed. The Owner stands behind it. Quality thresholds and a freshness commitment are published and monitored, lineage runs end to end, every field a consumer relies on is defined, access is reviewed on a schedule, and consumers get notice before a breaking change. Consumers, human or agent, can depend on it without checking with the Owner first.

A rising staircase from Uncertified (Registered, found but not yet vouched for), to Bronze (Known, someone answers for it), Silver (Described, explained and measured), and Gold (Committed, the Owner stands behind it). Each step up is more metadata, registered in the catalog.
Gold is where business-critical assets aim. Silver is the target for anything shared beyond its team. Bronze is the floor every registered asset should reach within a quarter.

The path from Bronze to Silver to Gold is the path of registering more metadata in the catalog. Once an asset's gaps are visible against the six dimensions, the next piece of work is concrete.

The six dimensions

Owners use the rubric below to certify each asset they own, at any grain, with the best fit: the level that best describes the asset as a whole. The rubric has six dimensions, each named for a question a consumer asks before relying on an asset.

Six boxes: Ownership (who answers for it?), Shared language (what does it mean?), Origin (where did it come from?), Quality (is it right?), Freshness (is it current?), and Protection, marked as the floor (who can use it, and how?).
Each dimension is named for a question a consumer asks before relying on an asset. No dimension outranks another, except Protection: a floor that no weighting can offset.

Best fit, not a mechanical score. An asset that is strongly Silver in most respects and Bronze in one dimension may still be certified Silver if the Owner judges that the overall state warrants it. The gap becomes the next thing to address.

The rubric

Each level assumes the one below it: Silver assumes Bronze is met, and Gold assumes Silver. Owners certify with the best fit.

DimensionBronzeSilverGold
Ownership
Who answers for it?
A named Owner on recordA Steward handles day-to-day questions; consumers know where to askAdvance notice to consumers before breaking changes; a defined escalation path
Shared language
What does it mean?
A concept name and a one-line statement of purposeThe fields that matter are defined in plain language and tied to the shared business glossary; a few example queries answer common questions correctlyEvery field consumers rely on is defined; known caveats, intended uses, and example queries with their expected results are written down
Origin
Where did it come from?
Source system namedUpstream sources and major transformations tracedEnd-to-end lineage, including downstream consumers, so the impact of a change is visible before it is made
Quality
Is it right?
Not yet measured; known issues noted where they existCompleteness, validity, and uniqueness checked on a regular scheduleThresholds published and monitored; consistency reconciled against the source; breaches reach the Owner
Freshness
Is it current?
Time of last update visibleExpected refresh cadence statedFreshness commitment published; late or missed loads raise an alert
Protection (floor)
Who can use it, and how?
Sensitivity classified and confirmed by the Owner. Required for Bronze.Sensitive fields flagged individually; access granted by roleApplicable regulatory obligations mapped; access reviewed on a schedule

Quality criteria follow DAMA's data quality dimensions; see Sources and Influences.

Protection is a floor, not a ladder. Best fit lets an Owner certify Silver while one dimension still sits at Bronze. Protection is the exception at the bottom of the scale: no asset holds any certification until its classification is confirmed. Sensitive data cannot wait for Silver to be handled properly, and the Restricted default for uncertified assets in new applications keeps discovery from outrunning protection.

How certification works

The Owner reviews each asset against the six dimensions and assigns the level that best describes its overall state. The certification is recorded in the catalog and visible to consumers.

Certification is a manager-level responsibility. An Owner may delegate the assessment to a trusted person, often the Steward, but still signs for the result. When a Steward registers metadata that meaningfully improves an asset, they raise it with the Owner for re-certification.

This is judgment-based and deliberately manual. The rubric tells the Owner what to look for; the Owner decides which level fits.

If your platform offers rule-driven certification. Some catalogs can evaluate a rule set and apply a level automatically. That capability reports dimension coverage for the Owner's review, and leaves the assignment to the Owner. An automated rule can confirm metadata is present; it cannot confirm an asset is trustworthy, and switching it on quietly moves accountability from a named person to a configuration. If you do adopt automated assignment, say so explicitly, because it changes what a certification means.

A four-step flow: the Steward registers metadata against the six dimensions, the catalog reports the current state but never decides, the Owner applies the rubric and signs for the level, and the catalog shows the certification to people and agents. A loop runs back from more metadata to re-certification.
The catalog reports; the Owner decides. A certification is a standing commitment by a named person, which is why it is never computed.

Tier, targets, and cadence

Certification says how mature an asset is. Tier says how much it matters. The two are separate, and Tier sets the target that certification is measured against.

TierWhat it meansTarget certification
Tier 1Business-critical. An outage or a wrong answer reaches customers, money, or regulators.Gold
Tier 2Shared beyond the team that owns it.Silver
Tier 3Local to one team.Bronze
  • The Owner sets Tier, and the Council settles disputes. Every registered asset should reach Bronze within a quarter of registration; beyond that, Tier decides where effort goes. A Tier 1 asset certified Bronze is a visible gap and goes to the top of its Owner's list.
  • Only Tier 1 is limited: one Tier 1 asset per concept name. Tier 2 and Tier 3 assets may overlap freely. For an agent, the two labels answer different questions. Tier finds the authoritative source for a concept; certification says how far to trust it.
  • Two cadences keep this current. Each quarter, the Owner reviews the domain: certify newly registered assets, re-certify those that have advanced, check each asset against its Tier target, and decide where to invest next. Twice a year, the Program Manager brings the program measures and the cross-domain patterns to the Council.

What gets certified

Each of the four grains carries a slightly different claim.

  • Column or field. Rarely certified on its own. Its metadata is assessed as part of the table that holds it, chiefly under Shared language and Protection.
  • Table or dataset. The most common grain.
  • Data product. The claim covers the product and its published contract.
  • Domain. The claim covers the domain's own metadata: charter, boundaries, ownership, glossary, and classification profile. At this grain, Shared language is the domain's glossary and Origin is its source systems.

Certifying a product or a domain

A product's or domain's certification is a judgment about its own metadata, informed by its members and never computed from them. It is neither the weakest member's level nor an average, so a domain can be Silver while some of its tables are still Bronze. A computed rule would hold every large domain at its weakest table indefinitely and turn best fit back into a mechanical score.

Consumers will still read a coarse certification as a promise about everything inside it. So the catalog shows the two together, for example Account domain: Silver (3 Gold, 8 Silver, 14 Bronze), and people and agents rely on the certification of the asset they use, never the one above it.

Concept names

Each business concept has one name in the catalog at Tier 1. When two sources overlap, they are different concepts, and they are named for what they are. CRM accounts and billing accounts are not the same thing (a billing system holds customers the CRM never saw), so the catalog holds crm_account and billing_account, and each declares how it relates to the other, usually through a shared key. Two assets named “accounts” is a collision, and it goes to the Council.

Where a presentation layer reconciles the two, that conformed asset belongs to an aggregate domain and carries the concept name. It is usually the Tier 1 asset for the concept, and it is where an agent asked “how many accounts do we have?” should look first.

An agent asks how many accounts we have. Tier finds the source in the aggregate domain's account asset, Tier 1 and Gold, which reconciles both the account domain's crm_account (Tier 2, Silver) and the billing domain's billing_account (Tier 2, Bronze) on a shared key.
Two overlapping sources are two concepts, named for what they are and related by a shared key. The conformed asset carries the concept name and is its one Tier 1 asset, so an agent finds it by Tier and decides how far to trust it by certification.

Certification is separate from lifecycle

Every registered asset also carries a lifecycle status, maintained by its Owner. Lifecycle answers is this still here, and should I build on it; certification answers how far can I rely on it. An active asset can be Bronze, and a Gold asset can be deprecated.

  • In development. Being built. Expect change; do not depend on it yet.
  • Active. In service and supported at its certified level.
  • Deprecated. Still works; do not build on it. A deprecated asset always carries a sunset date, and its known consumers are told.
  • Retired. No longer served. The catalog entry remains so lineage and history stay readable.

Four states are enough. Retention periods, archiving, legal holds, and disposal are records management, owned with the regulatory and compliance function, and sit outside this model.

How Success Is Measured

Most measures below can be counted from the catalog itself, so nobody has to run a survey or a time study. The headline tests the thesis directly, and it needs one more source: the data access layer's query log.

Headline: the share of AI agent queries answered from Silver-or-better assets. The data access layer logs which assets each agent query touched; the catalog says what each was certified at when it ran. This says how much of what agents actually used someone has vouched for. Before that log exists, sample it by hand.

Clear accountability: share of registered assets at Bronze or better, and share of Tier 1 and Tier 2 assets their Owner has reviewed in the last quarter.

Shared language: share of Tier 1 and Tier 2 assets linked to the business glossary, and the number of cross-domain term conflicts open at the Council.

Trusted data: the certification mix of shared assets, the share of Tier 1 assets certified Gold, and the share of Silver-or-better assets passing their quality checks.

Secure by default: uncertified assets in new applications (target: zero).

Explorable: catalog use, counting people and the AI tools and agents connected over MCP.

Lifecycle status: share of deprecated assets with a sunset date and notified consumers.

Scale: median time from registration to Bronze, and the share of certifications signed by Owners outside the central data team.

Where to Start

Start with one domain and one document.

Pick a domain whose data is well understood, used beyond the team that creates it, and matters to the business. Write data contracts as a formal, machine-readable spec: what each dataset is, who owns it, what each field means, how often it refreshes, what quality it promises, who may use it, and a few example queries that answer common questions correctly. An open standard such as the Open Data Contract Standard (ODCS) gives the spec a structure people and tools already recognize. No platform is required.

The contract is useful the day it's written. Added as context to any AI session, it lets people and agents work with that dataset right away: an hour, a day, or a week to value, depending on the dataset's complexity. It's also the first catalog entry, and writing it forces the decisions the rest of this model formalizes: a named Owner, a confirmed classification, and a shared definition.

Then work outward in three moves.

Register. Record the rest of the domain's datasets, each with an Owner, a classification, its source, and a line on what it is for. That is Bronze, and it is most of the value of the first quarter: an inventory people and agents can trust to be complete.

Certify. Take the one asset others depend on most, set its Tier, and bring it to Silver. It becomes the worked example every later domain copies.

Connect. Put the contracts where people already work with AI. Add them as context in an enterprise AI tool, have it generate queries, and run those queries against the dataset under the person's own read access. A Silver contract should produce correct queries without a subject-matter expert in the room. Where it doesn't, the gap is the next thing to write down. That is the thesis working at a small scale.

Keep the contracts in one shared, versioned repository so every session reads the current copy. Move them into a catalog exposed over MCP when files stop scaling: usually when agents need to choose between sources by Tier and certification, or when a second domain arrives.

Then add the next domain. The Council forms when a second domain produces the first cross-domain conflict worth resolving, usually two teams using the same word for different things.

What Changes

In practice, a lot changes.

Teams stop pulling time-constrained subject-matter experts into meetings to rediscover what has been known and discussed many times. The context is written down once, has an Owner, and is read as often as it's needed.

Every role works from that established context. Product ideation, architecture drafts, development work, and support all start from the same definitions, lineage, and contracts.

AI can self-serve answers that have historically required someone to run ad hoc BI or analysis. The agent answers from certified data, so the person asking can see what it relied on and who stands behind it. The headline measure (the share of agent queries answered from Silver-or-better assets) is how you watch that shift happen.

Adapting This to Your Organization

The document is built to be taken and made your own. It names no products: where the model needs a capability (a catalog, a quality framework, lineage capture), it describes the capability and leaves the tool to you. The PDF's Appendix C is a worksheet for mapping each capability onto whatever you run.

Everything specific is a placeholder. Organization names, domain names, role titles, named resources, field names, and example data are all generic. Substitute your own catalog or written contracts, your own domains, your own worked example in place of Account Health Summary, and your own governance body in place of the Enterprise Data Council. Name actual people and titles in the role table; an unnamed role is an unowned role.

The certification colors carry meaning. Bronze, silver, and gold are the only saturated colors in the source document. If you re-skin it for your own brand, keep that discipline and give the certification levels the loudest colors you have.

Keep, unless you have a deliberate reason: certification as a human judgment rather than an automated score; one named Owner per domain; the Council setting the bar while domains decide whether they meet it; one name per concept and one Tier 1 asset per name; certification and Tier staying separate; unclassified meaning Restricted in new applications; and a supporting layer that enables but does not decide.

Adjust freely: the wording of the six dimensions, which grains you certify at, your own review cadences, and the certification level names (Bronze, silver, and gold travel well, but Known, described, committed works too if your platform already uses medallion names for pipeline stages).

Continuing the Conversation

If you're working through any of this or think part of the program is wrong, I'd like to hear about it: jeremy.taylor@ossnetworks.com or linkedin.com/in/jeremytaylor00. The Enterprise AI Operating Model, which this program expands from Pillar 3, and additional articles are at ossnetworks.com.

About the Author

Jeremy Taylor

I lead global data and AI organizations with twenty years in technology. I've spent the past eleven years building and transforming cloud data platforms. Most recently, I've focused on adding an AI platform layer and agents in production, and I've had responsibility for standing up enterprise data governance.

I completed the Carnegie Mellon University Chief Data and AI Officer executive certificate in 2026. This program came out of that work.

Sources and Influences

This document is a synthesis. Where it builds on other people's work, that work is named here. Everything not named below is my own.

Governance philosophy. The stewardship philosophy in Part 1 and the idea that everyone is already a steward build on Robert S. Seiner's Non-Invasive Data Governance™ (KIK Consulting & Educational Services): that stewardship describes an existing relationship to data, that a program should formalize what people already do, and that accountability is layered across the organization. The five program layers follow his operating model of roles and responsibilities. Read the source directly: Seiner, Non-Invasive Data Governance: The Path of Least Resistance and Greatest Success (Technics Publications, 2014), and Non-Invasive Data Governance Strikes Again (2023).

Trademark notice. Non-Invasive Data Governance is a registered trademark of Robert S. Seiner and KIK Consulting & Educational Services. The minimally invasive framing in this document is the author's own adaptation and extension of those ideas. It says minimally because that is the more honest framing: any program asks something of the people in it, and the promise here is to keep that small. It is neither affiliated with nor endorsed by Robert S. Seiner or KIK Consulting, and it should not be read as a restatement of the Non-Invasive Data Governance method.

Data architecture. Federated ownership, data products, and aggregate domains follow Zhamak Dehghani, Data Mesh: Delivering Data-Driven Value at Scale (O'Reilly, 2022). This model does not require a mesh implementation.

Data quality. The measures in the Quality dimension (completeness, validity, uniqueness, and consistency) and the treatment of timeliness as a dimension of its own follow the data quality dimensions set out by DAMA: the DAMA UK Working Group, The Six Primary Dimensions for Data Quality Assessment (2013), and DAMA International, DAMA-DMBOK: Data Management Body of Knowledge, 2nd edition (Technics Publications, 2017).

Data contracts. The Open Data Contract Standard (ODCS) is an open standard from Bitol, a Linux Foundation AI & Data incubation project, published under the Apache 2.0 license.

Common practice. Grain is Ralph Kimball's dimensional-modeling term. Tier for importance and bronze, silver, and gold for certification are conventions found across enterprise catalog products. The names of the other dimensions (ownership, lineage, classification) are common practice. RACI and decision rights are ordinary management practice.

Agent context. The Model Context Protocol (MCP) is an open standard published by Anthropic.

What's original: a minimal governance program joined to a data catalog, so the subject-matter experts who know the data curate the context AI works from. They designate what is trusted and how much, and a named person's judgment is required where it counts. The mechanisms are the author's own: the six-question rubric with Protection as a floor, one Tier 1 asset per concept name, and certification at any grain without a roll-up formula.


© 2026 Jeremy Taylor CC BY 4.0 Version 1.0, published September 28, 2026

Download the PDF (37 pages)