Dispatch № 37Data governance7 min read

DMBOK Eats Your AI Roadmap.

Every ambitious AI milestone quietly assumes a data capability nobody funded, and the unglamorous work then consumes the time that was allocated to the interesting part.

The assumption inside every milestone

Take the retrieval assistant, the least ambitious item on the list.

It assumes a document set where the current version of a policy is identifiable, which is a records management problem. It assumes access rights that a machine can read at query time, which is a data security problem expressed as metadata rather than as a folder permission somebody set years ago and never revisited. It assumes a retired procedure is distinguishable from the one that replaced it, which is lifecycle and lineage. It assumes somebody can say which repository is authoritative when two of them disagree, which is architecture.

None of those are model work. All of them are prerequisites, and the roadmap treats them as ambient conditions rather than as deliverables with an owner and a budget line.

A roadmap is not a plan for building models. It is a bet that a set of data capabilities already exists. The bet is usually placed without anyone realising a bet was made.

Where the timeline actually goes

The interesting observation is not that these problems exist. Everyone who has shipped anything knows they exist. The observation is where they consume time.

They do not consume it at the start, when they would be visible and could be planned. They consume it in the middle, after the model works in a notebook and before it works in front of a user, in the stretch of the project that has no name and no line item. The demonstration is convincing early. The system is still not in production long after the date on the slide, and the intervening months went to reconciling two definitions of an account, chasing an owner for a field nobody can explain, and discovering that the historical extract used for training was filtered by a rule the pipeline no longer applies.

That is the sense in which DMBOK eats the roadmap. Not because data management is large, though it is. Because the schedule was drawn as though the data work were already done, so the data work has nowhere to go except into the space reserved for building the thing.

Lineage is nobody's job by design

In DMBOK, lineage does not stand as its own knowledge area. It lives inside metadata management, which is correct in the taxonomy and unhelpful in practice, because taxonomy is how organisations decide what gets a budget.

A capability that appears as a subsection of a knowledge area that most organisations have never staffed will not be funded, and it will not be owned. So when a regulator, an auditor, an internal risk function or a customer asks how a particular figure was produced, the answer is assembled by a person working backwards through a pipeline they did not write, over several weeks, with a confidence level nobody would put in writing.

Metadata is not documentation. It is the difference between an answer and an answer that survives being questioned. An AI system raises the frequency of that questioning sharply, because the outputs are novel, the reasoning is opaque, and the people affected by them are more inclined to ask.

The threshold nobody set

Data quality has a specific failure mode in AI programmes, and it is not dirty data.

It is the absence of a stated threshold. A regulated lender can run a customer table for years at a level of completeness that supports periodic reporting perfectly well, because reporting aggregates and aggregation forgives gaps. A model does not aggregate. It acts on the row. The same table that was fine for years becomes the reason a decision cannot be defended, and nothing about the table changed.

The question that was never asked is: what is good enough, for this use, and who decided. In a reporting context the answer stayed implicit and stayed adequate. In a decisioning context it has to be explicit, because it is now a statement about the conditions under which the organisation is prepared to be wrong.

Data quality is not cleanliness. It is a decision about the errors an institution is willing to accept and defend. That decision belongs to a business owner, not to an engineer, and it takes longer to obtain than any of the engineering it governs.

Master data decides what the model is even about

Reference and master data is the least glamorous area in the entire body of knowledge and the one most likely to stop a roadmap outright.

If an institution holds four systems that each maintain their own notion of a customer, a model trained across them is not learning about customers. It is learning about the union of four incompatible definitions, and every downstream statement about performance inherits that ambiguity. The metrics will look fine, because the metrics are computed over the same ambiguity.

This is the failure that does not announce itself. Quality problems produce visible errors. Master data problems produce a system that works, ships, and is quietly answering a different question than the one it was commissioned to answer.

Reading a roadmap backwards

The useful exercise takes an afternoon and is uncomfortable in a way that is worth the discomfort.

Take each milestone on the roadmap. For each one, write the data capabilities it assumes, in the vocabulary of the knowledge areas rather than in the vocabulary of the project. Then, for each capability, answer three questions.

Does it exist today, verified by someone looking rather than by someone remembering. Who owns it, by name, and does that person know they own it. If it does not exist, which milestone is currently funding the work to create it.

The third question is the one that produces the silence. In most roadmaps the answer is none of them, because the roadmap was built as a sequence of AI deliverables and the data work was assumed to be somebody else's programme, running in parallel, on a schedule nobody checked.

A roadmap that survives this exercise has fewer milestones and later dates. That is not a downgrade. It is the first version of the plan that describes what will actually happen.

The unglamorous half is the deliverable

There is a reasonable objection to all of this, which is that no organisation is going to pause AI delivery to complete a data management programme end to end, and none should. The knowledge areas are not a gate to be passed before AI begins.

They are a scope to be drawn per milestone. The retrieval assistant does not need enterprise metadata management. It needs metadata for one document corpus, with an owner, a currency rule and a machine-readable access model. That is a short set of decisions and a bounded piece of work, and it is the difference between a system that ships and a pilot that stalls in the unnamed middle.

What changes is where the work appears. Not as a dependency assumed to be satisfied, but as a funded line inside the milestone that requires it, with the same visibility as the model work and the same right to consume time.

DMBOK does not eat the roadmap because data management is enormous. It eats the roadmap because it was never written on it, and work that is not on the plan still takes the hours it takes. The organisations that ship governed AI on schedule are not the ones with better models. They are the ones whose roadmap admitted, in advance, what the model was standing on.

© 2026 Nabeel Khan. DMBOK Eats Your AI Roadmap is published under CC BY-NC-ND 4.0. Quote it, cite it, do not repackage it.

Keep readingMore dispatches2026
Fin · № 37