Week six is too late to meet your data
LUNTA · · 3 min read
Every enterprise believes it has the data, and most of them do — somewhere, in some state, under some ownership. The distance that matters is between ‘we have the data’ and ‘the data can support this decision, at this quality, for the people allowed to see it’. Organisations tend to measure that distance in week six of a build rather than week one of a diagnosis.
The schema describes an aspiration
The documented data model is what the system was designed to hold. What the tables hold is whatever a decade of workarounds put there. An optional field everyone actually depends on, populated often enough to look reliable and empty often enough to break the answer. A status column with far more values in its enum than in use, two of them meaning the same thing for reasons nobody still present can reconstruct. A free-text note carrying the real reason a decision was made — which is why the structured field beside it is not evidence of anything.
None of this is anyone’s fault. It is what production systems look like after being useful for fifteen years. It is also invisible to anyone reading the diagram rather than querying the tables, and the diagram is what gets shown in the scoping workshop.
Existing is not retrievable, and retrievable is not permitted
A corpus that lives as scanned attachments is not a corpus a retrieval layer can reach; making it one is an ingestion project, and nobody scoped an ingestion project. A source governed by row-level entitlements has to carry those entitlements through retrieval and into the answer, which is an architectural property rather than a feature added afterwards — and that requirement surfaces at security review, once the system exists and has a date attached.
Then ownership. Ground truth needs a maintainer, and a surprising amount of enterprise reference data is maintained by nobody in particular: a mapping spreadsheet nobody has owned since its author left, a policy document superseded in practice but not in the repository, a lookup table last edited two reorganisations ago. A system grounded in stale truth is not merely unhelpful. It is confidently wrong, which is the failure mode that costs an operation its willingness to use the thing at all.
Buy the finding early
This is why a diagnosis produces a data-reality assessment for every candidate before anything is built: what exists, what is retrievable, what is fiction, who owns it, and what the entitlement model will demand of the architecture. It is among the cheapest artifacts an engagement produces and routinely the most consequential, because it converts the most common causes of a failed build from risks into findings — and a finding has a date, an owner, and a price, while a risk has only a colour.
The verdict is not always green. Sometimes it is wait: the use case is sound, the corpus is not reachable until a records migration completes, and the correct move is to re-score the candidate afterwards rather than engineer around a gap that is about to close. Sometimes it is do not build, because no quantity of model quality manufactures a fact the business never recorded.
A readiness problem found in week one is a scoping decision. The same problem found in week six is a change request, a slipped date, and a meeting in which engineering explains something the business is certain it already mentioned. The work is identical. Only the price of knowing has changed.
Read next
Evaluation gates exist to kill workstreams. If yours can’t, they aren’t gates — they’re theatre.
This is how we deliver, not only how we write.
See the full delivery systemSee what an engagement producesStart a diagnosis