Data design · 2025-01-15 · Go Kyono
AI outcomes are decided by data design
Models change every few months. The structure of your data does not. That asymmetry is the whole argument for where a small company should spend its effort, and most spend it on the half that resets.
01
Two clocks, running at very different speeds
One clock is the model. It ticks fast: capability improves, prices fall, and whatever you built on last year’s model gets cheaper and better without you doing anything. Any advantage you constructed by picking the right model is temporary by construction, because your competitor can adopt the same model on the same afternoon.
The other clock is your data. It ticks slowly, or not at all. The way your organization names things, the relationships it records, the context it captures at the moment of a decision — that changes over years, and it changes only if someone deliberately changes it.
Work on the fast clock evaporates. Work on the slow clock compounds. This is not a subtle point, and it is still the most common misallocation we see.
02
Structure the low-volatility data first
Not all data deserves the same investment. The useful axis is volatility: how often the meaning changes, as distinct from how often the values change.
| Volatility | Examples | What to do |
|---|---|---|
| Very low | Product and part hierarchies, process definitions, organizational structure, account taxonomies, contract templates | Structure it properly. This is the asset. It pays back for a decade |
| Medium | Customer records, supplier terms, price lists | Keep terminology consistent and the relations explicit; the values will churn |
| High | Transactions, sensor readings, logs | Collect cleanly, but do not model it heavily. Its value comes from volume, not structure |
Companies systematically invert this. The high-volatility data is exciting because there is a lot of it and dashboards can be built on it quickly. The low-volatility data is boring, and it is the layer that determines whether any question can be answered at all.
03
What this looks like on a Tuesday
- Pick one process. Extract the nouns it uses. Decide which of them are the same thing.
- Write down the relationships you rely on but have never recorded — which part belongs to which product, which step uses which machine.
- Add the field that captures why at the moment of the decision, not in a monthly reconciliation.
- Name someone who owns the vocabulary, and a way for a new term to get added.
- Fix a set of questions the business needs answered, and measure whether you can answer them. Before and after.
None of that requires an AI budget, and all of it survives the next three model releases. When a model does arrive that could transform your operation, the companies positioned to use it will be the ones whose knowledge was already legible.
Next
Start on the slow clock.
Terminology control, relations, and a measured gold set — that is the ontology service, and the first consultation is free.