You cannot skip the "boring" part before the flashy AI automation
Generative AI cannot answer what it has no access to, and that limit is structural rather than a bug waiting to be fixed. It is why collecting and storing your data properly still comes first.
Every week there is a new generative AI model, and we enthusiastically talk about the same one thing every time; what it can do, what it beats, what we can build with it now.
We rarely talk about the opposite side- what these models cannot do, and specifically what generative AI is not able to handle by design, structurally. Not the things that get fixed or overcome in the next release. The things that are not going to be fixed, because they are structurally impossible.
They cannot answer what they do not know
A model can never tell you how many units are left in your warehouse if it does not have access to that data.
That is the whole limitation, and it sounds too obvious to be worth writing down.
What makes it worth writing down is that the model answers anyway. It does not stop and tell you it has no idea, even if they use words like “honestly” or “genuinely.” OpenAI published a paper in September 2025 called Why Language Models Hallucinate, and the argument in it is that the way these models are trained and scored rewards guessing over admitting uncertainty. Their example is a birthday. Guess a date and there is a one in 365 chance of being right. Say I do not know and the score is zero, every time. So it guesses.
That is why I call this structural. The model is not neglecting to tell you the data is missing. It was built to produce an answer regardless.
A surprising number of people forget this
This happens more often with upper management who want to use AI at the company level, and I understand why. From that seat you are looking at a capability, not at a data pipeline.
The numbers on this are not subtle. Gartner predicted in February 2025 that through 2026, organizations would abandon 60% of AI projects that were not supported by AI-ready data. In the survey behind that, 63% of organizations either did not have the right data management practices for AI, or were not sure whether they had them.
RAND interviewed 65 data scientists and engineers in 2024 about why AI projects fail. Two of the five root causes they came back with are the ones I am describing here. The organization does not have the data it needs, and the organization has not invested in the infrastructure to manage it.
So this is not a niche problem, and it is not a technical one. It is the most common way these projects die.
You cannot skip the boring part
What I tell clients when this comes up is that we can never skip the boring foundational work of collecting and storing data in the correct way, just like we used to do before AI, in order to get to the flashy part of using AI to automate the workflow.
Monica Rogati wrote this up in 2017 as a hierarchy of needs, borrowed from Maslow. Collection at the bottom, then moving and storing, then exploring and transforming, then aggregating and labeling, and AI sitting at the very top of the pyramid. The reason it is drawn as a pyramid is that you do not get to skip a layer.
Nine years later, with models that did not exist when she drew it, the bottom of that pyramid has not changed at all.
Citations
- Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate, OpenAI, September 2025
- Gartner, Lack of AI-Ready Data Puts AI Projects at Risk, February 2025
- Ryseff, De Bruhl and Newberry, The Root Causes of Failure for Artificial Intelligence Projects, RAND, August 2024
- Monica Rogati, The AI Hierarchy of Needs, 2017