Who Are You Optimizing For?
When it comes to the RFP process, the evaluation rubrics I see from the partner side of the table are remarkably consistent. Query performance gets a row. Ingestion throughput gets a row. Cost per compute unit usually gets its own tab. Somewhere near the end there is a slide about business value, and that slide is almost never scored.
That thin column is not a leadership failure. It is a population problem.
The people running a platform evaluation are, with few exceptions, the people who will live inside the tool. Engineers, architects, platform leads. They score what they can feel. A benchmark number changes their Tuesday. A migration timeline changes their quarter. What the platform does for whoever consumes its output is a return that shows up eighteen months later if it shows up at all, and no evaluation deadline is built to wait that long. So the bake-off gets decided by the room, and the room is never the population being asked to trust the answer.
The data confirms this disconnect. Recent surveys show that while 75% of C-suite leaders believe their teams are 'data proficient,' half of middle managers disagree. This gap is becoming even more expensive in the age of AI; Gartner warns that 60% of AI projects will be abandoned because they lack 'AI-ready' data. The prediction isn't just about accuracy—it’s about whether the system can actually function in a production loop.
The data confirms this disconnect. Recent
The operational version of the same gap is more recent and more expensive. Gartner's February 2025 research expects organizations to abandon 60 percent of AI projects that are unsupported by AI-ready data. That conditional matters and it gets dropped constantly. The prediction is not that 60 percent of AI projects fail. It is that 60 percent of the ones built on a foundation nobody scored for them will be walked away from. Same survey found 63 percent of organizations either lack the right data management practices for AI or are not sure whether they have them.
Many right now are discussing how to make the answer trustworthy. I want to pick up one step downstream of that. Once the answer is trustworthy, who is on the other end of it, and what do they actually need in order to do something with it?
Conversational interfaces made it tempting to treat the agent as one more consumer persona. Add a row to the rubric and move on. I think that move is worse than not scoring it at all.
Here is why. When a human consumer gets a platform that does not fit their day, you find out. They stop opening the tool. Adoption flattens, someone escalates it, and eventually a training budget appears. The failure is loud and slow, which is another way of saying it gives you time. When an agent gets state that does not fit its decision loop, nothing flattens. The agent acts anyway. It holds the transaction or releases it, escalates the claim or closes it, and it does all of that with exactly the confidence it would have had on fresh state. There is no adoption curve to watch. The signal you would have used to detect sthe problem is the thing the agent just removed.
That asymmetry is why an agent row on a builder-written rubric buys false comfort. The row will say something like "supports AI and agent workloads," someone will check it, and the platform will ship. Nobody scored the thing that actually breaks.
The thing that actually breaks is freshness at the moment of decision, and whether the action is governed on the way out.
"Can the agent reach the data" is a build question and it is easy to answer yes to. A gold table refreshed hourly is reachable. It is also wrong for a decision that has to resolve in the next four hundred milliseconds. Materialized state in an Eventhouse with a hot cache policy sized to the decision window is a genuinely different artifact from the same numbers arriving in a nightly semantic model. Both are available. Only one of them is true when the agent asks.
Governance is the other half and it degrades the same way. A human who acts on a wrong number leaves a trail that someone can question in a meeting. An agent acting a thousand times an hour needs that trail built into the action path. If you have to reconstruct it later from logs never designed for the question, it isn't an audit. It’s archaeology.
None of this makes builder-optimized scoring wrong. There are real cases where the builder is the consumer and scoring for them is scoring correctly. A data product with a handful of internal engineering consumers qualifies. So does a feature platform whose downstream consumer genuinely is another pipeline. The honest objection to my argument is not that builders should not matter. It is that there are more builders in the room than consumers, so of course the rubric reflects the room.
That objection is fair. Influence is still not headcount. A platform used by five hundred analysts and acted on by a growing fleet of agents does not get scored by the twelve engineers who showed up to the meeting.
So before the next evaluation, add one row that is not a capability question. Not whether the agent can access this. Ask whether the state is current enough to be true at the moment of the decision and whether the resulting action is governed on the way out. If you cannot answer that from the artifacts already in front of you, then you have not scored the platform. You have scored the build.
Think about the last evaluation you ran. Would the agent you are about to point at it have passed?