July 14, 2026·Architecture

Score It From the Other Chair: The RFP Section Nobody's Writing

data-strategyagentic-aigovernanceaction-systemsplatform-evaluationrfp

Yesterday's piece named the problem: platform evaluations get scored by builders, not by the people or agents who have to trust the output and act on it. If you read it and nodded along, you're probably already picturing your last RFP. Here's the part that piece couldn't give you: an actual template, a section you can drop into that RFP today, that scores the thing your current evaluation skips.

What's already in your RFP, and why it isn't enough

Pull up almost any enterprise data platform RFP and the shape is familiar. TEC's BI RFP template runs past 1,600 criteria across 14 areas: dashboards, self-service BI, mobile access, alerting, reporting, OLAP, data warehouse, ETL, integration, workflow. CDP.com's RFP guidance shows the same shape in newer language: weighted scorecards built around integration, identity resolution, governance, real-time processing, cost. Every one of these answers a builder's question: can we connect it, can we scale it, can we afford it, can we report out of it.

2026 added a genuinely new category on top: AI-readiness. Techment's seven-criteria framework scores data quality, interoperability, governance, scalability, observability, AI-native capabilities, and operationalization. Go further into agent-specific procurement and one 2026 AI procurement playbook weights "MCP and agentic AI governance" at 20 percent of the total score, explicitly flagged as the category most vendors will struggle to answer well.

That's real progress. It's also still the same blind spot in a new outfit. Reasoning-trace capture, tool access permissions, escalation controls, these are all "can we monitor and control the agent" questions. Not one of them asks whether the agent, or the analyst, on the other end actually gets something usable without leaving their workflow. The industry updated its RFP for the agentic era and reproduced the exact same blind spot one layer down.

The template: an RFP section your builders didn't write

Add this as its own scored section, run alongside your existing technical scorecard, not instead of it. Score every row 0 to 3: 0 means the criterion wasn't considered at all, 1 means it came up informally with no real weight, 2 means it was formally scored but weighted below your builder-side criteria, 3 means it was weighted equal to or above them.

Step 0, before you score anything: who is the primary consumer of this platform's output, a builder or pipeline, a business user, or an agent acting on business decisions? If the honest answer is "another pipeline built by the same team," stop. This template doesn't apply, score the platform on your standard technical criteria instead. That's not a loophole. A platform built for engineering consumers should be scored by engineers.

Consumer trust in the answer

| # | Question | What good looks like | What it looks like when it fails | |---|---|---|---| | 1 | Can the actual consumer interpret the output unaided? | An account manager reads the retention-risk score and knows what it means without calling the data team. | The number requires a Slack message to the analyst who built the model to explain what it's actually measuring. | | 2 | Does the platform surface its own uncertainty, or present every answer with equal confidence? | A forecast ships with a confidence band and a data-freshness timestamp. | A stale number and a fresh one look identical on the screen. |

Context switching

| # | Question | What good looks like | What it looks like when it fails | |---|---|---|---| | 3 | Does the answer arrive inside the tool the consumer already uses, or does it require opening a separate one? | The risk flag shows up directly on the account record inside the CRM. | The capability to embed exists on a slide. Nobody built the integration, so the analyst opens a second tool and looks it up by hand. | | 4 | Can the consumer ask in their own terms, or do they need to learn a schema? | "Which accounts are at risk this week" returns an answer. | The question only works if it's phrased in the platform's own field names. |

Agent readiness (score architecture, not just current usage, this applies even if no agent exists yet)

| # | Question | What good looks like | What it looks like when it fails | |---|---|---|---| | 5 | Could an agent reach this state programmatically, in a useful time budget, today or next cycle? | The same data an analyst sees on a dashboard is queryable by an agent in milliseconds. | The only path to that number is a nightly export or a scheduled report. | | 6 | Execution: does the platform natively support automated threshold-triggered adjustments (reorder points, price changes), escalation into an owned queue, and record creation, or does "action" just mean a notification? | A fraud score above threshold automatically holds the transaction and opens a case, no custom code required. | The vendor's answer to "can it act" turns out to be "yes, it can send an alert." | | 7 | Can you reconstruct why an action happened, without the builder team's help? | A director queries the decision trail and gets the answer directly. | Explaining a decision requires paging the engineer who built the pipeline. | | 8 | Does response speed match the actual decision cadence? | A claims-holding agent gets an answer in under a second, because that's the window the decision actually needs. | The platform is technically reachable, but the round-trip is minutes, and the decision needed milliseconds. |

The evaluation process itself (always applies, regardless of consumer type)

| # | Question | What good looks like | What it looks like when it fails | |---|---|---|---| | 9 | Did an actual consumer sit on the scoring committee, not just the builders? | A business stakeholder or a downstream ops lead scored the demo alongside engineering. | Every name on the scoring sheet reports to the same platform team. | | 10 | Does the written RFP weight consumer-facing criteria equal to or above builder-facing ones? | Rows 1 through 8 above carry real point value in the final score. | They're a debrief-meeting afterthought, if they're discussed at all. | | 11 | Can the consumer flag a wrong or unhelpful answer in a way that improves the platform, without routing through the builder team? | A thumbs-down on an answer creates a ticket the platform team actually triages. | The only feedback channel is an email to whoever built it, if anyone thinks to send one. |

Score on evidence, not claims

Ask for a reference customer already using each capability the way you intend to, not the vendor's own description of it. This is the same discipline the industry still owes itself on self-reported context-layer and ontology benchmarks: a claim on a slide and a capability running in production are different things, and an RFP response optimized to win the deal will blur that line for you if you let it.

What the score tells you

Mostly 0s and 1s: this is a builder-scored platform, which may be exactly correct if your primary consumer really is another pipeline. Mostly 2s: consumer-aware, someone's thought about it, but the evaluation still didn't weight it. Mostly 3s: consumer-scored and agent-ready, the platform that wins your technical bake-off is actually the one your business, and increasingly your agents, can use.

A platform can score well on your existing technical RFP and still fail this section entirely. That's not a contradiction. It's the same builder blind spot yesterday's piece named, just made checkable instead of implied.

Where this fits

This section runs alongside your existing RFP. It doesn't replace it. The core platform questions, integration, scale, cost, still matter and still deserve their own weight. What changes is that "will the consumer or agent on the other end actually use this" stops being an assumption and becomes a scored line item, the same as everything else in the document.


Sources