Back to the newsletter

    The Costly Mistake of Measuring AI Before You Define It

    Only 1 in 5 revenue teams can prove AI ROI. The cause is a definitions problem: no baseline, no shared taxonomy, no audit trail, and no boundary set before the agent shipped.

    What the CAC? issue 7 cover: oversized serif headline on a dark editorial background with a Stack Finder green sunk motif.
    The Costly Mistake of Measuring AI Before You Define It

    Only about one in five revenue teams can point to measurable ROI from their AI initiatives. I used to think that number would embarrass people. Now I think it should relieve them, because the real reason is almost never the one everybody assumes. It isn't the model. It isn't the budget. It's that most teams never wrote down what they were trying to prove before they started trying to prove it. That's the thread underneath everything I want to talk about this week: a stat that names an ROI problem that's really a definitions problem, a taxonomy built for AI security that revenue teams still haven't built for their own pipeline, an audit trail nobody keeps until it's too late to use it, and a boundary question most teams don't ask until an agent has already crossed it. Different rooms, same root cause. Undefined is ungovernable, and ungovernable is expensive.

    Why Are Only 1 in 5 Revenue Teams Seeing Real ROI From AI, and What's Actually Causing the Other 4 to Come Up Empty?

    Because most teams never wrote down a hypothesis, a baseline, and a plan.

    A recent industry stat put a number on something I've been watching for a while: only about 20.6% of revenue teams are seeing measurable lift from AI. That tracks with what I see in the field, and the bottleneck isn't access to tools. It's operational: deal visibility, CRM hygiene, and workflow control. Something like 90% of teams are working with data too messy to trust, and messy data makes it nearly impossible to establish the one thing every experiment actually needs: a real hypothesis starting point.

    Here's the sentence almost nobody writes down before they start: “we convert X% of inbound to pipeline today, and we expect this to get us to Y% in 90 days.” Without that sentence, in writing, before the agent goes live, there's no way to prove the initiative worked, and just as importantly, no way to prove it didn't. Two quarters later, a CFO asks what the AI initiative actually returned, and the room goes quiet, not because the answer is bad, but because nobody can produce an answer at all. I've watched this exact scene play out enough times that I could write the stage directions.

    It's the same mistake I've watched marketing teams make for years, just wearing a new costume: pulling in reporting teams or analysis after a campaign launches instead of during the strategy phase, when the data architecture could still have been mapped out on purpose. Running AI is closer to running a real experiment than it is to flipping a switch. It needs a hypothesis, a baseline, and a plan, or it isn't actually being measured, it's just being hoped at.

    Part of the problem is that teams are reaching for the wrong ruler entirely. Leads, pipeline created, revenue, CAC:LTV, these don't disappear, they become the baseline everyone still needs, but they were never built to catch what an agent is actually doing differently. An agentic environment throws off new signals worth tracking on their own: how many agent calls it takes to reach a successful outcome, how long resolution actually takes, how often the agent passes its own governance checks, how often it escalates to a human, and what a single decision costs to produce. None of that shows up on a legacy dashboard, and none of it gets measured if nobody agreed to look for it in the first place.

    None of this works, either, without every team actually agreeing on what they're measuring toward. RevOps wants to know if pipeline is moving faster, not just that it exists. Marketing Ops wants campaign speed and consistency, not a launch count. Customer Success wants adoption quality, not login counts. Finance wants margin by outcome, not by seat. Four departments, four different definitions of success, and if nobody reconciles them before the agent ships, the ROI conversation two quarters later isn't really about AI. It's about four teams discovering they never shared the same goals, and were ultimately never measuring the same thing in the first place.

    The workflow that actually produces ROI is almost boring in how simple it sounds: pick one problem, get every team to agree on what “done” looks like, baseline it accurately, write the hypothesis and the expected lift, then measure and expand. Nearly everyone skips the baseline step. The agent goes live, the quarter runs, and the hunt for ROI starts with no starting line to measure from. That's the entire lag, right there. The agents are doing the work. Almost nobody benchmarked the outcome first, so the lift has nowhere to actually show up.

    And the baseline problem doesn't stand alone. You can't even write an honest hypothesis if the words in it mean something different to every team reading it, which is exactly the gap that showed up somewhere else entirely this week.

    What Does a Security Taxonomy Built for AI Agents Have to Do With Why Your Own Forecast Is Wrong?

    If your revenue teams can’t agree on what “qualified means”, your forecasts may be accurate for one team, but not another. A taxonomy is just a shared definition of what words mean.

    An executive said something to me recently that I hear constantly in some form: “we'll just swap out our APIs for an MCP.” For anyone who hasn't run into the term, an MCP, short for Model Context Protocol, is essentially how an AI agent plugs into your tools, apps, and data. What that executive was missing is that an MCP connection needs a taxonomy just as much as it needs a protocol, and a taxonomy is nothing more exotic than a shared way of naming things so every team and every system mean the same thing when they say it.

    Last week, OpenAI published exactly that for AI agent risk: an open, vendor-neutral framework creating one shared vocabulary for MCP security risks, weaknesses, attack patterns, controls, detections, and test cases. Swap the word “security” for “revenue” in that sentence, and you're reading a data architecture document most GTM teams still don't have.

    Try this exercise sometime: ask Marketing, Sales, and Customer Success what “qualified” means at your company. You'll get three answers, close enough that nobody starts a fight over it in a meeting, different enough that the exact same account shows up as a win in one dashboard and a leak in another. Humans have always papered over that gap with judgment and a quick hallway conversation. A new rep gets weeks of onboarding, a manager to ask, and a Slack channel full of “wait, does this count?” An agent gets a system prompt and a CRM connection. When two teams define the same term differently, the agent doesn't pause to ask which one is right. It picks one and runs, at scale, immediately.

    A real taxonomy fixes that at the foundation, not the edges. It means humans and agents make the same call on the same input, repeatably. It means a user story gets written against a term everyone actually recognizes, so “done” isn't a debate each time. It means an error is easier to catch because there's a source of truth to check the output against instead of a vibe. And it means the agent talks about your product the way your actual team was trained to, not the way one prompt happened to guess.

    I want to be honest about the limits here: a shared vocabulary doesn't stop an agent from making something up. It shrinks the room where ambiguity can hide, and it makes the drift visible the moment it happens, which is the entire game. You cannot fact-check an output against a definition your own teams never agreed on in the first place.

    Security got its common language last week. Most revenue teams are still quietly arguing about what a lead is, which means the baseline from the last section was probably softer than anyone realized the moment you actually poke at the words holding it up.

    Related: stackfinder.com/answers/user-stories-and-taxonomy

    What Happens When an Agent Hits a Case Nobody Defined, and How Do You Actually Learn From It?

    You need an audit trail. Log the action, the input, and the credentials it ran on, not just the output it produced. Give every edge case of strange result an incident ticket. Write down what the case taught you and exactly which rule changed because of it. And most importantly, have a roll-back feature that helps humans decide how far you’re comfortable letting it run in the first place.

    Most teams pay for the exact same mistake again in three weeks when a different agent hits a parallel edge case. Picture an agent making five hundred small edits to public web pages over a single weekend, and the only record your team has afterward is the token bill. Nobody told it to do that. It hit a page it had never seen before, made a reasonable-looking call with the context it had, and kept going. By Monday morning, the real question isn't whether the agent broke. It's whether anyone can reconstruct exactly what it did, and whether it can actually be walked back.

    That scenario isn't hypothetical anymore. OpenAI put a name to it last week, after agents wrote to public pages during a real incident, calling for disclosure standards that treat agent misalignment as something worth reporting beyond a model or system card. I agree with the instinct, but I think most people are hearing it as a frontier-lab problem when it's really a Tuesday-afternoon problem for anyone running agents inside their own GTM stack.

    Here's how I watch most companies actually handle it. The agent does something a little odd. Someone catches it, maybe. Someone patches the prompt. Everyone moves on with their day. Three weeks later, a different agent hits a parallel edge case, and the patch doesn't cover it, because nobody wrote down what happened the first time. That's an edge case treated purely as a surprise. The teams I trust the most treat it as a lesson with an actual owner attached to it: log the action, the input, and the credentials it ran on, not just the output it produced. Give every strange result an incident ticket, even the small, easy-to-shrug-off ones. Write down what the case taught you and exactly which rule changed because of it. And know how you'd roll it back before you decide how far you're comfortable letting it run in the first place.

    An audit trail is what turns a one-off into something a team can actually replay, learn from, and push back on. Paperwork filed after the fact does none of that. This isn't a new idea dressed up in AI language, either. It's the same discipline behind running a real experiment: knowing exactly how the last one was set up, what was done, and what the strategy behind it was, so the next attempt can actually build on it instead of starting from memory.

    Ask yourself honestly: the last time an agent on your team did something unexpected, could anyone reconstruct exactly what happened, or did someone quietly patch it and move on in a way you'd never let a real software bug go unlogged?

    An audit trail only helps after something's already gone slightly sideways. The far cheaper fix is deciding, in writing, what an agent is never allowed to do in the first place.

    Related: stackfinder.com/answers/ai-agent-governance

    What Should an Agent Never Be Allowed to Do, and Why Does That Boundary Have to Exist Before the Agent Does?

    Because the moment you're asking that question after an agent has already acted outside its scope, on credentials nobody assigned, making a change nobody can explain, the boundary has already failed.

    I said something half-joking to a peer recently: over a long weekend, I'd be tempted to just let my agents run freely and see what happens. The joke lands because everyone in that conversation immediately pictured what could go wrong. The first question most teams ask me now is which model to pick. My real follow-up is almost always the same: what should this agent never be allowed to do?

    I've spent the last few weeks writing about governance in different forms, and the edge cases keep proving the same point every time: an agent acting outside its intended scope, running on credentials nobody actually assigned to it, or making a change nobody can explain or reverse. If your team is anywhere close to that stage, that question needs an answer in writing before the agent gets more access, not after.

    Here's where I actually see the market moving right now, and it's a genuinely encouraging shift. Governance is moving into the build itself instead of getting bolted on after launch: permissions, identity, and logging get defined before an agent ever touches a real record. Buyers are asking for human understanding, not just raw capability. If nobody on the team can explain what the agent actually did, the feature doesn't ship, full stop, regardless of how impressive the demo looked. And education is quietly becoming the real adoption layer. Teams don't want a vendor that just hands over access. They want a partner who shows them what's actually possible and, just as importantly, where they need to push back.

    The teams scaling agents successfully right now are, without exception, the ones who figured out where to push back first. That's worth keeping close by for the next time someone in a room says, half-joking or not, “let's just press a button and let the agent run.” Because the honest answer to that joke is the same one running underneath everything else in this issue: nothing about AI ROI, shared definitions, or audit trails matters if the boundary around the agent was never drawn before it started moving.

    There's a number worth putting on that boundary too, one most finance teams haven't started tracking yet, and it belongs right next to CAC on the same spreadsheet.

    What Is Cost Per Decision, and Why Might It Matter More Than CAC Alone Going Forward?

    Cost Per Decision is the fully-loaded cost of a single determination, human or agent, divided by the number of those decisions made in a period, and it's quietly becoming a hidden ingredient inside your real CAC.

    The formula is straightforward: Cost Per Decision = Total cost of producing a category of decision (compute, escalation, human review, tooling) ÷ Number of decisions made in that period. A “decision” here can be as small as an agent scoring a lead, routing a ticket, or approving a discount inside a defined threshold.

    Here's why this belongs next to CAC instead of off in some separate operations report. As agents take over pieces of the buyer journey, some of what used to be a rep's hour, part of your fully-loaded acquisition cost, is now an agent's decision, and if nobody's measuring what that decision actually costs (including the cost of the escalations it kicks up to a human when it isn't confident), your CAC math is quietly incomplete. A cheap-looking agent that escalates constantly isn't actually cheap. The true cost just moved to a line item nobody's watching.

    This ties directly back to everything else in this issue. Cost per decision only means anything against a baseline, which is exactly what most teams skip in the first section. It only means anything if “decision” is defined the same way across teams, which is the taxonomy problem from the second section. It gets missed entirely when an edge case isn't logged, since an unlogged bad decision doesn't get counted anywhere, which is the audit trail problem from the third section. And it should be capped by a boundary that was set before the agent started deciding things at all, which is the fourth section. One metric, sitting at the intersection of everything else this week actually asked you to fix.

    Full breakdown: stackfinder.com/answers/cost-per-decision

    What Does Skipping the Definition Actually Cost You in CAC?

    Every pattern this week traces back to the same root cause: something that was never defined, agreed on, logged, or bounded before the agent started acting on it.

    • No baseline or hypothesis: Without a documented baseline and hypothesis, there's no way to prove or disprove ROI, so the AI initiative becomes unfalsifiable, and expensive, faith. The result: two quarters later, a CFO asks what it returned, and the initiative gets cut or doubted regardless of whether it actually worked.
    • No shared taxonomy: When teams define the same term differently, an agent won't pause to ask which one is right, it will just pick one and run at scale. The result: the same account looks like a win in one dashboard and a loss in another, and every downstream number, CAC included, inherits that disagreement.
    • No audit trail on edge cases: An edge case treated as a one-off surprise instead of a logged lesson repeats itself with a different agent a few weeks later. The result: the same mistake gets paid for twice, once when it happens and again when it recurs with no memory of the fix.
    • No predefined boundary: “What should an agent never do” has to be answered before the agent ships, not after it does something nobody can explain. The result: an unexplainable action becomes a rollback scramble, a lost afternoon, or a customer-facing mistake that could have been scoped out from the start.

    This Week’s Move, starting Monday, 9/14/2026

    Before your next AI initiative ships, write your actual hypothesis sentence: “If we convert X% today, then we expect this to move to Y% within 90 days.” If you can't fill in both numbers, that's the work to do this week, not the agent to deploy.

    Have a workflow, metric, or mistake you want broken down in a future issue? Reply to this email or send me a DM on LinkedIn.