by Trent Foley
Your AI bill is the easiest number in the office to find. The value it produced is the hardest.
We’ve had the same conversation across our client work this year. The team can show AI spend to the dollar, but it can’t show what that spend bought. As Jeremy Newhouse, one of our AI leads, put it: CIOs are frustrated there’s no return because nobody defined the outcomes up front.
The pilots were interesting. Nobody measured them.
Why AI ROI breaks the usual math
There’s no baseline. Most teams launched pilots before writing down how long the work took or what it cost. Once AI is in the workflow, the “before” number is gone. When we’re asked to measure ROI across a set of pilots that never defined success, the first job is reconstructing a baseline nobody kept.
Value is diffuse. An analyst saves 40 minutes. Where does it land? Unless saved time converts to throughput, revenue, or avoided hiring, finance can’t see it. One fintech client came to us with a growing queue of AI tool requests and no way to measure the business process impact of any of them.
Cost hides. Tokens sit in a cloud bill, a software seat, a Snowflake invoice, and a developer’s personal card. The engineering time to build and review the work sits somewhere else entirely. Nobody owns the total.
Quality moves. A cheap wrong answer costs more than an expensive right one. Rework rarely gets counted.
What Token Yield means

Token Yield is the business value produced per dollar of fully loaded AI cost, measured at the level of one unit of work. A resolved ticket. A migrated stored procedure. A qualified lead.
Fully loaded means everything it took to produce that unit:
- Tokens and credits
- Platform and seat licenses
- Engineering time to build, integrate, and maintain
- Human review and rework
- Evaluation and monitoring
- Training and change management
So why call it Token Yield? Because tokens are the part of that cost you can steer every day. Licenses are set annually. Headcount is set quarterly. Model choice, prompt design, and caching change with every call.
It’s ROI rebuilt for metered AI. Traditional ROI asks whether a project paid off. Token Yield asks whether each workflow, agent, and model choice pays off every week.
Yield isn’t thrift. The goal of AI governance is spend that is predictable and proportional to value. An expensive model that produces a result worth far more than a cheap model’s is the better buy.
Six techniques we use to measure AI spend
- Inventory before you govern: Teams complain about the cost of AI before they understand what problems the spend is solving. The reflex is to put up a gateway and cap usage. Start instead by mapping what each team uses, for which use cases, and where spend differs between teams. Agentic workflows, code generation, and chat assistants have very different cost profiles.
- Pick the unit of work before you build: Cost per resolved ticket, per processed document, per merged pull request. Agree on it with finance first. Agreed in advance, it’s a measurement. Agreed afterward, it’s a negotiation.
- Price the value: Ask what getting there sooner is worth to them, such as a product in market a month early or a regulatory filing done before the deadline crunch. Then run the math. If a new capability is worth $50,000 a month once live, and AI pulls delivery forward six weeks, that’s $75,000 of value for the fully loaded AI cost to beat. Pick a number, write it down, defend it.
- Instrument the full cost, not just the bill: Tag every model call by app, team, environment, and use case. Then attach the human side: build hours, review hours, rework. Count only the people time the AI path itself consumes. The work people did before AI belongs on the other side of the comparison, as the manual path you’re measuring against. The token data is automatic. The people data takes discipline, and it’s usually the bigger number.
- Hold cost, performance, and quality together: Cheap but wrong fails. Accurate but unaffordable fails too. Build an evaluation harness that scores accuracy, groundedness, latency, and cost per task, and gate releases on it. It’s lighter than it sounds. A golden set of 20 to 50 real questions and a model that grades each run gets you most of the way, and the evolv team can show you how to stand up those graders. Systematic prompt evaluation has cut token use 20 to 40% in our work while holding answer quality.
- Tune the ratio: Output tokens cost five to eight times more than input, so output is the cost center. Cached context reads run at about a 90% discount, and session churn destroys the cache, so we target a 70% hit rate. Route by task: Shea Scott, one of our engineers, had a frontier model plan and delegate execution to smaller models. Results held. Token use dropped hard.
On Snowflake
Snowflake now meters Cortex AI Functions, Agents, Intelligence, and Cortex Code as token-based AI credits, separate from warehouse compute. That makes it one of the cleanest places to measure Token Yield today. Usage-history views give per-user, per-model, per-query cost attribution. Budgets can alert at 50, 75, 90, and 100% of spend.
Model access controls double as financial controls. Switching a user from a mid-tier model to a frontier model can raise their output cost 83% with one command.
What a good yield looks like
A team needed to migrate 30 legacy SQL stored procedures to dbt. Estimate: three weeks of manual work, about $25,000 of engineering time. With governed Snowflake Cortex Code, one engineer finished in four days.
Token Yield: about 3.5 to 1.
One catch. That yield is only real if the 11 days the engineer got back went to something that mattered.
AI making tasks easy can be a problem
The biggest drag on yield is a good model doing work nobody needed.
When a task took a day, people asked whether it was worth a day. When it takes five minutes, nobody asks. So, the low-value work gets done. The status deck nobody reads. The third version of an analysis that already answered the question. The internal tool for a problem two people have. Busyness hits a new high, and the few big rocks that actually move the business sit untouched.
The old prioritization tools still apply. An Eisenhower matrix sorts work by urgent and important, and AI changes neither. What AI changes is effort. In our client workshops we plot AI opportunities on impact versus feasibility. AI drags nearly everything toward easy, so every idea starts to look like a quick win. When effort approaches zero, impact is the only axis left. Rank by it.
The scarce resource moved too. Tokens are cheap. Human attention is not. Every AI output needs someone to review it, integrate it, and act on it. Generation got cheap. Attention didn’t.
A low-value task done cheaply still yields close to nothing. The problem is the numerator, and no amount of caching fixes it. The discipline that matters most hasn’t changed: say no to low-value work, even when AI makes yes easy.
Budget AI like headcount
Here’s the change we’re making internally for next year: No single central AI budget. Each team owns its own and justifies it the way it would justify a hire: what does this spend buy, and how will we know? A center of excellence sets the standards and helps each team measure.
Some organizations go further and capitalize part of the spend, like any other software asset. Under US accounting rules, AI model training, validation, and optimization costs may be capitalized during development when they resemble coding work.
What each token buys
Token Yield makes realized value visible. Your next board meeting won’t ask how many tokens you burned. It will ask what they bought. If you want a clear answer ready, connect with the evolv team. We help organizations build the AI governance and enablement that shows what every token yields.
Building on 20 years of experience as a technology leader and consultant, Trent Foley has delivered innovative and high-impact solutions in the finance, healthcare, and defense industries. He brings a cloud-first approach to designing, building, and deploying enterprise-scale solutions. He is passionate about helping clients conceptualize and deliver high-quality solutions with modern practices and architectures. Trent’s expertise includes leading application development efforts, application modernization, micro-service architecture, API development and integration, and data modeling and analytics. He is on a mission to build the dream team of technology consultants, pursuing creative and innovative techniques to attract, develop, and retain the top tier of like-minded professionals that relentlessly pursue excellence.




