Several frontier models now post similar benchmark scores at sharply different cost-per-task. That is competition working. It is not a strategy. Measure economic value per dollar in a named workload — not which logo is smartest or cheapest this week.
- Tied intelligence scores with a 40%+ gap in cost-per-task are a market fact, not a vendor win.
- Cost per task is an input. Value per task — ROI per inference dollar — is the decision metric.
- Fully loaded cost includes tokens, tools, retries, review, and error cleanup. The API bill is not the whole cost.
- As model capabilities commoditize, differentiate in workflows, data, and governance — not monthly model contests.
- If you cannot name the workload, you cannot name the return. No named job, no ROI story.
On 2 September 2026 an independent intelligence index put four frontier systems at the same score of 61. One of them, Meta’s Muse Spark 1.3, did that work at about $0.55 per index task. A peer sat near $0.95 for the same headline number — roughly 42% more. The reporting is straightforward; see Implicator’s write-up. The implication for operators is not “pick the cheap one.” It is that the frontier is crowding, and the remaining public spread is mostly price.
That crowding is the point. Model competition clusters capabilities, then vendors fight on cost, latency, and packaging. Boards hear cheaper inference and assume the P&L will follow. Cheaper tokens are real, and they will keep getting cheaper. Treating the cheapest tied score as a strategy is not. A lab that matches peers at a lower cost-per-task has won a procurement talking point. It has not yet created a dollar of margin in your company.
Scores are not outcomes
Intelligence indexes are laboratory instruments. They compress agentic work, coding, and knowledge into one number so the industry can argue in public. They do not see your cycle time, your error cost, or whether a customer stayed. Even inside one index, the story is messier than the headline: cost-per-task can rise when a “better” model uses more input tokens, and one eval can improve while another slips. A committee that only tracks the composite score will miss that trade-off. Same lesson we drew on evals in diligence and the gap between leaderboards and live work.
Cost-per-task is still a poor proxy for economic value. A cheaper call that drafts a worse contract, slows a ticket, or puts a wrong number in a customer email is not a saving. A more expensive call that cuts a week off a diligence memo, lifts win rate, or reduces rework can pay for itself in a quarter. Technical efficiency — fewer tokens, fewer tool calls, a lower sticker price — is not revenue, not bankable productivity, not margin, and not customer value. Those live in the workflow. ROI of AI adoption is still a business measurement problem, not a model-card problem.
Value per inference dollar
As model capabilities commoditize, differentiation moves up the stack. The scarce asset is not “we picked the smartest model this month.” It is the playbook: which jobs the model is allowed to do, what data it sees, which human owns the output, and how you measure the dollar effect. Firms that swap APIs every quarter without a value ledger will keep paying for demos. Buyers already have more routing options than a year ago; another paid frontier API is one more seat at a table we covered when Meta opened to enterprise procurement. Routing is cheap. Judgment is not.
A more useful unit is value per task, or ROI per inference dollar. Name the job: a support ticket, an IC memo, a forecast refresh, an outreach sequence. Measure the baseline cost of doing it without the model. Measure the fully loaded AI cost — tokens, tools, retries, review time, and the cost of errors. Then ask what changed: hours, quality, conversion, margin, risk. That is economic impact per AI workload. If two models tie on a public index and one is 42% cheaper there, you still have to prove the cheaper one does your job without leaking quality. Sometimes it will. Sometimes the “expensive” model is the only one that survives a real review gate. You will not know from a leaderboard.
Executives should stop asking only which model is smartest or cheapest. Ask what economic value this model creates per dollar in this workflow, this quarter, under this control environment. The labs will keep tying. Your P&L will not care which logo sat on the API key unless the work got better, faster, or safer in a way you can defend to a board.
Sources
- Implicator, Meta’s Muse Spark 1.3 Matches GPT-5.6 Sol at 42% Lower Cost per Task
- Agentivo Capital Labs, The ROI of AI Adoption: How to Measure It
- Agentivo Capital Labs, Benchmarks and Evals Are Now Core M&A Tech Diligence
- Agentivo Capital Labs, The Frontier Eval Gap
- Agentivo Capital Labs, Meta’s Model API Changes Enterprise Procurement