Who cares about your model if you don't have the right harness?
The model is only one part of the system. The harness determines what the model can see, what it can do, how it is governed, and what it costs to produce real work.
It seems like there's news about new models, capabilities, and costs every few days ever since 2023. Model efficiencies and cost per token have started to really hit the main stage as companies invest more in frontier labs' capabilities, while others try to operationalize workflows at lower costs using open-weight models.
That's useful, but not the whole story.
📌 THE POINT IS: Companies are spending too much time asking which AI model is best and not enough time asking which harness should govern the work. The same model can be cheap or expensive, useful or frustrating, depending on the environment around it. The harness determines what the model can see, which tools it can use, how it is reviewed, and whether the company is measuring token price or the cost of a completed task.
Harnesses are starting to get attention, but they still aren't getting enough of it. Recent experiments have shown that harnesses not only govern what models can actually do, but can significantly impact total costs.
Simply put: the model's harness is the system that gives the model permissions, capabilities and can turn a chat assistant into an autonomous worker.
A model inside a workbench that can inspect files, use tools, make changes, and verify results is another thing entirely.
Databricks Just Put Evidence Behind It
Ali Ghodsi, CEO of Databricks, recently described an internal evaluation across Databricks’ own tasks, codebase, and infrastructure, produced by more than 3,000 software engineers and spanning three hyperscale clouds. His key finding: for the same model,
“the choice of harness can significantly save costs (~2x).”
That is a significant, executive-level point. Token price is not the same as task cost. A cheaper model can consume more context, require more retries, or fail more often. A more expensive model can sometimes finish faster. And the harness can change the cost curve without changing the model at all.
The Difference Is Not Subtle
I experienced this moving from ChatGPT to Codex for my personal coding projects. ChatGPT is excellent for many knowledge questions. It is fast, conversational, and economically sensible when the task is simply to ask, think, draft, or understand.
But coding in a chat harness often leaves the human as the integration layer. The model writes code, but the user copies it into files, runs it, finds errors, returns to chat, and repeats the loop manually. This is very painful!
Codex changes the pattern. It can inspect the project, edit actual files, run checks, reason across steps, and iterate. The model is not just talking about the task. It is operating inside the full flow and acting more like a digital partner on your goals.
The Same Pattern Shows Up Outside Coding
As Treasurer for a nonprofit, I produce monthly and quarterly financial analyses and reports. I had tried to get ChatGPT CustomGPTs and Gemini Gems to learn my style, follow my template, and automate the reporting process from raw data. They helped, but the process remained manual and fragile.
With ChatGPT Work, I provided prior examples, the template, and the raw data and it produced a strong report in one pass. That was not just a model difference. It was a harness difference. Chat is for knowledge finding and research whereas Codex moves through technical workflows using your computer or the cloud. Work helps produce real business deliverables (also using your local computer and files).
The right question is no longer only, “Which model should we buy?” It is, “Which harness should govern this work?”
Model Capability Still Matters
The lesson is not that model choice no longer matters. It is that model choice increasingly becomes something the harness should manage. Anthropic described Claude Fable 5 as its most capable generally available model, especially strong on longer and more complex tasks. Anthropic also distinguished Fable 5 from Mythos 5, a limited-release variant, and noted that Fable includes safeguards that may route some requests away from the top model in fewer than 5% of sessions on average.
That's the power of Fable 5's harness: it acts as an enabler and a governance safeguard.
At the same time, Moonshot’s Kimi K3 shows how fast cost-performance is moving. Moonshot describes Kimi K3 as a 2.8T-parameter open model with native vision and a 1-million-token context window, built for long-horizon coding, knowledge work, and reasoning. Recently it was evaluated next to Fable 5 as a potential open-weight challenger to Anthropic’s most powerful, generally available model. Moonshot says it still trails the strongest proprietary models overall, but it's showing frontier-level performance across its evaluation suite.
What Executives Should Consider
The winning AI stack will likely be multi-model and multi-harness. In many companies, a proprietary or deeply customized harness may be the way to personalize AI around actual work. A custom harness could more accurately leverage company context, choose when to use a conversational AI, determine the tools and right model for different coding tasks, and quickly apply company assets to a document workflow. Harnesses could determine which governed enterprise agent may be required for a challenging or specific task to ensure consistency and business rule adherence. The harness can also determine when to route to frontier models versus cheaper or open-weight models.
The harness is what makes those decisions operational at scale.
It governs context, permissions, memory, tools, retries, verification, routing, and cost. It can turn model capability into leverage while removing the human guesswork and error, especially when employees are left to guess which tool, model, or workflow fits each task.
Ask yourself if you have these things figured out today:
Have I started looking at cost per completed, correct, reviewed task?
Am I still using weaker metrics like price per token, without task-level quality and completion data?
Do I benchmark model-harness combinations against real company work?
Am I making a common mistake such as buying a model and treating the harness as a fixed wrapper?
Moving Beyond Model Selection
The model matters. But the harness determines whether AI is an answer machine, an assistant, or an operating capability.
Executives who learn this distinction early will make better AI decisions. They will buy more intelligently, govern more effectively, and measure the economics of AI by the work it completes, not the tokens it consumes.
Like and Share this newsletter to spread the word about practical data and AI strategies for executive leaders!




