In short
Mid-year 2026 “reality check” on AI’s state and enterprise implications, arguing that capability has advanced, but ROI now depends on execution, data/workflow integration, and adoption.
Guests (co-authors/credits)
Brandon Powell, Matt Page, Omar Shanti (co-authors); Andy Smith (editor). Host: Matt Paige.
Guest backgrounds
The episode is produced by Hatchworks AI (AI consultancy). It references Hatchworks AI’s “forward deployed engineer” model and its status as an official Anthropic partner with certified engineers.
Key claims
- Models improved sharply since January (multi-step, long-horizon performance ~3x better than six months ago).
- ROI variance is mostly not the model; it’s data connection, workflow embedding, and adoption (MIT: 95% of pilots fail to hit P&L).
- “Trust-tiered AI” is arriving: identity verification and access tiers after the 18-day Fable 5 export-control ban.
- Enterprises should use blended portfolios: frontier models for hard work, open models for routine work.
Notable examples
Stripe’s 50M-line code migration with Fable; Coinbase cutting AI spend while token usage grew via open-model defaults and routing; Palantir’s air-gapped sovereign AI using NVIDIA Nemetron open models; Ramp AI Index spend gaps and adoption stats; DeepSeek/Quen open-model cost and download momentum.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOState of AI Overview
0:45 to 2:37
A comprehensive look at the changes in AI since January and predictions for the rest of the year.
“has fundamentally changed since January, and where the second half of the year is headed.”
Key Numbers and Trends
2:37 to 5:26
Discussion of important statistics and trends in AI spending and adoption.
“Fable 5, the most capable model ever released publicly, was suspended under a U.S.”
Models Not Plateauing
5:26 to 6:05
An analysis of the advancements in AI models and their implications for enterprise work.
“Cloud's helping employees move faster, but in many companies, the business itself hasn't changed.”
The Shift in AI Strategy
6:05 to 12:09
Exploration of the shift from token maxing to a focus on ROI and disciplined spending in AI.
“And here is where January's architectures are getting smarter.”
The Lab Landscape and Market Dynamics
12:09 to 14:01
Overview of the current players in the AI market and their competitive positions.
“Anthropic is the enterprise incumbent now.”
Going Public: A New Era for AI Companies
14:01 to 15:13
Learn about the confidential filings of Anthropic and OpenAI and their potential impact on public disclosure and enterprise revenue.
“Both leaders have quietly lined up the option to go public.”
The Distribution and Price Frontier in AI
15:14 to 17:36
Explore how Google, NVIDIA, and emerging players are shaping AI distribution and pricing strategies.
“From 7 % of enterprise share in 2023 to roughly 21 % now, Gemini is doing what Google does best, showing up everywhere its cloud and workspace footprint already is.”
The Rise of Trust-Tiered AI
17:37 to 21:46
Understand the implications of identity verification and trust tiers in AI access and usage.
“for over half of all global open model downloads.”
Implications of Trust in AI Development
21:47 to 24:14
Learn how trust impacts enterprise relationships and product development in the AI landscape.
“Access tiers are becoming the new pricing tiers.”
Sovereign AI and Control in Data Management
24:15 to 26:12
Discover how sovereign AI is reshaping data control and trust for enterprises and government agencies.
“Sovereign AI Gets Real The fable episode was the demand signal.”
Show all 21 chapters
Open Source as an Enterprise Hedge
26:13 to 28:00
Examine the strategic advantages of using open models for enterprises in the evolving AI market.
“The through line from the last chapter is that trust and control are becoming the product.”
The Portfolio Strategy in AI
28:00 to 30:17
Learn about the recommended portfolio strategy for AI models, emphasizing cost savings and efficiency.
“Roadmap risk Your product strategy inherits your provider's priorities, which are not your priorities.”
Coinbase Case Study and Tactics
30:17 to 33:16
Discover how Coinbase halved its AI spending using innovative tactics for model management.
“Coinbase and Five Tactics for Blended Intelligence.”
Building the New Enterprise AI Stack
33:16 to 35:57
Understand the key components of a new architecture for enterprise AI and their implications.
“build the evaluations that prove where the bar is, and cut the token bill without cutting the outcomes.”
Principles of AI-Native Companies
35:57 to 38:25
Explore the principles that define AI-native companies and their organizational structure.
“The intelligence layer started as an architecture decision and is quietly becoming an organizational design one.”
Agent Identity and Security Challenges
38:25 to 42:00
Learn about the security challenges posed by AI agents and strategies for managing them.
“Fourth, bring your own agent the newest and most disruptive of the four.”
Safe Deployment of AI Agents
42:00 to 43:28
Learn best practices for safely deploying AI agents in organizations.
“without inheriting anyone's permissions.”
The Jobs Question: AI and Employment
43:29 to 46:00
Understand the complex relationship between AI investment and job creation.
“The loudest narrative in AI is that it is coming for jobs.”
The Human Bottleneck in AI Adoption
46:01 to 48:23
Explore the challenges of human adoption in AI implementation.
“What it demolishes is the assumption that AI investment and hiring move in opposite directions.”
CEO Insights on AI Production
48:24 to 53:58
Gain insights from the CEO on the state of AI production and its challenges.
“The forward-deployed engineer is not new.”
CEO Insights on AI Production
53:59 to 55:54
Gain insights from the CEO on the state of AI production and its challenges.
“If the ROI question is sitting on your desk, We built our generative suite around exactly this journey.”
Transcript
Automatic transcript. May contain errors.0:00Matt Paige:Hey everyone, Matt Paige here, host of Talking AI, and we're doing something a little different with this special episode. So Haptrix AI just released our State of AI 2026 Mid-Year Reality Check, and it's a comprehensive look at what has fundamentally changed since January and where AI is headed in the second half of the year. And instead of only publishing it as a written report, we're using 11 laps to turn the full report into an audio experience. So the voice you're about to hear is AI-generated, but the research, the analysis, and the point of view came directly from our team at Haptrix AI.
0:26Matt Paige:The report covers everything from the latest jump in model capabilities to the growing pressure to prove ROI, the rise of open models, the new enterprise AI stack, agent security, the impact on jobs, and our predictions for the rest of 2026. You can also grab the complete written report through the link in the show notes. With that, here's the State of AI 2026 Mid-Year Reality Check. From the experts at Hatchworks AI, State of AI 2026, the mid-year reality check, a comprehensive roundup of industry stats, research, and critical insights on where artificial intelligence stands at the half, what has fundamentally changed since January, and where the second half of the year is headed.
1:05Co-authored by Brandon Powell, Matt Page, and Omar Shanti. Edited by Andy Smith. Copyright 2026. Hatchworks AI. All rights reserved. One line from the introduction frames everything that follows. The value is real. The spend is real. And the gap between the companies getting one in exchange for the other and the companies getting neither has never been wider. Chapter one. Mid-year by the numbers. Ten numbers. Ten storylines. One number per storyline. Three times. That is how much better the newest model generation holds long horizon, multi-step work compared with models from six months ago, according to Anthropic.
1:51680 times. That is the gap in monthly AI spend per employee between the top 1 % of firms at$7 ,450 and the median firm at$11. Source, the Ramp AI Index, June 2026. 41%. The share of U.S. businesses with a paid Anthropics subscription, now the most adopted AI provider, with OpenAI Flat. Also from the Ramp AI Index. $60 billion. What SpaceX paid for Cursor's parent company, AnySphere. All stock, four days after its NASDAQ debut. Reported by CNBC. Two. The number of frontier labs that confidentially filed for IPOs in June, Anthropic and OpenAI, both at valuations near$1 trillion. 18 days. How long?
2:46Fable 5, the most capable model ever released publicly, was suspended under a U.S. export control directive. We will come back to that story in Chapter 6. Number 1. DeepSeek's rank on Ramp's June trending vendor list, a live read on the appetite for price competitive open models, roughly 50%. How much Coinbase cut its AI spend while token usage kept growing through open model defaults, routing, and caching. Brian Armstrong, July 2026. 1.3 billion. The number of AI agents predicted to be operating by 2028, according to IDC, plus 10%. Headcount growth at heavy AI adopters in the two years after adoption.
3:35Entry-level roles grow 12%. Source. The Ramp Economics Lab. June, 2026. Chapter 2. The Plateau That Wasn't. In January, we ran a section titled Models Are Plateauing. Architectures Are Getting Smarter. It was the consensus view, and half of it aged very well. The other half got overtaken fast. The models that landed between late last year and the first half of this year represent a genuine step change rather than an increment. Anthropik's Opus series reset expectations for agentic work, and the fable generation went further. Anthropik's own language was uncharacteristically direct, capabilities that exceed those of any model they have ever made generally available, state-of-the-art on nearly every tested benchmark.
4:26OpenAI, for its part, has released its answer, GPT 5.6, signaling the next capability wave, and it is proving to be a truly formidable competitor to Fable. The frontier has found another gear. The receipts came fast. Stripe used Fable to run a 50 million line code migration, the kind of project measured in engineer months, in a single day. On long horizon work, the ability to hold the thread across hundreds of steps without losing the plot, the new generation performs roughly three times better than the models of six months ago. That last number matters more than any benchmark. The thing that held agents back was rarely raw intelligence.
5:08It was follow-through, memory, planning, and consistency across long tasks. A model that is three times better at not losing the plot is the difference between an assistant you supervise and a system you delegate to.
5:24Matt Paige:Quick break in the pod. I keep hearing the same pattern with companies I talk to. Cloud's helping employees move faster, but in many companies, the business itself hasn't changed. The value's still trapped in isolated chats and experiments. And that execution gap is why forward deployed engineers have become one of AI's most talked about deployment models. They embed with your team instead of advising from the outside. It's also why the FDE model is now central to every client engagement we lead at Hatchworks AI. As an official Anthropic partner, we embed Anthropic certified FDEs to identify high value business problems, build and deploy the solution and put governance and security around it, then transfer the capability back to your team.
5:56Matt Paige:If your cloud rollout is still mostly individual usage, check out how Hatchworks AI FDEs work at hatchworks.com slash cloud dash FDE. You can also find it in the show notes. Now back to the show. And here is where January's architectures are getting smarter. Call proved exactly right. Arguably more right than we knew. The capability jump is the model multiplied by everything now wrapped around it. The agent harness, meaning verification steps, sub-agents, permissions, and guardrails. Persistent memory, skills that encode how work should be done, and tool ecosystems that let models act instead of just answer.
6:35Better models made the scaffolding more valuable. Better scaffolding extracts more from every model. The compounding of the two is what moved the ceiling. which leads to the sentence we would put on the wall if we could only keep one. If your organization concluded the technology isn't ready, that conclusion has expired. The practical consequence, there are entire categories of work that simply were not automatable 12 months ago that now are. Multi-day engineering projects, end-to-end research and synthesis, workflows that chain dozens of tools and decisions. If your organization scoped its AI ambitions in 2024 or even in mid-2025 and concluded the technology wasn't ready, that conclusion has expired.
7:25For most enterprise work, the capability question is settled. The open question is who is positioned to turn that capability into a number ACFO will accept. Which brings us to the theme of the year. Chapter 3. From Token Maxing to Show Me the ROI. The first half of 2026 had an unofficial strategy, and it was token maxing. Scale through compute. Spend as strategy. Throw the new capability at everything and sort it out later. You could see it in the numbers. Enterprise AI investment hit$37 billion in 2025, tripling in a single year. And the freshest 2026, Reed shows where the momentum concentrated.
8:11Per Ramp's June AI Index, the top 1 % of firms now spend about$7 ,450 per employee per month on AI, while the median firm spends roughly$11, a 680-fold gap. Boards approved AI budgets the way they once approved cloud budgets, on the theory that underinvesting was the bigger risk and the most aggressive boards approved them at a scale the median company has not begun to imagine. The back half of the year is shaping up differently. The phrase we hear in every boardroom now is some version of show me the ROI. Notice how much that question concedes. Does AI work? Ended as a debate a while ago. Are we getting value?
8:58Almost everyone can point to something. The live question is sharper. Is the value worth the relative spend? Token bills are now a real line item. Co-pilot seats, API invoices, and platform fees compound quietly, and the CFO has noticed. The burden of proof moved from the skeptics to the spenders. Here is the uncomfortable truth underneath that question, and it is the central argument of this report. The ROI variance between companies has very little to do with the AI. Everyone is buying roughly the same models. The same frontier capability is available to your company and your competitor for the same per-token price.
9:43Yet one of you will clear the bar, and one of you will not. The difference lives in three places, none of which come in the box. First, data connection. A model that cannot see your systems produces generic output at premium prices. Companies clearing the ROI bar connect AI to their CRM, their warehouse, their documents, and their communications, operating directly on business data instead of the Internet's average. Second, workflow embedding. Value comes from AI wired into the actual flow of work, the ticket queue, the deal desk, the close process, The deploy pipeline, usage that lives off to the side in a chat window, evaporates.
10:29Usage inside the active workflow compounds over time. Third, and this is the stubborn one, adoption. Access does not equal adoption, and seats do not equal changed behavior. The MIT finding showing 95 % of pilots failing to hit the P &L highlights that AI success has always been a people, process, and integration challenge. The same MIT research contained the tell. Purchase tools and specialized partners succeed about 67 % of the time, while internally built solutions succeed 33 % of the time. The gap comes down to where the failure modes live. in integration and adoption, exactly where outside operators with pattern recognition across dozens of deployments have an advantage and where a solo internal team is learning everything for the first time.
11:27Call this the 2026 posture shift from token maxing to relentless execution. The second half of 2026 belongs to a disciplined approach. Instrument spend, connect data, embed workflows, drive adoption, and measure relentlessly. Companies that do this will find the ROI question easy. The others will find budget season very long. Chapter 4. The lab landscape found its new equilibrium. The who's winning conversation got stale precisely because the answer stabilized. The interesting story now is what the settled board looks like and what each player's position tells you about where this goes. Anthropic, the enterprise incumbent.
12:14Anthropic is the enterprise incumbent now. That sentence would have read as provocation 18 months ago. Today, it is just the data. Menlo Ventures' year-end 2025 study, still the most recent comprehensive read of the market, put Anthropic at 40 % of enterprise large language model spend, up from 24 % a year earlier and 12 % in 2023, plus a commanding 54 % of the AI coding market on the strength of Claude Code. The fresher 2026 signals say the lead is holding. Ramp's June AI Index shows Anthropic as the most adopted paid AI provider among U.S. businesses, present in 41 % of them. The Fable launch extended the capability lead at the exact moment enterprises were consolidating vendors.
13:07The story has shifted from an upset in progress to a leader defending a position, OpenAI. The agentic comeback. OpenAI staged the comeback nobody fully priced in. After watching its enterprise share slide from 50 % in 2023 to 27%, OpenAI turned the consumer flywheel into an enterprise engine. Enterprise now represents more than 40 % of OpenAI's revenue, on track to reach parity with consumer by the end of the year, on a run rate north of$25 billion annually. The pattern to watch. OpenAI is winning the agentic workflow purchase, selling outcomes and platforms to buyers who came for chat GPT and stayed for the stack.
13:55Counting them out was a mistake in 2024, and it is a mistake now. Going public. Confidential filings and market impact. Both leaders have quietly lined up the option to go public. Anthropic filed a confidential S1 on June 1st, on the heels of a May round at a$965 billion valuation, with reports pointing to a listing window as early as October. OpenAI filed confidentially in June as well, near a$1 trillion valuation, and is reportedly weighing whether to wait until 2027 after SpaceX's bumpy debut. A confidential filing only sets up the option to list, with no date committed, so timing stays flexible.
14:40Late this year or next is the likeliest window. The significance for enterprise buyers is bigger than the tickers, because public companies disclose. If these two list, the economics of the frontier labs, their margins, their compute obligations, and their customer concentration become visible quarterly, expect that transparency to reshape pricing conversations. And expect both labs to chase enterprise revenue even harder, because that is the number public investors reward. Chapter 5. The Distribution and Price Frontier. Google and Gemini. Quiet, compounding. Google is compounding quietly. From 7 % of enterprise share in 2023 to roughly 21 % now, Gemini is doing what Google does best, showing up everywhere its cloud and workspace footprint already is.
15:35Google wins deals nobody tweets about. In the enterprise, that is often what durable looks like. NVIDIA and Nematron. The quiet giant. And then there is NVIDIA's quiet giant. Nemetron, NVIDIA's family of open models, is remarkably capable and conspicuously under-marketed. The strategic logic is almost funny. NVIDIA's best customers are the frontier labs, and aggressively pushing models that compete with your customers is bad business. So Nemetron advances steadily, wins serious deployments, and stays out of the headlines. For enterprises, it is one of the best-kept non-secrets in the market. Frontier-adjacent capability, open weights, backed by the one company nobody in AI can afford to alienate.
16:26SpaceX and Cursor, the distribution game. SpaceX bought Cursor, and distribution became the game. On June 16th, four days after its NASDAQ debut, SpaceX, which absorbed XAI in February, exercised its option to acquire AnySphere, the maker of Cursor, for$60 billion in stock. Cursor brings roughly$2.6 billion in business-to-business revenue and a beachhead on millions of developer desktops. Read the move plainly. XAI's models were not winning the enterprise on merit, so SpaceX bought the Surface developers already live in. Expect more of this. When model quality converges, the fight moves to distribution, and the checkbooks in this industry are extraordinary.
17:17The Chinese Openweight Labs, DeepSeek, Quen, and the CostTrust Calculus. The Chinese Openweight Labs stopped being a footnote. DeepSeek previewed V4 in April to serious reviews, and Alibaba's Quen passed 1 billion hugging face downloads by March. faster than any open model family in history and now accounts for over half of all global open model downloads. The capability story is real. Near frontier coding and reasoning at prices that barely register. A coding session that runs about$10 on a frontier U.S. model can cost under 50 cents on DeepSeek. For enterprises, the calculus is uncomfortable.
18:04Open weights mean you can inspect them and run them entirely inside your own infrastructure, where no data ever leaves. But the trust questions are just as real. Provenance, alignment, and increasingly, politics. U.S. companies, including Airbnb and AnySphere, have drawn congressional scrutiny simply for disclosing that they use Chinese open models. Every enterprise will land on this trade-off in the next year, deliberately or by accident. Land on it deliberately. Self-hosting neutralizes most data exposure risk, but the reputational and regulatory exposure is a board-level call, not an engineering one.
18:46The composite picture. Capability is consolidating at the top. Price performance is collapsing underneath from two directions at once. And the competitive frontier has moved from benchmarks to distribution, trust, and price. Keep that trio in mind. The next two chapters are about the trust part. Chapter 6. Fable, the 18-day ban, and the arrival of trust-tiered AI. The defining AI story of the half-unfolded in the four weeks after a product launch. Here is the short version. Because the sequence matters. Early June. Anthropic releases Fable 5, the public mythos class model. It is the same underlying model as its sibling, wrapped in safety classifiers that root the most dangerous categories of request, offensive cyber, biological risk, and model distillation, to a safer fallback model.
19:44More than 95 % of sessions never touch the fallback at all. June 12th. A U.S. export control directive citing national security requires that foreign nationals be cut off from access. Unable to reliably determine which users are foreign nationals, Anthropic suspends both Fable 5 and Mythos 5 for everyone. The world's most powerful public model goes dark. July 1st. The Department of Commerce lifts the directive. Fable comes back globally after 18 days offline. July 8th. Anthropic begins requiring identity verification, a government-issued photo ID and a live selfie processed by a third-party provider for consumer accounts that want full Fable access.
20:35API customers are exempt. July 20th. Pressured by the capability of OpenAI's GPT 5.6 sole, Anthropic makes Fable 5 part of all Max plans, with the first 50 % of usage included at no additional cost. Seven weeks. A genuine turning point in the AI arms race. Perspective and context. Sit with how much is packed into one month. A frontier lab shipped its most capable model with a safety fence it spent a quarter building. A government demonstrated that it can and will switch off a frontier model by decree. And the resolution, the thing that got the model turned back on, was identity. Proving who is on the other end of the prompt.
21:22We think this is the arrival of something enterprises should internalize quickly. Trust tiered AI. The old model was one API, one price, capability gated by your credit card. The emerging model is capability gated by who you are. Anonymous users get one tier, verified individuals another, vetted organizations another still, and sovereign or air-gap deployments sit at the top of the trust ladder. Access tiers are becoming the new pricing tiers. The rest of the frontier is already adjusting to the lesson. OpenAI's approach to GPT 5.6 is a preview-first, staged rollout rather than a big-bang launch, easing its most capable model into the world with graduated access instead of daring regulators to react.
22:12Having watched a competitor's flagship go dark for 18 days, no lab wants to ship first and negotiate second. Measured, tiered, verified rollouts are becoming the frontier default. For enterprise buyers, four implications. One. Access is now a risk surface. If your workflows depend on a frontier model, a policy decision in Washington, or wherever your provider is domiciled, can interrupt your operations. The fable ban was 18 days. Your business continuity planning should now include the line, Our model got turned off. 2. Identity and compliance are becoming product features. Expect verification, usage, attestation, and audit requirements to spread across providers and expect your procurement and legal teams to care about them.
23:06Three, the trust ladder favors the prepared. Organizations that can demonstrate governance, access control, and auditability will get earlier and broader access to frontier capability. That is a new kind of competitive advantage, and it is buildable. 4. Trust runs both directions, and your intellectual property is part of the equation. The deeper your workflows, data, and fine-tuning flow through a model maker, the more of your business it understands. And the labs are demonstrably willing to move up the stack. Anthropics product expansion this year, with Claude Design being the sharpest example, landed squarely on territory held by its own ecosystem partners, Figma most visibly.
23:54So here is the uncomfortable question every company building on a frontier platform should ask. What happens when my model vendor ships my product as a feature? Defensibility now means being deliberate about what you send upstream, keeping your differentiating data and evaluation assets in your own hands, and structuring vendor relationships with the assumption that today's platform partner is tomorrow's competitor. Chapter 7. Sovereign AI Gets Real The fable episode was the demand signal. The supply side answered in the same quarter. The air-gapped stack. In June, Palantir announced an intelligent engine built on NVIDIA Nimitron Open Models to deliver AI for U.S.
24:40government agencies and critical infrastructure operators, deployable in fully air-gapped environments. The detail that matters. Agencies fine-tune the open models on their own data and own the resulting weights outright. No external API. No data leaving the boundary. No dependency on a commercial provider's uptime or policy posture. Palantir's stack enforces authorization and isolation. Nimitron supplies the intelligence. And the agency keeps the keys. Enterprise Adoption This is sovereign AI moving from conference keynote concept to procurement reality. And the appetite extends well beyond government.
25:23The same architecture, open models, private fine-tuning, owned weights, isolated deployment is exactly what a bank, a defense contractor, a healthcare system, or any regulated enterprise increasingly wants for its most sensitive workloads. NVIDIA notes that about two-thirds of companies already use open models in some form, and cost savings are a stated driver, national scale. Nations are running the same play at larger scale, standing up national compute, national models, and national data policies. Because intelligence infrastructure is being treated the way energy and telecommunications were treated in prior eras, too strategic to rent entirely from someone else.
26:12Trust and control. The through line from the last chapter is that trust and control are becoming the product. The Frontier Labs sell capability with guardrails and identity. The Sovereign Stack sells capability with ownership and isolation. Most enterprises will end up buying some of both, which raises the obvious question. How do you decide which workloads go where? That is a portfolio question, and it is the subject of the next chapter. Chapter 8. Open source is the enterprise hedge. Strategically, the most consequential shift of the half is bigger than any single model. A two-track market has matured, and it changes how enterprises should buy.
27:00Track 1 is Frontier Models, Fable, the GPT-5 class, Gemini, at premium prices, for work where the capability ceiling is the point. Track 2 is Open Models, Nemetron, the Llama class, the Quen class, and their descendants, which have crossed good enough for a large share of enterprise workloads, at a fraction of the cost, with weights you can own and deploy anywhere. The strategic argument for taking track too seriously is about dependence rather than ideology. Concentrating your AI stack on a single frontier provider is concentration risk in four flavors. Price risk. Your unit economics are hostage to someone else's pricing decisions.
27:47Policy risk. The fable ban made this concrete. Terms change, models deprecate, access rules shift, and, now demonstrably, governments intervene. Roadmap risk Your product strategy inherits your provider's priorities, which are not your priorities. Data risk The more of your proprietary context flows through someone else's API, the more your moat commutes. The numbers behind the hedge are hard to ignore. Open models now cover roughly 80 % of proprietary use cases at per token pricing 10 to 20 times lower than Frontier rates. And on input, the gap can run 35 times or more. Above tens of millions of tokens per day, self-hosting saves 40 to 60 % versus equivalent API spend.
28:40And a blended pattern, open models for routine work with frontier escalation for hard cases, cuts total costs 70 % to 80 % against running everything on a frontier model, at a quality most users cannot distinguish. So the hedge is a portfolio. Frontier models for the hard 20 % of work where the capability gap genuinely matters. Open models for the routine 80 % where it does. Not. and, critically, your own evaluation harness and data layer so you can measure the gap and reroute workloads as the models change. The companies that build that switching capability stop being price takers. One pattern we increasingly recommend within the portfolio is the advisor model strategy.
29:28Use the frontier model as the planner and the reviewer, the expensive intelligence you consult, and route the high-volume execution to cheaper open models that it supervises. The Frontier model designs the approach, decomposes the work, and grades the output. The open models do the bulk tokens. You get most of the quality at a fraction of the spend, and the ratio improves every time the open models take another step forward, which they reliably do. The operational side of this has gotten easy, too. Routing layers like OpenRouter make it close to trivial to send each request to the right model dynamically, by task, by cost ceiling, or by latency.
30:11So a multi-model portfolio no longer means multi-vendor integration pain. Chapter 9. A Case Study. Coinbase and Five Tactics for Blended Intelligence. The clearest proof point of the half came from Coinbase. In July, CEO Brian Armstrong laid out how the company cut its AI spend nearly in half while token usage kept growing. The playbook reads like the portfolio strategy running in production, and his five tactics are worth studying. Tactic 1. Better defaults, not usage caps. 91 % of Coinbase employees never hit their usage caps. So rather than lowering caps and driving up alerts, the company is moving its defaults to open-weight models, GLM 5.2 and Kimi 2.7, through its own gateway, while engineers stay free to pick the right model for the task.
Read the full transcript
31:06Code reviews deliberately run a diversity of models so they can check each other's work. Tactic 2. Better Routing Prompts are pre-processed and routed to the best model for the job, weighing cash state and price. Frontier models handle planning. Execution rarely needs them. Armstrong's version of the advisor pattern is blunt. Humans should not be choosing models at all, because routing is itself a task AI can automate. Tactic 3. Better Cashing Cash misses are the fastest way to inflate a bill, so every request is cash-aware. One internal implementation took cash hit rates from 5 % to 60%. Tactic 4.
31:53Lean Context. Fresh sessions when switching tasks. Narrowly scoped file context. Unused tools disconnected. In his words, the goal isn't fewer tokens used, it's fewer tokens wasted. Tactic five, visibility over suppression. Engineers can spend what they want on whatever model they want. Usage is simply visible and more spend carries the expectation of more impact. Armstrong's summary is the line worth keeping. The goal isn't to suppress usage. It's to build the infrastructure that makes exponential growth sustainable. That is the posture this entire report has been describing, stated by a public company CEO with the receipts.
32:39And note which models Coinbase now defaults to, GLM and. Kimi, both Chinese open-weight families. A live answer to the trust versus cost calculus from Chapter 5, made by one of the most heavily regulated companies in tech. This is also where we will be transparent about what we do. At Hatchworks AI, an increasing share of our engagements now includes exactly this work. Helping clients right-size their model portfolio, stand up open models where they clear the bar, build the evaluations that prove where the bar is, and cut the token bill without cutting the outcomes. It is some of the highest ROI work in AI right now, for a simple reason.
33:28It attacks the denominator of the ROI equation, spend, while everyone else is squinting at the numerator. The frontier labs will keep the capability crown, but the enterprise center of gravity is shifting toward owned, right-sized, blended intelligence. The winners of the next 18 months will run portfolios instead of monogamies. Chapter 10. The New Enterprise AI Stack Beneath the market noise, a reference architecture for enterprise AI quietly firmed up in the first half of 2026. Four ideas, new enough that most organizations have not operationalized any of them, are doing most of the work. Together, they describe how the companies clearing the ROI bar actually build.
34:14First, the intelligence layer. This is the anchor concept. Individual AI usage produces individual productivity, helpful and invisible on the P and L. The intelligence layer is what turns scattered usage into an operating capability. Connect the priority systems, the CRM, the warehouse, the knowledge base, the communications. Give agents governed access to that context and make the business itself queryable. It is the difference between employees who each have a subscription and an enterprise whose systems can be asked, reasoned over, and acted upon. Every other idea in this chapter assumes it, because agents without context are interns without onboarding.
35:02The most forward-leaning companies push the logic further. If an intelligence layer can root information and coordinate work, what is the org chart for? Block CEO Jack Dorsey published the sharpest version, a thesis he called, From Hierarchy to Intelligence, collapse the management layers that existed to relay information, push decision authority down, and rebuild internal tooling around agents that carry context natively. Y Combinator's playbook for the AI native company lands in the same place, with five principles, AI as operating system rather than tool. Closed loops everywhere. A queryable company.
35:45No human middleware. And scaling through compute rather than headcount. Whether or not an established enterprise adopts that shape wholesale, the direction of pressure is unmistakable. The intelligence layer started as an architecture decision and is quietly becoming an organizational design one. Second, skills. The standard operating procedures that actually run. Every company has standard operating procedures, and almost nobody follows them. They live in a wiki, and the work gets done from memory, differently by every person. Skills flip that. A skill is a company's way of doing one specific task, written down once in plain language, that an AI executes the same way every time.
36:32Encode how we write a proposal, how we review a contract, how we close the books as skills. and the. Best operator's process becomes everyone's process. Onboarding compresses. Institutional knowledge stops walking out the door. It is the first time process documentation has ever self-enforced. Third, loops and the agent harness. The shift from prompting to designing. The most important behavioral shift among advanced practitioners this year is that they stopped prompting and started building loops, systems that prompt the agent for them. You define a goal and the machinery around it. The agent finds the work, does it, checks it, and goes again until the goal is verifiably met.
37:23Boris Cherney, the creator of Claude Code, put it bluntly, I don't prompt Claude anymore. My job is to write loops. He estimates 30 % of his code is now written entirely by loops. Peter Steinberger, creator of OpenClaw, made it a rule. You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents. What makes a loop trustworthy is the harness around the model. Verification steps that check the work against the goal before calling it done. Sub-agents that separate the doer from the checker. Memory that persists across runs And permissions that bound the blast radius The harness is what separates a demo from a system you can leave running overnight Loops are how tokens spend shifts from foreground chat to background work And, not incidentally, how the ROI math starts to look like labor economics instead of software economics Fourth, bring your own agent the newest and most disruptive of the four.
38:33Employees and customers are beginning to arrive with their own agents, wired into their whole stack, and they want to use your systems through them rather than through your carefully designed interface. For software vendors, this inverts a generation of product thinking. Your most important user is increasingly an agent, And agents do not care how beautiful your dashboard is. They need clean APIs, structured context, and machine-legible workflows. The platforms that make themselves easy for a customer's agent to operate will keep the customer. The economics are attractive, too. When users bring their own agent, the tokens burn on their account instead of your margins.
39:20The four compound. The intelligence layer supplies context. Skills supply the standards. Loops supply the autonomy. And bring your own agent readiness supplies. The interface. That stack, more than any individual tool purchase, is what AI Native is coming to mean in practice. Chapter 11. Agent Identity and the Double Agent Problem. Now the part that keeps security leaders up at night and the natural sequel to January's AI is going rogue section. IDC projects 1.3 billion AI agents in operation by 2028. Every one of them is functionally a new worker with credentials. It reads data, calls tools, sends messages, and takes actions at machine speed around the clock.
40:11Which forces a question almost no org chart, identity system, or security model was designed to answer. Who, exactly, is this agent? And what is it allowed to do? The double agent threat. On a recent episode of Talking AI, our podcast, we sat down with Charlie Bell, Microsoft's Executive Vice President of Security, whose warning gives this chapter its name. Beware of double agents. An AI agent that can be manipulated through prompt injection, poisoned context, or a compromised tool stops working for you and starts working for whoever manipulated it while still wearing your credentials. Malice is optional.
40:57Persuadability is enough. Bell's prescription is what he calls agentic zero trust. Treat agents the way mature security organizations learn to treat everything else. Give each agent its own identity, never a borrowed human login. Scope its access to the minimum the task requires. Contain it so a compromise cannot spread. Observe everything it does. And assume breach. The first-class identity model. The industry is converging on the same answer from the product side. The clearest pattern of the half. Agents are getting their own accounts. Rather than acting as the user who summoned them, the emerging identity model gives an agent its own credentials in every system it touches, its own permissions set by an administrator, per compartment memory that does not leak across boundaries, and a full audit trail of every action.
41:54Anthropik's new team-facing, Claude products work exactly this way, and it is why they can be dropped into a shared channel with five humans without inheriting anyone's permissions. The pattern generalizes. The enterprises deploying agents safely are the ones treating agent as a first-class identity type in their access model, provisioned, scoped, monitored, and revocable, the way you would an employee rather than a browser extension. Practical guidance, distilled from Bell's framing and our own deployments. No shared credentials, ever. An agent acting as a person is an audit hole and a containment failure waiting to happen.
42:37Least privilege by default. Start narrow, read the logs, widen deliberately. Separate the doer from the checker. Verification agents with independent context are your first line against both error and manipulation. Instrument everything. If you cannot replay what an agent did and why, you are not operating it. it is operating you. Govern before you scale. The time to design agent identity is at agent number three, long before agent number 300. Treat the double agent problem as the price of admission for the autonomy described in the last chapter, rather than a reason to slow down. The organizations that solve identity and containment early will be the ones comfortable enough to let agents actually run.
43:28Chapter 12. The Jobs Question. Watch the net, not the headlines. The loudest narrative in AI is that it is coming for jobs. The most interesting data of the half says the story is more complicated and more hopeful than the headlines. In June, Ramps Economics Lab, working with the workforce data firm Reveglio Labs, published one of the first studies to pair observed corporate AI spending, actual purchases across 21 ,559 U.S. companies from 2021 through early 2026, with firm-level employment records. No surveys. No exposure estimates. Just what companies bought and who they employed. The headline finding runs straight against the doom narrative.
44:17Companies that invest heavily in AI grow headcount roughly 10 % over the two years following adoption. Entry-level headcount grows even faster, at 12%. The firms automating the most are hiring the most, and they are hiring at the bottom of the ladder, exactly where the displacement fears have been sharpest. Our read on where this nets out. Displacement is real, and so is creation, and the number that matters is the net. Some roles will compress, particularly work that is pure task execution with no judgment layer. At the same time, a suite of jobs is emerging that did not exist two years ago. Forward deployed engineers, agent operations and governance leads, loop and harness designers, AI change managers, evaluation engineers.
45:10Every prior platform shift produced job categories nobody forecast. The web gave us roles that would have sounded like nonsense in 1994, and this one is minting them faster than the last. Two nuances keep the finding honest. First, the gains accrue almost entirely to high-intensity adopters. Companies that dabble see no statistically significant change. Read that carefully, because it is this report's ROI argument wearing a labor market costume. Shallow adoption produces nothing measurable. in the P and L or in the org chart. Depth is what pays. Second, heavy adopters already skewed larger, more engineering heavy, and faster growing before adoption.
45:56So some of this is who they are, not just what they bought. The study leaves causality open. What it demolishes is the assumption that AI investment and hiring move in opposite directions. And there is a second-order effect hiding under the enterprise story. Democratization The same capability that lets a Fortune 500 team do more with less lets two people with a laptop do what used to take a funded team. Build the product Run the back office Generate the marketing Staff the support queue with agents The cost and capital required to start a real business have fallen further in the last 18 months than in the prior decade.
46:39If that holds, the labor story of this era may be less about incumbent headcount and more about a surge in new, small, AI-native businesses. Employment created beside existing org charts rather than inside them. The practical takeaway for leaders is to manage both sides of the ledger deliberately. Redesign roles around judgment and oversight rather than task execution. Invest in re-skilling toward the new categories. And treat the ramp finding as the pattern to emulate, because the companies hiring through AI adoption are the ones adopting deeply enough for it to pay. Chapter 13. The bottleneck is still human and the market's answer is the forward deployed engineer.
47:27Every thread in this report pulls in the same direction. The models cleared the capability bar. The spend is committed. The architecture is knowable. The governance patterns exist. The labor data says deep adopters come out ahead. And still, the MIT number looms over everything. 95 % of enterprise pilots fail to hit the P &L, with purchased and partnered approaches succeeding at twice the rate of go-it-alone internal builds. So the remaining constraint is the last mile. Connecting the systems, embedding the workflows, encoding the skills, standing up the harness, and, hardest of all, getting human beings to change how they work.
48:11January's conclusion holds at mid-year. The bottleneck is human because access does not equal adoption, and adoption is where the ROI lives. In six months, the market found its answer, and it is a role, the forward-deployed engineer. The forward-deployed engineer is not new. Palantir built its business on embedded engineers. but it has gone mainstream. The frontier labs now hire them aggressively because capability doesn't deploy itself. Someone senior has to sit inside the business, find the leverage points, build against the real systems, and stay until the behavior changes. Three verbs describe the job.
48:54Find the highest ROI opportunities in the actual workflows. ship production systems grounded in the client's data, agents, automations, intelligence layers, AI-native products, and multiply by transferring the patterns and training the teams so adoption outlives the engagement. That's how Hatchworks AI operates, and the first half of 2026 validated the model. Our forward-deployed engineers are senior AI strategists and builders embedded with client teams, with one mandate. Production outcomes, not slide decks. The winners. Bought the same models as everyone else, then put builders where the work happens and treated adoption as an engineering discipline.
49:40Two credentials. We are an official Anthropic partner with certified engineers who take Claude from seat licenses to production systems, which is a deliberate choice, given Anthropic's 40 % of enterprise language model spend. And everything this report describes is something we build, from the intelligence layer, to the loop and harness architecture, to agent governance. Chapter 14. The Second Half of 2026. Where This Goes. Nine calls for the second half, made with mid-year confidence and full awareness of how January's plateau call aged. 1. ROI becomes standard. AI line items get the scrutiny cloud got in 2019.
50:27Show me the dashboard replaces show me the demo. 2. The forward deployed engineer goes mainstream. Job postings with forward deployed in the title multiply across the fortune 1000, not just the labs. 3. Agent identity becomes a buying criterion. Requests for proposal start asking how agents are identified, scoped, and audited, and the answers start deciding deals. 4. Open weights rise. Open models move from experiments to owned production workloads, adopting advisor model patterns. 5. Trust tiers spread. Identity verification, usage attestation, and vetted access programs appear across major providers.
51:146. More sovereign deals More private fine-tuning, owned weights, and isolated deployments in public and regulated sectors 7. Compute shifts from chat to loops Token spend moves toward background, goal-directed systems, which become the proxy for AI doing actual work 8. Market consolidation continues Distribution remains the prize Expect more acquisitions of beloved AI surfaces by major players. 9. The Jobs Debate Reframes Displacement vs. Creation Powered by tiny AI-native teams We will grade ourselves in January. Chapter 15. CEO Commentary The View from the Field By Brandon Powell If more than one person is doing the same task the same way, that is usually a great place to deploy AI.
52:13Every January, somebody declares the year of AI production. Here is what I can tell you from the field at the half. Production arrived, and it split our client conversations into two camps. One camp automated their existing processes. They got a marginal lift, and now they are staring at a token bill, wondering why the needle barely moved. The other camp reimagined the work from first principles. They asked what the process should look like if intelligence were nearly free, then rebuilt around the answer. That camp has moved past, asking whether the ROI is there. They are asking how fast they can expand it.
52:55So when boards started asking, show me the ROI this spring, we welcomed it. Discipline is the best thing that could happen to this industry because it forces the honest diagnosis. AI fails in the last mile. Data that was never connected. Workflows that were never redesigned. People who were never brought along. You can put a treadmill in every home and you will not cure heart disease. The treadmill was never the issue. The habit is. Our simplest heuristic still holds. Start there. Ship something into production in weeks. Measure it. Then multiply what works. The companies pulling ahead treat adoption as an engineering discipline and put builders where the work actually happens.
53:42Budget size has surprisingly little to do with it. That is the bet we built this company on, and the first half of 2026 has only deepened my conviction. The technology is ready. The advantage now belongs to the operators. One last thing. If the ROI question is sitting on your desk, We built our generative suite around exactly this journey. Gen ROI to decide what to fund. Gen DD to build it. Right. And Gen EQ to drive adoption. Decide. Build. Scale. Wherever you are on that path, we would be glad to help at Hatchworks.com. Brandon Powell, CEO of Hatchworks AI. About Hatchworks AI. Less AI hype.
54:31More results. Hatchworks AI is a pure-play AI consultancy. AI is not a practice area. It is the entire business. We help enterprises turn AI into ROI through forward-deployed engineers, embedded inside your teams, agentic AI pods for larger initiatives, and end-to-end AI and data transformation. All of it is built on our generative suite. Gen ROI to decide. Gen DD to build. Gen EQ to scale, delivered by more than 100 certified engineers. We are an official Anthropic partner, ranked the number one AI services company by Clutch, and trusted by AT &T, Cox, and Stanley Black and Decker. Ready to turn AI into ROI.
55:22Start with our agent opportunity finder or talk to us about embedding a forward-deployed engineer at Hatchworks.com. This has been State of AI 2026, the mid-year reality check from Hatchworks AI. Full references and source links are available in the published edition of the report. Copyright 2026, Hatchworks AI.
55:45Matt Paige:Thanks for listening to the Talking AI Podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast platform. And don't forget to leave us a review. We love those. For more info on Talking AI, visit TalkingAIPodcast.com. quick break in the pod if you're listening to this podcast chances are you've been thinking about how to actually use ai inside your business and that's exactly why we built the ai opportunity finder it's a free tool that helps you uncover high impact tailored ai use cases based on your business your goals your pain points and your industry no fluff no generic use cases just real ideas that fit your business and the ranked by roi potential it takes about three minutes to run and it's like having your own personal AI strategist for free.
56:30Matt Paige:If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder.
From the publisher
The value is real. The spend is real. And the gap between the companies getting one in exchange for the other and the companies getting neither has never been wider. Six months into 2026, the top one percent of firms spend $7,450 per employee per month on AI while the median firm spends $11 — a 680x gap. The question in every boardroom has sharpened from “does AI work?” to “show me the ROI.”
In this special episode of Talking AI, host Matt Paige hands the mic to an AI. Hatchworks AI just released its State of AI 2026: Mid-Year Reality Check — a comprehensive look at what has fundamentally changed since January and where AI is headed in the second half of the year — and instead of publishing it only as a written report, the team used ElevenLabs to turn the full report into an audio experience. The voice is AI-generated. The research, analysis, and point of view come directly from co-authors Brandon Powell, Matt Paige, and Omar Shanti.
The report covers the step change in model capability that ended the plateau debate, the shift from token maxing to “show me the ROI,” the lab landscape’s new equilibrium, the 18-day Fable 5 ban and the arrival of trust-tiered AI, sovereign AI moving into procurement reality, open models as the enterprise hedge, Coinbase’s five tactics for blended intelligence, the new enterprise AI stack, the double agent problem, the jobs data that runs against the doom narrative, and nine calls for the second half of 2026.
In this episode, you’ll hear about:
- The ten numbers that define AI at mid-year — from a 3x jump in long-horizon capability to a 680x spend gap between the top 1% of firms and the median
- Why January’s “models are plateauing” consensus got overtaken — and why “the technology isn’t ready” has expired
- The three places ROI variance actually lives: data connection, workflow embedding, and adoption
- The lab landscape’s new equilibrium — Anthropic as the enterprise incumbent, OpenAI’s agentic comeback, and two confidential IPO filings near $1 trillion valuations
- SpaceX’s $60 billion all-stock acquisition of Cursor’s parent company, Anysphere, and why distribution is now the game
- The 18-day Fable 5 ban, identity verification, and what trust-tiered AI means for enterprise buyers
- Sovereign AI getting real — Palantir, NVIDIA Nemotron, and owned weights in air-gapped environments
- Open source as the enterprise hedge, and the advisor model pattern for blending frontier and open models
- Coinbase’s five tactics for cutting AI spend roughly in half while token usage kept growing
- The new enterprise AI stack: the intelligence layer, skills, loops and the agent harness, and bring your own agent
- The double agent problem, agentic zero trust, and why agents need first-class identity
- The jobs data — heavy AI adopters growing headcount 10%, entry-level roles 12% — plus the rise of the forward deployed engineer and nine predictions for H2 2026
Key Moments:
- 00:01:30 — Chapter 1: Mid-year by the numbers — ten numbers, ten storylines
- 00:03:35 — Chapter 2: The plateau that wasn’t — the step change in model capability
- 00:06:55 — Chapter 3: From token maxing to “show me the ROI”
- 00:11:10 — Chapter 4: The lab landscape’s new equilibrium — Anthropic, OpenAI, and the IPO filings
- 00:14:20 — Chapter 5: The distribution and price frontier — Google, Nemotron, SpaceX–Cursor, and the Chinese open weight labs
- 00:18:20 — Chapter 6: Fable, the 18-day ban, and the arrival of trust-tiered AI
- 00:23:30 — Chapter 7: Sovereign AI gets real
- 00:26:00 — Chapter 8: Open source is the enterprise hedge
- 00:29:30 — Chapter 9: Case study — Coinbase and five tactics for blended intelligence
- 00:33:05 — Chapter 10: The new enterprise AI stack
- 00:39:00 — Chapter 11: Agent identity and the double agent problem
- 00:42:35 — Chapter 12: The jobs question — watch the net, not the headlines
- 00:46:30 — Chapter 13: The bottleneck is still human — the forward deployed engineer
- 00:49:20 — Chapter 14: Nine calls for the second half of 2026
- 00:51:10 — Chapter 15: CEO commentary — the view from the field with Brandon Powell
Key Links:
Mentioned in this episode:
AI Opportunity Finder
Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/
