Jev Creator: System One models for Prod, not God

23 Sep 2026 · 22 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The “hidden crisis” of automation: today’s AI is brilliant at reasoning and benchmarks but fails at reliable, code-friendly tasks like classifying emails or producing strict outputs. The episode argues for TypeSafe’s System 1 / “large programmable model” Jev, designed for computer-to-computer automation rather than chat.

Guest backgrounds

Diogo Almeida (TypeSafe; founded TypeSafe; previously at OpenAI, helped deploy InstructGPT). The host(s) discuss his claims; no other named guests appear.

Key claims

RLHF makes models “people-pleasers,” causing over-explaining and output variability (mode collapse). Jev replaces text with API primitives: choice (enum-like), new (Bernoulli probability for thresholds), and score (ranking). TypeSafe rejects system prompts as “global variables” and avoids deterministic guarantees for cost/intelligence tradeoffs. It also refuses API-layer safety refusals, treating safety as application-level responsibility.

Notable examples

sorting thousands of support emails into billing/tech support/sales; deleting spam when new > 0.95; “dark data” processing (logs, PDFs); coding agents using Jev as a router to manage KV-cache/memory; computer-use/voice commands executed via primitives.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Paradox of AI Automation

0:45 to 1:30

Exploring the disparity between advanced AI capabilities and simple task execution.

“We've built this incredibly powerful, supercharged engine of reasoning.”

Introduction to JEV and System 1 AI

1:30 to 3:15

Discussing the new model JEV and its implications for the future of software.

“And this entire shift is being spearheaded by a new model called JEV.”

The Role of Diogo Almeida and InstructGPT

3:15 to 5:26

Understanding Diogo Almeida's impact at OpenAI and the limitations of InstructGPT.

“So if you ask an AI a yes or no question, a human rater will likely give five stars to a three-paragraph answer that, you know, outlines the nuances, uses bold text, adds bullet points, and ends with a friendly sign-off.”

The Flaws of Reinforcement Learning

5:26 to 7:40

Examining how reinforcement learning impacts AI's effectiveness in automation.

“It is built from the ground up entirely for computers to talk to computers.”

JEV's Architecture and API Primitives

7:40 to 10:35

Detailing JEV's new architecture built for reliable AI outputs in coding.

“It's like building with perfectly machined Lego blocks instead of trying to mold wet clay.”

Intelligence Per Dollar and the Jevons Paradox

10:35 to 13:00

Discussing TypeSafe's focus on efficiency and the Jevons paradox in AI.

“Are you saying TypeSafe is completely abandoning determinism?”

Safety Alignment vs. Capability Alignment

13:00 to 14:00

Debating the implications of unfiltered API access and safety alignment in AI infrastructure.

“It's when you ask a question and the AI says, I'm sorry, as an AI, I cannot help you with that topic.”

The Responsibility of AI Infrastructure

14:00 to 15:56

Explore the implications and responsibilities of providing unfiltered AI access.

“But wait, I have to push back on this impartially.”

Unlocking Dark Data with AI

15:56 to 17:08

Learn how AI is enabling the analysis of previously neglected unstructured data.

“Assuming this unfiltered, cheap, System 1 intelligence continues to roll out, what is it actually doing in the real world right now?”

Innovations in Coding Agents and AI Memory

17:08 to 19:24

Understand how advancements in AI memory management enhance coding agents.

“I read this line in the transcript and it threw me.”
Show all 12 chapters

The Future of AI in Everyday Software

19:24 to 20:30

Discover the potential of AI to transform everyday software into more intelligent tools.

“Jev is aiming to push AI into the background.”

The Silent Revolution of AI

20:30 to 21:46

Reflect on the implications of AI becoming an invisible part of technology.

“Your existing accounting software, your CRM, your email client, they all just get infused with this programmable intelligence.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I mean, think about it. We have artificial intelligence right now that can pass the bar exam. Right. It can diagnose incredibly rare diseases, solve millennium prize math problems, all of it. But if you take that exact same AI, right, and you ask it to just reliably categorize a thousand basic office emails without breaking a software pipeline, it completely falls apart. Yeah, it really does. And welcome to the hidden crisis of modern automation. For you listening, it's basically the technological equivalent of hiring a Michelin star chef who can flawlessly construct a 10-tier wedding cake blindfolded.

0:34But when you ask them to just chop a single onion, they freeze up, panic, and insist on having a 15-minute philosophical conversation with you about the onion's feelings instead. That is exactly it. It is this massive paradox that the tech world is kind of quietly wrestling with right now. We've built this incredibly powerful, supercharged engine of reasoning. But, you know, we haven't figured out how to build the right plugs to connect it to the actual economically valuable plumbing of our daily work. The day-to-day stuff. Exactly. I mean, we are using it for parlor tricks, right? And brainstorming.

1:08But it really struggles to be just a simple, reliable digital gear. Right. And solving that paradox is exactly what we are getting into today. because the notes we're looking at, they outline a shift that is going to fundamentally change how you see the future of software. Our mission for the Steep Dive is to explore a radically new paradigm. A really big one. Yeah, it's called Machine Native or System 1 AI. And this entire shift is being spearheaded by a new model called JEV. That's J-E-V, built by a company named TypeSafe, which, by the way, was founded by Diogo Almeida. Yeah, a lot of people know him.

1:41Listeners might definitely recognize that name from his time at OpenAI. Okay, let's unpack this. Where does this breakdown in basic automation actually start? Well, to really grasp why a model like Jev is such a massive departure, we have to look back at Diogo Elmita's time at OpenAI. Because he was instrumental there. Like, he fought incredibly hard to help deploy InstructGPT. Right, which was huge. Oh, at the time, InstructGPT was a monumental leap forward for AI usability. But after it was out in the wild, he started looking at the data, you know, and realized this kind of painful truth about how it was actually being utilized.

2:18Which was what? People were not using it to automate deep, valuable, rote background work. They were mostly using it to generate what he bluntly calls slop on web pages. Just endless SEO articles, like marketing copy, generic, polite emails. It just became a text generation machine rather than the background automation engine they originally envisioned. And the transcript we're unpacking lays out exactly why that happened. Right, the mechanics behind it. It comes down to the underlying training mechanism, which is RLHF, or reinforcement learning from human feedback. That's the magic ingredient, you know, that makes AI chatbots feel so remarkably polite and conversational and helpful when you type a prompt into them.

2:57Which is fantastic, honestly, if you want a digital buddy. Sure. But it is the absolute root of the problem for software automation. We really have to look at why that happens. Because RLHF fundamentally relies on human evaluators rating the AI's answers, right? And human beings naturally prefer comprehensive, confident, highly formatted, and polite responses. So if you ask an AI a yes or no question, a human rater will likely give five stars to a three-paragraph answer that, you know, outlines the nuances, uses bold text, adds bullet points, and ends with a friendly sign-off. Wow, yeah. So the model learns to optimize for that exact human preference.

3:37It learns to be a people pleaser. Precisely. And in doing so, it suffers from a phenomenon known as mode dropping or mode collapse. The AI's internal probability space literally warps. Warps how? Like, what does it do? It becomes terrified of giving a stark, unformatted answer or making an obvious mistake. So it over explains. It performs for you. It acts like a chatty co-worker who wants to impress the boss. Which, I mean, makes sense if I am asking for a recipe for banana bread. Exactly. I want the chatty co-worker for that. But it is an absolute nightmare if I'm a computer program looking for a simple, binary, true or false output.

4:17If my apps code is expecting a 1, right, and the AI outputs, I would be delighted to help you with that today. The answer is 1. My application crashes immediately. Oh, completely breaks. It is like we are treating artificial intelligence like a horseless carriage. Back when cars were invented, people just stuck a motor on a buggy and treated it like a horse. Today, we are swapping out a human co-worker for a digital one that talks like a human, instead of realizing its true potential is raw infrastructure. So if current AI is just trying to please us with conversational techs, how do we get an AI that actually just does the work reliably?

4:49What's fascinating here is that standard AI is trying to solve a human-to-computer communication problem. But true automation requires solving a computer-to-computer communication problem. Oh, interesting. Yeah. A Python script does not want a friendly greeting. It wants a reliable, predictable, mathematically sound output. That is the huge conceptual leap TypeSafe is making. They realized you cannot patch a chatbot to make it a good programmer tool. You have to build a totally new architecture. Enter the System 1 model. TypeSafe calls it a large programmable model. Right, as opposed to a large language model.

5:27Exactly. It is built from the ground up entirely for computers to talk to computers. They just stripped away the chatbot persona, the RLHF, the politeness, all of it. And this is where the architecture changes completely. Because instead of generating text paragraphs, JEV outputs three brand new API primitives. Okay, let's break that down. Yeah, primitive in computer science is a basic building block. These aren't standard text strings. They map directly to the foundational structures of programming languages. Let's walk through how this actually works because it's brilliant. Let's say I'm building an app that sorts a massive inbox of like thousands of customer support emails into different departments.

6:07A very common task, yeah. If I use a traditional AI, I have to beg it in plain English to read the email and reply with only the word billing or tech support. And I just cross my fingers hoping it doesn't add punctuation that breaks my code. So how does Jev handle that using these new primitives? Well, with Jev, you don't ask for a text string at all. The first primitive is called choice. This maps cleanly to switch or match statements or what programmers call enums. Okay. You mathematically force the model to pick from a strictly defined array of options. It cannot physically output a friendly greeting.

6:43It evaluates the email and returns a raw programmatic value pointing exclusively to billing, tech support, or sales. So it physically cannot go off script. Nope. It's locked in. And then there is the new primitive, spelled N-U, right? Named after the Bernoulli distribution. Yes. New is incredibly powerful. So new outputs raw probabilities that map perfectly to if statements in code. Right. So not words at all. Right. Instead of the AI outputting text saying, I am highly confident this email is spam, it outputs a mathematical probability, say 0.98. Got it. Then the developer just writes an if statement.

7:21If the new score is above 0.95, automatically delete the email. It turns a vibe into a strict mathematical threshold. That makes so much sense. And the final primitive is score, which maps directly to sorting or thresholding tasks. Like if you need to rank 100 documents by relevance, it just gives you the raw scores to sort them. It's like building with perfectly machined Lego blocks instead of trying to mold wet clay. Exactly. And to make sure developers use these primitives properly, Diogo Almeida strongly argues against a practice that is honestly standard in the AI industry right now, which is system prompts.

7:58He actually calls traditional system prompts disgusting global variables. I know. It's harsh but true. I love that phrasing. So a system prompt, for anyone who doesn't know, is usually that giant block of text developers hide in the background that says, You are a helpful assistant. Please follow all these complex rules. Never read this specific file. always format as JSON. Right. And Diogo argues that doing that is a terrible engineering practice. It's like a massive game of telephone. You are asking a statistical model to keep 10 different abstract rules in its head while evaluating a problem.

8:31Too much room for error. Way too much. Instead of one giant text prompt, Jeff expects you to pass it structured JSON objects natively. You don't ask it to do 10 things at once in a massive paragraph. You pass its structured data and ask multiple tiny isolated questions in parallel. Here's where it gets really interesting. By breaking a big task into a hundred tiny AI decisions using these primitives, the code becomes perfectly verifiable. Yes. Think about the alternative. If a standard AI writes a giant text block and makes a mistake in the middle of it, you just scratch your head. You don't know why it failed.

9:06But if you are using JEV and it fails on a specific no probability or a specific choice, you can look at the exact threshold that failed. You can adjust the math based on real data, write a unit test for it, and fix the bug forever. It brings true, rigorous software engineering to artificial intelligence. And because this model is designed to make millions of these tiny, isolated background decisions every single second across the software ecosystem, speed and cost become the entire ballgame. Which perfectly transitions into TypeSafe's core obsession, and actually where the model gets its name, right?

9:40Intelligence per dollar. Exactly. Jev is named after the Jevons paradox, which is this economic principle from the 1800s. William Stanley Jevons noticed that when the steam engine became more efficient with coal, people didn't use less coal. Because it was suddenly cheaper and more efficient, they found a million new uses for it, and coal consumption skyrocketed. And TypeSafe is banking on that exact paradox for AI. To stay on the absolute cutting edge, what they call the Pareto frontier of intelligence per dollar, they do what Diogo affectionately calls absolutely disgusting things behind the scenes to the model's architecture.

10:15Wait, like what? Well, they are so aggressively committed to keeping the compute costs down and the intelligence high that they don't even guarantee the model is strictly deterministic. Hold on. In traditional software engineering, determinism is a sacred rule. If I put 2 plus 2 into a system, I need 4 to come out every single time. Yes, normally. Are you saying TypeSafe is completely abandoning determinism? How do developers build reliable systems on top of an engine that might give a slightly different answer to the same prompt? I know it sounds chaotic, but it's a calculated tradeoff. TypeSafe realized that forcing strict character-for-character determinism at this massive scale would mean sacrificing either the raw intelligence of the model or running up the GPU power costs to an unsustainable level.

11:02Oh, I see. So they prioritize general robustness over strict determinism. As long as the semantic meaning of the output is consistently accurate, slight mathematical variances in the hidden layers are acceptable. That plays right into their stance on public benchmarks, which is wild to read in the transcript we're unpacking. Everyone in the tech world loves a good benchmark chart. They really do. Every time a new model drops, there's a graph showing it beating the competition by two percentage points on some standardized test. And TypeSafe despises those charts. They believe public benchmarks are profoundly broken and easily gamed.

11:37The industry has reached a point where companies literally build teams to curate training data that closely mimics the benchmark tests just to make their charts go up. Just teaching to the test. Exactly. Diogo argues that true intelligence has a je ne sais quoi. Like, it relies on vibes and trust that you can't capture in a sterile test. It only truly proves itself when a developer integrates it into a complex, messy, real-world workflow and sees if the application holds together. If we connect this to the bigger picture, when an AI is running as a silent background loop for an enterprise, shaving off fractions of a cent per token is what actually allows a company to deploy it at scale.

12:18Right. You aren't just calling the AI once to write a poem. You're calling it 10 ,000 times a minute to manage server state, sort incoming data, and verify micro actions. Precisely. But giving developers access to this incredibly cheap, highly efficient, and completely unfiltered intelligence brings up a massive controversy. And this is a crucial tension we need to explore. Yeah. If you build a model that just executes code flawlessly without any human persona, should that AI police what the developers are building with it? TypeSave has taken a highly controversial ideological stance on this. They absolutely refuse to implement safety alignment or refusals at the API layer.

12:57For anyone who has used a modern chatbot, you know what a refusal is. It's when you ask a question and the AI says, I'm sorry, as an AI, I cannot help you with that topic. And Diogo explains it from a purely architectural standpoint. An AI refusing a prompt is essentially a type error. Imagine a background script asking the AI to parse a file named DNA.py to check for formatting errors. And the AI suddenly hallucinates a safety violation and apologizes that it can't read files about genetics. That would just break everything. That refusal stochastically breaks the background software. You cannot have your core backend dependency suddenly deciding to give you a lecture on ethics.

13:38You know, your entire application pipeline will crash. He makes a very sharp distinction between capability alignment, which means doing exactly what the user asks, and safety alignment. He argues that safety alignment is strictly for end-user products, like a chatbot app a consumer downloads. But an API is just raw infrastructure. It is a database. And it shouldn't be the database's job to police its users. Right. But wait, I have to push back on this impartially. What if bad actors use this API for causing harm? Like, if you are providing the foundational, highly intelligent infrastructure of the future and you remove all the guardrails, you know malicious groups are going to flock to it to build automated systems for phishing, hacking or worse.

14:23Isn't there a massive responsibility there? I mean, this phrase is an important question. And honestly, it is the central debate in AI infrastructure right now, right? The notes detail exactly how Diogo fields it. And we really have to look at both sides. On one hand, the danger of providing unfiltered API access to malicious actors, it's very real. But on the other hand, Type State's technical counterargument is that putting your thumb on the scale at the base technological layer fundamentally fractures the model's overall intelligence. Like every time you force a model to overfit to a safety guardrail, you've degraded its ability to do general logic and reasoning.

15:02It makes the model dumber across the board just to patch a few edge cases. Exactly their point. They view this through the lens of Internet architecture. Think about the TCP protocol. That's the underlying infrastructure that moves data across the Internet. If you send an illegal file, the TCP protocol doesn't inspect the packet and refuse to send it. The internet service provider might block it. The application you were using might ban you. But the base protocol remains neutral. TypeSafe is arguing that JEV is the TCP of intelligence. By remaining neutral, it leaves application-level security, compliance, and safety strictly to the developers building the final product.

15:41So they are choosing to act purely as infrastructure, treating the developers like adults who must secure their own software. It is a bold stance, and obviously one that will continue to be heavily debated as these systems scale. So what does this all mean? Assuming this unfiltered, cheap, System 1 intelligence continues to roll out, what is it actually doing in the real world right now? How are developers using these permittives? Well, one of the most massive areas being unlocked right now is what the industry refers to as dark data. That sounds incredibly ominous. Like a sci-fi thriller about a rogue server farm.

16:22I know, sounds ominous, but it's actually a very mundane, practical enterprise problem. Dark data refers to the massive hordes of unstructured data. We're talking millions of customer service logs, decades of internal emails, messy PDF reports that companies have been sitting on in cloud storage for years. Just gathering digital dust. Yep, they've never analyzed it because paying a human to read it is impossible. And running it through a traditional expensive LLM would cost millions of dollars and take months. But with a model optimized strictly for intelligence per dollar, the cost plummets. They can finally unleash these cheap background primitives to process years of dark data overnight.

17:04That is going to unlock insights companies didn't even know they had. Exactly. And then there are coding agents. I read this line in the transcript and it threw me. KV Cash rules everything around me. Wait, explain the KV cache to me like I'm a developer who hasn't worked with LLMs. Why does that rule everything? Okay, the KV cache, or key value cache, is essentially the short-term memory or the scratch pad of the AI model. When you have a complex coding agent, like an AI trying to build an entire software application for you, it has to remember the context of hundreds of files. Currently, these agents are locked into a single, massive, expensive model.

17:42Because transferring that massive scratch pad of memory from one model to another model takes too much time and compute power, right? It's like trying to carry a giant filing cabinet from one office to another every time you need to ask a question. That's a perfect analogy, exactly. But JEV frees these agents. Because the intelligence is so cheap and fast, developers can use JEV simply as a router. Oh, wow. Yeah, the agent uses JEV to dynamically manage its own state and memory, looking at the filing cabinet and deciding, I only need to read these three lines, and I'll send just those three lines to a bigger model to get an answer.

18:16It prevents the system from getting bogged down in its own memory. That's incredible. They also talk about computer use and voice commands. Imagine whispering a command into your phone, and an AI silently and perfectly navigates a web browser in the background, clicking buttons, filling out forms, transferring data to execute it. And it does all of this without writing paragraphs of text explaining what it's doing. Just raw, programmatic execution using those Noom, Choice, and Score primitives. It really is the realization of background logic. The software operates entirely on mathematical probabilities and switch statements, moving through the internet seamlessly.

18:52So what does this all mean for you, listening right now? The average person going to work, using software, managing digital tasks. Imagine a future where the software you use every day actually updates itself. A world that operates on a do-what-I-mean philosophy rather than do exactly what I type. We're talking about a massive structural transition in computing. Up until now, AI has been a foreground novelty. It has been a neat little chat window sitting on your screen that you have to actively engage with. Diogo compares the current state of AI to the early days of the internet, when logging on was an event.

19:29Oh, the dial-up tones? Yeah. Exactly. Jev is aiming to push AI into the background. It wants to be the TCP or UDP protocols that run beneath it all. You don't see them, you don't interact with them, but they make everything you touch work flawlessly. And the economic implications of moving AI into the background are staggering. Diogo is predicting a massive 3 % bump in TFP, total factor productivity, within the next five years, just from this shift. To put that in perspective, TFP measures how efficiently an entire economy uses its resources. A 3 % jump across the entire global economy is astronomical.

20:07It's massive. That is the kind of leap you only see after something like the Industrial Revolution or the widespread adoption of the Internet. He calls this coming wave an inverse SaaS-pocalypse. Right, because everyone thought AI was going to kill traditional software companies. We thought AI would just write custom apps for everyone on the fly. But if TypeSafe is right, existing software doesn't die off. Instead, it gets suddenly and massively supercharged from the inside out. Your existing accounting software, your CRM, your email client, they all just get infused with this programmable intelligence.

20:42It's a complete reframing of how we deploy automation. Instead of trying to build a digital human to sit at a digital desk, TypeSafe is building a digital gear to run the machine. Which brings us all the way back to that Michelin star chef who couldn't chop an onion. For the last few years, we have been trying to force brilliant, conversational, philosophical AI models to do mundane data sorting. And they failed because they wanted to talk to us. Now we finally have an architecture that is perfectly happy to just be the prep cook. Quietly chopping a million digital onions a second in the background without ever needing to chat with you about it.

21:17The plumbing of our digital world is about to fundamentally change, and it's going to happen beneath the surface. Which leaves you with something fascinating to think about. If TypeSafe succeeds and AI completely disappears into the deep plumbing of our software, acting as perfect, invisible background logic that we never interact with directly, will the ultimate sign of artificial general intelligence be that we stop noticing it entirely? Thank you for joining us on this deep dive. Keep questioning the software you use every day. The silent revolution is already running in the background.

From the publisher

We discuss the Latent Space podcast's interview with Diogo Almeida, CEO of TypeSafe, about the launch of Jev, a specialized class of AI models designed for programmatic integration rather than human conversation. He argues that traditional models are "fractured" by safety alignments and chatbot optimizations, making them unreliable for economically valuable automation. Jev is presented as a machine-native "system one" model that prioritizes intelligence per dollar and reliable, structured outputs for software developers. Almeida explains that by moving away from human-centric "slop" and focusing on calibration and robustness, AI can finally automate basic administrative tasks and drive a technological revolution. Ultimately, he envisions a future where AI acts as a reliable utility deep within software infrastructure rather than just a visible digital coworker.

More from Best AI papers explained

All 475 episodes
Jev Creator: System One models for Prod, not GodBest AI papers explained · 22 min
Listen in VO