GLM 5.2 Clearly Explained (and how to set it up)

23 Jun 2026 · 23 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

GLM 5.2 for local AI—why it’s outperforming on benchmarks, how to set it up in Cursor/Codex via OpenRouter, and how to chain models to offset GLM 5.2’s lack of vision.

Guests

Amir (podcast host’s friend). Background: practical local-AI implementer; focuses on integrating models into daily workflows using Cursor, Codex, and OpenRouter; discusses cost/token economics and model chaining.

Key claims

GLM 5.2 has a 1M context window and strong long-horizon performance; it’s open source and can run locally or via cloud providers. OpenRouter can reduce token costs vs closed models. Model chaining (fusion/sequencing) improves results.

Notable examples

Using Opus 4.8 to describe screenshots (vision workaround), then GLM 5.2 to plan and implement UI changes (hero section redesign, carousel, bento grid). Token-cost example: ~44 cents vs ~$2.38 for similar output.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Learning Objectives with Amir

0:56 to 2:12

Discussion on the capabilities and setup of local models including GLM 5.2.

“Hit a like, comment, and subscribe for more of this sort of stuff in your feed.”

The Impact of GLM 5.2 Release

2:12 to 4:03

Exploration of how GLM 5.2 compares with previous models and its performance.

“Now, what I want people to take away from this session is one, how to actually get set up with it.”

Understanding GLM 5.2 Benchmarks

4:03 to 6:32

In-depth discussion on the benchmarks of GLM 5.2 and user experiences.

“So in this case, it's 62.1 % compared to OpenSys 69.2.”

Setting Up GLM 5.2

6:32 to 8:49

Instructions on how to set up GLM 5.2 using various platforms like Cursor and Codex.

“Codex does actually support open source models.”

Cost Analysis of Local vs Cloud Models

8:49 to 11:20

Discussion on the cost implications of using local models compared to cloud models.

“And I find that GLM 5.2 is a lot more refined and it's able to follow the instructions on what you want it to do.”

Future of Local Models and Investment

11:20 to 13:14

Exploration of future investments in local models and their long-term benefits.

The Economic Implications of AI Models

14:00 to 14:58

Learn about the economic factors influencing AI model usage and investment strategies.

“they're going to get you hooked into the workflows.”

Navigating Model Limitations in AI

14:58 to 18:20

Discover strategies to work around the limitations of AI models like GLM 5.2.

“and lets you kind of tap into that cost saving if you couple it with Composer 2.5.”

Effective Use of AI Tokens

18:20 to 20:40

Explore how to manage AI token usage effectively while maximizing output quality.

“In a way, you can have some sort of direct ROI between the tokens you're spending within the engineering team because you're like, okay, cool.”

The Shift in Token Management Philosophy

20:40 to 21:48

Understand the shift from token maxing to minimizing tokens while maximizing outputs.

“I can see that my usage limits are being hit faster and my cost is going up.”
Show all 11 chapters

Final Thoughts and Audience Engagement

21:48 to 22:42

Recap of the episode and encouragement for audience interaction and exploration.

“So my answer to that is if it works for you and you can directly have an ROI that you can show that, hey, I spent$200 and got a thousand out.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, you've probably heard of GLM 5.2 that's going viral everywhere on Twitter. Yes, it's this new open source local AI model that people are saying is the chat GPT moment for local AI. But no one's actually gone and shown you how to use it and how do you actually set it up. So I figured I'd bring on my friend Amir. He tells us exactly how you should think about running GLM 5.2, how you should think about running local models, how that integrates to something called Open Router, how you can use it with your codex or cursor or cloud code. And this episode in 20 minutes or less, you're going to get everything you need to know about local AI models, why GLM 5.2 is crushing benchmarks, and how you can set it up today so you can go and build your startup, build your business, and be more productive and be more efficient.

0:54Enjoy the episode and I'll see you at the end of it. Hit a like, comment, and subscribe for more of this sort of stuff in your feed. Enjoy.

1:09Welcome to the show, Amir. By the end of this episode, what are we going to learn? We're going to talk about, essentially learn about how local models are kind of keeping up now with the pace of these closed models as well. and how you can kind of use compounding models or fusion models, as OpenRouter calls it, to be able to do sequencing between a more extensive thinking model and a more execution-based model. We'll show you how GLM 5.2 actually compares and stacks up against other models and how you can effectively use it and get set up with it as well. Cool. I did an episode on local models.

1:43It was a hit. People wanted more tactical. How do I actually, you know, how should I think about local models? How do I implement local models? how I can actually make this a part of my daily workflow. So I brought on Amir. Welcome to the show. And let's, I give you two welcomes, by the way. That's how excited I am to have you share with everyone everything. So let's get into it. I'm super excited to be here. Let's jump right into it. So let's talk about what happened this week. ZAI came out with Jelon 5.2. And I think this was a big inflection point because we've typically seen with local models, it's either storage extensive that you can't essentially run it or install it on your computer, or you need a better GPU RAM performance to be able to actually run it locally as well.

2:30Now, JLM 5.2 is also resource extensive, but what we're seeing is with open source models, open source providers like Open Router or Llama, being able to help you run these models in the cloud and effectively being able to essentially pay slightly less for input and output tokens compared to the more closed models. Now, what I want people to take away from this session is one, how to actually get set up with it. We're not going to go through the detailed setup process, but I'm going to just cover how you can do it in cursor using open router or in codex, and then effectively talk about how GLM 5.2 stacks up against the other models, and then how you can effectively use it as well to do model chaining.

3:10And then we'll do model chaining. And then we'll do just a kind of a, maybe a quick walkthrough of how I'm currently using it. You know, I want to be very honest, like these local models still have a lot of work to do in terms of having tool capabilities to be able to, you know, have the modalities to be able to see images and conceptualize on what they're looking at. And I'm going to tell you how you can effectively circumvent some of that, where you can use other models to explain what the image is back to GLM and then have GLM work on it. and also just have a very live test on how this stacks up against other models you know benchmarks are great personally i'm not an expert in it i don't know what any of these benchmarks actually mean the way i do is off apps like let's build it out and see how this actually looks and how it stacks up against other models sound good yes sir let's do it okay so glon 5.2 came out and essentially it has a 1 million context window and it scores 81 points on the terminal bench 2.1 it's just about four points behind opus 4.8 and it does quite well on the long horizon task evaluation so this is essentially projects you have that you want to run long sequence tasks on and you know i think it's taking account like the the thinking parameters and how it can think through and plan through some of the the task it has at hand so you can see that across all these different kind of benchmark reporting or reports.

4:40GLM 5.2 actually does quite well. So in this case, it's 62.1 % compared to OpenSys 69.2. What's special about GLM is it's open source. You can run it locally on the machine if your device can support it, or you can run it in the cloud through the open model providers. It's a big leap from 5.1. I personally didn't test 5.1. I got straight into 5.2, but from what I'm seeing based off of like Twitter and conversations with people, it's performing quite well, especially on the front-end side of just execution-based tasks. I haven't really tested more on the back-end resource-extensive tasks, but just based on perception and what we're seeing in the reporting, it's stacking quite well.

5:22So when you say stacking well, are we talking like, because I look at benchmarks and honestly it goes through one ear, one eye and it goes out the other. I'm like, I glaze over it because I'm like, what does this really mean? like are we talking is it like 4.8 is it like 5.5 yeah what do we how should i think about this and yeah give it give it to me straight honestly man i'm the same i don't get it you know what i mean like i'm i'm not smart enough to understand how these benchmarks actually stack so i'll be honest for me it's like let's just let's just build it use it and see how we feel about you know it how it performs compared to the other models and for me i want to get the best out of it so if i feel like jln 5.2 is strong in one part but weak on the other then i think about how do i actually use other tools or other models to essentially now i think like almost like a fusion approach now i think open router uh which is one of the like the model providers uh coined this where it's like you're able to do like sequencing between two different models to get the best output so i'm totally game on if i can run a local model on my machine to do certain tasks but then call you know opus or codex do something else and have them work together by all means i want to be the most token and cost efficient and performance as well yeah so on the setup side i personally started using this through cursor using open routers api so how it works essentially is you got to go to zai which is the glm provider they created the glm 5.2 model you get an api key from them and then you take that key and you go into your cursor settings paste it into the open ai field and then from there you override the open ai endpoint with this api endpoint right here and essentially from there you go back to models add a custom model glm 5.2 and you're able to then actually call glm 5.2 directly so in essence instead of open ai key you put your glm api key from z.ai then you override the api endpoint for when you call openai chat completion with this one right here and then you go back into custom models and add the custom model protocol you can also alternatively do this using open router so if you want if you're using codex you can go to open router get your open router key and then go into the provider get the endpoint and then go into Codex, create a profile and say, hey, I want you to install this model, open source model.

7:57Codex does actually support open source models. So you're able to provide the details of what the model is, the context window. And then essentially when you're running Codex to the CLI, you can switch to GLM 5.2. Easy enough.

8:12Amir:Yeah, easy enough. And then maybe we can have a page or something to show later on where they can kind of follow these instructions. There's a lot of guides on Twitter and online you can follow those but essentially in my opinion the best way to get started with this is just go to open router and cursor and get that set up through and through so um let's talk about here the model we've talked about the benchmarks how it performs really if we want to just at least take you know put some weight to the benchmarks i'd say if they're scoring at 62 percent and opus 4.8 is at 69 you know that probably means something you know for for for the normal people we probably won't really know until we actually play around with it but i went in and was just looking at for example this website we have this is a small app and i built this in i think opus 4.8 i was just testing it around and i started refining the design using jill 5.2 so i was like hey redesign the hero section for me or refine it um there's this like section right here with all these images i was like why don't we just do a little like carousel style so it you know It's fascinating in a way because I personally don't think the local models had this kind of capability to be able to get it so refined and accurate previously.

9:30And I tested the models that we had. And I find that GLM 5.2 is a lot more refined and it's able to follow the instructions on what you want it to do. So in a couple of months, I was like, hey, let's do a carousel here. Let's make sure we are able to show the images. And then from there, I want you to build out like a bento grid style of all the features that we have. Now, this is all in one prompt. Obviously, you know, you can see it's a little bit vibe coded here. It has a little side like badge, the labels, you know, and you can tell. But at the end of the day, for a local model, for like, you know, a local model, if you're running this on a computer and not burning any tokens, it's doing quite well.

10:11And I think I can see how it's staffed, like internal reporting. I think so. What is the main benefit of using a local model versus something in the cloud is you don't burn tokens. The way to think about it is you're buying a machine. And correct me if I'm wrong, but you're buying a machine. And we should talk about some of the machines that people could potentially buy. But you're buying a machine. It's a cost. It could be$2 ,000,$5 ,000,$10 ,000. And then you can just run tasks. So you're building a startup. Maybe you want conversion rate optimization. So maybe you say, every day I'm going to feed you customer feedback and every single day I want you to work on the front end.

11:04Amir:And you just do that. So my question to you is, for local models specifically, for people who actually want to be building companies, shouldn't they be basically running it all the time on certain tasks and how should people be thinking about it yeah so i think that's a tough that's a little bit tough to answer and i'll say why because like this model specifically is really resource intensive so a lot of i think um existing consumer computers may not be able to run this from what i've seen i've been i've been running it on the cloud through open router directly um and what's the cost with that yeah so uh with so i actually was trying to map out the token cost of model training so if we had about 50 000 input tokens and 85 000 output tokens uh to get almost close to an opus 4.8 level of output it will cost us 44 cents whereas with opus 4.8 it costs you 2.38 cents so the you know there's a big almost like a big difference on the almost like 5x you know price difference between yeah which doesn't sound like a lot when you're like oh two dollars here 44 cents here but when you're actually using these things and you're running it all the time and you have pretty you know big tasks that you're going after and you don't want to be constrained by token costs so it's a big deal 5x is a big deal yeah yeah and this is based on kind of the averages on like the coding benchmarks and what what cursor charges you to the api pool but i want to i want to i want to be also teacher proofing yourself if we got this far with jill on 5.2 i wonder what next six months looks like right so would it be even worth i think it's also worth thinking about making an upfront investment in your machine right now to be able to potentially download and run these local models so that when we get to GL 5.3 or 5.5, we're essentially made the upfront investment in the compute and the equipment to now save a lot more in the long run with other feature models that are going to come out that are going to be much more extensive.

13:16Because I think this, you know, we're seeing the AI subsidy on tokens, right? Like we're getting a lot more output out of cloud, out of codecs. And I wonder, you know, I've personally seen it, and I'm sure you have too, where it's like, now we're hitting our usage a lot faster than we before. Especially when Fable came out, I remember I ran it and in the first day I hit my limit. Totally. So what you're saying is basically, if you look at the history of VC-backed startups, think about Uber. When Uber first came out, they actually subsidized rides and they got you hooked onto the app. And then over time, they started increasing prices, increasing prices.

13:56What you're saying is in the AI age with a lot of these LLMs, they're going to get you hooked into the workflows. You're going to build on top of it. And over time, those subsidies are going to go away as they go public and things like that. So what you're saying is maybe it's a good idea to actually invest in running this thing locally

14:22because the price of memory isn't getting cheaper. and the price of tokens aren't getting cheaper so building it now securing it while you can might be a good idea yeah and to tie this all together i'd say two things one um harnesses that are agnostic on models so like for example cursor where you're able to run multiple models across the same sequence of tasks are going to actually potentially benefit from this right so So I wouldn't be surprised if Cursor decides to directly support GLM 5.2 as a model provider and lets you kind of tap into that cost saving if you couple it with Composer 2.5.

15:04So this is where it goes into model training, right? Where I was, for example, looking at earlier in this task here, I wanted to refine the hero section. So what I did is I actually used Opus 4.8 to first import screenshots because GLM 5.2 doesn't support vision capabilities. So what I did is I actually used Opus 4.8 to import screenshots and explain back to me what it sees. I was like, tell me what you see specifically on the front end design for the hero section and lay it out. And then I switched to GLM 5.2 to study that layout and then actually act on making those changes. So it's kind of a way to circumvent the fact that you have limitations to what GLM 5.2 can do on image capabilities.

15:54But you're able to kind of now train the expensive model to think through the plan and then get the same level of frontier level quality, but at a much affordable price point. I mean that makes sense so your recommendation is basically it's almost like free trade versus protectionism not to be political this isn't political this is just an economic theory which is when people are trading amongst each other maybe it's in Canada where you are you might want to trade with Florida because we got good oranges here. We can't make maple syrup here, so we'll get your maple syrup, right? We'll get the maple syrup, yeah, exactly.

16:47So make the best use of it, yeah. Make the best use of it. You're saying is using cursor, and basically you don't have to use cursor, right? You can use whatever you'd like. Codex, cloud code, yeah. Use one of those to basically say, okay, for certain tasks I'm going to be using local models, for certain tasks you can be using the best in class cloud models. And then together, ultimately you're getting great results in terms of the output, but you're also not spending through the wazoo. And if you're a token maxi, like you and me are, in the sense of we're always pushing into the limit around anything we're building to get the most out of AI, because we don't want to hire 100 people, 500 people and stuff like that.

17:34It's helpful to do that. And exactly, anecdotally, on two parts, right? One internally within our company, you know, I think Satya at Microsoft, you know, mentioned how like human capital plus token usage is now a big factor into what they're doing, right? A lot of companies are now moving away from, you know, having direct access to the cloud code API to run the tokens because of how expensive it's become, right? So they're canceling subscriptions. So we're seeing this firsthand with a lot of companies now are saying, okay, cool. This first year was great. You know, we had the mandate, you know, AI adoption, token maxing, you know, that's how we're going to measure success and that's how we're going to become AI native.

18:12Now they're like, wait a minute. Okay, cool. We've done this, but we're spending way too much money on tokens. How can we now be more effective, right? And I'm seeing this firsthand too, where it's like, especially now, right? In a way, you can have some sort of direct ROI between the tokens you're spending within the engineering team because you're like, okay, cool. We're saving a lot of time. There's an output. it you know engineers expensive we get that but now you're providing the same level of harnesses and models to the non-engineering teams that you know are one-shotting like hey help me format this email and they're using opus 4.8 high thinking they're like maybe that's probably not the right model and that's a governance issue right that's a big thing and i'm having these conversations with companies right now where they're saying hey can you help us figure out like how to build governance and proper education on how to actually use the right models and this is where i think model changing is a big factor right by the way you know john at marketing maybe you shouldn't use opens 4.8 to run this like to just format this email for you and just helping them understand that and i think you know i wouldn't be surprised if in a year from now companies start you know we've been thinking about as well we're like hey why don't we just get our own machines and start running some local models because it's a lot more effective especially how much money we're spending on tokens what's the like just to play devil's advocate why wouldn't i just use open router and call it a day.

19:30Please do, yeah. Absolutely, I think they should. Because when I'm on X, I see a lot of people being like, buy a Mac Studio or buy these expensive devices for people listening. Do they need to buy a local piece of hardware if the price even goes up 2x or should they just use Open Router and Cursor or whatever harness they want? if I just have this right machine, I'll get to this result. That's not how it works. No, you don't need a Mac mini. You don't need this equipment. You can get started today. And what I love about Open Router and all these other tools is, again, they're so agnostic. They make it easy for you to be able to access this in the cloud.

20:14They run the models locally. And it's credit-based. Load$20, get it going, and easy to set up. I highly, highly recommend if you're starting to dabble with this and token usage is a thing for you, get one of these Asian hardinaces set up now that a lot of them are model agnostic run some tokens in OpenRouter get these OpenModels in there and start vibing I love to experiment to see how far I can take this what if I plan with Opus execute with 5.2 and then review with Composer 2.5 or Codex 5.5 there's a lot of ways and I think we can be really effective and I think that's what the smart people are going to be doing in the near future some people are saying i don't care how much tokens cost because i think there's so much opportunity in building startups and and optimizing and ai arbitrage that i don't even care if it costs me whatever what do you say to those people who are just basically ignoring this whole open source local ai movement i used to be the exact same person you know in our first episode you're like yeah how much does it all cost i'm like i don't know man i'm just vibe spending and And I think that mentality has changed now.

21:24I can see that my usage limits are being hit faster and my cost is going up. And now that our team is expanding internally as well. So I think as a solo person, it's a lot easier to build a case or rebuttal around why you should just token max as much as possible, which itself is kind of ironic. You shouldn't be token maxing. You should be token minimizing as much as possible and output maxing instead. So my answer to that is if it works for you and you can directly have an ROI that you can show that, hey, I spent$200 and got a thousand out. Great. Otherwise, sooner or later, the subsidy is going to run out.

22:00All right. Well, I think that's the episode, you know, unless there's anything else you want to add before we bounce. Yeah. I mean, for people that are trying to dabble with it, play around, have some, you know, see what it can do at least on the front end for you and start working back in tests. And yeah, I hope they learned something from this. I'll include links for where to follow Amir he's always one of my first calls whenever I'm trying out new stuff and so I'm happy that you were able to jump on I appreciate you, we appreciate you give him a follow like and comment this video let us know what you think we'll be in the comment section just out there trying to help and learn and thanks a lot Amir I'll catch you on the next one thanks for having me

From the publisher

In this episode I sit down with Amir to get tactical about running local AI models as part of a daily workflow. We center on GLM 5.2 from ZAI, how it stacks up against frontier models like Opus 4.8, and how a fusion approach lets you sequence a heavy thinking model with a lighter execution model for the best output at the lowest cost. Amir walks through setup in Cursor and Codex via OpenRouter, shares real token-cost math, and demos GLM 5.2 refining a live app. By the end you will know how to start today, where local models shine, and how model chaining keeps spend in check.

Timestamps

00:00 – Intro

02:09 – GLM 5.2 and Z AI

04:01 – Specs: 1M context and Terminal Bench 2.1

05:22 – Making sense of benchmark scores

06:42 – Setup in Cursor or Codex with OpenRouter

10:18 – Local model upside: buy a machine, run tasks

11:42 – Token cost: 44 cents versus $2.38

13:36 – Future-proofing with an upfront hardware bet & The Uber subsidy analogy

16:49 – Model chaining and the vision workaround

19:23 – Token maxing vs routing tasks to the right model

20:54 – Answering the "cost is irrelevant" crowd

21:59 – Closing thoughts

Key Points

GLM 5.2 ships with a 1M-token context window and scores 81 on Terminal Bench 2.1, landing about four points behind Opus 4.8.

A fusion approach (a term OpenRouter coined) sequences models: plan with Opus, execute with GLM 5.2, review with Composer 2.5 or Codex 5.5.

Running GLM 5.2 in the cloud through OpenRouter costs roughly 44 cents for a task that runs about $2.38 on Opus 4.8 — close to a 5X saving.

You can start today with credit-based access: load $20 in OpenRouter and route tasks to the right model.

For images, Amir uses Opus 4.8 to read screenshots and describe them, then hands the layout to GLM 5.2 to act on.

Teams are shifting from token-maxing to output-maxing, making model governance and chaining the smart play

The #1 tool to find startup ideas/trends - https://www.ideabrowser.com

LCA helps Fortune 500s and fast-growing startups build their future - from Warner Music to Fortnite to Dropbox. We turn 'what if' into reality with AI, apps, and next-gen products https://latecheckout.agency/

The Vibe Marketer - Resources for people into vibe marketing/marketing with AI: https://www.thevibemarketer.com/

FIND ME ON SOCIAL

X/Twitter: https://twitter.com/gregisenberg

Instagram: https://instagram.com/gregisenberg/

LinkedIn: https://www.linkedin.com/in/gisenberg/

FIND AMIR ON SOCIAL

Humblytics: https://humblytics.com/?via=community

X/Twitter: https://x.com/amirmxt

Youtube: https://www.youtube.com/@amirmxt

More from The Startup Ideas Podcast

All 140 episodes
GLM 5.2 Clearly Explained (and how to set it up)The Startup Ideas Podcast · 23 min
Listen in VO