325 | Maximize your AI ROI (ROAI 🤔) great output for less money with Isar Meitis

8 Sep 2026 · 22 min · 13 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How to maximize AI ROI by avoiding token/credit limits while keeping output quality—mainly by choosing the right model and “reasoning effort,” reducing token-heavy inputs in ChatGPT/Claude inside Excel, converting repeatable workflows into Excel scripts (near-zero token cost), using auto-upgrading settings, orchestrating multiple models, and leveraging Claude caching (including co-work/code/API).

Guests

No guest is mentioned; Isar Meitis is the host and speaker.

Key claims

Cheap/fast models (e.g., Claude Haiku, ChatGPT Instant) can match higher tiers if prompts are specific; Excel scripts replicate AI steps with zero tokens and ~half-second runtime; ChatGPT can auto-escalate reasoning for complex questions; Claude caching can cut costs (75% price drop mentioned).

Notable examples

ChatGPT in Excel generating multi-tab outputs and graphs; an Excel test where Instant/Medium/High produced identical table results except euro-sign formatting; Claude multi-step workflows assigning Haiku/Sonnet/Opus/Fable per task; using ChatGPT alongside Claude to offload steps while sharing the same folders.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Token Limitations

0:45 to 2:01

Discussion on the limitations of AI token usage and its impact on productivity.

“you can use per week, per month, or a part of a company pool where the company pool may drain if one person is using too many tokens for just one thing that is very heavy and so on.”

Choosing the Right AI Model

2:01 to 2:59

Insights on the importance of selecting the right AI model for efficiency.

“The reality is we overestimate that, or at least I do and most of the people that I know overestimate that.”

Leveraging ChatGPT in Excel

2:59 to 4:13

Exploration of how to effectively use ChatGPT within Excel for complex tasks.

“I tried it on this and tried it on that.”

Reducing Token Usage in Excel

4:13 to 6:06

Strategies for minimizing token consumption when using AI tools in Excel.

“from this past week, so I'm not going to use the file that has the last three years because it also includes the last week, that's fine.”

Creating Scripts for Efficiency

6:06 to 8:10

How to create scripts from AI interactions to enhance productivity without token costs.

“you will not understand and you shouldn't and you shouldn't care either.”

The Importance of Prompt Design

8:10 to 11:01

Understanding how well-designed prompts can lead to better AI outputs with fewer tokens.

“You're thinking, well, this sounds really, really great, but I don't want to use Haiku or I don't want to use Instant or I don't want to use Luna.”

Hidden Features for Token Savings

11:01 to 13:08

Discussion on lesser-known features in AI tools that can save tokens.

“prompt that explains exactly what you want, you will get that outcome even on the cheaper models, which is a huge benefit.”

Future of AI Model Efficiency

13:08 to 14:00

Predictions on how AI models will evolve to improve efficiency and save costs.

“If you've been listening to this podcast for a while, you know my thoughts about this.”

Auto-Writing Capabilities of Chachapit

14:00 to 14:48

Learn about the auto-writing capabilities of Chachapit and its advantages.

“And so Chachapit already does that, and it is very easy to turn on from the settings.”

Building Complex Processes with Claude

14:48 to 16:48

Discover how to build sophisticated processes using Claude and its task management.

“It's a skill that's doing it automatically right now and defines a multiple steps, well-designed process to complete anything that we're working on.”
Show all 13 chapters

Understanding Caching in AI Models

16:48 to 18:47

Gain insights into caching and its cost-saving benefits in AI models.

“The other thing that Claude has that now became significantly more powerful, when I say now is with the announcement of Astra, that's the other thing that they announced, is caching.”

Effective Use of Caching and Smart Routing

18:47 to 20:48

Learn how to effectively utilize caching and smart routing in AI applications.

“You need to flag specific reusable pieces of content or data that it needs to do to be cached.”

Maximizing AI Outputs and Budget

20:48 to 21:54

Explore strategies for maximizing AI outputs while managing costs effectively.

“before the end of the week, every single week, because I do multiple things in parallel in Cloud, even though I'm the$200 plan, I started using it in parallel to ChatGPT.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Isar Meitis:Hello and welcome to the Leveraging AI podcast, a podcast that shares practical, ethical ways to leverage AI to improve efficiency, grow your business and advance your career. This is Isar Meitis, your host, and I have a really important episode for you today. If you are a heavy AI user or just want to be a heavy AI user, one of the things that you're going to run into is running out of tokens. And it doesn't matter which plan you're on. I'm on the$200 plan on Claude as an example, and I ran out of tokens before the week. and I'm sure it's happening to you if you're trying to grow and you're in a$20 plan and you're not really sure how to continue to work without just waiting for your next round and the next cycle to start.

0:41Or you could be in an enterprise environment where you are capped by the number of tokens you can use per week, per month, or a part of a company pool where the company pool may drain if one person is using too many tokens for just one thing that is very heavy and so on. So in today's episode, I'm going to walk you through multiple ways and much deeper knowledge on how this whole concept works. And I'm going to show you lots of tricks on how you can still get the job done effectively while spending significantly less tokens. Now, I'm going to be sharing my screen, but those of you who cannot see my screen, that's perfectly fine.

1:19I'm going to explain everything that we're looking at. It just helps me to walk you through the right process in the right sequence. So the first concept is the easiest and most straightforward one, and yet most people don't apply it on the day-to-day, which is picking the right engine, picking the right model for the thing that you are doing. Now, the problem with that is that it is very tempting to use the better model. Nobody wants to use the cheap, crappy model. Everybody wants to use the better model. So when Fable came out, I dove straight in and started using Fable like crazy until I had to start paying for the real cost in tokens on Fable.

1:56And that was unacceptable because I would use all my tokens for the week in just half a day. And so the trick is to use the right model for the job. The reality is we overestimate that, or at least I do and most of the people that I know overestimate that. So the truth is the cheapest, slowest, crappiest, narrowest models that we have right now are better than most models we had six months ago and better than any model we had just a year ago. So if you are saying, oh, I cannot use Claude Haiku or I cannot use Chachi PT Luna, you don't understand how good these models are compared to anything else that you've seen before.

2:38The same thing with the level of effort that these tools have. So every time you switch the model, you can also pick how hard it is going to think. And the reality is for most tasks that you're doing in the office today, the instant or lowest level of thinking is actually going to be sufficient. Now, I know what a lot of you are thinking. Oh, that's not true. I tried it on this and tried it on that. If you have actually tried and did a perfect real test, then you know. But we're going to get to that on how to do that even better as well. So that's the number one thing. But let's look to a little more details.

3:13As an example, but before we dive into those questions, the next thing is, I know a lot of people started using ChatGPT in Excel. It's absolutely fantastic. It's really, really magical. It works really well. It allows you to do amazing things in Excel, anything from stuff you just don't know how to do, all the way to really complex things, including creating multiple tabs with one prompt that will create any output you can imagine that Excel can do, including graphs and charts and tables and comparisons and VLOOKUPs and statistical models, literally anything you can imagine, including very complex things across multiple tabs.

3:44It is really, really good. The problem with that is that ChatGPT in Excel runs as if, and not as if, it's actually how it's running, against the API, meaning there is no limit per how many tokens it is going to consume. It depends on the amount of data you're uploading in that one Excel and how complex is the thing you ask it to do. Now, there are several ways to reduce that. The first one is to be aware. What I mean by aware, but just by knowing what I'm telling you right now, if you have a choice between I need the data from this past week, so I'm not going to use the file that has the last three years because it also includes the last week, that's fine.

4:19You would know that using the last three years is going to cost you significantly more tokens than using just the last week. And you're going to just get the file from this past week from whatever tool you're exporting it from, or take the large file and just copy the components you need and use that. So that's very easy and straightforward. The other thing is, as of last week, inside of ChatGPT, inside of Excel, you can pick the level of thinking that it is going to do just like you can in the regular ChatGPT. So here is the same thing. You can choose instant, you can choose medium, you can choose high, and you can choose extra high, And that's going to have a very big impact on how many tokens or credits it is going to consume, regardless of the amount of data.

5:02Isar Meitis:So it's still going to grow with more data that you upload. But when it's going to think less, it is going to cost you significantly less tokens. But there's an even better way to save tokens when running ChatGPT or Claude inside of Excel. And that is creating scripts. So if you are now running a process that you're running regularly in Excel, and it's something that you're doing every week or every day or every month or every quarter, depending on whatever it is that you're doing in the file, you can go and open the extension for ChatGPT or Claude or Copilot for that matter, and ask it to do the thing you're trying to do, and it will do it.

5:38And if you build a skill, because you can run skills inside of ChatUpt and inside of Claude, inside of Excel, you can build a skill that will do the entire process and will give you exactly the output you want, exactly the same way every single time. It feels like magic. But once you create the response, once you had the conversation back and forth with whichever AI inside of Excel, and it ran properly, and it gave you exactly the output you want, all you have to do is ask it to create a script for Excel. and it will write a code that if you're like me, you will not understand and you shouldn't and you shouldn't care either.

6:10Because what that script will do, it will follow the same exact process that the AI did in multiple steps and a lot of minutes and a lot of tokens as a script that can run inside Excel itself. So what do you do with that script that it creates? I'm glad you asked. If you open Excel on the top menu, you have all the different options. You have home and insert and draw. And if you go all the way to the right, there's a tab called automate. And if you click on automate, the first button on the left says new script. So if you click on that, it opens a place where you can paste your script. It usually has a few existing rows in it that you can delete and just paste your script and save the script.

6:46And now you have a script that will do exactly what ChatGPT did or Claude did with two really big and important differences. One, it's going to cost you zero tokens to do. Two, it happens in about half a second instead of sometimes multiple minutes if it's a really complex process. So these two things are awesome and amazing. The third thing that is really, really cool is that you can give the script to people who do not have access to the license of whichever to redevelop the AI with. So they don't need ChatUPP, they don't need Copilot, they don't need Claude, they can just run it in a regular Excel just by getting the script and running the script and it will do the exact same output every single time.

7:25So this is a huge difference, both in means of time and in means of the amount of tokens and money it is going to cost you to do really large Excel analysis. Now, the thing you need to remember in all these tools, if you throttle them up, meaning let's say you take my advice and saying, oh, I'm going to run on Instant or I'm going to run on Luna or whatever the case may be, and you find a task that doesn't work well and you shift the gears upwards, it will not shift back on its own. Meaning if you forget it on High or Pro or Max or whatever the setting is for that particular AI, it will stay there until you throttle it back, which means if you forget or you start a new conversation and it wakes up on that as a default, then you have to move it backwards in order to get back to the lower settings.

8:12The other thing that I noticed, at least on ChatGPT, that it doesn't always wake up on the last thing you left the previous conversation, meaning you may leave a conversation on instant, start a new conversation, you'll be on medium or high, and you need to pay attention and lower this down. Now, I know what you're thinking. You're thinking, well, this sounds really, really great, but I don't want to use Haiku or I don't want to use Instant or I don't want to use Luna. It is crap. It cannot do the work. So I did a test. I took an Excel file that's actually pretty big and I'm sorry I'm picking up on Excel and you may say, oh no, Excel is not as complicated as writing for me in my voice, my tone.

8:45Try that as well. But what I've noticed, what I was looking for is something that is a very clear pass or fail. Writing a better essay or a better blog post or a better something is very hard to judge, right whether it's better or not so I was looking for something I can do a very clear pass fail so I uploaded a document again an excel file that has I don't know 20 30 columns and thousands of rows so it's a pretty big excel and I gave it a very detailed prompt and I'm not going to read all of it but it has four different segments and then it asking for three different very specific tables with specific columns in each table and exactly what the data and exactly what the currency is supposed to be and so on.

9:25So it's a pretty detailed prompt. By the way, I did not write the prompt. I asked AI, I told it exactly what I'm looking for. I told it exactly what the data is in the file. I told it exactly what I want to see in the results. And I asked it to write a prompt for me that will deliver consistent results every single time. And so what I did is I then ran it on ChatGPT on the three different tiers. I ran it on instant, I ran it on medium, and I ran it on high. And those of you who are watching the screen can see that the results are identical. call. Now, the only thing that is different between instant, medium, and high, that medium and high did not write the euro sign next to the numbers, and the instant one did.

10:04But other than that, the structure of the tables and the actual numbers in them running through a pretty complex analysis across tens of thousands of data points came out exactly the same. It's a 100 % pass without any miss running on the three different tiers with the differences being running on zero tokens to 10 tokens to, and I could have continued because the numbers would have stayed the same because these are the correct numbers. I could have spent 50 tokens to do exactly the same thing and I would have ended up with the same results. So there are many, many, many tasks that you're doing every single day for which you can use the cheapest model and still get consistent and accurate results as you need, which is going to save you a huge amount of tokens.

10:46Now, the trick is

10:48Isar Meitis:the prompt. The reason I got exactly the same results is because the prompt was very specific. I told it exactly what I'm looking to get. I told it exactly what the data is. I told it exactly the format in which I want the output. If you give it the right prompt, a highly detailed prompt that explains exactly what you want, you will get that outcome even on the cheaper models, which is a huge benefit. And again, you don't have to know how to prompt awesome. You just need to know how to explain in English what you want to get. You can run that on instant. So it's going to cost you zero tokens to create the prompt, and then you can run the prompt and still get really good consistent results.

11:23Isar Meitis:So that's another really good thing to know. And again, this is a real proof from a real test. I've done similar things with Claude, and I found the same exact thing. Now I'm going to talk about some quote unquote hidden features that exist in those tools that allow you to save a lot of tokens for yourself or for other people working with you in your company. Inside of ChatGPT's settings, in the section called model features, you can pick which level of reasoning effort is going to be available for you in the user interface. So you can't mistakenly leave it on max or even worse on ultra or even on extra high.

11:59You can choose to have only light, medium and high. And that's the only things that are going to show up. And then you cannot mistakenly pick the higher levels. Now, if you really need it for a specific thing, you can go back to the settings and add that in. But ChatGPT actually has a cooler and much better feature. That feature lives in the main general section of the ChatGPT settings. And it is called, it's like a flag, like an on-off radio button, and you can turn it into what's called higher intelligence. And what it says under that, it says ChatGPT can automatically use a higher intelligence settings when you ask a complex question.

12:34What does that tell you? It tells you that you can leave ChatGPT on the lowest settings, which is called instant, that again has unlimited number of tokens that you can use, and turn this on.

12:44Isar Meitis:And when ChatGPT itself runs into a situation where this is not good enough, it's not going to give me good enough results, it will upgrade on its own to a higher level to solve whatever section of the process it needs. And then it will throttle back automatically, which is really, really awesome because it's exactly what you want. And I wish it will exist in all the other models as well. And I assume it will exist in all the other models as well. If you've been listening to this podcast for a while, you know my thoughts about this. I think the labs will have no choice to figure out smart routers on their own.

13:17Isar Meitis:Otherwise, everybody will switch from using their model to choosing just a third-party router. And then it's not just choosing between their models. It is choosing between all the models, which is another thing that you can do, right? You can sign up to a tool that has routing capabilities and automatic routing capabilities. And there's many of them right now. And then it will pick between whatever chat GPT, whichever cloud, whichever open source models that you will allow it to play with. And it will use those in order to do the work in the most efficient way, which will save you even more money.

13:47Some companies are not going to do that

13:49Isar Meitis:because they are from a practical perspective, have an alignment with one lab. Some just won't do it for whatever other reasons, but the labs themselves won't have a choice if they want to stay competitive, they will have to auto route their own process. And so Chachapit already does that, and it is very easy to turn on from the settings. So Chachapit has this great auto-writing capability, and Cloud does not have that. But something that I've built for myself a long time ago that is working consistently is I build pretty sophisticated and complex processes inside of Cloud. Cloud uses multiple steps and multiple tasks to complete the process, and it builds a task list that tells it exactly what to execute at what sequence in order to complete the bigger job that I'm saying.

14:35Isar Meitis:And I'm not talking about the task list that Claude builds for itself as just from a agentic perspective where you sit in the top right corner where it gives itself tasks. It's much longer, more detailed planning that I'm doing with Claude. And again, I'm not doing it. It's a skill that's doing it automatically right now and defines a multiple steps, well-designed process to complete anything that we're working on. But in that process, when it is assigning the task, when it's defining which tasks needs to happen, it also defines which model it is going to use. And it is actually doing this for across the board.

15:08Isar Meitis:I would say the majority end up being sonnet, but there's definitely haiku and a little bit of opus. And every now and then there's a fable and it is really, really cool to watch. So it wrote the instructions itself. I'm going to tell you right now the quick summary of what it does. It says for haiku, mechanical, low judgment, sonnet, routine, knowledge work, which is the default, opus, output quality is the product, or the build is technically complex. And then for Fable, it said exploratory synthesis or pattern finding feeds the final output isn't the final output itself. So it is what Claude is doing, have defined on its own on how to use the different models, and it is assigning it for the different tasks on its own, and it is executing based on that.

15:55Isar Meitis:The other thing that you need to know is that when you're running multi-skill orchestration, you can not assign specific levels to specific tasks. But what you can do is when you're building really complex things is assign some of the steps to other models. So in the case that you're seeing right now, it is showing that I'm using NA10 and I'm using other third-party tools together with Claude to do some of the tasks. So when it's running the cloud skills, it is running the skills or whatever the default was when I was running it. But it can call other tools through the API and so on and use them to complete some of the steps at a much cheaper cost, including using open source models that could be significantly cheaper, which allows you to still stay within the cloud environment, but save a lot more money on things that do not require the cloud capabilities.

16:48Isar Meitis:The other thing that Claude has that now became significantly more powerful, when I say now is with the announcement of Astra, that's the other thing that they announced, is caching. Now, what the hell is caching? First of all, it's caching with CH, not caching with a SH, so it's not from the word like money. And it's actually spelled C-A-C-H-E in singular and then caching in plural. But what it basically does is it allows Claude, and by the way, the same thing in other models as well, to store information, large pieces of information online for you, which means they don't need to get processed again and again.

17:22Isar Meitis:It's going to be sent through the model less times, and hence it is significantly cheaper. And they just announced that their caching price is going to drop by 75 % compared to what it was before. So how and where does it work? So in a regular chat, you just go to Claude regular chat. It doesn't matter whether you're on the web or the desktop or the mobile version. There is no caching at all. Anything you upload, anything you write, anything you say is going to be charged at the regular rates, which is obviously the most costly option. If you go to co-work, it is going to be automatic, meaning Claude co-work knows how to cache information.

17:57Isar Meitis:And then you pay a lot less for the cached information. And it does it once every session. So if you started the session and it stored something for the rest of that session, it is going to have that information saved. You do not need to do anything. It runs automatically in the background as a service for you from Cloud. If you use Cloud Code, it is basically the same thing. It is automatic and it's on every turn of Cloud Code. It can reset the session and you're going to lose the thing. By the way, it gets saved for a limited amount of time. So if in that time you're not doing anything, that session expires, and then you have to upload information again, or it will upload information again, depending on how it is set up.

18:37Isar Meitis:And then in the API, if you're building anything with Cloud's API or the same thing with ChateBit or any other API, caching is available, but you have to set it up as the developer. You need to flag specific reusable pieces of content or data that it needs to do to be cached. And then it is going to sit in a section of its memory and you're going to get charged significantly less for that. So that's another little trick to know on how you can save tokens and effort of the AI and still get the same results. The other thing that you need to know in the caching stuff when it does it automatically, if the data is too small, and I'm not sure exactly what that means, but if it's not a big chunk of data, it is not going to cache it.

19:16Isar Meitis:It is going to go to it every single time. So from that perspective, there's actually a value in uploading more information to Claude in order to benefit from the caching pricing. So how do you know when to cache? How do you know what to cache? How do you know which model to use? How do you know how much to use it? The best way is to experiment in your particular use case. Do exactly what I did with the example in the Excel for whatever your use case is. And if you have licenses to more than one AI, if you have Chachapiti and Claude, try the same thing in both. Try it in all the different settings and see if you can tell a noticeable difference.

19:52Isar Meitis:Try it multiple times because the fact that it did it once doesn't mean it's going to do it consistently, good or bad. So run it multiple times and see if you're getting the same results, the same good consistent results at what level. And that's the level you should use AI at versus increase it. You can also use all the tricks that I said before. Two more things that you can do. One of them that I mentioned, you can use a tool that already has built-in smart routing into it. Many, many companies offer this right now. I've been using Open Router for probably two years now to test different things with different models at different times.

20:25Isar Meitis:And you can build smart routing capabilities in there. but also there are tools that are built to do exactly that. So this is another way to do this. And the last thing that I will say when it comes to not capping your tokens from a specific model, I use Cloud Cowork mostly. I use Cloud Code a lot. And what I started doing recently because I ran out of credits at the end of the week, before the end of the week, every single week, because I do multiple things in parallel in Cloud, even though I'm the$200 plan, I started using it in parallel to ChatGPT. So I do some of the steps in ChatGPT work, which is basically the same thing as co-work, while they're sharing the same folders and the same data.

21:06Isar Meitis:So I can offset and offload some of the tasks to ChatGPT to do them to save me tokens on Claude. And it's completely seamless because they're both reading and writing to the same files in the same folders. And they don't really know that the other model is working on stuff. They just check the current status. They see what it is and they just continue going. That's it for today. I hope you found this really useful. I can tell you for me, it is life-changing because I constantly ran out of tokens. And again, in my enterprise clients, they were paying a lot of money before we figure out all these things on how you can change the settings in order to save a lot of tokens.

21:40And now they can do a lot more work in the entire company with the same exact budget

21:46Isar Meitis:without giving up on quality. That's it for today. Have an awesome rest of your week. I will see you back this weekend with another.

From the publisher

Are you burning through ChatGPT or Claude tokens faster than your team can justify the cost?

The problem may not be how much you use AI. It may be how you use it. With the right model, reasoning level, prompts, scripts, caching, and routing, you can often get the same quality of work while consuming significantly fewer tokens.

In this episode of the Leveraging AI Podcast, Isar Meitis breaks down practical ways to make AI usage more efficient across individual and enterprise workflows. He shares tests, settings, and workflow strategies designed to help you accomplish more without automatically reaching for the most expensive model or highest reasoning setting.

In this session, you'll discover:

  • Why the most powerful AI model is often unnecessary for everyday business tasks
  • Why lower-cost models can still produce highly accurate, consistent outputs
  • How a detailed prompt can dramatically improve results from cheaper models
  • How Excel scripts can execute complex recurring processes in seconds
  • How model routing can distribute work between cheaper and more capable AI models
  • How caching reduces the need to repeatedly process the same large amounts of information
  • How running ChatGPT and Claude in parallel can help spread workloads and avoid hitting platform limits
  • How organizations can perform more AI-powered work within the same budget without necessarily sacrificing quality

The takeaway for business leaders is simple: AI efficiency isn’t about using less AI. It’s about using expensive intelligence only when expensive intelligence is actually required.

About Leveraging AI

If you’ve enjoyed or benefited from some of the insights of this episode, leave us a five-star review on your favorite podcast platform, and let us know what you learned, found helpful, or liked most about this show!

More from Leveraging AI

All 330 episodes
325 | Maximize your AI ROI (ROAI 🤔) great output for less money with Isar MeitisLeveraging AI · 22 min
Listen in VO