In short
Advanced language models and practical demos—Google’s experimental image generation for storybooks, and Hume AI’s expressive text-to-speech integrated via Anthropic’s MCP tool-use standard (Claude Desktop).
Key claims
Google’s new image model is a “huge change” in ease/quality; Hume’s TTS is more expressive because it uses an LLM-based approach (not just text-to-phonemes); reasoning models are best for step-by-step analytical tasks, while general models are usually sufficient.
Notable examples
Generating an 8-page “Clifford”-style soccer practice story using Google AI Studio (45 seconds per run); converting a poem (“I do not have long…”) into British-voiced audio through an MCP server; using a reasoning model to correctly answer the “last US presidential election not in a leap year” question (1900).
Guests
Richard Marmorstein, a software developer (CS and economics, Washington and Lee University) who demos Hume AI’s TTS workflow and uses Claude/Claude Desktop, MCP, and other LLM tools.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExploring Google's Image Generation
0:45 to 4:30
Discussion on the capabilities of Google's new image generation model and its personal applications.
“Well, I will start with something fun that I've done in my personal life.”
Demonstrating Text-to-Speech Technology
4:30 to 10:30
In-depth demonstration of a text-to-speech model and its various applications.
“Okay, and now I'll demo something a little bit more related to my job.”
Use Cases for Conversational AI
10:30 to 14:01
Exploration of the practical applications of conversational AI in business settings.
“Is Claude writing code that then goes back to your server that tells it what to do?”
Exploring Voice Technology Use Cases
14:01 to 14:56
Learn how advanced voice technology can enhance content creation and user experiences.
“and you want the voice to be better than what you're capable of speaking yourself and just a lot of control over it and to be able to iterate yourself without, like that's one use case, like people can use it for that.”
The Future of Blogging and Podcasting
15:29 to 16:02
Discover how to convert blog posts into audio formats for wider reach.
“You might, you know, have a voice that you want to read it to you.”
Understanding AI Models for Coding
16:03 to 17:38
Gain insights into different AI models and their applications in coding.
“Like you could just like read the blog post, but what I have is I have Claude sort of generate a little intro with an excuse why, you know, I'm not on the podcast.”
Using Reasoning in AI
17:39 to 21:59
Learn the importance of reasoning models in achieving accurate AI responses.
“this is a good way to tell like whether it's a reasoning model or not is like when was the last US presidential election not in a leap year?”
Integrating AI into Programming Workflows
22:00 to 23:50
Explore how to seamlessly integrate AI tools into your programming tasks.
“If I'm like actually working on like a hard project that the AI is not going to be able to do often, I will kind of use chat GPT as my pair programmer.”
Transcript
Automatic transcript. May contain errors.0:00This recent image model that Google released is a huge change. My company actually has two products. One is a conversational AI, which you can hook that up to Twilio and you can have it make calls or receive calls. Our other product is this TTS that we just released, where you could have a conversation, hook this up to Claude or whatever, and create a conversational experience. You can have an expressive conversation where you can customize the voice as much as you want. I have a blog. I'm a software developer. I'm going to just put my blog posts through this, put a little audio link on top. I'm going to publish my blog as a podcast on Spotify.
0:34I'm just going to go down that rabbit hole. That was freaking fascinating.
0:41Richard, in Las Vegas, I'm feeling lucky. Why don't you show me what you're working on, man? All right. Well, I will start with something fun that I've done in my personal life. I took him to soccer practice, two-year-old soccer practice the other day, and he didn't want to participate. He's really into Clifford. Like he watches the Clifford TV show and reads the book. So Gemini came out with this great image generation model. So I figured I would generate my own book about Clifford going to soccer practice and seeing if I could, you know, motivate him to participate in the exercises. So I guess I'll share my screen.
1:16Red dog named. Let's not name it Clifford because that would be copyright. Copyright. Copyright. Ifford. Yeah. Gifford. Who goes to soccer practice, initially doesn't want to participate, but discovers he really likes it. So this is Google's AI studio? Yep, AI studio. Don't go to Gemini. It's not there. It's in the AI studio. And you use their image generation experimental model. and just say like generate about one sentence per frame. Wow. And an image. Story should be about, I don't know, eight pages long. And then you run it. It obviously won't be production quality, but there we go. What are you talking about production quality?
2:09Look at that. Yeah. Oh my gosh. Dude, that's Clifford. I mean, he's not the big red dog, but he's the red dog. Yeah. Yeah, Iford, Iford. Iford, Iford, excuse me, excuse me, Iford. Dude, and it's written this soccer story for your son. I can't read any of the captions underneath, but. Yeah, it's, you know. He wasn't sure about this soccer thing. It looked like a lot of running. The little black and white ball rolled towards him and he just watched it go by. Come on, Iford, try kicking it. Hesitantly, Iford swung his big paw at the ball and to his surprise, it zoomed forward. what's crazy about this is is even the cartooning looks like clifford like it looks looks like the children's books i read growing up yeah it doesn't look exactly like him i think yeah i don't know i feel i'm not i'm not saying about the i'm not saying him just like the style yeah like the coloring and the anyway wow that's incredible yeah so like this takes 45 seconds i can print this out i can give it to my son and you know he's not picky about like the quality of the illustrations and he's just going to eat it up and maybe he'll he'll do a little bit better at soccer practice did you try this did you show it to him i not yet i need to i need to figure out how to export these into like i'm gonna have to individually save each image like it's crazy how just like downloading the images and putting them into the document is the part that like takes way longer than the illustration and the narration well i have four boys under 11 and i've told each one of them the same story like it started with my oldest and then there's a story that i tell about they see a creature in the distance they have to sneak up on it and it turns out the creature is bullseye from toy story and then they ride bullseye to their house and as they're as they're riding to the house they see all their friends they see spider-man and venom and batman and Iron Man and like every character you can think of.
4:03And they love this story every night. Obviously not my 11 year old anymore, but the two to five year old. And that would be cool to put in this and just have it generate a storybook for me. Obviously I couldn't use those names, but it could generate a storybook. Yeah, so I mean, I think just this recent image model that Google released is like a huge change versus like how easy and how good results you can get from what came before. I definitely recommend just checking it out and seeing what it can do. Okay, I love that. Okay, and now I'll demo something a little bit more related to my job. So I work at Hume AI, and we just released a few weeks ago a model for text-to-speech.
4:47And so my job is to demo the capabilities of our model. And so what I've been doing is I've taken, I went to college and took some English classes in college and wrote some poems and things. And I put those through the text-to-speech model just to kind of see what results you can get. And my workflow for doing this is I have made an MCP server that I use via Claude Desktop. Okay, I'm very stupid. So what is an MCP? Right. So I guess it's a standard around tool use. You know, most of the models support tool use. So you can, you know, you work for some software company, you have a software product, you want people to be able to just like chat with Claude or chat with OpenAPI and like use your tool.
5:38So MCP is the standard from Anthropic. You can use it from Cursor. You can use it from Claude Desktop. You can't do it from the website because like it all runs on your laptop or that's kind of the default. The servers run on your laptop. It essentially allows you to use your tools within the model. Yep. So what you do is, let me make this a little bigger. You edit this file called clod desktop config dot JSON. And then there's these commands that you can put in here to add new capabilities to clod. You can do this if you're not a developer. This makes a lot of sense if you are a coder. But like this one.
6:17I was going to ask, did you write this yourself? Were you using cursor to write this? So this one gives Claude access to the file system. I wrote this config. You could probably ask Cursor or Claude just to write this config for you and get some results. But so here I'm giving it access to the file system so it can see the contents of these directories where, you know, this is where I put my... Where your poems are located. My poems and stuff. And then, you know, this is where my Obsidian vault is. I use Obsidian, which is a note-taking software. And this is where I write my code. And then this is the server that I wrote, which is just a little wrapper around the software that my company makes.
6:55And so this is what will give Claude the ability to convert text to speech. So let's check it out. I will start a new chat with Claude. I will say, I have a poem in, let's see, users, twitchered, creative, out to get me. I would like you to help me design a voice to read the poem interactively and then produce audio files for the poem. So we'll do that. You went to school for computer science? Is that what your major was? CS and economics. Where'd you go to school? Washington and Lee University. So a small liberal arts school. I took a lot of classes outside my major. Yeah. Which is a good experience.
7:47I think I got looking at my file system and seeing that this is a directory. Okay, it's reading the poem. Oh, it likes the ending of the poem. And I'm going to create the voice. Hopefully you can hear what's going to happen because it's going to play it too. Oh. Claude is? Yeah. Well, Claude will call my little server and it'll play it through my space. Okay. Oh, I accidentally clicked to disallow it because Claude always prompts you to ask permission. Right, right. I'm going to try again. All right, and I'll click allow for this chat. And then it'll cook for a second, and then it'll speak. It'll cook.
8:28How long did it take you to build this? Not necessarily the total tool, but just your server. I do not have long, but they're out to get me. All of them, those knowing faces I considered friends, those awkward family arms that felt like love. Did you pick British? Like, was that just a, I think... Yeah, so it's coming up with... Well, this is the first line of my poem. I do not have long, but they're out to get me. After he's done going out. Those knowing faces I considered friends. Those awkward family arms that felt like love. Yeah, so my company's model, usually you want to... If you can describe the voice that you want, like you could say British accent, or you can say, you know, this person should sound really emotional, or this person's monotone and stuff like that.
9:19But also if you just like send it text, it'll try and find the voice that is appropriate to that text. So if you put like army hardies, it'll know that it's supposed to be a pirate or that sort of thing. So the first line that played, I was like, oh yeah, that sounds robotic. The second one that it played didn't sound as robotic. They were different. I mean, they were still the same British voice, but they were, I guess, different inflections. I don't even know the right terminology to use. Was that on focus? Like, how did that happen? Yeah, well, so that happened because it's not... Claude hasn't figured out how to use the model the right way.
9:53It can't hear the audio. Like, Claude doesn't support, like, audio input. So what I can do is I could say, you know, those voices sounded too monotone. I would like it to be more expressive. And then can you like emphasize parts of the sample text with capitalization and ellipses to add a little bit more stylistic flair and emotional variance to the voice? What? And then it'll... Is Claude writing code that then goes back to your server that tells it what to do? Claude isn't writing code it's just writing text and it's writing a voice description so the text this all goes to Hume's you know text to speech model but they're out to get me all of them those knowing faces I considered friends those awkward family arms that felt yeah so it's gonna try again how did you build the model so I didn't build the model my company has some researchers who do that so i just show off the cool work that they do so they did their machine learning stuff that i don't understand but all i did was build this connection kind of between clod and you know the text-to-speech you know you put in text and you get out speech you can try what my company does we got a demo well let me just ask what's the use case for this like why why do we need a another text-to-speech company?
11:30Yeah, I mean, there's lots of them. Ours is particularly expressive. So if you're a creative person that wants emotion in your voice or to be able to customize exactly what the voice is like, we're pretty good at that. Ours is the first model where it's an LLM trained. It's not like just a model that takes some words and converts them into the sounds that correspond to those words, it understands like what the speech means, right? Like most of the traditional text-to-speech models, they will go from, you'll type to them and then they might have an LLM like in there in between that converts the text to like phonemes, you know, just like some representation of the sounds that corresponds to those words.
12:19It's not like, whereas ours goes directly from, I'm not the researcher, but it's like a language model, which is a thing. Let me ask you this. What are the use cases for this? So a creative person who wants the voice to sound good, but are there other business applications? Like if I'm a small business owner and I want to have an after hours phone service, or if I want to have sales support, but I don't have enough money to hire a person, I want to hire AI. Would I go to your company and be like, hey, I want the voice to sound like this. And then I'm using this model in the background to answer the questions.
12:57Does that make sense what I'm asking? So it's like you're using kind of two models. One is the voice. One is the brain behind the way it operates. So my company actually has two products. One is a conversational AI, which, you know, you can hook that up to Twilio and you can have it make calls or receive calls. And then you'll have a conversation with the user. And then our other product is this TTS that we just released, where, you know, it's not like you could have a conversation, hook this up to Claude or whatever and create a conversational experience, but it's not going to, you know, be able to be interrupted by the user and like have a natural feeling conversation.
13:34And soon we're going to like, we're going to make our conversational AI model, use this model. This model is called Octave behind the scenes. So you can have like an expressive conversation where you can customize the voice as much as you want. So that's one use case is the conversational use case where you want to have phone calls that happen and deal with things with users. For just the TTS, it's more creatives, right? Just creatives. Yeah, well, I mean, if you want to put an ad on TikTok or whatever and you want the voice to be better than what you're capable of speaking yourself and just a lot of control over it and to be able to iterate yourself without, like that's one use case, like people can use it for that.
14:19Any kind of content creation where you need voices. We've seen people on Twitter, just like a little short animation that they're showing to people. This is so freaking cool because you can actually customize the voice. How would I use this in an application other than content? Because content's great. But if I'm a business owner, like what could I do? Is it better at cloning my voice than other models? We're going to release voice cloning soon. voice cloning is a thing. You could do that. I think one thing that people do is want to listen to things like on the web, you know, you're on the go and you don't have time to scroll your phone and you want to listen to something.
14:55All right. Now's the part of the show where I feel the most uncomfortable, but my therapist says I need to face my fears. So here we are. I've started a newsletter and I want you to subscribe. And what you're going to get every single week are the aggregated conversations from that week that I have on this podcast with an overview of what their business actually looks like. I'm also going to throw in a review of one or two businesses that are listed for sale. I'll give you my opinion on whether or not the EBITDA multiple is good, or there's customer concentration, or there's red flags or green flags.
15:19And then lastly, I'm going to give you one piece of actionable advice every single week on how to buy your first business. So click the link below, subscribe to my newsletter, and let's get back into the show. You might, you know, have a voice that you want to read it to you. So that's a use case for experiences like that where, you know, your users just want a voice to narrate content to them. Like something that I'm going to do is, you know, I have a blog, I'm a software developer. I'm going to just put my blog posts through this, put a little audio link on top so that you can click it. And then I'm going to actually, I'm going to publish my blog as a podcast on Spotify.
15:56Oh, cool. Who want to consume, you know, my written content can do so, you know, from their favorite podcast. That's a really, that's a really interesting use case. Yeah. Yeah. You can do things around it too. Like you could just like read the blog post, but what I have is I have Claude sort of generate a little intro with an excuse why, you know, I'm not on the podcast. And so the AI is just going to read my blog post and then I can have Claude also generate a little criticism of what I've written at the end. So it's a little bit more of a fun audio experience than just the text itself. Let me ask you this.
16:28This is what I like to ask people who are using AI every day, what models do you personally use and what are the use cases that you use them for? I use Cloud a lot. I use Cloud Code that they released, you know, a couple, a couple, a month ago now, almost. Was it at the same time they released 3.7? Yeah. Yeah. 3.7 Cloud Code. And so that can perform like pretty, pretty simple kind of tedious coding tests that I have to do. Like I use that. When I'm trying to solve kind of more complicated problems i use open ai's reasoning models they have a cool thing i can show you so let's just open up chat that's interesting desktop again when you say reasoning models which i don't know what that means yeah so they have like they have 40 4.5 0 1 0 1 mini 0 3 like to me they're all the same but are some of those reasoning and some of those aren't yes so 4-0 is just kind of like their normal model you type to it it types back and that's it these models have this other step where they reason about things and so like they'll they'll you can see kind of like the chain of thought that happens you know like a question i like to ask and this is a good way to tell like whether it's a reasoning model or not is like when was the last US presidential election not in a leap year?
17:53And the right answer is 1900 because, you know, leap years, like every like millennium or whatever, they get an exception to the rule that like 2000 wasn't a leap or was a leap year. Right. Whereas 1900 wasn't. Yeah. Right. Right. Right. Right. You have to like really know what the rules for when a, when a leap year happens is in order to get it. So if I ask this to 4.0, I think it'll get it wrong. Oops. I clicked update just instinctively. When was the last? Oh, here it is. Sorry, I can't help with that. Well, yeah, that's definitely sure. You can help with that. Come on, 4.0. I guess there's some safety thing that they've added to their model that makes it not want to answer this question.
18:39It must be like presidential election stay away from that so i'm gonna try like here's a history question for you when was the last presidential election not
18:59oh searching the web all right yeah that's all right it cheated you cheated yeah yeah yeah Yeah. Okay. But if you, if you ask, let's see, can I just, but if it had used its native intelligence, quote unquote, it wouldn't necessarily have been able to reason through, Oh, actually there was an exception. 1900 was the year. That's that's the answer. Don't use web search. Just figure it out. Yeah. We'll, we'll see if we can do this. yeah okay so it'll try and reason and it does this and it just decides that 1928 was not in the leap year because i don't know yeah and then oh interesting if you do a reasoning model and ask so which which model did you just pick i couldn't i couldn't see that was 4-0 and i had to i had to tell it that it was a history question so it didn't think that i was trying to like vote because it didn't want to get into like and then i had to web search so it didn't cheat and then it it did what I wanted it to do, which is try to find the answer, but get it wrong.
20:00And then if I start a new chat, I switch the model to a reasoning model. This is their most advanced reasoning model. Well, available to me, at least. Oh, so that's O1 Pro? This is O3 Mini High. O3 Mini, okay. Yeah, but O1's good too. I'm actually not sure on the differences, but it's a reasoning model. And you can see it's reasoning. You can usually see like the steps. It has like this internal monologue. the user is asking about this okay a leap year happens every four years with an extra day it does this chain of thought thing that that it hides from you by default but you can you can you can peek into it you can watch it and expand it and then eventually it'll spit out its answer and 1900 gets it right so let me ask you this i've seen all these models when i'm going into chat gpt i usually just use 4.0 i'm trying to write something i use 4.5 help me understand understand when would I use the reasoning models and when would I just use the general model?
21:01Yep. I mean, I would use, I almost always use the general model unless I'm coding or trying to solve just like some analytical problem where you have to reason step-by-step in order to come to the correct answer. Right. So great for coding. So I wouldn't use 4.0 for coding. Yeah. I mean, you can no i get what you're saying but like it would be better to use better results if you if you use a reasoning model usually why what's the difference do you know what the yeah the difference is just it does this chain of thought thing right like it instead of just you know all right i'm gonna write the answer now it it's explicitly trained to do this type of reasoning where it lays out its assumptions and and the steps that it is taking explicitly and that helps it you know it's still you know it's always trying to predict the next token right but like it's trying right right to predict tokens in a way that gives you this thing which is more likely to eventually get to like a right answer that follows from sound reasoning versus you know just something off the top of its head you know that's fascinating i like i had no idea what the difference was i was just like whatever bunch of models but that makes sense yeah so so i use i use quad for coding mostly when I use cloud code and I want it to like actually just like, Hey, here's this tedious coding thing that I don't want to do.
22:23If I'm like actually working on like a hard project that the AI is not going to be able to do often, I will kind of use chat GPT as my pair programmer. They, their desktop app has this feature where you can, you can like connect it to your code editor. Like I can connect it to terminal over here. And then, you know, I don't do that. Where did you just do that what oh you can connect apps yeah you can and then it can see this right and so you know i say this is javascript just you know today is all right what do i want to what do i want to do oh well i'll just like fibonacci and then i won't implement it and i'll pretend that i don't know how the fibonacci secrets works and you can i will like open up advanced voice mode and do hey Hey, ChatGPT, can you see my terminal?
23:17Nope. I can't actually see your terminal or interact with it directly. But if you describe what you're seeing or copy... Yeah, you're lying. Anyway, I have used ChatGPT successfully to like... You know, I don't have to like context switch in between like my chat window and my code window. I can just like... It's just there. I can use my voice to ask questions. Oh, that's smart. and a programmer, which Claude doesn't have an experience like that. It's not actually about the model. It's about - The UI. Yeah, the UI. That's a really cool use case. I like that a lot. Dude, Richard, this was awesome.
24:00Not just because your company's model is interesting. That was interesting in and of itself. But I think a couple of things that I really got out of this were when to use reasoning and when not. I'm just going to go down that rabbit hole. I think that's super interesting. And I kind of want to explore that. And I hadn't even thought of it. You're the first person to actually mention that. So that was freaking fascinating. The UI that you just showed at the very end there, which was speech to text, not even speech to text, you were talking, but also able to work in another setting and still have AI assisting you.
24:32You didn't have, like you said, context switch. I think that was fascinating. And then the third one was sort of in the very beginning with this Google AI studio, or I can't remember the name of it. Is that right? Yeah. Google AI studio and use the Gemini experimental model with image generation. Yeah. Yeah. So I have, I have that up. I'm excited to like try that and just make some stories for my kids. Like personally, that's a really cool one. I like that a lot. This was awesome, man. Thanks for talking to me. It's, it was fun to just demonstrate the way I use AI. I think, I think it is, you have to figure this stuff out, right?
25:03And it's not obvious from all the press releases, but if you're, you know, if you, spend some time with it you can figure out ways to make it work for you
From the publisher
MY NEWSLETTER - https://nikolas-newsletter-241a64.beehiiv.com/subscribe
Join me, Nik (https://x.com/CoFoundersNik), as I interview Richard Marmorstein (https://x.com/@twitchard). In this episode, Richard shares some fascinating ways he's leveraging AI in both his personal life, like creating a custom storybook using Google AI Studio for his son, and professionally, showcasing the expressive text-to-speech capabilities of Hume AI.
We also dive into practical business applications of AI voice technology and explore the crucial differences between general language models and reasoning models like those from OpenAI. Plus, Richard gives us a peek into how he uses AI tools like Claude and ChatGPT to enhance his workflow as a software developer.
Questions This Episode Answers:
• What is an MCP and how does it enable tool use with language models?
• What makes Hume AI's text-to-speech model different and what are its key use cases beyond creative content?
• What's the difference between using a general model like ChatGPT 4.0 and a reasoning model?
• How can AI assist with coding tasks and serve as a pair programmer?
• What are some unconventional ways to repurpose existing content, like turning a blog into a podcast using AI?
Enjoy the conversation!
__________________________
Love it or hate it, I'd love your feedback.
Please fill out this brief survey with your opinion or email me at nik@cofounders.com with your thoughts.
__________________________
MY NEWSLETTER: https://nikolas-newsletter-241a64.beehiiv.com/subscribe
Spotify: https://tinyurl.com/5avyu98y
Apple: https://tinyurl.com/bdxbr284
YouTube: https://tinyurl.com/nikonomicsYT
__________________________
This week we covered:
00:00 Revolutionizing Conversational AI and TTS
04:20 Exploring Google's AI Studio and Image Generation
11:32 Innovations in Text-to-Speech Technology
15:47 Practical Applications of AI in Content Creation
20:59 Understanding AI Models and Their Use Cases
