In short
Podcast Notes: A New Direction for AI Developer Tooling
Episode Overview Podcast Title: The Changelog: Software Development, Open Source Episode Title: A new direction for AI developer tooling (Friends) Description: In this episode, José Valim, creator of Elixir, discusses his new coding agent, Tidewave, designed to assist in full-stack web development. The episode covers various features of Tidewave, its integration with popular frameworks, and the evolving landscape of AI developer tooling.
---
Key Concepts and Discussions
Tidewave Introduction
- What is Tidewave?
- A coding agent designed for full-stack web development.
- Runs in the browser alongside your web application.
- Deeply integrated into frameworks like Rails and Phoenix.
Features of Tidewave
- Agent Flow and Verification:
- Tidewave interacts with the web application in real-time.
- Can verify features work by running tests in the browser, providing immediate feedback.
- Integration with Web Frameworks:
- Supports various frameworks including Phoenix, Rails, and plans for Django and Next.js.
- Understands the DOM and how it maps to templates, allowing for streamlined feature implementation.
- User Interaction:
- Offers multiple ways to interact:
- Via chat prompts.
- Through a browser inspector for direct manipulation of elements.
- Popup notifications for error handling.
Development Experience
- Current Challenges with GitHub Actions:
- Discussion on the shortcomings of GitHub Actions for debugging builds.
- Kyle Galbraith from Depot emphasizes a need for better observability solutions in CI/CD tooling.
- Concerns with AI and Developer Productivity:
- Debate about whether AI tools genuinely improve productivity.
- José argues that while there are benefits, the disruption caused by AI suggestions can negate productivity gains.
AI Agent Interaction
- Testing and Code Generation:
- The conversation highlights the limitations of current AI in generating tests and ensuring code quality.
- Proposes improvements in facilitating agent-driven coding while maintaining code quality and reducing redundancy.
- MCP (Model-Controlled Protocol) Servers:
- José critiques the reliance on MCP servers for AI coding tools, advocating for direct code execution capabilities.
- Suggests that the future of coding agents lies in optimizing existing tools rather than creating new layers of complexity.
Future Directions
- Cloud Code Support:
- Development of Cloud Code support as a key feature to enhance Tidewave's capabilities.
- Discussion around optimizing the user experience with AI by integrating better feedback mechanisms for code quality and productivity.
- Building a Sustainable Product:
- Tidewave aims to balance being a paid product while keeping up with rapidly evolving AI landscape.
- José highlights the need for a sustainable business model to continuously improve and support the tool.
---
Key Takeaways
- Real-time Development: Tidewave aims to enhance developer experience by integrating directly into the development workflow, allowing for immediate feedback and interaction.
- AI's Role in Coding: While AI tools offer potential benefits, they also introduce complexity that can hinder productivity if not implemented thoughtfully.
- Focus on Integration: The future of developer tooling lies in integrating existing capabilities rather than layering additional protocols that complicate workflows.
- Community Engagement: José emphasizes the importance of community feedback in refining tools like Tidewave to better serve developers' needs.
---
Conclusion This episode offers valuable insights into the development of AI-driven coding agents and the challenges faced in the evolving landscape of software development. José Valim's approach with Tidewave represents a promising direction for enhancing developer productivity and streamlining the coding process in real-time environments.
---
For further inquiries or feedback, listeners are encouraged to reach out via email at editors@changelog.com.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:14it's changelog and friends a weekly talk show about mcp hot takes thanks as always to our partners at fly.io, the public cloud built for developers who ship. We love Fly. You might too. Learn more about it at fly.io. Okay, let's talk.
0:41What's up, friends? I'm here with Kyle Galbraith, co-founder and CEO of Depot. Depot is the only build platform looking to make your builds as fast as possible. But Kyle, this is an issue because GitHub Actions is the number one CI provider out there, but not everyone's a fan. Explain that. I think when you're thinking about GitHub Actions, it's really quite jarring how you can have such a wildly popular CI provider, and yet it's lacking some of the basic functionality or tools that you need to actually be able to debug your builds or deployments. And so back in June, we essentially took a stab at that problem in particular with depot's github action runners what we've observed over time is effectively github actions when it comes to like actually debugging a build is pretty much useless the job logs in github actions ui is pretty much where your dreams go to die like they're collapsed by default they have no resource metrics when jobs fail you're essentially left playing detective like clicking each little drop down on each step in your job to figure out like, okay, where did this actually go wrong?
1:50And so what we set out to do with our own GitHub Actions observability is essentially we built a real observability solution around GitHub Actions. Okay, so how does it work? All of the logs by default for a job that runs on a Depot GitHub Action runner, they're uncollapsed. You can search them. You can detect if there's been out of memory errors. You can see all of the resource contention that was happening on the runner. So you can see your CPU metrics, your memory metrics, not just at the top level runner level, but all the way down to the individual processes running on the machine. And so for us, this is our take on the first step forward of actually building a real observability solution around GitHub Actions so that developers have real debugging tools to figure out what's going on in their builds.
2:35Okay, friends, you can learn more at depot.dev. Get a free trial, test it out, instantly make your builds faster. So cool. Again, depot.dev.
2:49I got like the news, like I got the update in the podcast that there was like an oxide event and you were there. Was that a thing? How do you do that where you go like to the place? Yeah, this was, this was a first for us, I guess. Cause that's like an internal conference for their company. And obviously we are not internal to their company. So that's a first for us, but you know, We've hit it off with Brian Cantrell and with Steve Tuck, the co-founders, and they wanted us to experience the team and meet everybody. And we had been kind of ogling and gushing over how cool their server racks are for years.
3:29And, of course, I don't have enough money to buy one of their racks, and neither does Adam. They're not going to make a Homelab version, despite Adam's incessant cries for them to have an affordable, maybe a half rack. and so we've always wanted to like see their hardware um in real life and so this opportunity presented itself so we went hung out met a bunch of cool people and saw kind of inside their company what they are up to which was it was cool it was different nice very nice first time to emeryville you've been to emeryville oakland area in the bay jose i don't think so you don't think so i don't Right across the street from Pixar.
4:10Like the Pixar headquarters. I actually didn't even know like Pixar was around that area. So I didn't either until I looked across the street and there was Pixar. You know, isn't like that area in Auckland where you're had like some startups starting to move there. Is that the thing that you're saying or I'm getting things mixed up? I don't know. Honestly, I think so, but we're outsiders, you know, we get invited to the Valley. from time to time. Yeah. And eventually we say yes. You know, it's a cool experience, I think. I think it depends on the, the different places we get invited to, but I think it's a chance to explore the world and meet some cool people, pull back the layers, tell some cool stories.
4:55So I favor the IRL. I think it's cool to do it a few times a year or, or as often as it makes sense. Some version of that. we usually go to all things open but our schedule conflicted this year what about you jose you get out and see the people over yes and i had kind of like i was used to do a lot of that uh especially at elixir like at the beginning go to a bunch of different conferences and just talk about elixir and then of course like at some point it gets very exhausting and then i ended up like just kind of like, okay, I'm going to focus on the Elixir community from now on. And even the Elixir community, like there is like the adjacent Erlang community as well that it's enough to keep a person busy.
5:43But now we've Tidewave and we are like supporting, now like we support Phoenix and Rails. We are working on like Django and other frameworks. I have started to like kind of go back, for example, to Ruby conferences. So last week I was at Uruko, which is one of my favorite conferences. I don't know if you are. You were Ruby folks, weren't you? Yeah. Yeah. So I'm familiar with that. I haven't been to it, but I'm familiar with most of the Ruby confs. Yeah. So just, I don't know if you're alive already, but for the listeners, I've talked about this in other places. What I really like about Uruko is that every year they, people say, look, I want to host on my CD.
6:29They do like a three minutes, five minutes presentation and people attending the event vote where it's going to be next year. And it's usually somebody with like no experience organizing an event now has to like organize an event for like 500 people, right? 600 people. And it's probably very daunting, but I think it keeps like, it keeps it fresh and keeps it always like community centric, right? because it's always moving around. So, yeah, so I was at Uruku and then the Elixir events, I'm going to, in two days, I'm going to the GoTo conferences in Copenhagen. So, yeah, I'm kind of back on traveling mode for now.
7:18You like that or do you just do it because you have to do it? I'm enjoying it right now because I think one of the things with everything that is happening around AI, right? And coding agents that they are like, nobody kind of knows where it's going. Some people tell you that it's going to go there, but nobody really knows, right? Like I think the CEO or CTO of Entropic made a prediction about like 90 % in six months, six months have passed, has not been 90 % of code being written by coding agents. So I think like I am enjoying a lot, like this opportunity with like talking to different people and you are getting like a bunch of different takes and different ideas and things to explore.
8:00So it has been really fun just going out and talking to people. But I'm sure that I'm going to do it enough that at some point, like maybe six months, I'll be like, okay, I had my fail. It's time to hibernate again and go back to the Elixir conferences. But right now it has been really fun. I agree. I think it's fun to step back for a while and become a recluse and enjoy your local world. And then to come out, you know, peek your head out from underneath the rock and see the people again. There's something invigorating and exciting about it. But when you're just constantly on that track of just like travel, travel, travel, conference, conference, conference, it can tend to burn out.
8:47So I think everyone needs to step back, but then also step out and see some people. because that's where the magic happens, isn't it, Adam? I mean, that's where the real relationships actually form. I think so. You know, the IRL is really where it's at. I heard that somewhere. I liked it. Then I experienced it, and I was like, you know, just give me more, please, nonstop. Well, trying to figure out where the AI thing is going. Cloud 4.5 dropped today. I don't know if anybody played with it yet. I have not. But, you know, better, stronger, faster. Uh, still not writing all the code for us. I think I did use it today.
9:27But they said something like it can go 30 hours on a coding bender. I just thought, well, that's really good marketing. Cause I had no idea if it's true or not, but I was like, that's a great way to describe what you're thinking do. Um, which is more than I can last. I don't know, Adam, how long you can bend or Jose. I'm sure you've been on some benders in your life, but 30 hours. Holy cow, man. Yeah. So that would be stronger. I'm trying to think of where it fits in the category of better, stronger, faster. stronger man 30 hours straight i think it doesn't lose context or something i don't know did you read this maybe they are doing more things so right now like they they do the auto compactation of the chat and context engineering is all the rage right now right so uh right but like they do the automatic compactation of the of of the the conversation which is summarizing it but But something that they do is also that when the context is getting filled, they just prune the tool's output from the beginning.
10:27So there is like some files or some searches or commands you ran at the very beginning of the conversation. They prune that. And that also allows them to go for long without having to summarize stuff. Because when you summarize, there's always a chance that you are losing some data. And something that is really something that I do it a lot is that you can actually have conversations with the agent about this stuff. So you can like, so people, so a lot of people say like, oh, like there is the agent, the agent doesn't tell me like, oh, I'm using this agent and this agent doesn't tell me which tools you have available.
11:12but you can always just ask the agent like which tools you have available and then you can you can you can like have a just say look invoke so something that i do is like invoke like echo with uh uh with you know a a hundred thousand times three times so i can force it to fill in the context name like what you see now on that tool that was invoked and then it's it's like oh the two output disappeared it's like it's surprised that stuff like itself like it now says this thing and you can you can trick it like to crash very easily as well say like so you're like oh why you have this tool like wouldn't you think like this tool would be better and then you just say like would this tool be better it imagines that the two exists and it's like yeah it would be better let me try calling it but it obviously doesn't exist and then the agent crashes so i like like having those metal conversations and they uh they get like surprised or tripped up yeah it's fun it's almost like talking to a kid you know it's just very easy to pull the wool over their kid's eyes and they'll just they're just gullible because they don't have life's experience that we do and you can have a lot of fun as long as you're you know keep it in good natured fun and not trying to actually trick a kid.
12:31But with your AI, who cares? It's a robot. Trick it all you want, Jose. You know, get it to do all kinds of stuff. Have you heard? We just learned this from Feroz that people are actually using prompts in their malware now. So if you can get arbitrary code execution on someone's computer, for instance, this was in the case of NX, which is like a monorepo command line tool. And they hacked the NX NPM package, distributed some malware, and if you're running the nx command and you're and you're infected in there was an actual prompt to ask cloud code to do stuff for it instead of like you know instead of coding it out yeah and and it was really kind of smart because what they asked it to do was the things that are kind of fuzzy finding for humans which is like for programs i should say which is like find all the interesting files on this computer which of course you could have a list of where the interesting files are and you could like search certain things but you can but cloud code can just go do tool calls and read stuff and just hand back a list of interesting hackable files you know like secrets and whatnot anyways i just found that to be amusing so even attackers are getting lazy to write their own code that's exactly like why like this was the process the promised land isn't it you know you don't have to code anymore when you're hacking someone's computer yeah i wonder why that was the best route was it because of their laziness or their lack of desire to write that script or just because they were just trying to leverage you know a cloud code enabled developer's machine like what do you think the the true psychology of that choice was i think they're just thinking this is the fastest way to the best result you know just like most programmers like what's the fastest way to the best result well i could write a program and besides i only have so much stuff i can shove in i'm assuming like the more stuff you put in the more likely to be found so maybe some compression is in there but it's like if i could just prompt something to to scour your computer for interesting files that's a lot that's pretty good at it that's a lot faster than me having to write a program that scours your computer for interesting files that's my guess i don't know jose why do you think somebody might do that maybe they're just showing off yeah i don't know they just want to trick a computer you know they want to trick an AI into their nefarious deeds.
14:53So, okay. You like to mess with them. How much value are you getting? Because a survey says that we're getting tons of value, but quantified research says that we're not. I don't know if you've read any of the research, but a lot of recent papers, a lot of, you know, at least more than one have come out and said, Developers think that they're more productive with AI coding tools, but it's actually slowing them down. What are your thoughts, Jose? Well, I think I have many thoughts on this. The first one is that to nobody's surprise, we're really useless at estimating stuff. Of course we are estimating.
15:37We've been proving that for years, haven't we? Yeah. So of course we are estimating things wrong. And I feel like the fact like people come like with this exaggerated, like, oh my God, like in three times more productive or even twice. For me, it's just kind of pointless because if actually, if you're even like a third more productive, 33%, that's like kind of like massive. That's huge. Yeah. Right. And then I think people fail to consider, like there are other studies where like developers, I don't remember the exact number. It's like we spend like 50 % of the time coding, let's say. And then if you're using agents for coding, of course, how more productive you're going to be is going to be where you're using agents.
16:23And if you're only using agents for coding, you can only optimize that 50 % and not all the other things. so all that said uh i yeah and then the other thing is that we a lot of people they don't estimate they don't consider the time that they lose when something doesn't work right everybody's happy like oh i use the agent it worked i was super productive but there are a lot of times where it's just it's not productive and then you ended up like trying to coerce it to do the right thing and then it doesn't do it and then you quit right and then you try it again and then it works and you completely forgot about that bad experience the the bad experience is actually one of the reasons why i never like the completion like the the ai completion suggestion because i would read it and if it's not what i want it would always throw me out of my loop and like that time where I'm like, I read it and then I'm like, oh damn, I lost my flow.
17:27Like, how do you measure that? Right. If you're only measuring like, oh, it was accepted, you know, two thirds of the time, but like the one third was so disruptive to me that, you know, like, so with all that said, I think that I, I get a, I get a benefit from it. You do. Right. I think. I think. There's my six caveats, but I still think I do. citation needed right uh i i like i i joke that i would love like the we could use the the the wikipedia citation need it should like be like an html feature we should just be able to put that everywhere you know like in conversation after every sentence that i say yes yeah because the other thing is it's rigby is that samsung's thing silicon valley oh gosh sorry oh yeah continue You know I was not going to get that.
18:21You were hoping Jose was going to get it, weren't you? No, yeah. All right, well, some people got it. Yeah, we can talk about Silicon Valley later, but yeah. So the, yeah, you just did the AI completion. He just autocompleted the wrong thing. Yes, perfect. In the reels. In the reels. He lost his limit. Okay, he's back. Context switching. So there are a couple of things that I do that I think, like you have to find where it works and where it doesn't work. And of course, it's going to change as those things improve. So for example, I tried it a couple of times to help me work with Elixir type system stuff, and it doesn't work.
19:00It's going to be useless. I'm not going to try again. Maybe in six months, maybe in a year, things change enough that it can help me with that kind of work. But I don't feel it's there. But for example, when working with TideWave, because it supports other web frameworks, I often implement the feature in Elixir. Tell the people what Tidewave is real quick so that the three of us know, nobody else knows. What's Tidewave? Yes. So Tidewave is a coding agent for full stack web applications. Okay. So I'm going to summarize it. We can jump into it later. But the idea is to have a coding agent that is tightly integrated with your web framework.
19:41So we understand like what is on the DOM and how that maps to a template. It can coordinate the browser. So it gives it a really strong verification loop. So as you ask it to build features, it can verify that features work. You can interact with the actual web page and ask changes on the page instead of asking for changes on the code. I have like this whole idea that I think we should run coding agents on top of what we produce. So if I'm working on a library, what I produce is like API docs, fine, run that in an editor. But if I'm building a web application, I want to run the code agent in the actual browser because I want it to understand what I produce, right?
20:28And I want to be able to interact with what I produce because if I can't do that, we are doing boring translation work all the time. Like looking what happens in the screen, go to the editor, ask it to change things, right? And then the agent says, I'm done. You reload the page. There's an exception. You have to copy and paste the exception back to the agent. Like, you don't want to do this boring stuff, right? And I say, like, the data science folks, they were the first ones to notice that because they were the first ones to put the coding agents inside notebooks, right? They're like, okay, let's run this thing inside a notebook because if it understands my variables, if it understands my cells, they're going to be more productive.
21:07But nobody caught up to that trend, right? We kind of regress. We first put in the editor and then we put it in the command line. Right. Like, you know, so we should be going up. Right. So that's TideWave. And I do think TideWave can help you like be more productive with AI because being able to allowing the agent to verify what it builds is going to make it so it builds better things, things that are guaranteed to work and are going to spend less time on that loop. Right. So when I'm working on TideWave, like we support Phoenix, we support Rails, we are working on Django, Next.js and a couple others.
21:50I usually implement the feature in Elixir, tell the agent like, hey, I implemented this feature in Elixir. Now do, then I go to the Rails project, implement the same thing. And I'm like, there are a couple of things that I tell it's like don't add tests don't use mocks right so there are some cracks in there uh but you say don't add tests and then you say don't use mocks i mean if it's not writing any tests why is it sorry don't add additional tests sorry than the ones i wrote because the lxrpr is like it's it's good i wrote it right it's right it has a test in there so it's copying those over it has like the the proper tasks.
22:29Yeah, because it tends, I think my experience with Coding Agent for coding is way better than testing because testing stands, not in Elixir, but because I'm doing a lot of Ruby and Python, it tends to use mocks a lot and just writes a bunch of redundant tasks. So it's a whole separate discussion, but it's really good. When I ask it, get this PR here, translate to this repository, a lot of the time is just perfect. It's done, it runs the task, it runs the linter, and I can just push it, right? I send a PR, people review. So that's really good. So I think that that's one of the things you have to figure out where it works and where it doesn't.
23:14Take notes of that, right? And find the loops and tools that make it work for you. And then you can be, it's like any other tool. And I think AI has this particular problem that some people say like, oh, it's just magical. It's kind of like a lottery, a lottery in some sense. Like some people go try AI and because it's, you know, it's probabilistic, they get a bad experience and they're like, oh, this sucks. And I'm going to try it again because people come with the expectation that it's just going to work. And then some people, again, like you're trying it for the first time by just the randomness of it.
23:50They have a good first experience and then they start investing on it and refining it, right? And that's the process. You do have like to figure out what is there and what isn't. And then the other thing that I like, I tell people to do, which works really well for me is to, I don't correct the agent. I just like, if it does something wrong or if it's like 70, 80 % good, I just go and finish it. That's fine. You code it for it. It depends. So I do two things. So imagine that I ask it, does this thing? And then I, I leave, I come back like, oh, this sucks. I'm not asking it like, oh, you were supposed to do this instead.
24:30Because often when it does something wrong, it's there in the context. It has a really. It's going to keep getting wrong. Yeah. And then when it fixes, it doesn't fix everything. So when it does something wrong, I usually go and like, I start a new chat. I just discard everything, right? No, nobody's going to be upset. I just discard everything. I was like, okay, start again, but do this, this, and don't do that. I add a little bit more of context and then if necessary, I start again. So you start fresh with additional little warnings or instructions. Yeah. Adam, you do the opposite, don't you?
25:03You never write the code. You just keep telling it to do stuff. Yeah, I don't really. Yeah, I think by and large it's writing code. I can't really write any, you know, myself anyway. So it's not, you can do a better job than I'm going to do. Jose is a better programmer than both of us. So he can just fix things. We're just different people, you know? Yeah. So what are you using it for? CLI tools, really. I'm having fun with like a Proxmox CLI where it instantiates like a virtual machine with a given cloud init image. And it's like a command line away, basically. It's cool. So I can spin up a new server immediately, essentially.
25:43I can package it as a server. You can share it with me as a Git repo. It's kind of cool. that and I would say 7zarch which is a uh you know 7z is the compression algorithm so I yeah I was just working on a version of that as a CLI that's just cooler basically because 7z's uh its existing command structure is just kind of like not a lot of fun it's hard to remember I can always forget it it's highly configurable and so I wrote something that was just more fun so does it wrap it and then call it underneath the hood with specific flags yeah or is it essentially it did that for a while until then it was like uh we essentially just like we rebuilt uh something called lib 7z which is a wraparound i think it is like rust 7z2 is a crate out there okay and so it's like it actually acts as a library around 7z essentially and then you can write a CLL layer on top of that because it's a library.
26:46So that's where it's currently at right now. That's cool. So you don't have to actually shell out. You're actually re-implementing the functionality with a Rust crate. And you get a lot more data in that API as well. Like you get a lot more granularity around like files and process and progress. And you can control a lot of the UX around the CLI that way. We deal with a lot of large files and folders. Yeah. So I'm just sort of enamored by archiving them very well. Archiving it to the best of his ability. Yeah.
27:30well friends it is time to let go of the old way of exploring your data it's holding you back but what exactly is the old way well i'm here with mark de poe co-founder and ceo of fabi a collaborative analytics platform designed to help data explorers like yourself. So Mark, tell me about this old way. So the old way, Adam, if you're a product manager or a founder and you're trying to get insights from your data, you're wrestling with your Postgres instance or Snowflake or your spreadsheets. Or if you are and you don't maybe even have the support of a data analyst or data scientist to help you with that work.
28:04Or if you are, for example, a data scientist or engineer or analyst, you're wrestling with a bunch of different tools, local Jupyter notebooks, Google CoLab, or even your legacy BI to try to build these dashboards that someone may or may not go and look at. And in this new way that we're building at Babby, we are creating this all-in-one environment where product managers and founders can very quickly go and explore data regardless of where it is, right? So it can be in a spreadsheet, it can be in Airtable, it can be in Postgres, Snowflake, really easy to do everything from an ad hoc analysis to much more advanced state analysis if, again, you're more experienced.
28:41So with Python built in, you know, Python built in right there in our AI assistant, you can move very quickly through advanced state analysis. And a really cool part is that you can go from ad hoc analysis and data science to publishing these as interactive data apps and dashboards, or better yet, at delivering insights as automated workflows to meet your stakeholders where they are in, say, Slack or email or spreadsheet. So, you know, if this is something that you're experiencing, if you're a founder or product manager trying to get more from your data or for your data team today, you're just underwater and feel like you're wrestling with your legacy, you know, BI tools and notebooks, come check out the new way and come try out Fabi.
29:20There you go. Well, friends, if you're trying to get more insights from your data, stop resting with it, start exploring it the new way with Fabi. Learn more, get started for free at fabi.ai. That's F-A-B-I dot A-I. Again, fabi.ai.
29:40I also use coding agents for things that I'm not reviewing, particularly for like prototypes. And that part has been really fun because if you're working on a product, you have ideas of, wait, what could this project, which directions we could go in the future. but usually before you would like think about it, put on some notes, right? And then maybe if you're lucky in two, three months, like somebody from the team can take a look at it, give feedback, right? And now with agents, you can just say, okay, go for it, implement this thing, right? And so as I was saying, I have this idea that coding agents should run on top of the thing that we produce, right?
30:26And we talked about Tidewave web that works for our applications. I talk about notebooks. Well, but if I'm working on a game, right, I want to have Tidewave running in the game engine. If I am building a mobile app, I need to know about the mobile device, simulators, right, and all this kind of stuff. So I was able to, I think for four or five weekends straight, what I would do during the weekend is to come to the computer from time to time, see if the agent was working, and just have it build a different proof of concept of embedding TideWave somewhere completely different. Like, oh, what would a Tidewave browser extension look like?
Read the full transcript
31:14Which capabilities we got from this? And, you know, like this, doing that when I was, I had to do this kind of things for other products. Like we were doing that for Lifebook. You know, it would take a really long time to validate all those things. And I could very quickly explore something different, get the lessons learned, and provide a way better blueprint for the team to work on. Do you run it in YOLO mode or whatever is the equivalent where it's just doing whatever it wants to do and you come back every once in a while? Yeah. Yeah, totally. So do you have on it? Have you considered like a notifier and text me when you're finished kind of a thing?
31:54Otherwise you're going to keep coming back. Are you done yet? No, it's not finished. Oh, it's been done for two and a half hours, but I was watching TV. like yeah so in this case because it's the weekend i don't care in the sense that i don't want to be also interrupted right like it's not my priority gotcha so when you feel like it you go over and check it yeah otherwise it's like uh yeah other yeah otherwise i'm using the notifications right like uh i use that a lot and tide wave and they all have notifications and then i'm kind of like listening i'm waiting for them oh they don't like push your phone or anything i don't think that tide wave doesn't i don't think that then you don't have to wait and listen you can be out on a walk or whatever and be like oh it's done maybe yeah but if i yeah but if i want to give it his next task you know yeah walking yeah like so it's funny because i talked to chris mccord about this right and it's like oh maybe i am on a coffee i'm i'm out to get a coffee and then i'm like look if i'm out to get a coffee i'm out to get a coffee you know it's like if it's not i don't care i'm out you know it's right yeah same but i do three maybe you lose three hours of productivity man i mean all you gotta do is sell it to keep going you know these trade-offs i get it it's the weekend i like to unplug as well but i don't do any of the stuff that you're talking about i don't have anything coding for me over the weekend so if i was i'd at least want to like you know be a good babysitter not a neglected not a neglecting babysitter but to each their own I guess so you're talking to Chris it sounds like you and Chris are you guys competitors now I mean doesn't Chris have Phoenix.new and isn't this like there can be only one Jose right so I
33:43and that's why he's not coming on the show anymore no I'm kidding so no we do talk talk a lot about those things. And like, we are still balancing many ideas of each other. So the way I think about this is that I think there's a very easy way to separate those things is that Phoenix.new is remote. And I did not, so maybe we should go deeper in TideWave because there is a bunch of additional context here. So as I was saying, like TideWave is a coding agent for fully stack web applications, but the thing is that it runs on your machine, you know? So it's not, so my whole, so one of my ideas that we are looking at like bolt.new, lovable.dev, right?
34:30And they have all those things where you can click around, ask it to do changes. But it's like, they want to kind of own your code. They want to be responsible for your code. And most of the time it's like for its front end or for like React apps. And then I'm like, I want that for my Phoenix app that I run on my machine, right? Like there are, so there's like this, like a lot of people are pushing, oh, AI and those app builders that are running on the cloud, right? TideWave is, you know, you are accessing localhost. So the way you would install it is that you would add the TideWave package. So today it's like for Phoenix or Rails or in the future for next Django.
35:12So you just install the package. After you install the package, you go to your application, localhost, whatever, 4000, and you do slash TideWave. And then the agent is running there, like in the browser and your web app running on the side. And now you can do all those things. You can go to the inspector, click it and say, hey, I want you to, on top of this element, I want you to add a chart of the most listened podcasts in the last month. Right. So you can be very UI driven and everything is running locally. So for the people who are fine, look, I want to have Phoenix.new be responsible for my code, for my deployment.
35:54I don't care about that. And I want that thing to do everything for me. Then go use Phoenix.new. I still think it also owns the getting started experience. It is the best way of getting started with a Phoenix app, right? Just go put things in the prompt. It's going to build something for you that you can throw it out, right? And for me, I'm like, look, okay, I have my own thing. I already have my own infrastructure, my own development cycle. And I want to incorporate all those tools into what I do every day. For me, it's like, you know, when there was like a trend, everybody was saying like, oh, you're all going to be developing on remote machines.
36:35And then there was like those developer containers and that never really happened, right? I remember we did the show, didn't we, Adam? We did the show with whoever it was, like GitHub Codespaces. Cloud development environments, essentially. Yes. Gitcontainer.devcontainers. Yeah, I know some people use it. It is used. Yeah, people do do it, but they're on the fringe. You can use it locally, but it's not like everybody, because people would say that everybody would use it, right? Like, why would you have a local machine, right? So I see it the same way. I want those tools for my framework and running on my machine.
37:17Okay, I am with you. I actually, when I heard TideWave runs in the browser, I was like, another browser thing, Jose? Like they're all running in the browser, but actually it's different than that, right? It's in the browser because that's what your output of your web app goes, but it's in your local browser running against your local web server with your local environment and helping you build cool stuff right there, which is kind of how I develop now anyways. Whereas phoenix.new was making me go into the browser and have a remote browser session, which I've always got excited about for the hour that we do the show.
37:55And then when I go back to my real life, I just don't want to do that. I want to be on my local machine. I always have. Maybe I always will. I'm getting old, so I'm getting stuck in my ways. So that makes me like Tidewave a little bit more than when I first thought, oh, it's, it's, because one of my questions for you is going to be like, why the browser, you know, but it's because I didn't understand. Yeah. And, and, and the thing is, so we actually went through many possible designs. So over the last month you had, we already had like, let's talk a little bit about the browser design. So like we already had for some time, the Playwright MCP.
38:32So somebody may be listening to this and say, well, I can use VS Code with Copilot and install the Playwright MCP. recently, Chrome DevTools, MCP. Chrome, yeah, they released there. And I think yesterday, Coursor Browse came out. What's that? It's just controlling Chrome. It's like the Playwright Puppeteer MCP. It's just building, right? And the issue with those tools is that it is a separate browser session. It's not the one that you are developing. So imagine, for example, that you are working on a project manager and what you need to do is that you need to implement a feature for transferring a project between two organizations, right?
39:17So in order for to implement this feature, you need to create a user, create two organizations, probably make sure that the admin is, the user is admin on both organizations, create the project, and then you can transfer it. And then a lot of the times the MCP is going to get stuck just in this process. Like a lot of the times MCP cannot create an account because create an account requires sending something to an email that the MCP doesn't have. So now we start writing like those backdoors for tests. Like, so there's such, there's a big amount of work, right? And then the fact that we run in the browser, we are literally running in your browser session.
39:58So because when you're going to develop the feature, right, you open up the, you are already logged in your development version. You go to the page already. And when you're going to validate that the feature works, you already have all that set up. So because it's running there in your session, everything that you do for development, the agent can do. And the agent is going to verify the things in front of you, not in a separate session. and then you can actually have a back and forth. Like if you're using the MCP, imagine like the agent's like, okay, let me test that it works. And then the MCP with the separate browser is running and then you see a bug right when the thing is testing.
40:39How you're going to debug that? Because it's a separate browser. How are you going to click things and say, hey, you would have to go around, say around this page, maybe there is a bug here. With TideWave, it is your browser. I think that's the most important thing. You can stop the testing, go with the TideWave Inspector. There is a bug here. Fix it. And we also go the next step, which is that we integrate with the web framework. We understand, like when you inspect like a DOM element, we know the DOM element and send it to the agent. But we also know where in the template or which React component that thing came from.
41:20and we send that to the agent. So we don't have to do the manual working of figuring that out and do it to the agent. When there's an error page, we detect the error page of all the web frameworks we support, automatically feed that to the agent. So it's really meant to be like, look, it's like you, the agent, the browser, the web framework, like in a shared context, kind of everybody can see what the other is doing because otherwise it becomes your responsibility. You are the ones who are like getting information from all those places and pass it around. sounds pretty cool man it does sound pretty cool uh bypassing a lot of that stuff i mean the fact that it has you know that i mean that's something i just uh i've always wanted i guess i mean you're going back and forth like that it's better to do it right there real time i i haven't played with it to know the ux really of that like when you're filling with let's say is it a button maybe it's not working properly what is the back and forth with the the experience can you speak to it?
42:18Can you type to it? What are some of the interfaces you can think of? I think there are three ways that you are interacting with it. One is the usual chat prompt. With the difference that we know what is the page that you're currently looking at. So you can talk to the page in the sense that, for example, imagine that you just boot up your dev instance, your database is empty, you can go to a page that is listing all the podcasts, like for change login, you can say, oh, this page is empty. Add some podcasts. It knows which page it is at. So it can find information from the controller or from the live view and then say, okay, that's the data I need to, like, it gives you an entry point.
43:05So that's the chat. So you can kind of, it has the context of the page. The other one is the inspector. It's like the browser inspector, but so you can click it and then you can mouse over elements. We show the DOM element. We also show like which templates on which Phoenix template it came from. And then you can like click it to open your editor or you can click it to ask the agent to do something. And the other way that we interact with it is when we detect that something goes wrong. we just show up a pop-up like oh you want to fix it right and then you can just click a button and have it fixed for you uh so i think as a human those are the i may be missing some those are like the three so it's a very classic like chat experience with a few like things on top like inspector the error but i think a lot of the part that we shine is in giving more tools to the agent, right?
44:09So the agent can do everything that the coding agent can do, but it can also run JavaScript on the page. And that's how the agent can test that it implements something. So for example, one of the coolest features that we use TideWave to implement, like if you go to tidewave.ai today, we have videos in the homepage. page. And so I added like the YouTube URLs or I got YouTube, like the URLs for the, I added the video tags, right. And then I wanted to make it. So as I was scrolling through the page, the videos started to autoplay. So I asked a Tidewave to implement this, which it can do is like, it's a, it's a, it's a straightforward feature.
44:53You know, I can't do it, but I assume it's a straightforward feature for me. I didn't look at the code, but I'm sure it was pretty easy. Yeah. So it implemented the thing, right? It implemented the thing. And then in order to test that it worked, it actually like reloaded the page. So this Tidewave, it wrote JavaScript code to reload the page and scroll to the first video. and then it runs on JavaScript to validate that the first video was playing, but not the other two, then it automatically scrolled a little bit more. So the second video started playing and then it ran some JavaScript to make sure that the second video was playing and not the other two, right?
45:37And I think that's the important part because if you can see the agent doing that, because if the agent doesn't do that, there is a chance they get it wrong, right? And then if they get it wrong, who is paying the price to fix it? It's you because you are going to be the one who test it. And then you have to go and tell it. I thought you were going to say your users because you're going to push it out live. Awesome. Could be. Could be. And then your users will have to tell you if it's broken or not. Yeah. Yeah. How do you limit it to the viewport? Like I assume the scrolling is either simulated or it's real or it's simulating it.
46:11So you think it's only scanning with you. It just runs JavaScript on the page. So what's in the viewport? So it's, it's looking at what you're seeing essentially. Yes. Yes. It's running in your browser. Like there is no, there are a lot of complexities in there, but like this part is like as straightforward as it can be. It can control the page. So, so the same way, because that's the thing, like people are coming up with all those different APIs to have the agent, like there's an MCP with 30 different commands to control the page. And I'm like, it knows JavaScript. It knows the DOM API. just have it run things on the DOM.
46:48It knows what is, I don't know what is the command, but it knows what is the command to say, hey, scroll a little bit, right? It knows. So the only things like we had like to intervene like very little, like, so there's one of the things that it can't do, like resize, resize the browser window. I think it's because browsers don't allow you to do that because of security concerns or something like that. So there are some things where we have to intervene and add like extra capabilities, but it's just running things on the page.
47:36Well, friends, you don't have to be an AI expert to build something great with it. The reality is AI is here. And for a lot of teams, that brings uncertainty. And our friends at Miro recently surveyed over 8 ,000 knowledge workers. And while 76 % believe AI can improve their role, most, more than half, still aren't sure when to use it. That is the exact gap that Miro is filling. And I've been using Miro from mapping out episode ideas to building out an entire new thesis. It's become one of the things I use to build out a creative engine. And now with Miro AI built in, it's even faster. We've turned brainstorms into structured plans, screenshots into wireframes and sticky notes, chaos into clarity, all on the same canvas.
48:25Now, you don't have to master prompts or add one more tool to your stack. The work you're already doing is the prompt. You can help your teams get great done with Miro. Check out Miro.com and find out how. That is miro.com, M-I-R-O dot com.
48:46Can it take a, I wouldn't say a fixed width, but a desktop designed website and implement it? So it looks where you want to look on a desktop. And can you say, make this a progressively, what's it called? Not progressive web enhancement, responsive web design. There we go. Can you make this responsive for these six viewports or something? Right. Not yet. I knew it. I knew it. I got you. Because the resize thing that I told you. Not because it's his fault. Not because it's Jose's fault. Because I don't think these things can do that anymore. It's hilarious though. I love it. This pursuit of rightness.
49:29I'll try. I'll actually try. Because I've been using, I've been doing a lot of front end lately and I'm not good at it anymore. I'm learning. All the new tools are fancy and they're hard to use. I can't figure Clamp out. I mean, I've been using Clamp wrong for weeks now, finally starting to get it to work. And none of these tools can do it either. So I play what I call LLM Russian roulette. So I take the same prompt and I'm usually I'm like, hey, can you do this thing in SVG or whatever? Like I'm trying to accomplish stuff that I don't think is possible. And I thought it should be possible. It's the modern web, you know?
50:07And so I ask chat GPT, I ask Claude, I ask, I'll even ask Grok if I get too angry and then I'll ask Gemini and they all give me different responses that are all wrong. None of them can do it. I want one that just tells me, actually, Jared, that's not a thing that you can do. You know, like you can't do that with web technology. They're not going to do that because they want to make me happy. But I know that like that kind of stuff, we're not there yet, man i'm just doing way too much work in the browser as a human night now here's how i would try to implement it okay okay and let me know if that's an approach that you try because if that's an approach that you tried then my solution obviously is not going to work i'll let you know trust me yeah so i hope i was talking about the resize that's something that we we identified recently so we haven't implemented it so it doesn't have currently the ability to resize which means that it cannot validate responsive designs.
51:06Right? As simple as that. But so I would try doing would be is add the feature to resize and the feature to take a screenshot of the page which there are some other complications because the browsers don't allow you to do it for security reasons as well. We need, I know how to solve it. It's just going to take, I'm just explaining why you're not going to have this feature tomorrow. There is some work. And then have it look at the screenshots and see if you can see things are good or bad. How do you think about this approach? In my experience, their ability to look at screenshots and decipher things is really bad.
51:46Like it's not there. Like they have vision, but it's not precise enough, you know? And so I haven't tried that specifically. I also don't think it's going to work. I would love you to try and prove me wrong. I would love to be wrong. but in my experience when you pass a screenshot or you say take a screenshot and then inspect the visual nine times out of ten they're wrong all of them so i wonder if we could if we could use accessibility apis wait what this is just a the guy is such i love the exchanges on he's such a problem solver that he like can't help himself right now he's like let's let's debug this thing.
52:28So you already, you already saw me like getting off track with the AI suggestion. We saw it live. So this is a, also a real life nerd sniping happening right here. Yeah. We're shaving the yak. So yeah. What were you going to say? Accessibility APIs. If we could use accessibility API somehow to measure like size of elements and what is visible, what is not, but I maybe, maybe not. Yeah. Right. Yeah. I don't know about that. I just get angry and I just do it myself. Because it goes back to like, to, to what I was saying, the sense that the way for us to eliminate the AI guessing is adding more verification tools.
53:13So if browser had had a way, if the browsers could tell me like, oh, the phones here are too small. These things are clipping. That's why I was thinking about accessibility APIs. because if the browser tells me that then i can get that thing which is going to be better than a screenshot and send it to the agent that might actually work right but i don't know i don't know if this accessibility api exists right so that's why that's why i'm well don't ask it online though i'll tell you that it does exist right i'll give you the code radically i love when i have it they produce svg and i'm trying to get like a tapered border and all this kind of stuff and they're like, here you go.
53:49And then they like tell me all the reasons why it's going to look good. And I put it in there and I put it in there and I'm like, dude, it looks like a bow tie. Like you just drew a bow tie. And it's, it's so far off that I have to laugh because I'm otherwise I'm just going to cry and just be like, why am I even wasting my time with you guys? So there's certain things where they just have these inadequacies and they're all inadequate at this point. In my experience, I haven't done 4.5 yet. So maybe after this call, I'll go see if Claude can do this. But I don't know. I don't feel like I'm pushing the envelope.
54:22I feel like I'm a kind of an, just an intrepid person trying to get something done and thinking that you can do things that maybe you just can't even do in the browser right now. But I think being able to develop out a simple, I'm not going to say fixed width desktop styled website and say, make this responsive, like that should just be a thing. Don't you think? If you build that in the highway of Jose, I mean, people are going to line up with their money. I think so. I mean, cause that's the thing. Like, so I, I don't want to do that work. Right. So if I hope that I can do it, that's the thing.
55:02Like I can also do it, but it's just slower and tedious. And that's what, that's what the promise is. We don't have to do this stuff anymore. And I'm not a good at it anymore. It's just, it's guess and check. I got to guess and check. Oh, it's still too big. Now it's too little. All right. I'm done complaining. Adam, take us somewhere else. I'm just, I'm airing my grievances. Well, you know, one thing I was going to go back to was, Josie, I think one thing you were mentioning was how, when you scroll, uh, tie wave dot AI, as you see these movies come in, I'm actually, you know, back, I think it would be 15 minutes potentially, but you were describing this page here and now that i've actually caught up and i'm scrolled it maybe that's where we can go is like uh what was the aha moment here when you did this because you said you were kind of going back and forth did you not do any of this design yourself did you just sort of prompt it what was the experience like for getting this page to be like this yeah no for this page in so for this page in particular we were just doing the design of the page and then we knew we wanted it should add auto-scrolling uh and then we just ask it to do it and then it did it right and and i think the i think what was surprising about that is because i mean it's obvious but that's exactly how the autoplay of the videos was key right autoplay video but it wasn't the autoplay it was how it tested itself to know that it got the autoplay right yes and that's exactly how we would test i mean it's obvious that the way you test the autoplay scrolling is by scrolling you watch it autoplay and you make sure the other ones aren't, but it's just running JavaScript.
56:32But it's really nice to see it happening by itself. Right. And then it goes back to, to other stuff. Like, like Tidewave has access to everything. Another way that I like to phrase this is we imagine like you're working with somebody and somebody sends up a request and then you open up like the work they did in the browser. And then they're like, wait, this looks bad. And then you go back to the person like, did you look at it in the browser? Did you try it out? And then like the person says like, no. And then we'll be like, what? Like you have to test things in the browser, right? And that, or like I use the repo all the time as well, right?
57:13It helps me develop a lot, but we are asking coding agents to develop without a proper browser, without a repo. So TideWave gives all those things as well, right? So, oh yeah, I was talking about like the, you asked about what are the user tools and I started talking about the agentic tools. So one is coordinating the browser, but the other one is that we also give access to a repo running inside your web application because we use the repo for development. Why are we not giving one to the agent? Like I would be a worse developer if I didn't have a repo, right? And then we have MCPs for like, oh, you can install an MCP to talk to Postgres.
57:53But I'm like, my web application already knows how to talk to the database. It already has all the credentials in there. Why are you asking me to configure a separate thing? So a lot of the times it builds a feature and then it tests the feature in the browser. And then it does a database query to make sure that the change also happened in the database. So that's kind of, yeah, so we're going back 15 minutes, but that's closing the loop of like, what are the tools that the agent have? And the whole purpose is to make sure they are producing something that is really good. And I'm not going to waste my time telling it obvious things like, oh, the video actually doesn't play.
58:34Oh, the change was not actually saved to the database. You mentioned, but I think one thing you mentioned there was MCP servers. Have you, Jared, messed with MCP servers at all? I really haven't. Personally, just Figma's. Yeah. And I told you my experience with that, which was not. Right. hit or miss. There's nobody's fault except for the state of the art. It's not quite what it's needed. I imagine you're probably playing with them heavily. How exactly does that fit into your flow? Because from what I understand it just adds more tooling to the context window, which is already kind of small and so we're always battling that auto compression or just having to refresh the entire chat whenever you feel like it, I suppose.
59:16How do you work in those kind of tools into your workflows those without, I guess, bloating the context. I actually have a hot take in here that it's not a, it's not a unique hot take, but, so to answer your question, which is going to kind of reveal the hot take is that He's just teasing it. He was not going to reveal this. He's just teasing the hot take. Come on, man. So, cop setting it up, Jose, and give us the hot take. So, like, almost all of our APIs is write code. Oh, you can execute code in the context of the web application. You can execute code in the context of the web page. That's it.
1:00:00Because like we are doing all this dance. Like, oh, I said about the database, right? Like, oh, I'm going to have an MCP for the database. No, my web application already knows how to talk to the database. Just use that. Oh, I want to have an MCP to talk to GitHub. And I'm like, well, I already have like, Like I already logged in on GitHub in the browser. I already have the GitHub command line. Use that, right? Like for coding agents, we are even going as far as adding MCPs for documentation. And then I'm like, why am I going to a separate website to get documentation? You are a coding agent. The code is in your machine.
1:00:38And usually with the code, you have documentation. Why don't you use the documentation that is there already on your machine with the exact version that you're using? Because sometimes you go to the remote server and then we get the documentation for Phoenix 1.8, but we are still on 1.7, right? So for me, like the answer for, oh, they are like too many, like the context thing is like, I'm going to have just a small amount of tools and what those tools are going to do is that they can run code, right? And I'll let them do whatever they need to do. So trying to keep the amount of tools minimal and powerful.
1:01:18And like this, this take that, you know, like, oh, MCPs are too much. You probably just need code, right? That's kind of, I'm not the first one to say it, but I also think like MCPs, the user experience, like the developer experience around MCPs for coding agents is really poor. Like, I mean, to be fair, it's new, it's still evolving, right? It's probably six months old at this point. But like, so we have issue where, so one of the MCP2 that we're using, it works, it was working for GPT-5, but not for Gemini. And then we fixed it for Gemini and it broke for GPT-5. Like it, if the server disconnects, they cannot reconnect again.
1:02:04Like there are all those sort of like annoying issues there. And then do you know about the Figma dev mode thing? Mm-hmm. So like, so there is an MCP in Figma dev mode. So you can run like Figma on your machine. There's a desktop client. And then I can go to Figma, inspect an element, click like a component that I wanted to implement. Right. And then, you know, what is the workflow today? It's like, I have to go to Figma, click on the component. And then I have to go to the agent and say, I have selected a component. Please implement it. And then like, when it's done, you have to redo it yourself.
1:02:47Yeah. That's my experience. Oh, good job. Not good. Not good. It doesn't look like it's supposed to look like. I already clicked the thing. Why do I have to go back and tell you that I already, like, you know? Because they're separate tools, right? They're distinctly different tools. You're meant to have a protocol for those things to communicate. I know you know the answer. I'm just saying it out loud. whereas it was TideWave it's all integrated it's all it's all integrated and I actually I want to I want to hold as much as possible with actually adding MCP support to TideWave because I think we will have better integrations if we do it by hand so for example when I do Figma for TideWave when you click on the Figma thing we will know and we just tell it oh you You want to implement this?
1:03:40Oh yeah, just click the button, right? Like don't have to type anything. Just click it. Now, is it going to do it right? Well. That's up to your model, right? Yeah, maybe we can give more. TideWave is just using whatever model you bring to it, basically. Yes. What we can do is that we can help it. Like you'll be able to click something on Figma and then click something on TideWave. And then we'll be able to say, oh, you should implement this. And this is exactly where it is. So we can improve the experience there, right? But when we send all the information to the agent, if that's going to be ultimately better, and the agent will be able to validate that some things look good, like it did with the video scrolling.
1:04:23So we are giving it more tools to verify that it did a better job than it just working blind. But ultimately, yeah, right? And I think that's going to be true for a lot of things. So there is a tool called Conductor that, for example, they added a GitHub integration. And one of the things they do is that they know which GitHub, they know which Git branch you're using. So in their GitHub integration, they know the comments dropping a PR for that branch and they automatically surface that in the UI. So you can ask the agent to solve a comment as somebody's commenting on GitHub. So like those sort of experiences doing through MCP, it's just, oh, you know, it's like, oh, get all the comments for me.
1:05:15Right. It's like, so I really think that for coding agents, I don't want to generalize this too much, but I think like for coding agents, for a lot of the things you can build, like MCPs doesn't allow you to push information. Right. So that's what I'm complaining about. Like GitHub should be able to push information for the MCP. like, oh, there's this comment. Oh, I clicked this on Figma. It doesn't support that. And we're not talking about even the security issues. So I feel like, yeah, like I want to give you like a good package with everything at the point you're telling users to go like, oh, just go and install those different MCPs.
1:05:56You kind of like gave up on the developer experience, right? Because it's not there yet. I would tend to agree, I think, with that. How hot was that? Was way too much teasing for not too spicy? It's a lot of spice in there. A lot of spice. It's a variety of spices. There was some hedging around. I just feel like you could have dropped it a little hotter. I could, yeah. Kanye could have gone ghost. And also, I just tend to agree. I think MCP server seems to be like a builder. Builder-driven technology right now versus user-driven. Like, I feel like it's, it was like so quickly adopted by all the builders.
1:06:37And as users were kind of like, were we asking for this necessarily? And can you, could you do it so that it was made, I would say more transparent perhaps, or like maybe just user friendly for us as end users. But man, I've never seen an API or specification or a protocol get built out across the entire tech industry so fast. And we're talking like less than a year. well I mean from their first announcement back in November I believe was when MCP was announced by the Anthropic team like less than a year ago and nobody paid much attention to it for three months and then all of a sudden in the spring ish it's like everybody just started building MCP servers like everybody that's true yeah and then as end users we're kind of like did we ask for this or I don't know I'm not sure what it's like you want to be first I don't know you don't want to be left out I'm not sure why everybody just thought immediately we got to do this but it was pretty interesting to behold that you know I don't know I don't have a lot of contexts around MCP servers but I don't really use any of them but when I think of them it's more like a CLI tool that's on the system already rather than like you know pollute my context window with a tool that's an MCP server why not just have a tool on the system that you can use, not have to be an MCP server.
1:08:05Instead of using the GitHub MCP server, you might use the GitHub CLI to access data from GitHub? Is that what you're saying? Right. If I already have GH installed and it's already authenticated, why not just use the tool versus some sort of MCP server that just is in my context? Why does it have to be in my session and configured? It's also instrumentation of tooling. It's a lot of ceremony. You know, it's a, it's a lot. But even if you say like, look, well, what if, what if you, I, you don't have the GitHub command line tool. Yeah. Have it right. Like it can write Elixir, it can write JavaScript, it can write Python to talk to the API.
1:08:43Then of course, so like, I think like the authentication part of the MCP is interesting because if it had to write like a tool, it would have to ask for your credentials somehow. So that's part is good, but it feels like that's in certain ways, that's probably all we needed, a way for the agents to ask, which we have all of and other things, and a way for the agents to ask for your permission to use, to talk to some API on your behalf. Because if it has the code, it could also do things like, oh, it can ask information, it can get the raw data from GitHub, then use whatever library to compute the information that you want, right?
1:09:29And give you a better result than trying to do with the MCP, getting plain text, and then maybe doing something interesting with that, right? Like we have those things that are really good at coding and we are sometimes like dumping things down to a text interface while they could write code. I like to think one of the questions that I ask myself, like people are talking a lot about like personal devices, right? You're going to have our personal devices that are AI augmented and this kind of things. And I say like that thing needs to know how to run code, right? It's like, because how can you have like some like generic, like personal assistant can do everything and that thing cannot run code.
1:10:12Any assistant in mind has got to be able to run code. That's right. You better run code. I'm telling you, get out of here. It's like the first thing on the resume. Can this person run code?
1:10:26Well, yeah, I tend to agree. I think MSCP is an interesting phenomenon and most widely adopted by builders. technology in that I can think of in history. So there it is. It's there now, but not necessarily, it didn't necessarily have to be there. And, uh, there you have it. Spicy, spicy Jose. What else? I mean, Tidewave, you're trying to make a business out of this thing. You're trying to make a living. What are you trying to do? Yes. So it is a paid product. We, um, consider a little bit, but then I realized, well, everybody, this is an AI thing. It's going, it's a very rapidly changing landscape.
1:11:08So if we want to be able to keep up and feel invested on this and continue improving it and also support different kinds of frameworks, we need to find a sustainable way of doing that. So yeah, we'll see where the launch was pretty good. We got a lot of people excited but it also pointed out like so today's like bring your own api key and the the feature that people ask the most for is cloud code support so being able to bring like codecs and cloud code um and yeah let's see so right now i like to i think in the email i sent you folks like uh my product uh history my history of building products have has been like catalogued by uh by the changelog yeah pretty much we're we're doing our job there like the changelog of jose's products you know yes yes it really is so yeah so uh there is livebook which is also running right and now Tidewave and what about Elixir man is it done or are you done with it or still working on it still working on it I still it changes so around now that there are a good amount of Tidewave things happening it's fresh I think we are about five weeks since we launched so we just launched it and you know when you launch something like that there's a lot of work feedback and prioritizing.
1:12:46I think it's kind of like about half, half of my time on Elixir, half on TideWave. But otherwise, like most of my works is still going to the Elixir type system and Elixir work. The other thing that I want to do, one of the things that going outside, TideWave is an example, but going outside of TideWave, one of the things that makes me excited about AI is because we can, look at the tools and find ways to improve and build new developer tools. And I've been like exploring some ideas around those areas. Like, well, so I was saying like tests, the tests that the coding agents write, I usually don't like them.
1:13:32They are redundant. And I think a lot of people don't pay attention to like, or they use too many mocks. A lot of people don't pay attention to code quality in tests, right? Test is test, right? So I'm trying to figure out ways of improving that. So for example, when the agent's writing tests, can we measure coverage and guide the agent to write tests based on coverage, but also give information about, oh, those tests, they are redundant. They are pretty much checking the same lines of code. You can try and define them. And the cool thing is that we are thinking about those things because we want to automate the agent, but a lot of it translates to better developer tools, right?
1:14:13We released this for the agent, but developers can also use it. So I think a lot of the work that we are doing right now will go back, will feed back into like better tools. Even when I'm working on TideWave, a good amount of the work will eventually feed back into better tools for like Elixir and the community too. Right on, man. Well, keep fighting the good fight. Always love talking to you. I always love hearing what you are working on I am going to give Tidewave a ride in earnest I have a Rails app now you know we have an Elixir Phoenix app so I can use it in both contexts and let you know what I think give you some feedback let me know and then right now it's either bring your own key or you can use your GitHub your GitHub co-pilot integration and then uh hopefully in about a month what are the tools that you use today so i use i use cloud code i have chat gpt pro but i don't actually use codex i'm not sure if i get codex with pro um i have gemini cli okay i don't know it's very confusing like when you get cloud code do you also get tokens for the api i don't think so right so you buy those separately i'd rather not buy more of those i'd rather is this why people want to bring their cloud code subscription because exactly yeah i would love to do that and not have another toll bridge toll road as adam calls them toll booth toll booths so that's my current setup adam what are you and you got some amp subscription maybe about five bucks left in amp it's uh i still can't hold right it's always expensive for me it's really great though it's it's so cool how it works it's really uh one of the best but i haven't found a way to hold it in a way that isn't expensive and uh okay so cloud code primarily uh same i have an anthropic key but only because of i think one thing had to have it and i think i got like the the trial balance they give you i'm still on that So there's nothing past that.
1:16:25But Cloud Code, Augmented Code I like as well. They're cool. AMP, I still like AMP. It's just I haven't found a way to make it that expensive for me. I just, I don't know. But it is really, really, really good when it does this thing. So is there anything in particular that you like about it? It seems to be just, it's got this Oracle. So speaking to AMP Code, it's got an Oracle where it can go back and consult. It's kind of like UltraThink now that I think about it, Jared. It's not quite that, but it's a bit more where it'll go into a deeper understanding of coding patterns and like a learned behavior across, you know, let's just say like a Rust CLI ecosystem.
1:17:03Like how do those work generally? What are good patterns? And it will come back and tell you stuff like that. So I find that it's research and its ability to execute in a, you know, hands-free YOLO environment is just, it's really good at that. Like you wind it up on the right thing with the right research, the right context, the right everything. It just plows through it for hours and just does amazing work. So, but it gets expensive if you don't like work with it and babysit it. Do you prompt for the, for the Oracle or it automatically figures out that it's like the plan mode and like the code or the, or it automatically figures out, Oh, now's the time for like some Oracle.
1:17:48yeah you know i think it does it on its own desire but you can also say hey in this exercise go ahead and prompt the oracle as well tap them get them involved i don't know it feels cool it does it and uh good results come i guess but you can either prompt it yourself or it just kind of does it when it needs to i am not an amp code expert by any means but that's how i experienced Yeah. So wrapping it up, yeah, right now is bring your own API key and you cannot, yeah. And Cloud Code, like your OpenAI subscription or Cloud subscription does not give an API token. And we cannot actually, because, you know, Cloud Code is just using a Cloud API, right?
1:18:37But we cannot use that API. It's actually not according, it's not legal according to Claude to entropic terms. So we decided to not do that. That's why we're working on the whole Claude code, codex integration kind of things. And so either bring your own key, but I really, at the end here, I really would recommend giving the GitHub Copilot a try because it's confusing because Microsoft calls everything Copilot. Right. But you can, there is a GitHub co-pilot plan that gives you access to kind of like a bunch of different models. And it's actually like, it's a predictable plan in the sense that, because the thing with paying for tokens is that it's very hard for you to predict how much it's going to be.
1:19:26and paying and GitHub, the GitHub Compilite subscription is per messages, which at least improves the visibility a little bit. And it has like a basic plan, quite affordable. So that's a good way to try it out for now to get some feedback. And yeah, we are hopefully launching Cloud Code. We are, Zed released something called ACP. I don't know if you saw those news. like with the Asian client protocol. So you can talk to Codex, Cloud Code, Gemini CLI, right? So we are building on top of that, but it's a bunch, it's work because you're running on the browser, right? And ACP is an IO protocol. So you can figure out all the hopes that we have to jump to make those things talk to each other.
1:20:18But yeah, hopefully we'll be launching that soon alongside Django, Next.js, and so on. I wish Anthropic would just give you, when you get some sort of subscription, they'll just give you a token or a key that you can use against that subscription at the same pace that you use CloudCode, right? Like, I guess they're just subsidizing that to death and then don't want to subsidize their API. Because I can pay 20 bucks a month or whatever it is and use the dog do out of CloudCode, but I got to pay 200 bucks or 500 bucks equivalent to use this API the same amount. I just made those numbers up, but you can see like the discrepancy is there.
1:21:01It doesn't make sense to me. I guess they just want you using their, their CLI a lot. It's, it's not even, I would say it's not even about the, the, how much the cost is just the previous, like the, how predictable it is. Right. Because. Yeah. because you don't want, you don't want to get dinged for making the bad prompt, you know, like you want to set, I'm fine with 20 bucks, 40 bucks, 50 bucks a month, but just because I use it a lot, don't give me a, I tell it to ultra think and it's like, well, that's$17 for that ultra think. And I was still wrong. But I'm still wrong. Like, can I get my money back too?
1:21:39Are there returns? I know, right? Service degradation is a real problem for me. Yeah. But here's the thing. I actually think pushing people towards using cloud code more or codecs and building on top of those tools like with acp is not actually a bad idea because because here's okay let me tell you a story quick story i know you're going um so when we first implement a tide wave we focus on entropic and the cloud models and like if you go to cloud code prompt it has things like you should be concise, you know, don't use too many words, use four words. I think it even said at some point, one answer word is best.
1:22:25And of course, like it doesn't, it doesn't listen to that. Right. It's like, it finishes the feature. It just like dumps like four pages of text about the thing that it implemented that nobody ever reads. Right. So when we did our prompt, we tested with those things as well. It does improve a little bit. Right. But it also say things like don't write a code comment. And it always writes a code comment. So anyway, so we wrote the prompt. And then when GPT-5 came out, we decided to give it a try and started supporting OpenAI. And it was very curious because it would say like, hey, implement this feature.
1:22:59It would do all those things. And then at the end, done. And then you would ask something and it would say, good. And then we realized that the prompt we had for Entropy that was saying, be concise. Otherwise, GPT-5 was actually listening to that prompt and it was being concise. That's why I was just saying, done, good. It was not doing any fluff or anything. And that's when you realize that you actually have to come with a... If you're building a coding agent like I am, you have to actually build a prompt per model. Right? And now GPT-5 codecs came up with its own prompt that it's different from the GPT-5.
1:23:40And then not only, so just doing that, like fine tuning the prompt per model, that's a pain. That's like, that's boring work. That's not something I want to do. Right. And then it gets the other thing, which is then they have the tools. And at this point, those coding models, they are becoming so important for those companies that they're actually fine tuning, like how we should send edits to a file. They are fine-tuned the models for that. So when the GPT-5 codex model came out, they also said like, look, this model is best at sending these kind of diffs and edits over the wire. So now I have to implement the specific editing tools per model that I support.
1:24:29And then each of those models come with like their own context engineering techniques. So at that point, like it's, you know, If you're like me, you're building a coding agent, you want to be able to get that infrastructure and build on top, right? And then it comes with the nice thing. So like going back to the hot take, if you're building like your agentic tooling for coding, instead of doing the MCP, don't do the MCP. Build on top of ACP and have control of the agent and use all those things and extend that instead. And with the announcement with Cloud Sonnet 4.5 today, yesterday, they actually recognized that.
1:25:18They renamed the Cloud Code SDK to Cloud Agents SDK. They moved a couple of things around for it to be a better SDK for people to build on top. Because I think that there is a lot to gain for leveraging everything. They are tightening the model of those tools. and we want to be able to leverage that. That makes sense. They're putting a lot of work in to take that model and make it an agent. And there's no reason why everybody else needs to do that work as well. What could you, how would you go about building on that right now? Like, where would you go? What's the starting point to building that right now?
1:25:56I don't know. What's the website? Oh, I don't know. Zen.dev slash ACP or something? No, I would just search for agent client protocol on Google and see where that. Agentclientprotocol.com. all right all together right no all spell it out yep all right of course when you google that it will be your first hit and i believe there is an sdk typescript i don't know other languages right now but that's the only language that you know
1:26:32you heard it now there's a hot take you heard it here first jose only knows typescript that's when you know, it's getting late. I just auto-completed you. So, yeah, but I can see it becoming like more and more important and we are going to see like more SDKs. But the protocol is also relatively straightforward as well. So, yeah, I'm hoping that I can really see a lot of value in there and I hope it's going to catch up to the point where, because a lot of the, I think Gemini CLI supported built-in in the CLI. But when Zed released support for Cloud Code, for example, it's because they have a wrapper.
1:27:17So I hope it grows to the point where more like the CLIs, they are coming with built-in support for it, right? And then I hope it grows to the point that, like is it who which of the big providers have their CLI version right now is Gemini it's OpenAI and and it's Entropic right they do like Grox they don't have theirs I think ZAI they they don't have theirs so I actually hope like those other companies they start providing those CLIs as well with with all those things we have been talking about in the sense It's like, look, here are the optimized diffs. Those are the things we improved for. So we can move to the point where we are all building on top and not particularly inventing that wheel.
1:28:10So I really hope it grows. More CLIs. Give them to me. There you go. Thanks for hanging with us, Jose. It's always a pleasure, man. Pleasure. Yeah. All right. Bye, friends. Ooh, synchronized. All right. that is your changelog for this week. We hope you enjoyed Monday's news episode about exiting the feed for Cell vs. Cloudflare and why overengineering happens. We hope you enjoyed Wednesday's interview with Evan Yu and we hope you enjoyed this episode with Jose. Because after all, we're here for your enjoyment. We also want you to learn and to keep up the easy way and we want you to level up your own work and to feel connected to this worldwide community of hackers.
1:28:52But we'd love for you to do all those things while enjoying the process. because after all, the process, that's our life, isn't it? If you do enjoy our work, please tell a friend or three or send us an email, editors at changelog.com. We absolutely love hearing from you all. Thanks once again to our partners at Fly.io and to our sponsors of this episode, depot.dev, fabi.ai, and miro.com. Thanks also to the one, the only, the mysterious Breakmaster Cylinder. Have yourself a great weekend. Let someone else praise you and not your own mouth. And let's talk again real soon.
From the publisher
Elixir creator, José Valim, is throwing his hat into the coding agent ring with Tidewave –a coding agent for full-stack web development. Tidewave runs in the browser alongside your app, but it's also deeply integrated into Rails and Phoenix. On this episode, José tells us all about it. Also: his agent flow, YOLO mode, an MCP hot take, and more.

