In short
Emmy-winning nonfiction filmmakers use AI to automate documentary post-production toil—turning massive, messy media archives into searchable, fact-checkable databases.
Guest backgrounds
Tim McLear is a producer at Ken Burns Florentine Films, responsible for technology and processes that support documentary production and research.
Key claims
The biggest AI win for documentaries is not generative video, but automating media management (metadata extraction, transcription, and semantic search). Guardrails (embedded metadata + web scraping) reduce hallucinations. Automating “logging” frees researchers to do more research and improves asset quality and discoverability.
Notable examples
Shooting ratio for a Muhammad Ali series: 20,000 stills, 100+ hours of footage, 35 interviews. “Autolog” pipeline: 5-step REST API that parses metadata, scrapes sources, generates descriptions for images/video/audio (Whisper + frame sampling), then creates CLIP + text embeddings for semantic search and reverse image search. “Flip Flop” iOS app: captures archive photo fronts/backs, transcribes back text, embeds it into EXIF, and renames files for clean imports. “OCR Party” macOS app: crops specific document regions for OCR/translation with editor location tracking.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOTim McLear's Role in Filmmaking
0:52 to 1:32
Discussion on Tim's contributions to technology in filmmaking.
“I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools.”
Tim McLear's Role in Filmmaking
1:36 to 2:23
Discussion on Tim's contributions to technology in filmmaking.
“Brex is bringing that same power to finance.”
Challenges in Post-Production
2:23 to 2:52
Tim discusses the complexities of post-production in documentaries.
“What I love about what we're going to talk about today is you work in a very interesting and creative industry putting out amazing content.”
The Shooting Ratio in Documentaries
2:52 to 4:08
Exploration of shooting ratios and their implications in documentary filmmaking.
“how did you think about what problems there were to solve in AI relative to your job and the people that you work with and why did you start where you started?”
Using AI for Data Management
4:08 to 5:22
Tim explains how he leverages AI for effective data management in filmmaking.
“Because that will maybe give us a sense of how much of this you have to grapple with to get a good piece of content on the end.”
Demonstration of AI in Action
5:22 to 7:20
Live demonstration of AI tools for enhancing image metadata extraction.
“So I'm going to start by kind of just showing you the like end result before I go right to like how I got here.”
Improving AI's Accuracy with Metadata
7:20 to 8:30
Discussion on enhancing AI outputs by incorporating metadata for accuracy.
“Write me a script that submits the JPEG at the root of this workspace to OpenAI for description.”
Building a REST API for Image Management
8:30 to 11:28
Tim describes the technical architecture of a REST API for managing images and metadata.
“It's running, submitting this image to OpenAI for analysis.”
Automating Metadata Extraction for Video and Images
14:00 to 23:30
Learn about a five-step process for gathering and describing image and video metadata using AI.
“If I pop into the jobs folder here for a second, we could zero in on basically what we were just doing, but the current iteration of it.”
Automating Metadata Extraction for Video and Images
23:34 to 24:21
Learn about a five-step process for gathering and describing image and video metadata using AI.
“Brex is bringing that same power to finance.”
Show all 20 chapters
Field Research and Archival App Development
24:21 to 28:00
Explore the development of an app for capturing and organizing archival images in the field.
“Okay, so this is more of your archival and footage data, but you capture a lot of stuff in the field where people are not sitting in front of cursor or their desktop.”
Automating Image Metadata with AI
28:00 to 30:20
Learn how AI can automate the tedious process of embedding metadata into images.
“But for now, we're just going to create a collection, tap into that collection and capture.”
AI-Powered Document Processing
30:20 to 32:20
Discover how AI enhances document transcription and translation workflows.
“But I think Flipflop is certainly making the process easier since they've gotten back.”
Building Custom AI Applications
32:20 to 35:50
Explore the creation of specific AI applications to streamline workflows.
“So you can imagine in our films, we work with a lot of documents.”
Learning and Adapting to New Technologies
35:50 to 37:50
Understand the importance of embracing new technologies for creative work.
“I mean, I think I could sell it to like two colleagues.”
AI in the Film Industry: Opportunities and Concerns
37:50 to 42:00
Examine the practical applications and concerns of AI in filmmaking.
“Okay, well, we're going to do a couple of lightning round questions.”
Embracing AI Tools in Filmmaking
42:00 to 43:09
Learn how filmmakers can benefit from utilizing AI in their creative processes.
“But my approach has certainly just been like jump in and learn the tools.”
Using AI for Production Efficiency
43:10 to 44:02
Discover how AI can streamline production without compromising authenticity.
“And I think that stands for people in your industry.”
Effective AI Interaction Techniques
44:03 to 45:51
Understand how to communicate effectively with AI for better outcomes.
“Well, last question have to ask you when, you know, you're on your dog walk with chat GPT doing voice mode and it's not listening to you or not giving you what you want.”
Connecting with Tim MacLeod
45:52 to 46:24
Find out more about Tim MacLeod's work and where to connect with him.
“Tim, where can we find you and how can we be helpful?”
Transcript
Automatic transcript. May contain errors.0:00Tim McAleer:How did you think about what problems there were to solve in AI relative to your job and the people that you work with? And why did you start where you started? Post-production is like a technical mess of media management. You have many different file types. You have images, you have archival footage that you're gathering, live footage that you may have filmed out in the field, interviews, transcripts. So it ends up being hundreds of hours of footage, tens of thousands of photos. The data management piece when you're dealing with all that different stuff is the mess that I have used AI to tackle.
0:31My goal was to automate this. For years, this has been manual data entry.
0:36Tim McAleer:Automate away toil. That's what you want to do. No one was going to make me this app. And so the ability to make an extremely specific app that makes a workflow on my team and my company easier, it's been an unbelievable moment.
0:52Tim McAleer:Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today, we have Tim McLear, a producer at Ken Burns Florentine Films, who's responsible for the technology and processes that bring these amazing films to life. Instead of focusing on how AI can create creative for these films, We're actually going to talk about how Tim uses AI to build software products that make his post-production and research team's lives a lot better. If you're working with images, video, sound, or just a lot of data, this episode is a great one for you.
1:31Tim McAleer:Let's get to it. This episode is brought to you by Brex. If you're listening to this show, you already know AI is changing how we work in real, practical ways. Brex is bringing that same power to finance. Brex is the intelligent finance platform built for founders. With autonomous agents running in the background, your finance stack basically runs itself. Cards are issues, expenses are filed, and fraud is stopped in real time without you having to think about it. Add Brex's banking solution with a high-yield treasury account, and you've got a system that helps you spend smarter, move faster, and scale with confidence.
2:12Tim McAleer:One in three startups in the U.S. already runs on Brex. You can too at brex.com slash howiai. Tim, welcome to How I AI. I'm excited to have you here. Thank you for having me. What I love about what we're going to talk about today is you work in a very interesting and creative industry putting out amazing content. And we're going to talk a little bit about how AI is impacting the creation side of things. But you've actually used AI to smooth out some of the challenges you've had on the production and post-production side of things. So I'm curious, how did you think about what problems there were to solve in AI relative to your job and the people that you work with and why did you start where you started?
3:01Yeah, I think most of the flashiest use cases of AI in creation or media and entertainment right now are often in like generating full video content or images or whatever it is. But post-production specifically is like a technical mess of media management, especially in nonfiction. You have like many different file types, right? And you have images, you have archival footage that you're gathering, live footage that you may have filmed out in the field, interviews, transcripts. And so like the data management piece when you're dealing with all that different stuff is the mess that I have used AI to tackle.
3:40And I think that the sort of like AI as a tool versus AI for generation is even more immediately applicable in our field at the moment.
3:49Tim McAleer:Well, and I have a very, you know, very simple, humble little podcast, but even for us, we create a lot of research and longer content and we're editing it down. I'm just curious with documentaries and nonfiction work, what do you think the ratio is of media captured, researched and archived to actually publish? Because that will maybe give us a sense of how much of this you have to grapple with to get a good piece of content on the end. We have a thing in our industry called a shooting ratio. And so you can imagine in like a fiction series or you know like a sitcom on air i don't quite know what those shooting ratios would be but you're working with a script and so you're going to have a slightly lower ratio in documentary it can get quite high like i can tell you that we made a series about muhammad ali a few years ago it was an eight hour show we gathered 20 000 still images in the database of just stills i think it was over 100 hours of footage because he had a lot of fights and that kind of thing news news footage.
4:48And then we also filmed, I want to say like 35 interviews for the piece. So it ends up being like hundreds of hours of footage, tens of thousands of photos. And that's just like, that's one example of, you know, a particularly famous individual, but that tends to be what it looks like for our shows.
5:04Tim McAleer:So that's what you have to manage, make searchable, make usable by the entire production team. And you got inspired by ChatGPT and some of these early AI tools to do some of that. So you want to hop in and show us what, you know, the first use case is? Absolutely. So I'm going to start by kind of just showing you the like end result before I go right to like how I got here. So on any film that we work on, we end up having some kind of database, right? So this is a database where you can see the still images we've gathered. You can see there's a footage section, a music section, anything that might go into the film and all the kind of stuff you might expect to see, right?
5:44Descriptions, tags, a date on the thing where we got it from. Some more technical detail is also going to appear over here. In any event, my goal was to automate this. For years, this has been manual data entry. And so I remember vividly, I'm going to jump into cursor now, but I do remember like when I first started doing this, it was ChatGPT. I remember ChatGPT added image upload and it was this insane day for us. I was like in the office with my colleague Clark and we were just like throwing images at it and seeing kind of the quality of the output. Like it was this, an aha moment where it was like, oh my God, this thing can see.
6:22And how could we harness this text generation, right? To, to use it for our database entry. So I'm going to simulate that, like the starting point, and then we'll jump to where we're at today. But essentially what it looked like at the beginning was we would throw something into GPT and we would say like, Hey, can you describe this? and it would hallucinate a little bit but it was so tempting to figure out a way to harness that that I started essentially like writing little python scripts with chat gpt and at that time it was like vs code on one monitor and gpt on another and I'm gonna all right I'm just gonna go ahead and demo what that kind of looked like I'm gonna speak my prompts if that's okay I use this tool called Super Whisper because it kind of cleans up my off-the-cuff dictation.
7:11So I have an image here of a nice street in somewhere America, maybe mid-20th century. We're going to see what kind of description we get from AI. All right. Write me a script that submits the JPEG at the root of this workspace to OpenAI for description. I want just a general visual description of what we can see in the image. Any API credentials you need are in a text file at the root of the folder. And what we can see here is that like everything I just said got funneled through this app called Super Whisper. So it got funneled through a prompt that itself is cleaning up my like messy vibe coding.
7:54I think it's clean enough. So we're going to go ahead and submit it.
7:56Tim McAleer:And I see you're using Claude 4.5 Sonnet. Is that by choice or by default? or that is because I'm on a podcast right now, to be honest, like, I think this is a very easy task for AI. I could keep it on auto for this, right? I will say I switch between various cloud models depending upon the like difficulty. And I do try and be cheap and stay on auto if I know that I'm asking for easy stuff, you know? Okay, so you're just you're you're giving us a little bit of quality control here. Yeah, I don't want it to mess up. We're live on air, you know? Yeah. All right. So it's telling me that I need to install some requirements.
8:33My guess is I have those requirements. It's got a submit image script. Let's see what it did. Here we go. It's running, submitting this image to OpenAI for analysis. What kind of description will we get? There we go. This image depicts a small rural main street from what appears to be the mid-20th century. We had guessed that. There are a series of wooden storefronts, each with signs indicating there are local businesses. Okay, so this is great. And this is kind of what we were getting in those early days of GPT image upload. But the problem here is like, you're making a film, you want to know what rural Main Street, what town are we in?
9:13What is the exact year? And you can't really just go with this kind of generic description. So a lot of times we happen to know that images come with embedded metadata. And you know, if you're using your iPhone camera today, you know that maybe there's some metadata like GPS data, that kind of stuff. But archival images will often come with whatever notes people have scribbled onto them over time. And so I'm going to now I'm going to iterate on this one time and say, I want you to add a step to this script. I want to scrape any available metadata from the file first and append that to the prompt.
9:48The goal here is that we are using any available metadata as like a source of truth for what this image actually is and not just guessing.
9:57Tim McAleer:And so just repeating that while this is running, what you're saying is, yeah, for this particular use case, you're working with a set of archival photos from sources that have embedded probably additional layers of metadata into it that you can read that give more information, which is different than, you know, scanning something or taking something off your off your phone, which I think we're going to look at a bit later. And so you're trying to harness the structured metadata off this file, which if you go back to the tab that shows the image, we can't see with our human eyes, but our agent friends can read with its robot brain.
10:37Tim McAleer:And you're using that information to then upgrade this script that is going to do all this AI analysis for you. That's exactly right. And so in this case, it's going to be embedded metadata. I happen to know this is an image from Library of Congress. There's going to be some metadata on it, but it could also be something on the web. Like where this eventually goes to is like, okay, I know that there's a website with information, may not be in the file, but hey, how about you go and scrape the web, gather anything you can know about this. Because ultimately, like this is a journalistic endeavor.
11:10These shows get fact-checked. We want everything going into our database to be, you know, true and verifiable information. All right, so let's see how it did when it added that metadata check. So let's see, we can see in the console, it did a little bit of a scrape. It looks messy as hell. But somewhere in here, we can see stuff like, yeah, archival information. And it's now going to use that. And what we've generally found is that when you add those guardrails, when you give it information you know to be true about the image, it relies on that so much more than just what it can see. Like, you know, AI really wants to perform for us.
11:49It really wants to do a good job. And so when you give it the tools and information to kind of write a better description, it's going to be able to get there.
11:57Tim McAleer:And I want to call out something. So we talked about using the anthropic clawed models in particular for the actual coding of the script, but you're relying on the open aim AI models for the image analysis. Why open AI versus any other models that like stick with the one that you love, or it was the first one that did a good job for you, or do you feel like it's particularly good at image analysis? I'm curious why you select those different models for different use cases. Yeah. It's mostly that it's the first one. Like they were the first one who had a, they had a vision preview on their API. They did it before Claude.
12:30and like I had built up enough of an infrastructure using that API call that it was like the switching costs were too much, you know?
12:38Tim McAleer:Yep. All right, so let's see what we got this time. It's much more detailed. It is, it's much more detailed. So the image shows a street scene on the main street of Cascade, Idaho. There we go, we know where it is now. Captured in 1941 by photographer Russell Lee. We've got photo credits. All right, so this is a great example of like, you add the guardrails and you're going to get more detail, but you're also just going to get facts, right? Before, I don't know if it's still up here somewhere yet. Before it was a small rural main street. Now it is the main street of Cascade, Idaho. And so we can imagine this getting duplicated in various ways, right?
13:15This image has embedded metadata. Maybe it's a website that we're going and gathering it from. But effectively, like this is where it all started. It started with a single Python script that I was running on my computer. And I was like, this is awesome. My database software is like, it's advanced enough. to call external scripts. You can kind of use any database to do this, you know, Airtable, whatever. But you just need something that has an API and that can call an external script or webhook or something. So this is where we started. And now I'm going to switch my screen share to a remote machine, like a little Mac mini that I have running in my office.
13:50And what this, you know, it's hard to, at this moment, it's a more complex cursor workspace you can see. Maybe I'll bop into the rules. Basically, what this is, is a REST API so that every image file, video file, music file, anything that ends up in that database that we looked at at the beginning, pings off of this REST API for all kinds of different like metadata tasks. If I pop into the jobs folder here for a second, we could zero in on basically what we were just doing, but the current iteration of it. So I call it autolog because the process of writing this in for years, the manual data entry is called logging.
14:32So it's not the cleverest name, but it fits. And you got a five-step process here. Basically, first, we're going to gather the info, meaning file specs, how big the image is, is it a JPEG, is it a TIFF? We're going to copy the file to our server. We're going to name it our ID number. We're going to parse it for metadata. Is there any metadata? If there is, great. But either way, we're going to look for more information on the web in this step four here, scrape URL. And then once we know everything we could possibly know about that image, we're going to generate a description for it. And when you imagine how this might work for video, well, like video is itself, it's just 24 images in a second, plus some audio.
15:11And so basically this just gets scaled up to deal with video files too.
15:16Tim McAleer:Are you using the same model for video files? Are you taking them, extracting the stills and putting them through OpenAI? Are you using a different model? I use a different model for, so I have to the, the video files requires like two levels. Most video like AI models out there seem to do basically some version of frame sampling. so it could be extremely expensive if you were sending all 24 images every second to an API right so I pull at five second intervals because I'm cheap some others maybe pull in a more in a smarter way maybe at like lighting changes or something like that like there's different ways of thinking about the frame sampling so for the frame captions themselves I will use a cheap model I'll use like a nano gpt5 nano but then for the and I can go in and show you a prompt here which maybe illustrates this.
16:06I have frame prompts, which basically ask for just like a prompt of an individual still image extracted from video. But then I have a larger parent prompt. You can see that my prompts have gotten slightly more sophisticated over time. Basically, what this does is it sends every single frame that we've extracted from a video file, it extends anything like any of the audio we've transcribed from that video file, it packages it up into this elaborate prompt and it sends it to a reasoning model. And the purpose of that is to say like, these are all the video events that we have observed in this video.
16:46Here is like a massive text file of data. Tell me what you think is happening in the video. Got it. Yeah.
16:53Tim McAleer:Yeah. I, you know, maybe, maybe tip from one of our how other how i ai guests but i found that the gemini um the gemini models are quite good with video it's actually what we use to do our podcast raw recording to both highlight stills and a blog post that i put out i process them through the the gemini models and have had a lot of success and it just pulls out like the stills that might be just it automatically pulls interesting stills. It actually gives me interesting stills plus five seconds or like plus five seconds plus minus five or minus five seconds because sometimes the guest and I are looking ridiculous.
17:32Yeah, yeah, yeah, of course.
17:34Tim McAleer:So tip to anybody out there with video who hasn't tried the Gemini models, I find those particularly good for this use case. You might have just, you know, added something to our little roadmap here. Well, and so and then I'm curious about the audio side of things. So I kind of, you know, I play with the Gemini models for video. The still makes tons of sense to me. Tell us a little bit about the audio side of things. So the audio is also I now I feel like I'm an open AI shill. Everything I'm using is open AI. And I think except for the coding, which is interesting, but I think it's just habit.
18:08I use Whisper for audio. So like Whisper is an incredible open source model for speech to text detection even the like medium size model does a pretty good job and what i do is and i can pop back into the database software maybe to like illustrate this what i do is i extract you can
18:26Tim McAleer:see like frames pulled every five seconds and there's a caption associated with each frame and then there's this is a shot of an alligator in a swamp so he doesn't have any audio he wasn't talking but i basically pull audio at five second increments so that when we send those like video events up to the reasoning model, we are sending a full transcript, but we're sending it like kind of like pegged to the moment in the video that it happened, if that makes sense. Yep. So the transcription is all happening, you know, on my back end over here. Everything like I think I could probably open up the console and see like, there we go, like someone just sent a job through not that long ago.
19:04Like I can kind of come in here and see what my colleagues are doing as they ping my API all day long.
19:10Tim McAleer:Great. And so you're pairing a snapshot image every five seconds from a video, the five second transcript of the audio speech to text via whisper, metadata, if you have it, parsing that all together, and then getting a very robust description and analysis of the content that you have available in back in this tool that you're using to archive log manage all your assets. Yeah. And like I said, that tool could be kind of agnostic. Like you could do it in a Google Sheet if that's, you know, if that's what you like. But I like this. We've been using it for a while. Everything we just talked about is how we kind of get to like metadata that we can read, right?
Read the full transcript
19:52Like generative metadata that is a, we know it's accurate because it's kind of been put on these guardrails by our metadata extraction steps. And then also it provides this like nice visual for us. We can see what this thing is at a glance. but the next step of this now that you have this like api running in the background is you can generate something that maybe i can't read but the ai can read pretty well which is vector embeddings so i'll jump back to stills for this because i think it's a maybe an easier illustration of it every asset in our database gets put through two modes of embedding so we'll send the thumbnail through and run it against an open source model.
20:34I use clip for this. And I'll generate an embedding off of that. And then we'll send the description through. I use, again, an open AI text model for this and get an embedding for that. And then we'll fuse them. And the purposes of that is that so now we have like the ability to discover things semantically, like prior to this, and I think in a lot of film production today, you're working with exact text search, you know, like if that description says dog, but you know, somebody wrote in puppy, you're not finding that image. And so this has been like, kind of the most exciting part of it, not necessarily where I knew it was going when it started.
21:11Like I was just excited to generate a description, right. But now the ability to discover semantically is I think, you know, the most the most robust part of the system.
21:21Tim McAleer:So what I love about this, I mean, a couple things is one, you've really pushed every step of the way you know you could have stopped at like we got good descriptions or we got like the structure metadata out and now i have a script that runs it you could have stopped at images only but you took it to video and video and audio you could have stopped at structured data only but you went to embeddings to get semantic search so i love just the breadth of applicability of the ai in this process but what i probably love more is i doubt this was anybody's favorite part of their job. Like, I doubt it was anybody's favorite part of their job to be like, I'm going to go read some Library of Congress metadata.
22:01It used to be my job. So I can tell you firsthand, not my favorite part. And it's also like, I think the best argument I have for all the work I've done creating this system is that like the same people who used to write this data were the ones who are responsible for doing the research. So you've now freed them up to just look more, right? Like maybe now we could gather 25 ,000 still images for the Muhammad Ali project because you have that much more time. You're not just like copy and pasting stuff off a website to put it in this form, you know?
22:30Tim McAleer:Well, and you probably get to select from this big archive of data, better assets to use in your content because they're more discoverable because you have more confidence in the source and the content of that data. So I bet it up levels at the end of the day, the quality at at the end because you have just much more data to work off of. 100%. I mean, like a real quick example of that too is like, I'm going to use Abe Lincoln here, which is maybe not the best use of this image. But embeddings enable us to find things in ways we never would have thought to find them before. So like I have a button down here where when I click it, what it basically is going to do is reverse image search within our own collection.
23:10So if I'm an editor and I like an image, and this is going to take a while because I'm not on site. But if I like an image, I can click the find similar button and it's just going to go and find every image that kind of has that vibe. You can see here we have a duplicate of this one. But then there you go. It recognized the man and it started pulling in other portraits.
23:30Tim McAleer:This episode is brought to you by Brex. If you're listening to the show, you already know AI is changing how we work in real, practical ways. Brex is bringing that same power to finance. brex is the intelligent finance platform built for founders with autonomous agents running in the background your finance stack basically runs itself cards are issues expenses are filed and fraud is stopped in real time without you having to think about it add brex's banking solution with a high yield treasury account and you've got a system that helps you spend smarter move faster and scale with confidence. One in three startups in the US already runs on Brex.
24:15Tim McAleer:You can too at brex.com slash how I AI. I love this. Okay, so this is more of your archival and footage data, but you capture a lot of stuff in the field where people are not sitting in front of cursor or their desktop. I'm looking through these assets and I know that you use some vibe coding and a creative approach to get more information about those assets. Could you walk us through that? Yeah. So the next use case is an app that I developed for archival research in the field. So I think that we really pride ourselves on like turning over every rock on not just relying on what's digitized and available online and going and visiting physical archives.
25:01And so the process of visiting a physical archive is basically you have a bunch of folders that you pull ahead of time you arrive there and your goal is just to snap like low res resolution iphone snaps of everything you can possibly get and so you're snapping the front of the image and you're snapping the back of the image because the back is typically where there's going to be like a scrolled description or maybe like an accession number an id number that the archive is out of themselves and so this process used to look like you show up at the archive you take iphone snaps for two days you get back to the office you have the messiest camera roll you've ever had you cannot actually pair your fronts to your backs because it just got out somehow it got out of order along the way and so the goal was basically to make that process like a little better so i i vibe coded this ios app to deal with this problem.
25:56And I, I tend to just like speak in screens, like the way I maybe it's because I'm a visual person, like the way I deal with it is I just think like, okay, I see a screen that does this and a screen that does this, I imagine a button that does this. And the purpose of this was basically like, I want people to be able to create collections for each folder they're capturing, I want them to be able to snap a front and a back, like the flip side of the image, so that they can easily associate those, so the file names associate them. And I want to immediately transcribe any information on the back and embed it into the original image.
26:29So now I have this app called Flip Flop. I ask ChatGPT at the end of my dog walk to generate some kind of specs doc or requirement doc. It pretty much does it in one go. If you chat with it for 30 minutes, you know, you can get a lot done. And then I fed this PRD to Claude Code. and it this one it like it it didn't build it in one shot but it certainly built the ui in one shot and so i guess maybe we should just jump into like the actual app yeah let's do it so flip flop which is my cute little name for it is uh basically designed to capture those fronts and backs that i was talking about so you have three screens here you've got a collection screen where you're going to create your folders you've got a capture screen where you're going to take your images.
27:14And I'll just quickly highlight this part, which is where you kind of have your AI processing options. So I allow people to define a separate prompt for what I call the flip side of the image, the front and the flop side of the image, the back. And so in this example, I'm going to show you some photos of my dog now. And the flop side of the image is going to have some text on it. So our prompts here are really just designed to get a decent caption from the image and to transcribe any text that we see on the back end.
27:43Tim McAleer:So let's create a new collection. We're going to call it how I, AI, that's good enough. There's also an option here to add more context. You know, the AI loves context. And so maybe if you're, you know, you could imagine if you're digitizing an entire collection of, you know, someone's personal letters or someone's portrait photographs, you would add that kind of thing here. But for now, we're just going to create a collection, tap into that collection and capture. So here we go. It's a screen share within a screen share. We're going to not care about the glare too much. I'm going to capture the front side of this image of my dog Tony's third birthday.
28:24I now have the option to add notes if that's what I want to do, or I could just add a flop side of the image right here. And when I complete that, it will have, because it's lightning fast already, sent it up to OpenAI for a description and embedded it. And this is the really crucial thing because you just saw the first system I had, embedded it in the image metadata itself. So the flop details have the transcription, Tony's third birthday, and all of that will show up in the, what we call EXIF metadata, which is just the image metadata standard.
28:59Tim McAleer:Got it. And just for people that may be passed by, instead of simply generating kind of the text description and storing that in a database, relative to the original image you took, you actually now have this structured metadata on the image file itself, which again, like what a pain to do. Oh, a giant pain, yeah. A pain to do manually. And so now anytime anybody uses one of these images, even if they don't have access to this app, even now that image is embedded with that metadata. 100%. So you could pull this onto any computer or any app, anything that can read underlying metadata, and it's going to be able to see that this was Tony's third birthday.
29:41And so that's structured metadata in the sense that we've now structured the actual information about the image. But the other thing that's really crucial, honestly, is that we've structured the files themselves, right? So you can see they're getting named in a particular way. And so we've moved from like camera roll mess to like files that are going to sort in your computer that you're going to be able to import cleanly. You're going to be able to distinguish easily what's the front of the image, what's the back of the image. And that has, I think, been the other unlock. Like I had two colleagues out in the field a couple weeks ago, and they came back with 1400 images.
30:16And I don't think that's only because they were able to use Flipflop to capture it. But I think Flipflop is certainly making the process easier since they've gotten back.
30:25Tim McAleer:The thing that I want to call out for folks, maybe a general takeaway here is these AI models are so good with files and code can do a lot of stuff with files. And a lot of the people we talk to, you know, Markdown is the file type du jour these days, which is, you know, just like a specially formatted text document. But if you start to look at other file types and really understand what can be put in a particular file type, you can actually discover some pretty interesting things you can do with a combination of AI and coding to make those files much more useful for your use case. So this is one of these takeaways where I'm like, I haven't thought about like what can be embedded in an image file or what can be embedded in a video file.
31:17Tim McAleer:And even just having, you know, ChatGPT or one of your general models say, hey, I'm working with an image. How can I load it up with as much context and specificity as possible? What's available to me. And then using that as a jumping off point for what you do is a pretty interesting use case of AI. I didn't even know, like, I'm very familiar with stills underlying metadata fields, but I didn't really know what was available in audio or what was available in video files. And I just sort of, I go into cursor and I ask, right? Like, now we have a music workflow, which we're not going to look at, but like where we embed artist album kind of like licensing data into any music we consider for a film.
31:56And I didn't know that there was in the metadata field we could just store that in. But of course there is, you know, somebody thought of this a long time ago.
32:03Tim McAleer:Yep. Amazing. Okay, we have one last use case, which, mom, if you're listening, I think you're gonna like this one. My mom's a genealogist. So I think she's gonna like this, this use case. But let's show it first. And then I'll call out mama where I think you can use it. Okay. All right. So you can imagine in our films, we work with a lot of documents. And we're not always interested in the entire document. Sometimes, like, we just want to transcribe maybe part of it. Maybe we want to translate and transcribe part of it. Like, take this newspaper document, for instance. Like, maybe the Arkansas State News is the article we're interested in.
32:43That's the transcript we want to be searchable. That's what our editor might want to consider for the film. We can't just, like, put this in Adobe Acrobat and OCR the whole thing. It's, like, it's not going to work. and even more than that like the quality of the image would not work with most OCR engines you know so AI is really good at OCR of old documents it's really good at handwriting it's pretty good at translation too so I built and we're not going to get into the building necessarily but this is this is one of the few like xcode builds I had to do so this is a swift build a little mac menu bar app.
33:18It's called OCR Party, which stems from the fact that we're just OCRing part of the image. You got to have fun with these things. And let's see, we're going to open up that newspaper in OCR Party. We're going to get like a little preview window. So let's say actually what we want is Coolidge seeks peace in the world. So let's zoom in a little bit. Let's open up our cropping tool. this little thing down here is basically a choice between mac os vision and uh an ai api call and the purpose of that is because sometimes people don't sometimes people don't trust ai you might have heard and so i i built that in as an option essentially i would i would think the ai option gets used more but nevertheless now you're going to select just the part of this article you care about or this paper that you care about and you can see there's like a crease in the paper there's a weird black mark here.
34:13But you can imagine we submit this for OCR. Now we have just that text that we pulled. We're also calling out for our editors, like where on the page, they're going to be able to find it if they want to sort of zoom in on it, crop to that particular article. And I can't exactly remember what text we were looking at, but it certainly completed those sentences where there was a black marker. Yeah. Right. So AI was able to kind of infer to the best of our ability what that sentence might have said. And, you know, if this ends up in a film, I could guarantee it would get fact checked later. But for the purposes of gathering documents, thousands of documents, this ability to kind of like precisely OCR is, is, has been a nice little unlock for us.
34:56Tim McAleer:One thing I also want to make sure people take away from this episode is we've seen basically three form factors of apps. So yes, they've all used AI, but you've been able to swap between sort of like a Python API service that gets called by another software application or database, a iOS app that, you know, you can run on your phone, and then like a little desktop toolbar widget. And what I like, what I love about this moment in AI with, with regards to software engineering, is like, if you have basic software engineering practices, and then you know, enough to be dangerous like yeah you can you can vibe code uh and you know a swift swift app to run on on your local desktop hyper specific app yeah no one was gonna make me this app and so the ability to make like an extremely specific app that makes a workflow you know on my team and my company easier it's been it's been an unbelievable moment yeah i would say the tam for this app is like you.
35:58Yeah. Yeah. I mean, I think I could sell it to like two colleagues.
36:01Tim McAleer:Well, and then my mom. So what I was going to tell you is my mom is a genealogist for the Daughters of the American Revolution, of which I am one. Fun fact on Claire. Oh, no way. And she does the lineage tracing. And do you know how many times she screenshot something and is like, can you read this cursive? Like, what in the world is this name? And it's like, you know, one name and a big, a big image. And so I do think AI is and I'm like, yeah, I'm gonna drop this in a chat GPT. And I'll tell you what I think it says. And I think its ability to read handwriting, old typefaces, kind of understand the nuances of spelling and things like that are just really, really interesting for these sort of research use cases.
36:44Yeah, we didn't look at a handwritten doc here. But that is definitely something happening at our company, like the ability to read letters that we could not read before and also just other languages. Right. And then we immediately have that text to you have letters written in some kind of cursive scrawl from the 17th century that is now translated to English and made legible for you.
37:05Tim McAleer:Amazing. Well, we've seen three great use cases. I am sure you are the hero on the team for this kind of stuff, because I can imagine again, people might be tired of hearing me talk about AI, but thank you. Yeah, but I mean, this is hard stuff. It's tedious work to do. It requires a lot of time, a lot of detail orientation. And I'm sure people love using this information to produce amazing things, but probably is not their favorite thing, like zooming in and squinting at the text to try to get it the most accurate as possible. Trying to automate away painful processes, right? Not the things people liked.
37:46Tim McAleer:Automate away toil. That's what we want to do. Okay, well, we're going to do a couple of lightning round questions. I'm going to get you out of here to go digitize a thousand more or more images. So the first thing I want to ask you about is just your approach to learning. It seems like from what I'm seeing, you're pretty fearless about new technologies, new things. I think this moment is such a critical moment for upskilling and learning. How do you think about learning in this moment? I think that one of the reasons that I find like tools like cursor or cloud code kind of intuitive is to me, there's a parallel with creative software.
38:24So like at various moments in my career, I have been deep in Photoshop or deep in Adobe Premiere, Avid Media Composer, whatever it is. And those softwares are so complex. They are like a maze of tool menus. And you end up on Reddit and on YouTube, doing your research, trying to just like figure out how to accomplish the thing. And I think that that's essentially what a lot of these tools are today too. Like I've been on cursor YouTube and cursor Reddit and learned tips and tricks from the vibe coding people of the internet. And I think it sort of starts from knowing what could be done or what's possible.
39:01And the like path to get there is swifter than ever before.
39:06Tim McAleer:What I like about this, I started sort of my fascination with technology in these creative tools. I will like this is like pre Photoshop where I would go and how can I make my text look like liquid gold and I would follow these like five step you know graphics tools tutorials and what I love about this moment in vibe coding or AI assisted engineering is coding feels such so much more creative than technical where these tools feel really like creation engines to me more than functional tools to write code. And so I love that parallel because it's what's made me so excited about technology my entire career.
39:50Tim McAleer:And I think it's why I'm so leaned in this moment, like activates that same feeling of like, oh, now I can do, can make this thing that I didn't think I could make before. I think that there are a lot of people too in my industry who have a kind of creative brain and creative approach to these things that would, you know, maybe like looking at a cursor window right now when you have no idea what it is, is a little scary. But I actually think that they are more well suited for the work than they might know. Well, let's talk a little bit about your industry, because I know that the film and creative world is deeply skeptical of AI.
40:23Tim McAleer:Sometimes we wade into the waters of AI video generation on this podcast and get a little feedback. And I totally understand. I have family that's in the creative industry. I'm curious, you know, what's your point of view of AI, particularly in the film world? What are you excited about? And where do you think these kind of concerns are really warranted? And then where do you think the most practical applications are? I think today it's like sort of where we started at the top. The practical applications are more in like tooling than they are in creation. But I do you think that like the creation is going to get there?
40:57Like today I play with, I play with all the generative video models. Like how can I not? They're, they're super fun. Um, they are not like at professional grade quality yet. Like the amount of time you spend throwing tokens at even the highest end video models, you're not going to be able to match your shots that well. You're not going to be able to match the footage you shot yourself that well. And so I don't think they're there yet, but like, I'll be honest, they're going to get there. I think that like, they are still exciting to me but i would separate a couple things like in the non-fiction world i think i think people should be careful like i think we should not be generating archival footage we should not be trying to fool our viewers into thinking that there was video in 1750 you know and i think that that's the part that's like a little scary and then of course there's the display like job displacement aspect of things i think people are scared if you film stuff for a living, you're definitely scared that you're going to be able to just use text to generate that same video you used to shoot.
41:56So I don't know how to like, I don't think anybody has good answers to that part of it. But my approach has certainly just been like jump in and learn the tools. Like they are, they are gonna be here, whether we want them to be or not. And I think that they have a lot of practical benefits today that are less scary.
42:16Tim McAleer:Yeah, the best advice I can give to people, and I have, of all the spaces, and I'll say this honestly, of all the spaces I have the most job displacement concern, it's in video generation for non-archival, non-documentary. Commercial. commercial use cases. You just see how it could be very applicable. And the best advice that I can give to people in this moment is the more you learn the tools, the better off you will be. Whether or not you love where the tools are taking us as an industry or as a culture, knowledge is power. And so the more you learn and understand, one, you can identify opportunities where it does add value even in your creative process.
43:02Tim McAleer:And two, you're going to be differentiated in the market from a job perspective because you're going to have a more robust sense of what's available in your industry. And I think that stands for people in your industry. I think it stands for people in my industry and technology. So I just say there is no harm in learning this stuff. Yeah, absolutely. I also think that like there's a place in the process for it which allows you like a place to learn without thinking it needs to end up in the final product, right? You can use video models for storyboarding all day. You can maybe prove whether or not that shoot is worth spending that money on.
43:33Now you've learned how to use the video models a little bit. And you haven't necessarily displaced anyone. But you've made your production a little bit more efficient, a little smarter. Maybe you've shot better footage as a result of it. Yes.
43:45Tim McAleer:But we're not generating fake archival footage of Genghis Khan. We are not doing that. Definitely not doing that. And I'm like, PBS, which is where most of our films end up, have a lot of guidelines around that. And I think that's a good thing, but it's the other stuff. It's commercial, it's visual effects. Like a lot of stuff's going to get easier. And so it's coming one way or another. Great. Well, last question have to ask you when, you know, you're on your dog walk with chat GPT doing voice mode and it's not listening to you or not giving you what you want. What is your personal prompting technique, especially because you use voice?
44:20Tim McAleer:Like I'm willing to type things to AI. I don't know if I'd be willing to say them. So what's your technique here? It definitely is different when you have to say it out loud. I am super nice to the AI. I can vividly remember the one time I was mean to it. I'm nice to the AI. I don't know where this is going. I'm going to be nice to all the models. What I do is like, for lack of a better way of describing it, I just start over. Like I will, I know that a lot of these things have ways of like consolidating the context window now and sort of summarizing, but I will ask for what I call like a resume work prompt.
44:53I'll be like, this isn't working. I want to resume work later with another AI dev. Can you give me a prompt with everything they'll need to know? And typically what you'll find is that that prompt shows you where it was off, you know, like in its summarization of what it was doing. I'll be like, oh, see, like I wasn't asking for that. That's that's why we were not communicating. And then I'll take that resume work prompt. I'll prune it a little bit, pop it into another chat. And then, you know, you'll find that you wish you hadn't beat your head against the wall with the previous chat for 20 minutes.
45:24Tim McAleer:You know, I am also team be polite to your AI. But then again, like you hurt the one you love the most. And I've found myself occasionally getting testy. And you know, when I stopped being mean to AI is when reasoning really started to show and I could see it reasoning how upset I was. It'll be like the user is mad at me right now. The user is really frustrated with me right now. I need to totally rethink. Go sweet, sweet baby AI. I'm sorry. Apologize. I'm not that mad at you. okay so create a you know return to progress prompt really get the summary take that to understand if there was some misunderstanding improve that and then just start fresh that's great well tim this has been super fun so much for me to learn i have tons of ideas even just for my day-to-day life about how i could use i have kids so i probably have 30 year thousand let me know if your mom wants the ocr party i will she'll love it okay mom i have gotten you your first vibe coded app direct from the podcast source.
46:22Tim McAleer:Tim, where can we find you and how can we be helpful? Yeah, I'm not that active on social, to be honest, but I am on LinkedIn, you can find me on there. I have a website that is itself a fun vibe code project. So you find me timmacle.com. I have a little chat bot there that GP Tim, you can go chat with him, learn a little bit more more about me and my work. And then other than that, I would say tune in to Florentine Films upcoming production. We have a series about the American Revolution coming out in November. So on your local PBS station. My kids are obsessed with the American Revolution. So everybody sounds like it's in the family.
46:58Tim McAleer:Yeah, we will we will be big fans. Tim, this has been great. Thank you so much. And thanks for joining How I AI. Thank you for having me. Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube or even better, leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.
From the publisher
Tim McAleer is a producer at Ken Burns’s Florentine Films who is responsible for the technology and processes that power their documentary production. Rather than using AI to generate creative content, Tim has built custom AI-powered tools that automate the most tedious parts of documentary filmmaking: organizing and extracting metadata from tens of thousands of archival images, videos, and audio files. In this episode, Tim demonstrates how he’s transformed post-production workflows using AI to make vast archives of historical material actually usable and searchable.
What you’ll learn:
- How Tim built an AI system that automatically extracts and embeds metadata into archival images and footage
- The custom iOS app he created that transforms chaotic archival research into structured, searchable data
- How AI-powered OCR is making previously illegible historical documents accessible
- Why Tim uses different AI models for different tasks (Claude for coding, OpenAI for images, Whisper for audio)
- How vector embeddings enable semantic search across massive documentary archives
- A practical approach to building custom AI tools that solve specific workflow problems
- Why AI is most valuable for automating tedious tasks rather than replacing creative work
—
Brought to you by:
Brex—The intelligent finance platform built for founders
—
Where to find Tim McAleer:
Website: https://timmcaleer.com/
LinkedIn: https://www.linkedin.com/in/timmcaleer/
—
Where to find Claire Vo:
ChatPRD: https://www.chatprd.ai/
Website: https://clairevo.com/
LinkedIn: https://www.linkedin.com/in/clairevo/
—
In this episode, we cover:
(00:00) Introduction to Tim McAleer
(02:23) The scale of media management in documentary filmmaking
(04:16) Building a database system for archival assets
(06:02) Early experiments with AI image description
(08:59) Adding metadata extraction to improve accuracy
(12:54) Scaling from single scripts to a complete REST API
(15:16) Processing video with frame sampling and audio transcription
(19:10) Implementing vector embeddings for semantic search
(21:22) How AI frees up researchers to focus on content discovery
(24:21) Demo of “Flip Flop” iOS app for field research
(29:33) How structured file naming improves workflow efficiency
(32:20) “OCR Party” app for processing historical documents
(34:56) The versatility of different app form factors for specific workflows
(40:34) Learning approach and parallels with creative software
(42:00) Perspectives on AI in the film industry
(44:05) Prompting techniques and troubleshooting AI workflows
—
Tools referenced:
• Claude: https://claude.ai/
• ChatGPT: https://chat.openai.com/
• OpenAI Vision API: https://platform.openai.com/docs/guides/vision
• Whisper: https://github.com/openai/whisper
• Cursor: https://cursor.sh/
• Superwhisper: https://superwhisper.com/
• CLIP: https://github.com/openai/CLIP
• Gemini: https://deepmind.google/technologies/gemini/
—
Other references:
• Florentine Films: https://www.florentinefilms.com/
• Ken Burns: https://www.pbs.org/kenburns/
• Muhammad Ali documentary: https://www.pbs.org/kenburns/muhammad-ali/
• The American Revolution series: https://www.pbs.org/kenburns/the-american-revolution/
• Archival Producers Alliance: https://www.archivalproducersalliance.com/genai-guidelines
• Exif metadata standard: https://en.wikipedia.org/wiki/Exif
• Library of Congress: https://www.loc.gov/
—
Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.




