Small AI Models with Yoeven Khemlani

24 Jul 2025 · 41 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Small AI Models with Yoeven Khemlani

Episode Overview In this episode of Software Engineering Daily, host Gregor Vand interviews Yoeven Khemlani, the founder of JigsawStack, a startup focused on developing small custom AI models for various tasks like web scraping, forecasting, VOCR (visual optical character recognition), and translation. The discussion delves into Khemlani's journey, the inception of JigsawStack, and the innovation behind their technology.

---

Key Players

  • Yoeven Khemlani: Founder of JigsawStack, who has a background in game development and startups.
  • Gregor Vand: Host of the podcast, CTO, and founder of Wintik.ai.

---

Background of Yoeven Khemlani

  • Started as a game developer, transitioned through banking, and eventually pursued a master's degree, which he dropped out of due to COVID-19.
  • Founded Stay, a hotel aggregation service in Southeast Asia, which transitioned into co-working spaces and was later sold.
  • Khemlani's interest shifted towards developing technology for developers, particularly in the AI space during the rise of GPT-3.

---

Introduction to JigsawStack

  • Mission: To provide a suite of small models that automate backend tasks.
  • Focus: Initially centered on web scraping but expanded to include multiple tasks designed for efficiency and accuracy.
  • Unique Selling Proposition: Emphasizes on fine-tuning models for specific use cases rather than generic applications, leading to higher accuracy and lower costs.

---

Key Technologies and Innovations Small Models

  • JigsawStack utilizes small models (around 70B parameters) instead of larger models (400B), which are costly and inefficient for specific backend tasks.
  • The focus on scaling down models allows for cheaper and more efficient deployment.

Applications of JigsawStack

  • Web Scraping: Automated extraction of structured data from web pages without requiring extensive coding.
  • Optical Character Recognition (OCR): Combines traditional OCR methods with modern vision language models for enhanced accuracy.
  • Speech-to-Text: Optimized processes to offer faster transcription services.
  • Data Transformation: Focus on translation and transforming data between formats, such as translating text in images.

Prompt Engine

  • A robust engine designed to manage prompts, route them to the correct model, and apply techniques to improve output quality and reduce costs.
  • Incorporates a mixture of agents approach to validate outputs and ensure accuracy.

---

Developer Experience

  • JigsawStack aims for simplicity in setup (e.g., NPM install), focusing on intuitive libraries that minimize documentation needs.
  • The company offers a consistent API structure to enhance the developer experience.

Future Directions

  • Planned expansions into embedding models and enhanced data extraction capabilities.
  • Continuous improvement of existing products based on user feedback from the startup community.

---

Business Model

  • Transitioning to a token-based pricing model to align with industry standards and provide flexibility for users.
  • Initial pricing set around $1.40 per million tokens with a million free tokens offered monthly.

---

Community Engagement

  • JigsawStack actively engages with developer communities and participates in hackathons to iterate on feedback and improve product offerings.
  • The user base primarily consists of startups and indie hackers looking for efficient tools for their AI needs.

---

Conclusion Yoeven Khemlani shares insights into JigsawStack's mission to deliver specialized AI solutions, focusing on backend automation tasks. He discusses the technological innovations behind their small models and the company's strategic direction to continually refine their offerings while providing a seamless developer experience.

---

Links

  • [JigsawStack](https://jigsawstack.com)
  • [Software Engineering Daily Episode](https://softwareengineeringdaily.com/2025/07/24/small-ai-models-with-yoeven-khemlani)

---

This summary encapsulates the key discussions from the podcast episode, providing insights into the innovative approaches taken by JigsawStack and the visionary leadership of Yoeven Khemlani.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jigsaw Stack is a startup that develops a suite of custom small models for tasks such as scraping, forecasting, VOCR, and translation. The platform is designed to support collaborative knowledge work, especially in research-heavy or strategy-driven environments. Jovan Kemlani is the founder of Jigsaw Stack, and he joins the podcast with Gregor Vand to talk about making use of small models for diverse applications. Gregor Vand is a CTO and founder, currently working at the intersection of communication, security, and AI, and is based in Singapore. His latest venture, Wintik.ai, reimagines what email can be in the AI era.

0:41For more on Gregor, find him at van.hk or on LinkedIn.

0:58Hi, welcome to Software Engineering Daily. My guest today is Jovan Kemlani. Welcome, Jovan. Hey, hey, good to be here. So Yvonne, you are the founder of Jigsaw Stack, a ultimately Singapore company, but you're now in the US, which is a pretty well-trodden path, I think, for this part of the world. So before we get into Jigsaw Stack, which I cannot wait to get into, it's a very interesting product and very timely. I'd love to just hear a bit more about your background. I know this is, as they say, not your first rodeo when it comes to having founded a company. So yeah, just tell us about your journey to Jigsaw Stack.

1:33yeah i think it didn't start very exciting right it's like every other engineer so i'm an engineer like everyone else i love building products i love exploring new technologies and that's kind of like where i started so i started as a game developer which many choose not to take that path because it's one of the hardest industries to kind of break into from a revenue standpoint or like a salary standpoint like your initial career but something i enjoyed did that for a bit love building didn't love the industry went into like the banking industry didn't love the corporate industry decided to leave that and as I went to study my master's at imperial covid hit right at the right moment so I decided to drop out I didn't want to pay over two hundred thousand dollars just to sit in my dorm room so I decided to drop out and eventually that's when I got inspired to start this company called stay are way back I think it's like four years ago it was a hotel aggregation company in Southeast Asia that basically did short term bookings for like the pool, the gym and different things within the hotel because domestic markets were growing at back then.

2:35Then that grew to a point we went into co-working spaces and eventually we sold it. And I think that kind of kicked off my path of like, okay, startups is like where I want to be. I love building and I love building for people, right? It's like, how can I make money from the things that I built? And that's when I realized that I specifically love building tech for developers and tools and software and that's when i went into this rabbit hole of what should i do next and this was like during the gpt3 when it came out i was like ai space is getting interesting and a lot of technology being built is for front-end human in the loop applications right chatbots or tools that people can use in the loop and i was like can we bring this technology to backend applications where there's no humans in the loop.

3:22It just works. It takes away processes that used to require a lot of manual intervention. The most common thing right now is like web scraping, right? Like everybody's trying to scrape the web. Can we structure it in a way where you don't need to write puppeteer code anymore or playwright code anymore? And basically you just give a URL and you prompt fields that you want to extract and it does the work for you at a 98 % accuracy, not just marked down, like actual fields that you can pull out. So that's the pain point that we realized that can we automate these backend tasks? And we went down this rabbit hole of like fine tuning models and training and like how can we increase the accuracy to 97, 98 and get it up there.

3:58So that's kind of like how Jigsaw Stack started and then we started narrowing our field there. Nice. So I think you've kind of hinted at it there, but Jigsaw Stack, what is the, in sort of, I don't know, one sentence, what is Jigsaw Stack? So Jigsaw Stack is a suite of small models to automate your backend task. That's the way I like to phrase it. Awesome. Okay, so let's kind of dive just into that. You kind of already touched on there was a pain point around pure web scraping. But then I guess when did that morph into small models versus just going down like, oh, we're just going to make a web scraping model?

4:31basically yeah so the first challenge that we tried or the way we try to tackle this problem like can we use gpd3 or gpd4 one of the big large llm the best in the class at a point to solve this problem most of the times we get good solutions for very human in a loop things you get marked down generated from a website that you can use to pass into a chatbot to then you know talk with that markdown but you don't get structured data that you can actually use for example if you want to pricing comparisons between two products on amazon typically a developer would need to write some code to kind of extract that specific price field that they need to.

5:05And OpenAI or with any kind of scraper tool can't achieve that at that same accuracy a developer would be able to. So that's when we realized that you had to fine tune. So when we started fine tuning, we tried 400B models, like the Lama 3.1 at that time. And then we started scaling down. We said, can we take the same and bring it to a 70B model? Because then we reduce our cost, increase efficiency. And then we started going down even further. Like, can we bring that down to a 13B model. 13B didn't work really well. So now we're kind of on average at like the 70B scale. And that's where we see it's easy to deploy, it's cheap.

5:38And since we're so specialized in that one use case, we can get rid of a lot of the other generic use cases built into the model, right? And kind of read and train a lot of it specifically for that one thing. So that's where the small model kind of utility really came in. Nice. So let's sort of maybe go through a few, I guess, well, you have many small models, is my understanding. So could you kind of go through a few of those? You've touched on kind of web scraping, that which is still very much in the product today, I believe. What are some of the other ones? Yeah. So when we launched the web scraping that blew up and it kind of opened this door of like, where should we go next?

6:12We didn't want to dive super deep into web scraping because we realized the technology that we built for web scraping, it was 50 % the model, 50 % the infrastructure, because we trained the model specifically to Puppeteer, for example, that controls Papathia and injects JavaScript code. So now the question is like, how can we broaden the scope to do other forms of data extraction? And that's when we exploit data extraction as one pillar, right? So OCR is an example of pulling out text from things. So traditional OCR, like optical character recognition, uses machine learning to like kind of recognize the characters.

6:45A very traditional way of doing it and very inaccurate way of doing it. Now we combined it with today's tool of like vision LLMs, right? That gives us a whole new realm of OCR with bounding boxes and this kind of tool sets. And then we went into things like speech-to-text, is basically extracting text from an audio. That's where we took Whisper3, optimized the shit out of it, and basically made it one of the fastest speech-to-text. And we're actually faster than Grok at this point. So that's the direction we're taking. So that data extraction is just one example. Eventually, we realized that generative or data transformation was another big pillar.

7:21A lot of the users were asking like, hey, I scraped this chunk of data from, let's say, a Mexican passport. Like they were doing passport verifications with our OCR. And then now they need to translate some of the information. Translation became a huge pillar by itself, where people were moving away from Google Translates and the typical translation providers because the quality of the product, LLMs that exist today. So can we take a model, train it specifically for translation, and then add languages as we go, right? And then become a specialized translation model. So that's what we basically started to do for data transformation as well.

7:55And I saw you give a presentation a couple weeks ago, and you were talking also about translating text in images. That's a kind of interesting challenge, right? It is. I think people have seen this in consumer apps like Google, where you point the app at a match and it kind of overlays a bunch of text in a different language. which they have this in Apple, Microsoft. And we were looking into this and we're like, surprise, there's no API for that. Literally all these cloud providers don't provide an API for that specific service. I think one reason it could be because it's not high quality. They're actually just overlaying like a blurred text on top of the current text.

8:26So what we are exploring is that we saw that image ads and image generated content is becoming a big thing in the market, right? Is there a way that we can take an existing image, understand the text on that image, then translate it and diffuse it back with the same style, right? So basically don't affect the style and the way the shape of the text is. And that's something to explore. So with diffusion, you can really do a lot of that. So you can train a model on glyphs, on text glyphs, to then basically translate the thing. So I think we're like two weeks away from like version one. And then version two is going to be the full diffusion model that we're going to launch.

9:02So right now we're going for like English and Chinese as the two major languages to kind of like diffuse against. And then we're going to include Hindi and Arabic and like different languages here. Awesome. I'd like to touch on another sort of, I guess, part of the product offering. And you can correct me whether this is sort of an amalgamation of bits of the product or totally standalone, but the prompt engine, that's something that we've talked a bit about before previously, outside the podcast. And I had a sort of specific use case, which I can always talk about. But could you just describe what is prompt engine and like, what is that solving?

9:33So Prompt Engine, it solves three big problems. First is prompt management, which there's a lot of providers in the market that do that for you. The second is prompt model routing. Again, another suite of products in the market that does that for you. And the last is prompt techniques, right? Techniques that you can apply to obviously increase the quality of your output, reduce your token costs and all these other things. so we decided and we see in the market there's so many challenges for like developers using three different products just to solve this like these lang chain plus they use another product has to track their tokens and another product has to track evaluate their model outputs and stuff so what we managed to do is like can we train a really small model like 1b 13b size to basically make these decisions for you at runtime so we took different data sets in different industries, law, history, English, mathematics.

10:26And we trained a really small model that can make decisions on which model to pick at runtime based on your input of your prompt. So whenever you give an input, say generate a poem, it will understand, okay, it's a writing focused product. It needs to generate poems. Then we have a data set that has been trained. So poems are really good with GPT 4.5. Like they're really good at writing. thing it'll automatically pick that model for you but it will also run the same prompt against like five other models and when you get the output every single time you run that same prompt it narrows down the best model based on this concept called mixture of agents have you heard of that idea where it basically runs it and then uses llms as a judge to basically say hey three out of five this thing gives the same output typically this is going to be the correct output and then we can merge all these three into one output and then it gives that so it's a chain of techniques that kind of basically gives you this thing so from a user perspective you don't pick claw 3.7 or gpd 4.5 or gemini 2 you basically get everything and the selector just the model picks that for you and you can also store the prompts and then right rerun it and stuff yeah i mean the example i spoke with you about previously was it was around time zones and i I think we all know in programming, time in time zones is a hilariously difficult thing to still achieve various scenarios on.

11:50So that was kind of what I at least ran through Prompt Engine initially to kind of get an understanding of what is this product and how can it actually help us. And it was super interesting because it was far and away better from a results perspective. if I put into just pure GPT or I'm talking about GPT over the API or just go to Claude and just put it in there or go to prompt engine and say like hey we've got three people in three different time zones here's a few rules though like don't start a meeting before 6 a.m for any of these participants strangely the other two really struggle with it it's like you have to keep correcting them like when they come up with these three like oh here's the three times yeah one's at 5am it's like i explicitly said not before 6am oh yes you're right and then whereas prompt engine came back pretty reliably covering that so i think that's a very i mean can you sort of help explain why then it's able to produce that result so when you send your request what happened behind the scenes is basically it sent the request across five models right and typically what would happen is that maybe two or three out of that five models gave the right answer.

13:01And what happened to be the smaller model that we trained, X as that judge, right? So models, small models are really good at consuming data, but are horrible at generating. So they can make a decision if you give it a lot of information to be like, yeah, this is the right one based on its context. But it's really bad at saying, hey, I can give you the right answer. So if you tell it, okay, these are five of the answers and pick the best answer, it automatically uses a weighted average and be like okay this too has these answers coming up more often plus from my understanding this is the prompt the initial prompt this is the answer and it will automatically start merging those top two answers together so what happens is that you now you always guarantee the best answer and you're not going to get the worst answer because two out of three or three out of five of the models are basically giving you that same answer or similar answers and it merges the output so even if one of the model gives let's say they add AM and the other model SPM, when you combine it, it overwrites based on the initial prompt.

13:59So it's a mixture of agent is just one technique that we apply on it. Recently, you know, using the O1 concept of chain of thought is another thing, another idea, but implementing an entire model like R1 or O1 mini or O3 mini into from tangent is going to take tons of token costs and like take a really long time process. But we applied the same technique of chain of thought into our models from scratch way before like the whole deep CR1 came out, but obviously at the smaller scale. So it runs a lot faster and cheaper for the user. So that's why it tends to work more accurately, especially even for structured data.

14:31So if you ask it for a structured response, you almost guarantee get structured response. Very similar. I think one thing I'm curious though, like if a lot of users who are relying on LLMs, they're not actually tuning models themselves, but it's kind of usually a sort of secondary exercise once you have approved product market fit or so forth. And clearly the next step is to then tune a model, and become more deterministic for their use case. So at least where I sit, I kind of get used to understanding the likelihood of what's going to come back from a specific model. For example, if I look at our use case we have, doesn't really matter the details of it, but if I put that use case to a Lama model versus to GPT, I get very different responses back.

15:11And I'm like, okay, we're always going to go with GPT for this one because I'm just confident that what's coming back. In the prompt engine example, how does that work in the sense like, The result might be, quote, better. But I guess with the model behind the scenes have like flipped. I'm just trying to get my head around like, how much can the same prompt going to prompt engine end up then changing from sort of, I don't want to say quality perspective, because quality clearly is what the product's about, but change from sort of structure or like, oh, that's completely different to what I was shown.

15:44Expecting it. Yeah. A hundred percent. That's an interesting question. So the way we do that, that's why we always suggest a user to create. with two parts to this system. You create a prompt first and then you execute that prompt. That's where the prompt management layer comes in. The reason we suggest to the user to always create the prompt and run it, like the typical LLM, you just execute the prompt, right? Like you just generate the prompt. But in our approach, you create, because you're storing this thing, there's a level of memory that takes place. Meaning every time that you run your prompt, we store a generated version of that prompt of like how it outputted every single time.

16:20So let's say if you run this prompt six, seven, eight, 10 times, it basically will always store the version of the output previous to the next run in a database. And then we use that as a baseline as well. So if you keep changing your prompt and you keep tuning your prompt, we know clearly something is wrong. But if you kind of run the same execution of the same prompt, then we know, okay, hey, we say, hey, this is positive, this is positive, this is positive. This is how the structure should kind of look like. This is the output that it should kind of look like. so we give it some form of point system at the back where we're saying this is good this is bad this is good this is bad and as the more you run the more accurate it gets actually and the less models that we run against so initially it runs against five by the time you're on your fifth run is running against two models by your tenth run you're running just one model and we stick to that model so you don't change your model by your tenth run so if you're using let's say it automatically picks gpt4 row and gpt4 keeps being the best model for your answer it will just stick to GPT-4O throughout the lifetime of that prompt that you created so that you don't have changes and like breaking changes as you scale.

17:21Unless the model fails, right? So that's where we fall back. So unless the model is down and then we fall back to another one. Nice. APIs are the foundation of reliable AI and reliable APIs start with Postman. Trusted by 98 % of the Fortune 500, Postman is the platform that helps over 40 million developers build and scale the APIs behind their most critical business workflows. With Postman, teams get centralized access to the latest LLMs and APIs, MCP support, and no-code workflows all in one platform. Quickly integrate critical tools and build multi-step agents without writing a single line of code.

17:55Start building smarter, more reliable agents today. Visit postman.com slash SED to learn more. Something you touched on earlier, and I know is sort of quite a big piece of the product when it comes to i guess advertising this to developers is the speed can you talk a bit about how does that work let's just talk about small models versus large models maybe specifically first and then then move on to like jigsaw stacks infrastructure and how generally how is that making things so much faster so we started with the idea of gpu poor i mean we don't have a lot of gps right even have a h100 physically a lot of ai companies you see today the first thing is like let's buy one gpus for training so we came in with the idea of like hey we're going to be gpu poor as much as possible we need to work on a100s a10gs any of the smaller gpus i mean at that time it was not small but now it's considered small we have to kind of be restricted to that scale that was one of the ideas or like the methodology that we stick with the second is deployability the biggest issue with i think llms or big providers like open ai right now is that distribution has been a big blocker for them they are distributed with azure or people have to use open ai directly and that's how you see a bunch of these enterprise ai's companies coming out trying to train open source models because you can't self-host open ai's models anywhere being the best in class and that was one of the biggest blockers currently so we looked at this problem like we don't want to be in that realm right we cannot afford to basically be we have that as a blocker where enterprises want to use our ocr model and we're like oh we can't self-host because we rely on 30 different proprietary APIs that require like this much work.

19:32And that's when the small model idea came in because most enterprises don't have the resources to run large language models at 400 B scale. So if we can focus on training the model to be deployable and cheap to run anywhere, it became an easier distribution thing for us from the get-go. It's hard to build for sure, but in the long run, the cost kind of pays off because we own this proprietary model that we can eventually then go out to the market and be like, we can distribute on your systems easily, no restrictions. If you want AWS, you want Azure, you want GCP, you want your own on-house GPU.

20:08Eventually when we get big enough, we can even host on Grok to be even faster. And that's like the kind of opportunity that we have with smaller models. Cost efficiency is a big part of it. Yeah. So just so I kind of read that back, Grok is behind the scenes from an infrastructure perspective or multiple or? For Prompt Engine, we use a bunch of Grok. under the hood as well llama guard is a feature that they have that's pretty interesting that there's a lot of prompt guarding and we didn't want to add llama guard into our platform without grog because it's one of the fastest and we didn't want that bottleneck eventually so we do speak to grog and the goal is like you know if we can get big enough volumes we can also host the cost of hosting on grog is huge right you need millions and millions of users to kind of like be using you to even host one small model on grog so we want to obviously reach that scale where we can start hosting other models on infrastructure dedicated for small models.

20:58Okay, so let's, I guess, then talk about just general devics here. Obviously, we've got a very technical listener base, and this is precisely who should be using Jigsaw Stack. So what's the kind of likely first kind of what does the product look like to a developer on day zero when they're getting up and running? Yeah, right now, it looks like an NPM install, right? So you just honestly NPM install Jigsaw Stack, get your API key, that's it, everything else works. When you build the libraries in a way where you don't need the documentation so as long as you install the library you should have enough typings that kind of like guide you in what you need to do so from days you're over like every field needs to be a named field everything needs to be descriptive everything needs to be typed even our python library is fully typed so that we as a developer my best experience is using stripe or a library like super base where you literally like stripe dot you get your options customer dot subscribe we wanted to be that intuitive right where you don't need a documentation in most scenarios but obviously if you need advanced conflicts you go there but you should be able to get started with just the library and that's the kind of the developer experience thought that we had like can the library just be self-sufficient enough that the mp install they'll figure it out they pip install they'll figure it out and that's kind of like the design structure that we went with yeah i like that you mentioned both stripe and Superbase, I think they're both great reference points, Stripe especially on the API side and Superbase just for pure, I think, DevX when it comes to, yeah, what that product is.

22:27My understanding is the API itself, it's a pretty consistent structure across all the services. Is that the way to think about it? Exactly. So we kind of have a system, at least where we kind of structure every API to be pretty consistent in the way you call it, in the way you structure your body, JSON to kind of pass that data over. all the keys we use the same key across all the api so you don't get confused everything is just url or file slot key so there's no confusion in that sense yeah nice and i mean i guess we kind of have to talk about pricing and how you do charge for this and i think most of these products we're talking just in ai and gen ai especially you know it's a usage-based model could you just talk a bit about like how have you determined at the moment any of the like the best balance for like accessibility for developers versus then sustainable business economics that kind of thing like how have you looked at it so far and like how do you see that evolving yeah i think it's a good time to announce as well so we're changing our pricing right okay so we're recording on 12th of march today so just 12th of march yeah exactly so we'll be changing it in like a week or two ish the reason is that pricing i think is an ongoing change we have been learning a lot since the time we launched, we've removed so many products as well.

23:44And we've added a lot of new products. The reason we've done that is because we kind of like speak to customers, we develop and we understand truly what they want. When we started the product, we were like, we built on things, what we wanted. And then now it's like based on what developers wanted the whole time. So we've reached a point where we realized that the pricing that we have doesn't make sense for a lot of the products that we are offering. Big issue is that, for example, we charge right now 0.05 cents per API call. It made sense for things like the AI scraper, the OCR, because you have like a fixed cost that you can scale with, and then obviously get discount pricing the more you use.

24:19And it worked in competition to providers like AWS and like, you know, GCP, right? This is a similar pricing model. You have a fixed price to an API cost. What happened is that we started releasing products like speech-to-text and text-to-speech. And what happened is that if you're charging 0.05 cents for one hour of transcribing a video you're still getting charged the same for five seconds it didn't make sense to a lot of users at scale right and managing discounts manually for each customer didn't make sense a lot of the users came back to us like hey why not just use token-based pricing i'm like yeah why didn't we not just use token-based pricing like every like every llm provider is using it every developer is used to token-based pricing and if we can just shift to that kind of idea it gives one the flexibility to us to now charge for more usage in a more minute way meaning if you use less we charge you less you use more you based on the output we'll charge you more accordingly secondly it lets us explore more technologies because now we can allow you to configure specific apis to run for a longer period because we're not kept by that 0.05 cost underlying cost behind that scene so now we're shifting to this token based pricing where we're estimating it to be around$1.40 per million tokens, which is actually pretty good because we're trying to keep ourselves as cheap as possible to most users for scale, right?

25:38So$140, I think, like Claude is at$15. So$140 for like infra plus like a specialized model is something that we think it's fair. Now, obviously, we're still going to test more. We've been getting a lot of feedback. We haven't got a lot of experiments yet because we've not launched it. But when we launch it, we'll get more feedback, figure it out, and we'll still keep adjusting pricing to what developers will come to comfortable with you know at the end of the day so and i guess both well i mean today but i guess with this future pricing model there'll still be you know intro free tier of some description or yeah yeah oh yeah for sure 100 so it'll be like a million free tokens every month yeah this is sort of not pricing exactly but it is related like i guess context window i'm trying to get my head around how does context size play with small models because i guess if i think about without you saying anything further i think well context window small model surely it's just a really small context window but what is it yeah yeah in this small context window but we don't really rely on the context window as much it's not all language models a lot of them are trained models to control infrastructure so like the ai script for example i gave which basically takes in fields right like i need the price i need the description of this players i need the image on the left side of the whatever it's a very small information context input huge output on the output level that's where you see basically the model generates js puppeteer code then gets injected into puppeteer and then the puppeteer outputs a huge context but it doesn't need to go back to the model right and that that goes to the user and then that context gets basically post process so context is not a big problem for us because we are more of infrasite problem if you're obviously building a chatbot or like some form of a chat system or that it requires huge amount of context to be in then yes that i think context makes a bit different typically for most chat applications or like human in the loop applications but not so much on like where we actually never had like a really context problem at this point yeah yeah okay that's super interesting so i want to talk about the future of the product as well as currently with the developer community that you have so let's let's start actually with current which is like developer community who's using it i'm aware you guys did like a hackathon maybe a couple weeks ago which looked very interesting so maybe what was kind of some of the results of that and like what is the current kind of community that are using jigsaw stack and how do they talk to each other i think there's two big communities one is the typical startup indie hacker is kind of you know there's people who love the product let's play with it try it out but they are i think for dev tool like any like super base for sale or any product on the market it's always we face this kind of problem where it's when you need the product versus when you discover the product right so you a lot of times you discover the product but then you don't need it at a point and then you when you need it you forget about that product so kind of having us exist constantly in front of developer to the point where they need it and then get reminded of us is a key thing for us so that's why we target a lot of startups while they might not need us today they enjoy the product they eventually when they build their startup scales or they build a product that do does need jigsaw stack we're the first one to give a try give a shot or compare us against like gcp and aws the competitors that we're going against and then basically be like hey yeah this is a better product let me try it out and let me go for it so startups is our key thing so series a companies are smallest at this point we do do one or two enterprises but enterprises is something that are a bit more challenging for us at a scale we're happy to do it because we are still a small team it's like three to five of us at this point and as we scale we'll take on more and more larger companies that are series a and bigger but they have special needs that like you know they needed to be deployed only in australia or only in specific parts of the u.s and stuff like that which we wanted kind of do at a latest scale but right now it's mostly startups and it's and we get better feedback to be honest right like at this point because it's so high like like the hackathon because of the hackathon on that day i was not even building i was building jack saw stack Like everybody came to me and they were like, hey, I have this bug over here or I didn't even know this field existed over here.

29:44Like for the AI script, right? I'm like, okay, updating the documentation and like fixing this field. So the feedback loop is really good, like in real time, right? When you do stuff like that from the startup community. And that's what I love for, yeah. The startup community both have, I think, quite a very high bar versus some people and then like quite a low bar other places. It's like, you know, when you look at the product, very high bar for like, this should just work. You know, I should be able to just get running in like 10 seconds and then a fairly low bar when it comes to like oh maybe this piece of documentation isn't great but that's okay yeah is that kind of how it feels or oh for sure depends on where you talk about it then when reddit the heart of bar is very high reddit everyone's anonymous they're like this product is shit and i'm like oh my god you speak to the same guy in real life he's like oh yeah i know i love your products sometimes you don't know but yeah i think you're right basically right developers are very intuitive i think the more senior of a developer you speak or a developer that's comfortable with their skill, they're very intuitive.

30:41So if the product works or they can figure out, they love to fix things, right? So if they can fix the product for you, they would, right? And like, especially if they can. But that's why if the documents are broken or something is small, it's not big of a problem for most engineers, right? And that's what they used to. And like, so they will go and they'll take screenshots and be like, can you fix this? Can you update this? And even the biggest of companies like Google Google, Vertex AI doesn't work half the time. So the expectation from a startup is... I have experience of that and just thinking like, no, surely it's not actually Google.

31:17Just the API doesn't work right now. No, it doesn't. Exactly. It needs to happen. If Google can do that, like, you know, the expectation from like Jigsaw starting a startup is like, surprisingly from companies, it's actually a lot higher. Like in the US, when I, a lot of the startups that we meet there, the expectation from a startup is like, we expect you to be way better than GCP, right? I think in Southeast Asia, it's a bit different where they're like, yeah, I can expect you to be worse than GCP because you're a small guy, which is good. I like the expectation in the US where the expectation is like, hey, I expect you to be better than GCP.

31:49And that's why I'm going with you in the first place. I'm not going with GCP. So that kind of downtime and these things are very important to us. And the clarity of feedback that we get from a lot of these companies and engineers are very upfront, very clear. And they only complain about the things that have real problems. If it's actually down, they will complain about it. If the docs are wrong, it's not a big issue. They will just come and screenshot, can you fix it? And they solve the problem already. So it's a forgiving community when it's the right thing that goes down. And I guess now just looking kind of future and obviously like roadmap and obviously super early days for you guys.

32:24So roadmap is always one of these very difficult things sometimes I think to kind of pin down. But I mean, for example, I saw a post I think you guys put out very recently about sort of how you compare to say Mistral when it comes to OCR. Because like Mistral had this big splash with how their OCR is like in theory well in a way better than anyone else. And you sort of made some good comparisons of why in many cases I think you argue that you're actually better. So is the future direction sort of let's take these cases that the big guys are doing and try and perfect them? Or are there just like completely other use cases you're adding, thinking of adding?

32:57or yeah so we want to focus on two big pillars the first is data extraction the second is data transformation so we're going to stick in this realm so as long as it falls within this realm we're competing with any big guy that comes in this split market right michael was never in this market they were always in the llm market and then one day they dropped an ocr model like oh okay cool they want to get into the small model kind of like space as well and that's when it got exciting so i'm like let's try it out and trying after trying it out it wasn't as good as they claim to be based on their title of their article well they literally they use the word world's best i don't like shitting on on certain things but like i love mistral i love the pro they love especially that they have their first few models that they launched it was like they kind of kind of initiated a lot of the open source scenes like way before the llama and you know and the guys did so i love them for what they did in the open source world but when they did this mistral ocr and I'm like, world's best, like really.

Read the full transcript

33:53Just like, so I re-benchmarked it and it was like very clear, far apart, like from the world's best. And so I just had to put it out there and kind of show like, okay, hey, like there's a lot of room for growth and I'm happy that they're doing it. And it kind of ups the OCR market. So let's actually be the world's best and then claim that title, right? So this time we did the benchmark, we're like, hey, yeah, you know, it took us a lot of time to get our OCR model out. And we were surprised that there's a world's best, especially when Google can, with Gemini, level of resources couldn't be the world's best ever.

34:27Like I'm surprised I misruled it. So we were kind of like benchmarking and we were like, oh yeah, a lot of room, a lot of room for improvements. And I mean, in terms of like future of just the company team, very recent announcement a couple of days before us recording this, funding has increased, which is awesome. And like how sort of what is that going to enable you to do? and like how do you see the team going up or not going up? I mean, obviously we're in the golden age of small teams now. So, which I love because I've actually always been a small team person and was very frustrated through the 2010 to 20 period when you had to keep defending why you had a small team.

35:01So how does it look for you guys for the next couple of years? So we raised the one and a half million. The goal is to grow the team to like a five-man team, including myself, keep it very lean. And then the goal is to get the product to a solid standpoint, right? We're in beta still. We want to get ourselves out of beta. we are launching two significant products the first one is the embedding model it's a multi-model embedding model we realized that embedding is only tag space which is like but everybody embeds pds images and a bunch of different documents like we need an embedding model that can support all these document tags natively so we're launching that and the image to image translation that we spoke about one new idea that we had is in the data extraction space is i think microsoft released a segmentation model where you can segment buttons and different fields in a in a site can we combine that object detection and other forms of detection into a single model right so segmentation object detection and a few things into into one and that's something that we're exploring so we're going super deep into some of the the detection space and the embedding space so this the next i think one year is diving way deeper into the technologies that we've built improving the quality of each thing rather than scaling into more products we kind of like this is where we are stopping the product roadmap, but now we're just going to go deeper.

36:14How can we improve the developer experience? How can we make it even more seamless? Can we improve our AI scraper to make it even cheaper and make it faster by updating the engine under the hood? There's so many new property alternative engines that are coming out that's faster to run. Can we scale that, make that faster? So I think the next one year is really about the quality of the product, the developer experience, and how deep we can go and distribution rather than scaling the kind of like the product roadmap. So the product I shared with you right now, like basically going to be like the last three additions.

36:46And after that, we're just going deep. That's a great strategy. I think, I mean, you know, for anyone listening and kind of getting going with the product, that's nice to know in the sense that what you're using is just going to improve and sort of, you've already been through that phase of like shaking out, like what bits actually make up the product and what kind of gets people excited and also useful. Are you going to be hiring? We've got a great technical listener base because it's great if you say yes or no, because then you either don't get flooded with emails or you do, but... Just, yeah, go to jigsawstack.com, like, slash careers.

37:15We're always hiring at this point, and the significant role that we're hiring for right now is a founding full stack AI engineer. We only hire A-star players. We have three questions. One of the questions is, like, do you have a side project? That's my benchmark. Like, if you, as an engineer, if you don't have a side project that you're working on for yourself or for fun, then don't need to apply. Just don't apply. I think that's great. probably a caveat there is if you're already working on a startup but then realize this is a better opportunity that can that effectively is your side project right yeah but yeah if you're perhaps in a sort of regular role but itching to kind of get into something more interesting side project is always kind of looked upon very favorably by people like yourselves just a kind of i guess closing question generally was i think you founded this like effectively solo is that correct yeah i'm a solo founder it's not that i wanted to be a solo founder i think being the technicals because in my previous company both my co-founders were non-technical when i typically go for a non-technical product that i'm building then it's easy to find a lot of co-founders that in that space that specialize in a specific space i think the same way non-technical founders find it challenging to find a technical founder it's difficult for a technical founder to find another technical founder it's even more difficult because the issue is that you have a belief system of like how things are built right and then you're always kind of like in this realm of like what the way you see things and finding another technical founder is even more challenging.

38:37I've tried, couldn't find a perfect fit. And that's why I'm like, let's just find a specific roles. And then, you know, eventually those people become co-founders, right? And that's why I'm kind of hiring like the founding team that you eventually scale, you get equity and your equity grows. And I can give away honestly more equity because I'm a solo founder. So that's a good benefit there. I think it's great to call out. I mean, there's just, there's so much out there basically saying like almost don't even try unless you find your co-founder and i think there's quite a few examples in the last couple years where i've seen solo founders doing incredibly well so i just kind of thought it was good to call out so 100 like i don't think i'm not really so like a really great team and i don't feel alone in the company and that's the thing right like i think solo founders just have to build their team better and that's like the only challenge and so it's not a pain point i think it's not a big pain point for me yeah yeah absolutely and yeah i mean you sort of founded it effectively in singapore i believe kind of moving well is it west or east it's kind of equal distance pretty much well let's just say we're going you're in east of the states is that right yeah yeah so moving to san francisco one big reason is because when we launched in singapore naturally i'm here so it's easy to launch here we put on hacker news we're done reddit we put it and we could just majority of our customers and users are from the u.s at this point at at least at the initial stage, we want to focus on selling the US customers or in that region.

40:02Our second biggest market is actually the UK. And so we just think being in that region makes more sense for me to speak to customers, get a better feedback loop, just be in that same energy and like at the forefront of technology in that space, right? Southeast Asia is still a big market, obviously, but we require a lot more capital to tackle Southeast Asia. And that's something that we want to come back down the line, you know? Yeah, for sure. Well, Evan, it's been so great to catch up and to hear all about Jigsaw Stack at this stage in the journey. I'm sure we'll catch up again in time when Jigsaw Stack is probably 10 times the size in only a couple of years or something like that.

40:36So yeah, just wishing you guys all the best and we'll be following along. 100%. Thanks for having me on.

From the publisher

JigsawStack is a startup that develops a suite of custom small models for tasks such as scraping, forecasting, vOCR, and translation. The platform is designed to support collaborative knowledge work, especially in research-heavy or strategy-driven environments. Yoeven Khemlani is the Founder of JigsawStack and he joins the podcast with Gregor Vand to talk about making

The post Small AI Models with Yoeven Khemlani appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Small AI Models with Yoeven KhemlaniSoftware Engineering Daily · 41 min
Listen in VO