SED News: Data Land Grabs, Copyright Fights, and the Great AI Talent War

8 Jul 2025 · 46 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Software Engineering Daily - Episode Summary: SED News: Data Land Grabs, Copyright Fights, and the Great AI Talent War

Episode Overview In this episode of Software Engineering Daily, hosts Gregor Vand and Sean Falconer discuss significant recent stories in the tech industry, focusing on Meta's legal challenges regarding AI training data, the competitive landscape of AI, and data ownership issues. The episode provides insights on the current state of the tech industry, particularly the ongoing battles among major players in AI and data acquisition.

---

Key Topics Covered

  1. Meta's Copyright Battle
  2. Lawsuit Overview:
  3. Meta has been involved in a lawsuit concerning the use of copyrighted material for AI training, particularly related to books.
  4. The initial ruling favored Meta, citing fair use similar to Google’s precedent with Google Books.
  • Fair Use Argument:
  • Meta argued that the AI-generated outputs do not constitute market dilution of the original works.
  • There’s irony in how different content types (e.g., books vs. online articles) have varying levels of protection against data scraping for training purposes.
  1. Data Land Grab in AI
  2. Importance of Data:
  3. The competitive edge in AI currently hinges on access to quality training data.
  4. Companies are racing to secure data sources to improve their AI models, leading to strategic partnerships and acquisitions.
  • Meta's Investment in Scale AI:
  • Meta invested $14.3 billion in Scale AI, acquiring a 49% stake, indicating a push to secure high-quality training data.
  • The investment prompted other companies, including Google and OpenAI, to halt projects with Scale AI due to concerns over competition.
  1. AI Talent Wars
  2. Talent Acquisition Strategies:
  3. Meta reportedly offered signing bonuses of up to $100 million to attract AI talent from competitors.
  4. This reflects a broader trend where firms are competing aggressively for skilled professionals in the AI domain.
  1. Inter-company Struggles
  2. OpenAI vs. Microsoft:
  3. Tensions are rising between OpenAI and Microsoft regarding their partnership terms, particularly concerning clauses that could affect their long-term collaboration.
  4. Allegations of anti-competitive behavior have surfaced, reminiscent of past regulatory scrutiny faced by Microsoft.
  • Salesforce vs. Glean:
  • Salesforce’s limiting of API access for Slack messaging poses challenges for Glean, a startup providing internal search solutions.
  • This highlights the growing importance of data ownership and access in corporate environments.
  1. Emerging Trends and Predictions
  2. Market Predictions:
  3. A slowdown in major announcements is expected due to the summer holidays, with fewer conferences and events scheduled.
  4. Potential challenges for companies like Manus AI regarding user satisfaction and clarity of billing practices are anticipated.

---

Key Takeaways

  • Legal Battles: The legal landscape for AI and copyright is complex and evolving, with major implications for how companies use data for training models.
  • Strategic Data Acquisition: Companies are heavily investing in data sources to gain competitive advantages in AI.
  • Talent Competition: The fierce competition for AI talent is leading to unprecedented offers and potential shifts in the workforce landscape.
  • Corporate Dynamics: Relationships between major tech firms are becoming increasingly strained as they vie for dominance in the AI sector.

---

Conclusion The episode provides a rich analysis of the current state of the tech industry, focusing on the strategic maneuvers of major players in response to new challenges and opportunities in AI and data management. With ongoing legal and competitive developments, the landscape will continue to evolve, making future discussions essential for those in the software engineering field.

---

For more insights, you can listen to the full episode of [SED News: Data Land Grabs, Copyright Fights, and the Great AI Talent War](https://softwareengineeringdaily.com/2025/07/08/sed-news-data-land-grabs-copyright-fights-and-the-great-ai-talent-war/).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:12Hello and welcome to SED News. I'm Gregor Van. And I'm Sean Faulconer. And this is a certainly different format of SE Daily Podcast, where we basically take a spin across the last few weeks in terms of big headlines. We have a main topic in the middle. We look at hacker news highlights. And we just kind of give our little thoughts on what's been going on in the tech and predominantly software world. So how has your few weeks been, Sean? What's been going on? It's been good. I mean, I think last time we chatted, we were kind of, I was in the thick of coming off of Snowflakes, a big conference, and Databricks, a big conference.

0:48And then the challenge with going back-to-back conferences is that then I come back and I have like a mountain of work to actually catch up on because it's hard to do the work while you're at the conferences. So I think the last chunk of time has been essentially just trying to recover from a work perspective on all those things that were piling up. And then getting ready for a personal trip to Hawaii soon. so oh awesome nice yeah like why what have you been up to yeah it sounds like we've had sort of opposite schedules then i was sort of heads down for a while and then yeah the last couple weeks have been quite busy from a sort of events perspective yeah we had super ai which is like this big conference that got put on over here in singapore it was a bit weird i'm not gonna lie it was a sort of like going into a nightclub for like two days straight so i had to sort of take breaks but it was interesting just a lot of agentic agentic stuff i would say it was more sort of corporate leaning perhaps but good keynotes there was dorkish patel another ai podcast that's very popular these days and so had a quick chat with him and edward snowden was also a keynote speaker but by video link for obvious reasons yeah i haven't heard edward snowden and it's been a while yeah quite yeah so yeah it was a very interesting event and lots of people had flown in from all corners of the earth for this thing so it's always nice just to meet a bunch of new people from around the world as well yeah and think forward asia is happening in singapore i think next week oh yeah very soon i guess singapore is like a hot bed for tech conferences so yeah in this part of the world for sure it's kind of the place and yeah facilities here kind of you can just go down to this place called marina bay sands i believe it's actually owned by sands of las vegas and yeah you just go down there any day of the week and there's something going on down there so it's kind of fun so but yeah thanks to the Singapore government for the ticket I did not pay a thousand bucks for that ticket so I am a real startup founder I don't just blow a thousand bucks on these things anyway let's get on to the headlines the big one we're going to talk about to begin with is around meta there's a lot to talk about meta this week but meta and copyright basically and they've been in a lawsuit and at least the first sort of round of that is that they effectively won their argument that this is all around books.

3:04So it's the idea that publishers are not very happy, that basically chunks of text come out in responses. Obviously, this means the model has been trained on these books. And yes, the publishers are not very happy about this, but it seems like Meta has won on the basis that I think the TLDR here was simply that their argument wasn't strong enough yet. And Meta, I imagine, had gone to town on their lawyer side. But yeah, what was your take on this one, Sean? Yeah, I mean, my understanding was they were kind of able to argue it was under sort of fair use, which it was the same kind of argument that Google made over a decade ago about Google Books, which was also determined fair use.

3:47And I guess it's a very complex issue. I'm certainly not a lawyer. Some of the fair use arguments kind of make sense to me. You know, if I can't prompt the model essentially to regurgitate the entire novel word for word, then are you really getting anything more than perhaps you would be able to get from Wikipedia or Quif Notes version that also falls under presumably fair use? Yeah, I think that's a good example. Yeah. But then at the same time, if I'm a online content creator, you know, if I own Reddit, for example, I can block third parties from crawling and ingesting that content and training models on it.

4:23So it's kind of strange in some ways that if you're in the digital world and your thing exists digitally, I can prevent essentially model providers from essentially getting that data and using it for training. But then if I am an author of a book, I don't have the same easy control over it. And I have to, even if they didn't make the right argument, they have to essentially take this thing to court in order to be able to make that argument. So it's very complex. I don't pretend to be the person who has expertise on it. I do find it kind of, there's some irony in it though that there's protections basically put in place from a technical control perspective.

5:00But if you're not in that world, suddenly you don't have the same control to protect your IP. Yeah, I think this is the point. Like a lot of us in tech, I have a, technically I do have a lot of background. I'm maybe unusual, but I don't ever sort of pretend like I'm more knowledgeable in law these days than anybody else in tech who's not like an in-house counsel. So I think the point is that we're not lawyers, but we still have to understand on what side is sort of right or wrong in some ways. I believe what was said here is that there was no meaningful evidence of market dilution. That's a fancy way of saying they don't believe that this is so the judge saying like, I don't believe an LLM is going to stop people from buying the book.

5:38It's kind of like a translation of that, i think right yeah like if you want to read harry potter are you going to chat gpt and saying like you know i'm already paying for this i want to save some money so i'm just gonna have it tell me the story of chat gpt i certainly am not but i can't speak for everyone i mean i think that that's a fair argument right because i think to what you mentioned just before that you can't sort of prompt to get the whole book out and exactly you can't just say oh i want to read chapter one right now and then chapter one pops out it doesn't work like that to my understanding and that's i think and i think yeah cliff notes for those that remember cliff notes yeah it's a dated reference so cliff notes were you know kind of like cheat sheets for when you were studying like literature and things yeah it's like if you couldn't understand i don't know the shakespeare play that you're supposed to read in high school english class you could get the cliff notes version where it explain it to you in plain english exactly i definitely use those yeah so here we are we're going to see how this plays out and obviously we're touching on it here because this does affect all of us in software you know if suddenly oh the model has to stop referencing any books tomorrow well a lot of products just don't have that kind of concept that that might not come out so then what are the cascading effects of people's platforms or products where suddenly this sort of content that was just assumed that would be available is gone for example Yeah, and I think this kind of talk is one thing that points to a larger sort of theme and some of these things we're going to be touching on through the course of our conversation, but it's just there's such a land grab right now around data.

7:12Like the big competitive advantage for people building models or even those building applications on the model is really what data do I have access to? What data can I either use for training purposes or if I'm building applications on these models? Like how do I get the right contextual data into the prompt in order to get something relevant for my business application? And the people who are the most successful are the ones that are going to kind of win the market. That's really where the competitive advantage is. It's less, at least currently, around new, truly innovative techniques in terms of how these models are constructed.

7:49It's a lot to do with essentially the data and how organized it is for training purposes or prompt assembly purposes. Yeah, exactly. And we've been seeing a few things like this pop up, especially in the last few weeks. We're going to get onto that, as you mentioned, Sean, and the main topic around basically walls that are going up. So we'll get to that shortly. The other kind of main, I guess, announcement headline that made mainstream news as much as tech news was Meta again, making a 14.3 billion investment in scale AI, which gives them a 49 % stake in that business. And that's always a great number.

8:29If you see 49 or 51, you know, it's effectively saying this is effectively equal ownership. It's just that someone's decided there's a reason to take one off one side and put it on the other. Is that the same for Microsoft and OpenAI? Is it 49 %? That's a good question. I'm not super sure. I don't know. It sounds familiar, but don't quote me on it. Yeah, exactly. It could be that kind of similar. So in this case, ScaleAI, they have high quality training data and they are a vendor to all the big players, to my understanding. And this is just a blatant land grab by Meta. But crucially, because they haven't gone over that 50%, it's not an acquisition.

9:09so again let's just put the legal hat on for five seconds it's the idea that this is not going to be scrutinized sort of from antitrust and i'm sure just a whole bunch of like time and effort that would be needed to like fully acquire or majority acquire a business and this is no no we're just 49 lots of money in your bank account but you know you can still supply other people if you want but so yeah i mean you're kind of on the ground over on that side sean like what do you make of this It's a similar structure as I think Microsoft's investment in open AI, whether that's 49 % or not, I can't remember the exact percentage breakdown, or Amazon has a stake in Anthropic as well.

9:46And a lot of these giant tech companies are investing in these companies to access the AI capabilities without necessarily triggering any sort of antitrust reviews. It's a bit of a hack around that system. But as a consequence to the move that Meta made, Google, I believe, paused their scale AI projects within hours of the announcement. OpenAI is also winding down their relationship. Elon Musk's, you know, XAI project also halted some of their projects as well. So a lot of these companies are pulling out of this. And this kind of goes back to even the earlier conversation that we were talking about.

10:24Like, it's all about data. That's what Scale AI does is it gets data ready for training these massive models. They have all kinds of people deployed around the world that are involved in the cleanup process and the labeling process. There's a lot of human labor that goes into preparing the data. And that Scale AI has been able to address that. And they were working with all the biggest companies in the world to help prepare the data for training. and now of course meta gets to strategically kind of own that data funnel this is though maybe still the bit that maybe some of the audience as well are like scratching their heads on a little bit like i'm still just trying to fully understand okay meta own let's just say effectively own scalei but at the same time and scalei is providing the data so what changes when is around like how they're going to structure the data is it going to be very skewed towards like say llama models versus something else like why is there such an immediate pullout from these companies like i'm still trying to get that concern yeah i'm not 100 sure there either like why the immediate reaction was punished scale ai in some fashion because i mean who are the alternatives at this point yeah i don't know who else is in that i mean there's a couple other companies like label box and stuff like that that are in sort of the there's a bunch of data labeling companies.

11:43But as far as I know, there's no one sort of operating at the scale of scale AI. But there's a bunch of companies kind of focused on that problem set. So I don't know if the plan for the Googles of the world is to go and leverage those competitors in the space, or maybe they're going to build out known some of this themselves because they realize that they don't want to be dependent on a third party vendor to provide this. I don't fully understand what the impact would be if all those companies were using, they're already using scale AI. But presumably, when i give data to say aws and amazon it's not like just because it's running in an amazon server anybody at amazon can go and just like look at that data so i'm not sure why they had the feel that compulsion to pull out yeah it could also be well it's a pure economic thing you know if you sort of you know that that thing effectively is your competitors and then you say well you know these big multi-million dollar contracts yep they're potentially even billions at this point Like, oh, they're gone.

12:40It could be a kind of power play where they're trying to maybe have Meta double check their decision on that one. But who knows? Yeah, I mean, it could be also just because the space is so competitive. And presumably, all those companies using scale AI have some proprietary data that's part of that process. They do have concerns of something nefarious going on. We talked last time about the corporate sort of spying and espionage. So just because you sign a contract doesn't mean people aren't going to necessarily break it if the right incentives are put in place, especially if they feel like they can get away with it.

13:13So maybe there's just enough potential risk there that they want to go and seek somebody else to do that job for them. Yeah. And I think this is maybe a good time to move kind of onto our main topic, which is just the idea that big tech, the walls are going up. And why is this sort of significant? Well, I think it's fair to say maybe pre-AI or pre-GBT, we just didn't maybe see so much of this where the big tech was like really going at each other. they didn't really have like a specific piece of land so to speak to like fight over they were kind of like well you know google has its like ad thing and so does meta but meta also has this other thing and we're kind of all doing our dances around different things whereas this is i think like cloud is maybe the closest between the major like cloud providers but all the biggest companies are kind of multi-cloud anyway so there's lots to go around they're certainly competitive but i agree like i don't think we've seen this kind of level of competition since maybe the early days of social when Facebook or slash meta became really big, there was a real existential threat to Google's business.

14:21And they really put a lot of resources and time and effort behind like Google plus and the other failed Google social projects and stuff like that. I forget what the circle, maybe that was Google plus, but Google buzz was another one. They had social social experiments. And none of them really took off. And then eventually they kind of gave up on that as a business and went after other things. But, you know, since then, I don't think I've seen, I think AI, at least in my lifetime, since I've been working in the industry, is probably the thing that I feel like has the highest competition and companies are just like throwing crazy amounts of money at it.

15:00You know, they're trying to steal each other's talent with offering massive amounts of money incentives for people to come over, for talented people to come over. And it's like this land grab where I think that they see this as the future. There's going to be winners and losers and they want to make sure that they're on the winning side. Yeah, for sure. So we're going to dive into a few of these specific battles going on. So this is sort of in the realms of the walls are going up, like who have we got against each other and on what grounds? So the one that's also touched the big headlines, main headlines very recently, OpenAI versus Microsoft.

15:33and this is around legacy agreements that they had around Microsoft as you've touched on Sean owning a significant chunk of OpenAI for a long time however there was this interesting clause in there called like the AGI clause and this is around like at what stage does technology get to that stage and the thing is it's a sort of ironic problem because achieving AGI would automatically terminate the partnership but surely you want to reach that stage if the idea is just to advance technology and there's been something mentioned where like many executives at microsoft back in 2019 they thought this clause was nonsense but again reportedly satya nadella was like no we're too far behind on this i imagine like transformer etc we just need to do this deal like we'll figure out later and well here we are later six years later almost and yeah so yeah how does this look i mean i think the challenge just on the agi clause thing is there's not a clear definition of even like what agi is so let's just make sure agi is artificial general intelligence so essentially we've reached the place where we have human level intelligence and some people argue like hey we can have a chatbot pass the Turing test, which was sort of the original idea of a test created by Alan Turing many, many years ago of where essentially the idea is you have somebody that's interacting behind like a closed door asking questions to either a human or some sort of computer.

17:09And if it's a computer and the human can't tell the difference between the answers coming from another human or from a computer, then essentially the computer passes the Turing test. And for certain types of question answers, certainly you could argue that something like ChatGPT could pass the Turing test now. And I think people have kind of tried to prove that. But at the same time, there's things that these models are like incredibly stupid. Like there's ways of tricking the model that would never ever trick a human. And even just like basic arithmetic and this kind of thing. Like there's still, you would expect to ask a human, most humans, hey, what's one plus one and get the right answer.

17:44and there's enough evidence to show you might not always get that answer from a model at the moment. Yeah, so there's these types of challenges. So how do you even prove that you've reached AGI? It would probably become some sort of legal thing again. Like, how do you enforce that? There's no clear set mathematical definition of what that is. So I think that's a challenge. But the relationship between OpenAI and Microsoft has continued to get, I think, more and more contentious over the last couple of years or last 18 months, certainly. Like OpenAI has accused Microsoft of anti-competitive behavior multiple times.

18:19In a lot of ways, it kind of reminds me of the old browser war days where there was a lot of accusations against Microsoft in terms of like forcing OEMs to have Internet Explorer installed. And Microsoft certainly has a history of anti-competitive behavior. Famously, they got brought up in the early 2000s on antitrust charges where they tried to essentially break apart Microsoft as a company. And eventually those charges went away. but it really damaged Microsoft's, their sort of reputation. Yeah, their reputation, sorry, during that time. And we've seen the Windsurf acquisition, of course, which sort of plays into this whole thing as well.

18:54Yeah, so OpenAI acquired Windsurf and then Microsoft has GitHub and Copilot and all their IDEs and suddenly you have this competition that's happening between these two companies where there's a significant amount of stake. You know, from Microsoft's perspective and sort of investing in OpenAI and OpenAI has been running on Azure Cloud. I think they've been trying to pull back some of their dependencies there to be less vendor dependent. So there's all these things that are happening sort of behind the scene. And then I think also OpenAI over the last year or so has started to really pay attention to the enterprise and be less just about a consumer facing application.

19:33And I think enterprise is where Microsoft historically has really thrived as well. And that's where a lot of the dollars are. So that of course creates more tension between the two companies. Yeah. So we're going to move on to the next battle. This is Salesforce versus Glean. Now, what does this even look like? So Salesforce owns Slack. That's kind of where we're going to go with this one. Glean is, let's say, there's more of a startup. So this is interesting. We're actually seeing a kind of incumbent go against one of the more early guys. We did have an episode on Glean back in April with yourself, Sean.

20:07And this is funny because I remember listening to that episode. I happened to be in London And listening to this episode, I went straight into a meeting with a friend who works at a very large PE firm. And we got straight on to the topic of AI. And they said they just rolled out this white labeled search all our internal information. And I said, oh, who is it? Oh, it's Glean. So suddenly, you know, all kind of made sense. That's the context here. We don't need to name the name here, but the largest PE firm in the world, effectively, is running on Glean. And here comes Salesforce or Slack saying, oh, you're not going to actually be able to, with any sort of practical means, now catalog Slack messages via the API.

20:47They've put in like a pretty onerous rate limit on that. So it's not like full block, but it is incredibly hampering. And it seems like Glean was the main target for this. Yeah. What do you make of that? Yeah, I think it's unfortunate because Glean, I think, you know, was born out of Google. Google had created internally a product called MoMA, which was to try to solve this kind of heterogeneous search problem of internal documentation across Google. Like Google has been a company over 20 years, 100 ,000 plus employees. There's stuff everywhere. Like how do you make it findable, essentially? And then a bunch of smart engineers from Google left, started Glean to take that idea and build a product around it.

21:27And it solves a real pain point for enterprise business. Anybody who's worked in a large company can identify with this challenge of like, oh, someone said something to me, like, where the heck is it? You know, was it in a doc, an email, Slack message? The enterprise search problem is really challenging for most businesses. And Glean did a really good job of addressing it, where it can index essentially all these different disparate data sources, give you one interface in the search into it. And then they've done a lot of stuff since then. And now they have a bunch of AI tools. You can chat to it so I can ask questions and I'll go use sort of a rag based application behind the scenes to pull in the context and provide a proper response.

22:04So I don't have to necessarily just search it and click on links. And then they also have the agent support now where I can build my own workflow and use my own internal docs and stuff like that. So really, really cool stuff. Cool company. And I think it's kind of unfortunate from a competition standpoint that Salesforce is penalizing their ability to do that. But it kind of all goes back to the data problem. Like all these companies want to own the data because they own the data, then that's like where the AI serving is going to have to be dependent on. And Salesforce is making a huge play around becoming a data cloud company.

22:39They're going after lots of snowflakes of the world. And they just bought Informatica. They're trying to get more and more data sort of in their gravitational pull as a company. And most of these companies are not interested in that data flowing out. They're only really interested in the data coming in. And for their products to work, they kind of need to own the whole world to make it work. So I think that's unfortunate, but hopefully there's some resolve to that where Glean figures that work around or the other products in the same space. I think Notion is also looking to go after it. Notion, yeah.

23:12I went to Notion more trying to make big inroads over here. I believe there are APAC offices in Australia but they put on this very on-brand event over here rented out this very nice space in the National Gallery and it was called Cafe Notion and had like jazz music and fancy breakfast canapes which I thought was just a very nice spin on this kind of event where you're not pushing it to the evening with drinks and things you actually push it to the morning and just make it very nice. So all on-brand but I think the big standout for me with them during that whole presentation was a we're going for enterprise and they make a big case about open ai in theory all running on notion now and then b this data integration this is their big push it was you know oh you can integrate all these platforms and these ones are coming soon and but you can already integrate slack and so on and i'm just wondering like where do they go when their whole offering is saying look we know you've got disparate data you might not be all in notion yet but what can we do about that well we can integrate and like well now one of the biggest providers of the integration that would probably help you because i think most companies if they're using notion it's not like they don't use notion instead of slack that's just not a thing it's actually slack was trying to kind of recreate notiony bits inside slack and yeah yeah i don't think it really worked very well but so what does like a notion do in this in this case as well i'm assuming they have the same throttling challenges right all the companies that are integrating there besides Salesforce, which probably has a, you know, workaround through some internal API.

24:39I'm not sure what will happen. Or you have to get into a place where you have to end up with like more of a strategic partnership with Salesforce where they unblock you. You're not just using sort of a public facing API endpoint. There's an API that has higher limits or it's a different API and stuff like that. It's interesting. I'm just on the basis that, you know, Notion roll in. They're pushing very much, you know, we can clean up these 10 SaaS platforms that you use and even if you need to keep paying for them you won't actually interface with them all the data will just be pulled into us and as we're going through right now data seems to just be the thing that actually companies used to be kind of more open to exchanging because it seemed like a sort of fair trade like well this data for this purpose and you need it for that person sure like let's have apis that's the name of the game but as companies like salesforce clearly i guess looking at this and saying but why can't we be the people to give you the best context from our own data so and this context is the value here seems to be it's what we're being told so much data context is the value so the next one we're going to move to it's more of a rumor but it has been strongly i guess reported and you touched on it earlier sean meta versus open ai and this is in the sense of people.

25:54So in theory, Sam Altman came out in public saying that Meta had been offering$100 million signing bonuses, which is obviously, I don't think we've ever seen a number like that in terms of, at least publicly stated, for signing bonuses. Meta haven't commented on this, and I don't know if that sometimes means it might be true, but I don't think we have seen sort of this kind of level of for a while at least or in terms of competition yeah i mean even if it's not 100 million i'm sure it's a lot to poach talent from some of these other companies and certainly i think there's been three fairly prominent researchers that moved from open ai to meta recently and i'm sure there's a movement all over the place it's almost like you're in the space of professional sports or something like that where people are getting these like huge contracts i wonder if there's like a multi-year contingency to that you know it's like 100 million dollars but over 10 years to join or something like that are based on your performance and things like that so you get into a place where you're not just acquiring companies you're acquiring the talent to build the company that you want yeah exactly it really goes back to this arms race that's happening in ai right now whether it's it's for data or it's for essentially the talent to make things happen yeah and as you call out it's 100 million okay sure that's like a headline these things are never structured that way it's not just like oh you join and then on a monday and on Tuesday, 100 million lands in your account.

27:16There's many ways they sort of structure this in terms of options. And I just say like kind of over a certain amount of years and performance basis and so on and so forth. So, but I obviously it's not unusual to see someone like Sam Altman just take the number just to kind of make a splash. And he's obviously trying to say, look, we have the best engineers. This is what people are trying to pay for them. It's marketing effectively. Yeah, that's another angle, right? In terms of the competition, like the competitions, I think going, it's all the way up the stack, essentially all the way at the hardware level.

27:50Like if you look at the hardware level, NVIDIA, AMD to some extent are essentially providing all the chips, but a lot of the cloud providers and also big model companies are also looking now to figure out they're investing essentially in their own chip designs because they don't want to have this vendor dependency. And a lot of AI startups are going multi-cloud because they don't want to have too much dependency on a single vendor. It's really about either it's protective moves. They're trying to diversify their stock portfolio to some extent so that they're not beholden to any one company. And you have also come across the fact that Google has donated A2A to the Linux Foundation.

28:31So how does this sort of play into the walls going up? Or are these actually walls coming down? Or like, what is this? Yeah, it's kind of the opposite in some ascent, at least like from a surface level where we talked about Google agent to agent a couple times on here, you know, came out a handful of months ago. It's kind of a similar idea to Anthropics MCP, but focused on interagent communication. And just recently at the open Source Summit North America, the Linux Foundation announced the formation of the agent-agent project, which involves AWS, Cisco, Google, Microsoft, Salesforce, SAP, and ServiceNow, perhaps others.

29:08I think it's good from we're kind of moving towards consolidation. Even Cisco, which had a somewhat competitive product, is integrating A2A support into their agency's core. So I think that's good from a standards perspective. Overall, it's good that companies that are looking to invest in using something like agent to agent, now you have neutral governance, remain vendor agnostic. It can be a community-driven project. But strategically, for Google, there's probably multiple reasons, I'm sure, to do this. It's not necessarily something they're going to directly make money from. And more adoption of agent to agent and other such standards overall is better for the companies that are already succeeding in the agent market because it essentially becomes another thing that people could build against less barriers to entry you don't want to reduce as much as that as possible make it as easy as possible for people to build stuff and if you're you own the gpus you own the cloud serving or you own the models and the tokens being generated then you want more people essentially building yeah absolutely and yeah sort of on the same vein we've got a episode in the coming up in the future around mcp security actually and actually during that episode we talked a lot about there is no sort of one place at the moment to go for validated mcp servers for example and maybe this is sort of the same thing we're not the same thing but you know if google are kind of handing this off to the linux foundation like they maybe hope they're going to be a good steward of that protocol for example more so than say a google doesn't look neutral unfortunately linux kind of does so yeah yeah exactly i think like one of the challenges with a lot of the MCP servers right now too is that the majority of them aren't managed experiences from like a vendor that you trust.

30:56It's source code that's available from the vendor that you have to run yourself essentially. So you're taking on the burden of doing it or you go to one of these MCP aggregators that are running that on your behalf and then you have to know like do I trust this aggregator or not? How are they running this? All that type of stuff. So I think those are just signs of it being such a new thing it's like a year old so it's going to take a little while like there's it's surprisingly few companies that have an actual managed mcp server that you can just use yeah that's precisely what we touch on in that episode where a lot of it's around do you know who provided this server effectively so yeah look out for that in the future yeah so i think we've hit the high notes on walls in tech obviously we could have covered at least double that but i think we've kind of covered what the last three to four weeks have shown us in terms of what's going on I'm no doubt in next month's news episode we'll see some developments here as well as yeah who knows let's move on to I think it's almost my favorite part of SED news where we just get to look at hacker news you know throughout the weeks we just sort of have a scratch pad of things that we've maybe seen and kind of interest us so I might kick off with one that I just love these little projects that people do for no other reason than it's cool and fun but actually they have interesting uses you know at the end of the day and this one was called my iphone 8 refuses to die now it's a solar powered vision ocr server and i just was like i have to click this and see what this is so thank you to the user that submitted that this is basically someone who has rigged up i think it's an iphone 8 and it their argument was this thing was sitting doing nothing but it's actually incredibly powerful especially with like vision ocr from from apple which is a on device vision model they've rigged this thing up like a power unit which is powered by solar so this whole thing runs effectively from its own power and this person says they've got like a lot of image processing needs so they don't explain what that is but they say you know like thousands of images i think at least per week that need to be analyzed and labeled and this person claims is doing it fantastically well it's all self-powered the iphone 8 is remarkably stable and powerful for this purpose and i also like he says it's also a great conversation starter this thing just sits on his windowsill and people are like what is this thing so yeah very fun very fun yeah i also love to get a kick out of these little you know side projects and stuff like that and a lot of this like side project stuff is kind of like how i started my interest in computer science software engineering and stuff like that you know many many years ago i have less time unfortunately to do those things now but i loved how the he talks about the amount of over engineering that he went oh yeah it's like you know a little side project and stuff like that so yeah it's fun yeah super fun i almost got it confused that he was using it for like bird watching because of the fact that this the phone was like on the windowsill and i thought oh is this like you're using the camera as well but it's not that i don't think it's like he just happens to place it on the windowsill he's clearly feeding images to it from i think he's got like a micro server also sitting there which is also powered by the solar yeah i think it's like a laptop or something like that i think it's a laptop he's yeah it's like not connected to anything it's all like one internal network there's actually bird feeders now that have like cameras yeah they'll tell you which bird that's what i was thinking about it because i was like well i'm unfortunately a bit of a closet bird watcher so so as i was sort of looking at this thing going oh could i do this could i set up my i mean i imagine you could but i say to be clear i don't think he was using the camera piece of this it was actually just using this as a pure piece of hardware that's surprisingly good at this purpose and as he points out he's feeding thousands of images for assessment and he doesn't actually want these going to a cloud provider so this is all on device so yeah pretty cool pretty use case of that he points out the fact that the vision ocr side of things was pushed by apple like pretty quietly.

35:05So maybe we'll see more with that. And obviously Apple's not having a great time of it from the AI standpoint right now. People don't really associate them with any sort of strong AI offerings at the moment. So it's kind of interesting that there's maybe some like little hidden gems right now in terms of what's actually possible on device for Apple. Sean, what did you find in your travails of Hacker News? I want to talk about this article or highlight this article. It's actually written by a colleague of mine, Gunnar Morling, who is pretty well known from the Billion Row Challenge from a couple of years ago.

35:38And he's been involved in a number of open source projects. Now he's a principal technologist at Confluent. And he wrote an article called This AI Agent Should Have Been a SQL Query, which was like the number one article on Hacker News for some period of time. It was inspired by this talk by Seth Wiseman, which was titled, That Microservice Should Have Been a SQL Query, where he made the case for implementing microservices as SQL queries on top of stream processors. And then in Gunnar's article, he kind of explores that similar idea, but for AI agents. And he goes through this use case of processing research documents and being able to do agentic workflows, all using Apache Flink and Flink SQL running on top of the stream processor.

36:20And he highlights a bunch of the various open source flips that have been contributed to Apache Flink over the last year or so or year plus that bring in AI functionality, including one called Flip 531, which is about Flink agents, which I actually co-authored. It's really interesting. I think he did a really, really good job of just kind of explaining where maybe you don't want to do all agents this way. But there's a lot of agents that are really just like, hey, give me input. I'm going to go do my, you know, AI magic box and spit out an output. it. So like, why do I need to stand up like a whole bunch of infrastructure to do that, if I can just run that essentially as a stream job?

Read the full transcript

36:56Yeah. And I'm sure there's many examples of that where people are running things now through LLMs, and unfortunately, still RegEx, or just an SQL query could have done it at least as fast, if not certainly more efficiently. Or even traditional models, like you see a lot of people doing things like basic classification or sentiment analysis, there's other models that have existed that can perform those tasks quite effectively. And you know i work in ai i love you know i love building stuff on foundation models but you don't have to do everything on them you don't always need this massive power of this model to do things that have been kind of solved problems using wider weight techniques yeah that's a great call out yeah we got an episode coming up in the future with jigsaw stack which is a small model specialty company so yeah that's exactly what you've just mentioned sean these are small models trained for very specific purposes faster more accurate for those purposes look out for that that's a fun one very very young company so great to sort of have a chat with someone thick in the trenches with that one okay so this is obviously security leaning i always like to try and bring something security in if i can so this was ultimately on the tail scale block it was submitted by user ingve and this is an article by i believe it's the ceo avery pennerun and it's called frequent re-auth doesn't make you more secure and i had to go into this one because i've worked a lot in this space and there's a product that we have under mailpass which is pretty much exactly what he's talking about which is we sort of took a look at credential management and was like why are we constantly asking people to log in bouncing them out of a service changing passwords like more frequently than is needed and you know it also highlights the fact around device like devices now can do a lot of effectively indirect checking this is super interesting because yeah he's just talking about like frequent logins are the wrong answer like you don't you shouldn't need to he says it's kind of from a bygone era of like internet cafes which i think is like pretty astute way of looking at it you know we used to often back then use like shared computers like in a school or in an internet cafe if you can remember about that far i think maybe schools and universities are maybe like the more sort of likely case there where you're like logging into your gmail whilst on the school network or something and of course everyone just has a laptop now yeah i think the last time i was in an internet cafe and the last time i was in internet cafe was i traveled to europe for a conference in graduate school and the computer i brought the wi-fi wouldn't work for some reason so i had to go to an internet cafe just to in southern france just and paid you know whatever it was to be able to use it for 20 minutes to check my email yeah i think for me it's also europe it's croatia it was just out of high school inter railing and arriving in a town and not having anywhere to stay and using it to you know go on one of these like room share websites and like trying to find someone at 6 a.m in croatia who will like accept me in their home which worked that's a whole other story so anyway let's go back to the re-auth which is the fact that yes device possession which is kind of what we're what we're talking about here that you actually tend to now just always have the device it's your device you've probably unlocked it yourself so why are you now also re-unlocking a platform why are you logging back into it you know these very short session timings like a day or even somewhere like 30 minutes and Avery points out like okay for banking sure like there's probably a bunch of reasons like just that's a nice extra extra safeguard to have but for most platforms this this doesn't make sense and the other point he makes is something around passwords being being asked to change your passwords on a specific schedule usually by enterprise that's kind of at least the classic case of that it's something i've been absolutely against from the early days i remember when we were running our startup and my co-founder was like oh we need to change our passwords on back then it was last pass and i was like but this is a really really really good password and it hasn't been breached and if i'm just being forced to change it like i'm going to forget the password or it's going to become less secure.

41:05And that's his point, you know, like as soon as you ask someone to change their password every couple of weeks. So yeah, I don't know. What's your kind of like experience with all this re-off nonsense? I mean, I think that when you put all this friction behind using a product in the goal of increasing security, what it does is actually create additional security holes because people figure out ways of working around it because they have to get work done or they need to access whatever it is. And maybe in certain circumstances where it's not mission critical that they get access to it, they just give up with it and use something else that doesn't have that friction.

41:40And you end up creating a situation where I think people are, they run out of passwords, like they're good passwords. They start using passwords that are just easy to remember, or they just start writing them down and putting them in places where maybe they shouldn't. And I think the other point that he makes in the article too, is about how, you know, attacks aren't due to physical access to devices for the most part. Attacks of getting in your email happen remotely. So logging you out on your laptop isn't necessarily helping prevent an attack that's happening remotely. It's kind of like people being concerned over going back to the cloud providers and AWS having access to my data.

42:23Breaches don't happen because somebody in a data center walks up to the computer that has the hard disk with the actual physical data in it and grabs it. They happen because someone left an API key in their GitHub repo and someone found it and they use that to get access to the data or they find someone gets access to a server that has log files in it that haven't been encrypted and they grab the log files and the log files happen to have a bunch of passwords dumped in it and stuff like that. So it's more of these remote use cases and these kind of like password protection stuff of forcing people to log out, re-log in really doesn't solve sort of the root of the problem exactly and you know any modern platform a they're most likely using some version or derivative of jwts and most likely now they're also using jwts with refresh tokens which basically means that behind the scenes this token is being refreshed so if someone is lurking around on your system like every hour that thing's being changed anyway so this bizarre idea that logging you out logging you back in has any effect on that is kind of nonsense, quite frankly.

43:30So yeah, I think it was just, it's a pretty short article. I think he makes his points in record time. So I really liked that one from Avery. Yeah. It's the illusion of security. Yeah. We didn't touch on this, but yeah, he also mentions like MFA fatigue and that's like an attack in itself now where basically do something to make your MFA pop up. And eventually, eventually you just like accept it because you're like, Oh, I don't know why this thing's popping up, but I guess I should just like use my touch ID because that's what I'm being told to do. And then And Bing, they've got the access into the platform that they're looking for.

44:01So yeah, it's a very good one to look at. Anything else from your side on Hacker News, Sean? Nope, that's all. Cool. All right. So that was a nice little tour around Hacker News. And in terms of looking ahead, we're going to obviously be back next month. This is a purely hypothetical. Do we have any predictions for what we might see playing out through July, I guess? i actually think things are going to cool off in july because in the u.s you in canada as well you have first week of july is kind of gone due to holiday you know national holidays then a lot of people are on vacation and stuff like there's not a lot of conferencing going on in july so you have a sort of slowdown in like the announcement cycles and stuff like that so i guess my prediction is there's not gonna be a major major announcement in july and we'll see if i'm probably gonna be 100 % wrong on that but that's what I'm gonna say my what's my prediction so interesting the company Manus AI has been really blowing up I would say like over here and I do feel it's just something though where users are finding that they're spending too much money on Manus and they don't know what they're spending the money on for so I'm going to make a prediction that something comes out about Manus that is in that realm that like user frustration with Manus I think it's an interesting product I'm not trying to knock Manus here but it was just seemed like an underlying theme where people are not kind of clear on the on how they're spending their credits there so so yeah that's a i'm going to go out on a limb there and just say something about manis or if it's about something we talked about today again i'm just going to go out on a limb and say maybe the scale ai pseudo acquisition hits a roadblock so let's see about that awesome as always great to catch up sean i hope we've also given the audience just a nice little tour around what's been going on over the last three weeks there was a lot to cover so again i hope we've hit some main areas for everyone and some fun things as well so hope to see everyone next time on sed news next month

From the publisher

Welcome back to SED News, a podcast series from Software Engineering Daily where hosts Gregor Vand and Sean Falconer break down the latest stories in software engineering, Silicon Valley, and the wider tech industry. In this episode, Gregor and Sean dig into Meta’s legal battle over AI training data, discuss the strategic implications of Meta’s

The post SED News: Data Land Grabs, Copyright Fights, and the Great AI Talent War appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
SED News: Data Land Grabs, Copyright Fights, and the Great AI Talent WarSoftware Engineering Daily · 46 min
Listen in VO