In short
Eamonn Maguire discusses how data profiling starts before birth and how Proton’s “Born Private” aims to reduce that by anchoring family communications in privacy-preserving email. He also critiques AI/data practices, arguing that proprietary model providers rely on user data and that “open” models often hide training data provenance.
Guest backgrounds
Eamonn Maguire is a bioinformatics and computer science researcher focused on data visualization. He has worked in security (insider threat/detection engineering at Facebook), digital humanities (visualizing language/sound and document evolution), postdoc at CERN, then finance, and is now at Proton for ~6 years.
Key claims
Email and ad ecosystems infer age, politics, religion, and more via clicks and engagement; trust in AI providers is difficult given past copyright/piracy controversies. Privacy should be easy, not just default.
Notable examples
visualization of tongue-position grids for poem reading; document-change analysis (e.g., Bible/codes/files); Molly Russell’s Instagram-driven suicide narrative; GPT-2 training-data “0.3% Wikipedia” claim; Proton Born Private reserves a child’s email at birth via a symbolic $1 donation/voucher.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOEamonn Maguire's Background in Data Science
0:45 to 2:53
Eamonn Maguire shares his diverse background in bioinformatics, machine learning, and data visualization.
“So my major interest for a lot of my life was working on visualization, data visualization.”
Proton's Origins and Mission
2:53 to 5:43
Eamonn discusses the founding story of Proton and its commitment to privacy in email communication.
“I like everything, which is interesting.”
Proton's Freemium Model Explained
5:43 to 6:48
An explanation of Proton's freemium business model and how it supports free users through subscriptions.
“they're able to also help bring privacy to other users, which are not, who are not paying.”
Privacy Risks with AI Models
6:48 to 11:52
A deep dive into the privacy risks associated with using AI models and the implications of data ownership.
“And AI infrastructure in particular is quite expensive.”
The Importance of Data in AI Development
11:52 to 14:00
Discussion on how data acquisition is crucial for developing effective AI models and the ethical considerations involved.
“So only 0.3 % of the training data is basically everything on Wikipedia in English language.”
The Investment in AI Data
14:00 to 15:00
Learn about the significance of data acquisition in AI development.
Digitizing Paper Archives
15:00 to 18:00
Explore the potential for digitizing paper archives for AI training.
Google Books and Knowledge Access
18:00 to 20:00
Understand how Google Books revolutionized access to research materials.
“I know that some of them are quite quick.”
The Evolution of AI Models
20:00 to 23:00
Discuss the advancements in AI models and their implications.
“But basically any model that comes along which looks to be performant, like the next frontier in performance, we will use that.”
Understanding Data Privacy for Children
23:00 to 27:40
Examine the concept of data privacy for children before birth.
“So we do, even if I think doing things properly in end-to-end encrypted environments is hard, but just because it's hard doesn't mean we can't do it.”
Show all 17 chapters
The Ecosystem of Data Collection
27:40 to 28:00
Learn how data is collected across different platforms and its implications.
“Or generally, that experience that we've all had of talking to your spouse and then suddenly getting ads served related to that.”
The Creation of Digital Profiles
28:00 to 31:20
Learn how digital profiles are created and manipulated based on user behavior.
“I think there's a lot of conspiracy theories about what's done.”
Born Private and Digital Identity
31:20 to 34:30
Discover how 'Born Private' addresses the privacy of children's digital identities.
“There's no real safeguards to stop people having a negative impact on their child's psychology.”
Social Media and Privacy Challenges
34:30 to 39:40
Explore the tension between social media convenience and privacy risks.
“Reserving an email address is one part of that, but it gets the conversation going in people's heads about how you want to deal with your kid's privacy online.”
The Need for Regulation in Digital Spaces
39:40 to 42:00
Understand the importance of regulations to protect users from social media manipulation.
“But on the other hand, there is a positive side to this profiling because it does make you aware of things.”
The Need for Regulation in the Digital World
42:00 to 42:44
Learn about the necessity of regulation as digital lives evolve.
“is diverging from the platform goals, then you need regulation to bring them back together.”
Building a Private Digital Ecosystem
43:51 to 45:26
Understand how Proton products can create a secure workspace.
“And so you can build, I mean, as you were saying, you're not going to create a social media platform, but you can build a digital ecosystem through Proton.”
Transcript
Automatic transcript. May contain errors.0:00Can you introduce yourself and give that background? My background is varied, I would say. I started most of my professional type of career was in bioinformatics. It was basically using computational techniques to analyze the genetic sequences or protein sequences and so on. And that very much got me into the academic mindset of things. and I went and did a master's in bioinformatics and then a PhD in computer science. Largely focused on also quite a bit of bioinformatics, but I also worked in security applications, a lot of stuff in insider threat, for example. Worked in things with digital humanities, also like analyzing poetry for the way things are spoken and how you can visualize that type of information.
0:46So my major interest for a lot of my life was working on visualization, data visualization. So the intersection of computer graphics, mathematics, statistics, machine learning later, I would say, and then the computer science aspect of it. And visualization opened up a lot of different avenues, in fact. So I worked in biology, I worked in security, I worked in digital humanities. I ended up doing my postdoc at CERN. I spent a couple of years there. Then I went and worked in finance for three years, taking a lot of what I learned from, I would say, my previous career cycles into finance also. So I did a lot of work applying techniques typically used in biology or bioinformatics, but doing it using those techniques in financial time series analysis, for example, or looking at different ways of correlating different signals together.
1:39And after that, I ended up working at Facebook where I worked a lot and more the, there's more machine learning and data science for all external and internal threats. So everyone trying to get into Facebook's network, preventing that happening. Anyone trying to exfiltrate data out of Facebook, from within Facebook, also stopping that as well. So more detection engineering. And then after that, I came to Proton. So I've been at Proton now for six years. okay i gotta ask the digital humanities is fascinating what you're creating visualizations of language patterns in poetry is that essentially what you're doing one project in fact was looking at how the tongue moves or how there's basically different ways there's a representation about how you create sounds like from the back of the throat from the front from the tongue and so on and basically had a grid representing that and then you're able to visualize the changes of this over time so you could see how different poems were supposed to be read in a more smooth way versus those which are more harsh tongues where you have big dynamic changes in positions which manifested themselves then as being very large changes in tonality for example so it was cool I like everything, which is interesting.
2:59And we worked on lots of stuff looking at, which is relevant also today for computing side of things, like looking how documents change over time. If you look at things like the Bible, for example, it's many manifestations of copies over many hundreds or thousands of years that things were added, things were removed. There's different languages, so you've got different translations. And looking at how these are done also in other types of texts as well. So like the additions and deletions, which then, if you think in computing terms also, it's the same. Like looking how files change over time is a complex thing or how code changes over time.
3:38And using that in security is also interesting. So we ended up using a lot of those things in security, visualization problems as well. Wow, that's fascinating. And how did Proton get started? And what was its original mission? The story goes that Proton got started in the CERN cafeteria. And so Andy was doing his PhD at CERN. I think many of the founders were also still at the time doing their PhD at CERN in particle physics. And it was the time of the Snowden revelations. And Andy is, of course, he's Taiwanese-American descent. So for him, what was happening in Taiwan, for example, the threat of China, for example.
4:20but then also in America where it's a country where it was purporting itself to be the free democratic flag bearer in some way and then to understand like just how deep the surveillance operations were going I think it scared a lot of people at the time and they decided at that point in time to create something from mail because in the day-to-day especially in science I would say It doesn't seem very scientific, but a lot of the day-to-day operations go around sending emails. And all of this IP that is being sent around is suddenly globally available to players like Google or Yahoo, for example, or Microsoft.
5:01And a lot of your personal life is attached to these emails. So starting it at email seemed to be a logical step. And that's when he started it in 2014. And it was crowdfunded. So there was a Kickstarter project. that raised, I think, around half a million dollars. And that went a long way to setting up the company, bootstrapping everything to make it have the servers and have the capacity to handle the users. And then since then, it's funded by our subscribers. We don't have investors or we don't have venture capital. It's basically just funded by the fact that we've had paying users and they really believe in the product.
5:40They believe in what we're trying to do. And they, by paying for those subscriptions, they're able to also help bring privacy to other users, which are not, who are not paying. Because of course, Proton is also available for free to users as well. So across all of our products. I didn't realize that. So the business model is a freemium offering, and then you pay for premium services or something. For instance, for VPN, if you're a free user, you get access to certain countries. You don't have very much choice. But if you're a paying user, you've got access to hundreds of countries. I think something it's a lot of different gateways you can connect to.
6:21And for mail, for example, you'll have more storage. You have the ability to have custom addresses or custom domains for Lumo for the product that I control or I run. Basically, you've got increased limits, better models and so on by paying. So we believe in the need to be able to make privacy available to all, but we also need to run a business as well, which is actually paying for all the people who are working on these things and all the infrastructure it requires to run it. And AI infrastructure in particular is quite expensive. So it's difficult to run that on a free model. Unless, of course, you're getting into things like advertising and so on, which is the way that most players are to monetize.
7:01And we're going to talk about Born Private, this new product or project. That fascinates me, but which is that parents can register an email address for their children at birth and that follows them throughout their lives. Before we get to that, what are the risks particularly in using AI? Before we got on, we were talking about proprietary models and I'm going to start paying more attention, but my understanding was that proprietary models, most of them say that they don't use your data for training. I know I've seen that notice on some of them. What happens with your models? And can you talk about, is it LUMO of your model, which is open source, but how do you protect privacy?
7:54What makes that different from other models? And if you don't have persistent memory, doesn't that limit the capability of the model? can you talk about that most models so if you go to as a consumer let's say if you go on to many of the big platforms like chat gpt or cloud and so on you'll by default typically you give them permission to train on your data that's their if you're a free user basically they need to make something from your free users right otherwise why give you access to the platform unless you're not paying for it. If you're a paid user, then it's a bit different. If you're an enterprise user, typically they'll have it in the document in the contract saying, we do not use your data to train models.
8:44But then it's also a trust problem, right? Because they do have access to all of your data and not only all of the conversations that are happening back and forth between the user and the model, but also all the context that is provided as well. And the business practice of these at least at this point in time for all these providers the real valuable thing is data that hasn't seen before a text that has not seen before because this is what makes the model better right it's not you can have all these fancy derivations of algorithms or mixture of expert type of architectures and and small changes to how things are being processed or attention is being given but at the end the only thing that really matters is compute so how much you're wanting to throw at it but also how much data you have and for them getting access to data is the most important thing so that's why opening up to enterprise and education platforms is particularly of interest to these platforms because they want the data that hasn't seen and within a corporate infrastructure inside proton for example we have probably hundreds of thousands of documents sitting around that have never been exposed to the public web.
9:56And education is the same. So if you have a partnership between, for example, ChatGPT and Oxford University, what do ChatGPT get in return, right? Because they give favorable terms to Oxford. And Oxford are presumably also giving back some data to them. But even if they didn't do it directly, all the students are still uploading all of their files into OpenAI servers that they have access to. And then you have to trust that they're not going to use your data. And I think that's a difficult thing to trust whenever you've seen that they basically threw away all sorts of regard for copyright law when they decided just to start scanning lots of books.
10:36And that with Claude Anthropic as well, where there is$1.5 billion lawsuit against them, because they basically went along and bought thousands and thousands of books, scanned each of the pages and threw away all the books. so that there'd be no paper trail because there'd be no digital trace. But at that time, they were going to buy all these books just in the shop. At least they bought the books, I would say. But then what they did afterwards is quite bad. And then you also have the same thing with Meta, for example, where they were found to have used this big archive of pirated books as well in order to train their models.
11:14So I think trusting ChatGPT or OpenAI or Anthropic with copyright, they haven't really shown themselves to have much regard for the law, I would say, at this point. Not to say that others are much different, but there are open, there are providers of models such as the Allen Institute, for example, so AI2, developing the open models, which are called ALMO. There's a Swiss AI initiative also, which has developed the Pertus, which is all based on open data too. NVIDIA are also creating their nematron models which are all based on open data as well so there are people who are doing creating open models but truly open not just open weights but also you can see where the data comes from you can see the code you can see the process behind it and i think i think open models has been the term has been degraded and somewhat there's open washing where people they say something is open source or whatever something is open but actually it doesn't really satisfy the requirements of being open because just one small facet of it that's like llama mother's is that right tips they don't open weights open weights but you don't know where the data comes from in the end no one wants people to know that because we know for example for gpt2 which is a long time ago now there is like some crazy stat on the wired article from around that time which stated that 0.3 of of the training data used for GPT-2 can be accounted for the entire English language section of Wikipedia.
12:51So only 0.3 % of the training data is basically everything on Wikipedia in English language. The rest, script web pages, social media profiles, whatever else before everything got locked down by Reddit and X and so on. I remember hearing and repeating in the early days of these pre-trained transformer models that they had read the entire English language internet. And I never knew whether that was really true. It sounded unlikely. But what's your view? Is that possible even? I think so. With the databases that you have, you've basically got OpenCrawl, for example, that you can go and use. you can figure out which pages are interesting we have to filter out all the garbage of course because there's a lot of garbage there too and scam websites and so on but knowing which pages are english language is pretty easy from mostly from the domain name i would say and then going off and having giant crawlers getting all that content is also not a big deal the most impressive thing at the time for companies like open ai for example where the whole transformer model thing was around for quite some time but no one really thought it would actually give decent outputs and the fact that they went off and generated gpt2 with there's a lot of investment behind it there was still like hundreds of millions of investment so spending the money to go off and build cheap crawlers to go and get the content from the internet was not a big deal so if you know that the data is the king and this data is the oil for your machine you make sure that you go and secure the oil and that's what all these companies did and all the companies that do the best in terms of having the best models are the ones that manage to acquire the most data it's so fascinating i want to get get back to born private but let me ask one more question i do a lot of research and there's so much that's not on the end of this that's in on paper that's not digitized is do you think at some point all paper archives all over the world will be digitized and then available for training i think probably yes one project which is surprising and mocked at one point was open ai's pen they had this pen thing so you could when you write it takes down your notes and so on this is an interesting concept because in some way it's seen as being backwards it's like why would you create a pen and the other side a lot of people still write a lot of their notes on paper i'm the same like even to do lists and stuff like this i still there's something to be said about writing something to formulate your ideas better than typing it in the computer and sometimes I still do it by via pen or paper and pencil and I think that's interesting I think a lot of these collections for example the Bodleian Library in Oxford it's one of the biggest archives of research material I have not confirmed it but I have no doubt that OpenAI are looking at trying to get access to that because it's it is like literally a treasure trove of information and books and knowledge that many people haven't seen which in that case is a shame as well right the fact that it is locked away it's a shame google books when it came out google books was amazing is on one side it's google so i have to maybe i have to say it's not great but on the other side what they did was incredible because there is this you know having all this knowledge but in a book that you had to go to a library and that was nice in itself but to find the information was really laborious and you had to go takes a long time to be able to find the right book and the right information and then laborious to go to each shelf and take off each book and now just having this tool where you could go and search across all this knowledge and find the references that you're looking for i remember at the time when i was at oxford i in the colleges you sit beside like everyone at the table like different profiles and so on and there was one girl who was doing her PhD on English literature, but looking at how electricity, the advent of electricity became evident in literature over time.
17:17And she was spending a lot of time in the library and it was Google Books was around, but it wasn't super well known. And I told her, have you looked at Google Books to see, type in electricity and see when the trends are, but when it first exists first is mentioned and so on and she said no i never heard of it and then i showed her it showed it to her she said wow this is amazing and okay it doesn't have everything inside it but it opens up the ability to type a word and then see how that word was trending over time it would have taken her years maybe to go through the same collective knowledge to be able to have something similar are there machines that have been produced that can turn pages and scan It must become robotic.
18:03I know that some of them are quite quick. So you've got ones which are basically like a magic wand. You pass over the page. So you pass over the page, you pass over the next page and so on. And then I think there's a way then that they can control, like to understand the context of which page it was. If you read it properly, of course, you have the page number as well. This is something that's been done automated by Google already for Google Books. And then existing systems, I don't know exactly what they were doing at Anthropic when they did the same thing. What I understood was that they basically tore out the pages, but that seems a little bit brutalist for my liking.
18:43I would have thought that this tech company would have come up with something more interesting, but maybe Occam's Razor came to effect here and the simple solution wins. That's right. Literally a razor, right? We get again to born private. The model that you're using, that you've built, or the proton has built, if there's no persistent memory, doesn't that make it very limited in its use? So just to clarify, we don't use our own model. We don't build our own models. The reason is primarily financial. It would cost us so much money to do that. Instead, we use open models, different flavors of open, I would say.
19:25So Almo was one that we were using at the start, Almo 2. We're not using it anymore because it's not since July last year where we launched Lumo. The expectations about what these models can do is changed somewhat, I think over that time period even. So we're basically keeping up with the frontier models by deploying the best current open models that we can deploy. So that includes things like GLM 5.1, Kimi 2.6, which is released just today, QN 3.5. We're using GPT OSS 120 billion. But basically any model that comes along which looks to be performant, like the next frontier in performance, we will use that.
20:11Nematron is another example, Nematron by NVIDIA. They have released two of the three flavors. So the first flavor was, I think it's, I can't remember exactly, if it's 20 billion parameter model, 30 billion parameter model. They have 120 billion parameter model, which they just released recently. And they will have a 500 billion parameter model as well. They're using fully open code, open data, it aligns very much with what we want to do as well. If that model is close enough in terms of performance, we'll use that too. Then there's Apertus, which is the Swiss AI one. The version one of that was probably not at the level where we could deploy to users yet.
20:51It wasn't created for that purpose, I should say also. It wasn't created as a general purpose chat GPT type of competitor. It was created as a base model that was going to be used for different applications. So in sciences and in legal context and financial context and so on, that could be fine-tuned towards those different application cases, but not as an out-of-the-box ChatGPT competitor. But their upcoming releases of their plans will also improve that as well. So if it's close enough to something that we can deploy to production that users like, we will be able to run that as well. And for the context side of things, it's more complicated, let's say.
21:32So we don't, just because we're end-to-end encrypted for the chat history, for example, doesn't mean that we cannot use the information that's already there, right? There is no real persistent context. There's different flavors of that. So, for example, memory itself is, it can be implemented in many ways. It can be like a global database that we are pushing to all the time. It can be something which is local to the user that you can just look up as you're making calls. It can be something which is hybrid. we keep it locally and then we sync it every so often, which is basically what we do. So when we sync, we keep it locally, it's encrypted.
22:08When we sync it, it's encrypted with the user's key, so we can't read anything on our servers. For Drive, for example, in projects, so Lumo has this projects feature, you can link your project with a Drive folder, and then that Drive folder can have all the information that you need to be able to answer questions in that domain. So if you're an enterprise and you've got one team, which is focused on some financial stuff, you can have all your financial documents loaded there. What Lumo does then is it goes and looks, that folder is now synced. It downloads all those files. It transforms them into text.
22:44We index them all locally. So then we know for, we can basically do two types of search. One is just basic keyword type of search. but the second is that you can type in full prompts and it finds the most relevant documents based on the most relevant extracted terms and then we inject those documents automatically into the context for the prompt and then send those back to the GPU, get the result back and then the user has their context from their businesses within the response because we provided it at a runtime. So we do, even if I think doing things properly in end-to-end encrypted environments is hard, but just because it's hard doesn't mean we can't do it.
23:27So we've already been doing it for mail. I think search is particularly hard because you need the data on the client and not all clients are at the same level and the ability to be able to take all that information. With projects, we made the decision that if someone links their drive folder, they would link something which is more around the topic that they're looking at, rather than say link all my drive files, which is basically could be everything. It would be a lot to be able to sync onto your machine. Whereas linking a folder, which is very narrow or much more narrow, it gives us the ability that we can do that quite well without overloading the user's machine.
24:08So we do have ways. And then you've got the whole search thing. Our search is pretty good. So we have web search, we have financial searches with APIs in the background. users can enable or disable those as they see fit depending on their threat model so if they think that they i don't want the risk of anyone seeing anything i'm typing and any query going to some third party you turn it off and if you do want it on you can actually see what it's searched for and what results it got back and so on as well born private describe born private to me it's basically the basic premise is more that parents can go and reserve their email address for their kids.
24:47I would say the process is that you choose your email address, you donate $1 or whatever you want to the Proton Foundation to support the privacy mission. It's more of a symbolic donation in some way. You get a secure voucher back to your, basically a link. And if you want to unlock your account for your child immediately, you can. If you want to wait for 15 years, you can as well. That's it. It's basically, I would say, born private more than being that particular feature, which maybe it sounds underwhelming at this point, is more about the act of thinking about your child's privacy. Where does the privacy journey start?
25:28Or where does the data collection journey start? Depending on your point of view. Some of the things that we like to talk about is that once a parent starts having emails with their gynecologist for example, or perhaps for a fertility clinic or for something else, suddenly these systems outside are able to know, okay, these people and I want to be parents. We should already start sending them advertising on maybe some hospitals that are good for birth or some advertisements for some gynecologist or whatever else or pediatricians and so on. And as soon as people start thinking more about how their profiles are being created and how the profiles of your child is being created even before they're even born.
26:15I think that part is the most powerful part of the initiative, really. It's more about thinking about your child's privacy, thinking about the privacy of not just yourself, but also the people around you, and also highlighting some, I think, important things about what tech companies and what companies in general are doing, which is not at the benefit of you at all, but only at the benefit of them in advertising. Maybe you can explain how that data is collected and merged because we've all had the experience. My wife is convinced that her iPhone listens to the conversations because we'll have a conversation at dinner and then she'll get served an ad.
26:55My argument is, no, your profile is very complicated. You interact with digital systems all the time, and it's all companies, data brokers collect all of that and create a profile. But how it surprises me that Google, who, of course, if you have Gmail, has access to your emails, uses that data, the contents of the emails. I know they use the metadata, but can you talk about that ecosystem and where all of that data gets collected and where all of it goes through email? Through email, specifically through email? I mean, from... Or generally, that experience that we've all had of talking to your spouse and then suddenly getting ads served related to that.
27:49Google is a complex machine because not only do you have your Gmail, you have your Google home system, perhaps that's listening for keywords that are going to come up. Like Amazon Alexa has plenty of examples of them listening also to conversations of people as well. I think there's a lot of conspiracy theories about what's done. Those conspiracy theories happen because there's a lot of gap in knowledge, which is also intentional, right? There's no real transparency over to say how exactly these profiles are created in the first place. But let's say from a very simplistic nature and only focusing on email, you create an email address.
28:28And the first thing you do is go off and sign up for Instagram. That tells me already something about your age, probably. Then you go and sign up for some newsletter, some political newsletter. right so now i know basically your age and probably also your political leanings then you sign up for some book or newsletter on ai so at this point now i've got three data points and then from that i can infer a whole bunch of stuff right just by connecting those things like which people are interested in ai and also have right-wing ideology for example and then you find a whole bunch of stuff which people are using instagram or the age of instagram users that say also interested in ai and also are interested in this right-wing commentator or left-wing commentator whatever else and then you have a whole you start off with those three data points and then you end up with a whole bunch of suggested data points and then if you're a place like google you can basically say oh let's send a an advert for this thing that's tangential right we think you might be interested in this and then the user clicks on it and say okay the user was interested in that now you've got a new data point that helps you build a different profile or you might create a send another advert and they don't click on that for things which are not super clear for example maybe you don't have any information about their political leanings or say religion let's say and you sent an advert for some catholic church thing whatever maybe you're not going to click on it anyway i don't know but maybe you don't click on it but then you send another one for a protestant one and you say are they clicking the protestant one okay now we know their religion too so google and other providers like this which you know it's not just about creating the profile with what data you give them directly but also how they change they're able to interrogate certain questions that they might be having their own model gaps by basically proposing content and seeing what you click on or say putting a link in a google search result higher up or lower down and seeing if you click on it or don't click on it so changing all these visual cues and changing everything around the version of the internet which is being tailored for you also can have an impact on how you perceive the world so it's a you didn't have politically right leaning ideologies but then there was a an advertisement came up which was super interesting and something you agreed with and all of a sudden you're pushed down a rabbit hole of believing in this ideology but you didn't have that inclination to start off with and that's even a more interesting thing because you're starting to change people's behavior and change not only in changing who those people are based on what you serve them and so on and I think there is an interesting film I met the director in fact last month it was the director is called Mark Silver and the film is called molly versus the machines i don't know if you heard the story about molly russell he was 14 years old she killed herself things severe depression caused by basically she had suicidal thoughts and the instagram feed kept propagating more and more adverts and more and more content let's say on suicide and in the end she killed herself and the whole thing is about You think your child is safe in their bedroom, but all these companies are basically changing the behavior of your child, even if you're not in the room with them and they're not physically in the room.
32:00There's no real safeguards to stop people having a negative impact on their child's psychology. That's fascinating. And the way you describe it, which hadn't occurred to me, this profile that different companies are creating of you is a cloud that's constantly morphing through time. It reminds me a little of your visualization work. So how does Born Private work? Born Private starts by saying a lot of people's digital identities are anchored in the email address, right? So where you're getting information from at the start. So if the parents already are using Proton for all the communications with their gynecologist and hospitals and whatever else, basically none of that information is ever going to be used to create any of those profiles.
32:49Now, that doesn't stop people from, if you use Proton, but then, for example, use Google Chrome and you use Google Search, you've basically negated a lot of the benefits, right, from using something like Proton. Because on one side of things, yes, you don't have all these things that are coming into your inbox that are tracking you, making sure that you're seeing what advertisements you click on. We're not profiling at all, right? Google already knows that you're going to see a gynecologist, and maybe you've got some apartment for some fertility clinic. Then it already knows that you're planning to do those things, so it can start creating advertising profiles.
33:28If you start going and doing that also on Google or on Google Chrome and so on, then they will know that information too. Born Private is more about this is one step in the process, but it's more to highlight that we should be trying to increase the sphere, the privacy sphere for individual users. But also within that individual user, a lot of people make decisions which are also going to impact their entire family tree. right a classic example is if you go to 23andme a few years ago and got a genetic test not only have you basically made a privacy decision on your behalf to give away your data which can be used for something in the future but also for all your kids but also your family your brothers and sisters and aunts and uncles you've basically given away part of their privacy as well so So understanding, basically if born private does anything, it's about trying to increase the understanding of what privacy means, where it starts.
34:32Reserving an email address is one part of that, but it gets the conversation going in people's heads about how you want to deal with your kid's privacy online. So I have a daughter. There are no pictures of her online. Every time the school asks, the school does ask at least, do you want the photographs to appear and are even internally? I say no. my family tried to my sister for example will try and take photographs and put them on instagram or whatever i say no i say you can remove her from the picture we might appear in the pictures because i think our day is gone probably but for for her we've made the explicit decision that we don't want anyone that we don't know seeing pictures of our daughter and if you see what happens like later on with like defects of the potential futures for how all that information can be misused i think people would take a second look at their decisions and maybe not do that do you think that as awareness of privacy and it's the negative impacts of lack of privacy spread through society that there will develop an ecosystem around maybe proton will be the kernel around an ecosystem of privacy to counter the Googles and metas of the world.
35:53So that someone with a born private email address could then participate in social media or participate in search without having somebody build this elaborate, detailed profile of them. At the minute, I would say maybe there are private social media systems, but I'm not sure if they're any good or if anyone's using them. Platforms like Facebook and Instagram and these types of tools, they're basically a TikTok or whatever. They're basically the dominant force. and many people have made the decision that either they don't recognize the risk and their decision making has been basically zero. They just fell into the crowd of peer pressure.
36:45No, your friend does it, therefore I have to do it too. If you don't do it, then you're left out. I think a lot of people are in that boat. I think maybe social media as it stands today will not be the social media that stands in a year from now or five years from now, especially with all the pressures that are coming on regulation for how people are treating people's data. And the example I give for Molly Russell, for example, is a big part of that type of drive, right? But people typically gloss over danger for convenience. But that's not something which is as unique to privacy space. It's probably for almost everything, right?
37:24If something is easy, people are probably just going to do it. If something is a bit harder, then there's going to be a bit more resistance to doing it. It's easy to follow the crowd. It's harder to reject. I think with Proton, before ProtonMail came along, it was difficult to use PGP for creating secure email. It wasn't that easy at all. And Proton came along and then made that whole process easy and transparent. The idea is to get to a point where privacy is not just the default, but it's also easy for people to do. And sometimes it's a long process because there will be unique cases, unfortunately, which show what can really go wrong.
Read the full transcript
38:08It's the same what you get from data breaches, right? It's easy to set up a website or to vibe code a project these days or whatever else. It's still hard to do security, to do security well. and so people just ignore security as the first stage then they're breached and they suddenly all the records for all their customers or everyone that's been on their platform is now leaked online and they say oh that was a mistake i should have invested more time on the security side people there's drift thread i would say data leaks in the general population and people are more wary of it for identity theft in the in the u.s for example is super easy if you have the a small piece of information less easy in europe because you can't just create bank accounts so easily in europe with someone else's name but all these things are if you make it easy for people to have a private ecosystem where they don't feel like they're left out because they don't use a certain tool they're overall safer for it they have no manipulation you don't have cases where companies are trying to modify people's psychological profiles in real time.
39:17If we can get people to think more like this, that would be a success. Ideally, they'll come to Prouton and use our products where possible. I don't see a day where we create a social media platform. Email is kind of like social media, let's say. You send messages, you've got connections and groups and so on. But I also don't think maybe social media is not going to be super popular in a decade from now. Because of all of the risks. But on the other hand, there is a positive side to this profiling because it does make you aware of things. You are served ads of things that you are interested in that you wouldn't have found before.
40:04Is there a way of managing that so that it's not abused? Well, transparency is the important part of it. You do something on this, like you say, why did I see this ad type of thing does happen on particular platforms. The regulation call as it stands is happening because we've long treated social media as some sort of utility function rather than being a platform that was basically abusing the users right keeping them hooked the same way as you have for cigarettes or alcohol or gambling i think gambling is a good is the best example and i brought this up a while back and now it's i think now i've read more that people are using that example also in public discourse but gambling systems have regulation in place because otherwise what's to stop the gambling company from just taking everyone's money if you don't have these regulations that people abide by then then the population would be miserable.
41:04You'd have a bunch of people just taken for granted. They would probably be higher divorces and mental health problems and so on. But if you look at the effects of social media, it's largely the same, right? You have people who are addicted to it. They scroll all day because of the things that have been engineered to make people stay on the platform. The whole thing is about increasing active users on the platform. So increase engagement. the more time they spend on your platform, the more opportunity there is for shareholders and they will invest more so stocks will go up. So basically the reward system is broken.
41:45There is no real incentive for social media companies to try and limit usage because it would damage themselves. So I think when these types of things happen where the user health and the user focus is diverging from the platform goals, then you need regulation to bring them back together. And I think that it's fair that people are talking about it now because for a long time, people just said, let them self-regulate and so on. But clearly that hasn't worked. And also our digital lives have overtaken in so much of our real lives that in the beginning, that you didn't spend that much time in the digital world.
42:31But now we're, unfortunately, we're all doing what we're doing right now, speaking through a digital medium. So Proton's kind of a Proton, born private is a starting point, at least with a Proton email address that the content of your emails emails are not contributing to this profiling. Although, how does somebody get, and born private is a great concept, but presumably anybody could get a Proton email address. Just go to protonmail.com. Is it protonmail.com? I'm going to check just to make sure. Protonmail.com. It's been a long time since I actually had to go to the website. Protonmail.com and create a free account.
43:16Or pay also. Paying is good. It gives us jobs and so on. But yeah, protonmail.com for create your free account. And that gives you then the access into the rest of the ecosystem. You've got access to VPN, to Proton Drive for your document storage or also photo storage. We have Proton Pass for your password manager. We have Docs and Sheets. We have Lumo for the AI assistant. Meet, which we launched, I think, two weeks ago, two weeks ago, three weeks ago, which is our video conferencing solution. So you can keep everything private there too. And more. So it's pretty easy to set up. Yeah, that's fascinating.
43:58And so you can build, I mean, as you were saying, you're not going to create a social media platform, but you can build a digital ecosystem through Proton. And we just released a Proton workspace, which is also our plan. And more for B2B focus. but the idea is that basically companies can go along and they will have everything that they need to be able to function as a company by using Proton products so all of your IP stays with you you don't have to worry about confidential data getting lost somewhere if you want to use AI assistance you can use it within ProtonMail we also have Scribe which is the first product we first AI integration we did which is similar to what you have from Gemini now for Gmail And you can use Lumo for everything else and connect it to your documents.
44:48You can ask questions about it and use it for your productivity gains, hopefully. And that's the whole idea is to create this ecosystem or workspace where you don't have to compromise on functionalities anymore to be able to use something which is fully private. And for born private, that's also what parents can find out on the Proton website. Is that right? Every time you select search for born private in your favorite privacy friendly browser, yeah, then you'll be able to find it. Yeah. But basically proton.me slash mail slash born private. Okay. Why don't we leave it there? That's all fascinating.
From the publisher
What if your child already has a data profile, and they haven't even been born yet?
In this episode of Eye on AI, Craig Smith sits down with Eamonn Maguire, Director of Engineering for AI and ML at Proton, to explore one of the most urgent and underappreciated questions in the age of AI: who owns your data, who is building a profile on you, and what can actually be done about it?
Eamonn brings a rare combination of depth and range to this conversation. With a PhD from Oxford, a postdoc at CERN, and years at Facebook engineering ML systems to detect internal and external threats, he now leads Proton's AI efforts, including Lumo, their end-to-end encrypted alternative to ChatGPT. He makes a compelling case that the surveillance economy is not just a privacy problem but a behavioral one, where the systems profiling you are not only observing who you are but actively shaping who you become.
We get into how just three data points are enough for advertisers to infer your age, political leanings, religion, and spending habits. We discuss why trusting mainstream AI platforms with sensitive data is a structural problem, not just a policy one, and why the AI labs with the best models got there by acquiring the most data, often with little regard for copyright law. Eamonn also breaks down the difference between truly open models and open washing, and explains how Proton builds AI that is genuinely private by design, with local indexing, encrypted memory, and user-controlled data sharing.
Then there is Born Private, Proton's initiative to give children a private digital identity from birth. It sounds simple on the surface, but the conversation it opens up is anything but. Data collection on your child begins before they are born, the moment a parent emails a gynecologist or a fertility clinic. Eamonn argues that until we start thinking about privacy the way we think about other rights, from the very beginning, the surveillance machine will always have a head start.
Subscribe for more conversations with the people building the future of AI and emerging technology.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on AI on X: https://x.com/EyeOn_AI
Timestamp:
(00:00) Introduction and Meet Eamonn Maguire
(00:38) From Bioinformatics to CERN to Facebook: Eamonn's Career Arc
(05:23) How Proton Started in the CERN Cafeteria
(09:23) What Mainstream AI Platforms Actually Do With Your Data
(13:00) Copyright, Training Data, and Why Big Labs Can't Be Trusted
(15:10) Open Models vs Open Washing: What Truly Open AI Looks Like
(24:22) How Lumo Works: Encrypted Memory and No Data Leakage
(31:18) Born Private: Reserving a Private Email Address at Birth
(33:00) How Data Profiling Starts Before Your Child Is Born
(34:26) How Three Data Points Become a Complete Profile
(39:07) Molly Russell and the Consequences of Algorithmic Profiling
(53:55) The Full Proton Ecosystem: Mail, VPN, Drive, Lumo, and Workspace




