In short
How children’s “data profiles” can be created before birth through adtech and data brokers, and how Proton’s “Born Private” aims to start a privacy journey by anchoring families to a private email address.
Guest background
Eamonn Maguire, PhD from Oxford; worked at CERN; career spans bioinformatics, security (insider threat, detection engineering), digital humanities/visualization, finance time-series analysis, and machine learning/data science at Facebook; now at Proton for ~6 years.
Key claims
Major platforms can infer sensitive traits from minimal signals (e.g., newsletter signups, Instagram use), then refine profiles by testing what ads/search rankings users click. There’s little transparency about profiling. AI providers may seek training data because “data is the king,” and “open” models can still be “open-washing” if training data provenance isn’t clear.
Notable examples
Molly Russell’s case (Instagram content amplification tied to suicide); Google Books digitization; 23andMe genetic testing affecting relatives’ privacy; GPT-2 training-data accounting anecdote; claims about companies using pirated book archives for training.
Guests
Eamonn Maguire of Proton (single guest).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChildren's Data Profiles
0:09 to 0:23
Exploring how children's data profiles are created pre-birth.
“There's no real transparency over to say how exactly these profiles are created in the first place.”
Proton and Privacy Ecosystems
0:23 to 0:51
Considering the development of a privacy-focused ecosystem.
“We're going to talk about Born Private, but Proton has a bunch of privacy-preserving products.”
Eamonn's Background in Bioinformatics
1:16 to 2:24
Eamonn shares his diverse background in bioinformatics and security.
“So, yeah, can you introduce yourself and give that background?”
Career Journey to Proton
2:24 to 3:50
Eamonn discusses his career path leading to Proton.
“So the intersection of computer graphics, mathematics, statistics, machine learning later, I would say, and then the computer science aspect of it.”
Digital Humanities and Visualization
3:50 to 4:47
Exploring Eamonn's work in digital humanities and language visualization.
“You're creating visualizations of language patterns in poetry.”
The Founding Story of Proton
4:47 to 5:55
How Proton was founded in response to privacy concerns.
“I like everything, which is interesting.”
Proton's Business Model Explained
5:55 to 8:06
Discussion on Proton's freemium model and user funding.
“Well, the story goes that Proton got started in the CERN cafeteria.”
Introduction to Born Private
8:06 to 9:40
Introducing Proton's project for children's email privacy.
“they're able to also help bring privacy to other users who are not paying because, of course, Proton is also available for free to users as well so across all of our products.”
Risks of AI in Data Usage
9:40 to 11:05
Discussing the risks associated with AI and data usage.
“And we're going to talk about Born Private, this new product or project.”
Trust Issues with Proprietary Models
11:05 to 14:01
Examining trust issues related to proprietary AI models and data use.
“So first, most models, so if you go to, as a consumer, let's say, if you go on to many of the big platforms like ChatGPT or Cloud and so on, you'll, by default, typically you give them permission to train on your data.”
Show all 26 chapters
Copyright Concerns in AI Training
14:01 to 16:18
Explore the ethical implications of using copyrighted materials for AI training.
“they basically went along and bought thousands and thousands of books, scanned each of the pages and threw away all the books.”
The Role of Open Models in AI Development
16:19 to 19:17
Learn about the importance of open models and the quality of data in AI.
“data, so 0.3 % of the training data used for GPT-2 can be accounted for the entire English language section of Wikipedia.”
Digitizing Knowledge: The Future of Archives
19:18 to 22:09
Discuss the potential for digitizing paper archives and its impact on research.
“And the other side, a lot of people still write a lot of their notes on paper.”
Advancements in Page Scanning Technology
22:10 to 23:18
Discover the technology behind automated page scanning and its applications.
“And okay, it doesn't have everything inside it but it opens up the ability to type a word and then see how that word was trending over time.”
Proton's Approach to AI Models and Memory
23:19 to 28:00
Understand Proton's strategy for using open models and their memory management.
“and existing systems, I don't know exactly what they were doing at Anthropic when they did the same thing.”
Lumo's Innovative Document Syncing
28:00 to 30:40
Learn how Lumo syncs, indexes, and retrieves relevant documents for users.
“and then that dry folder can have all the information that you need to be able to answer questions in that domain.”
Understanding Born Private
30:40 to 33:00
Explore the concept of 'Born Private' and its implications for children's data privacy.
“Describe born private to me, to listeners.”
Data Collection Begins Before Birth
33:00 to 33:40
Discover how data profiles for children begin even before they are born.
“Well, maybe you can explain how that data is collected and merged because we've all had the experience.”
The Ecosystem of Data Collection
33:40 to 38:40
Gain insights into how companies collect and utilize personal data through various online interactions.
“But how, I mean, it surprises me that Google, who, of course, if you have Gmail, has access to your emails, uses that data, the contents of the emails.”
The Impact of Algorithms on Behavior
38:40 to 39:20
Learn how targeted algorithms can influence and change individual behavior.
“I don't know if you heard the story about Molly Russell.”
Privacy in the Digital Age
39:20 to 42:00
Discuss the importance of privacy decisions and their effects on families and future generations.
“It reminds me a little of your visualization work.”
Understanding Privacy for Children
42:00 to 43:31
Explore the importance of privacy for children's online presence and the decisions parents make.
The Future of Social Media and Privacy
43:31 to 45:58
Discuss the potential changes in social media and the importance of privacy awareness.
“so that someone with a born private email address could then participate in social media or participate in search without having somebody build this elaborate, detailed profile of them.”
The Need for Regulation in Digital Spaces
45:58 to 50:09
Understand the role of regulation in protecting users in the digital landscape.
“I think with Proton, before ProtonMail came along, it was kind of difficult to use BGP I mean, for creating secure email, it wasn't that easy at all.”
Proton's Vision for a Private Ecosystem
50:09 to 54:00
Learn about Proton's initiatives to create a secure digital ecosystem for users.
“The whole thing is about increasing active users on the platform.”
BornPrivate: A Resource for Privacy-Conscious Parents
54:47 to 55:42
Learn how parents can navigate privacy for their children with Proton's BornPrivate.
“And the whole idea is to create this ecosystem or workspace where you don't have to compromise on functionalities anymore to be able to use something which is fully private.”
Transcript
Automatic transcript. May contain errors.0:00My wife is convinced that her iPhone listens to the conversations because we'll have a conversation at dinner and then she'll get served an ad. The profiles of your child is being created even before they're even born. There's no real transparency over to say how exactly these profiles are created in the first place. You think your child is safe in their bedroom, but all these companies are basically changing the behavior of your child. Do you think that there will develop an ecosystem around, maybe Proton will be the kernel, around an ecosystem of privacy to counter the Googles and Metas of the world so that someone with a born private email address could then participate in social media or participate in search without having them, having somebody build this elaborate, detailed profile of them?
0:50We're going to talk about Born Private, but Proton has a bunch of privacy-preserving products. I thought maybe you could start by introducing yourself. I know that you've been working in this space for a long time. I believe you have a PhD from Oxford and you worked at CERN. Is that right? Am I wrong on that? Yeah. No, you're not wrong. It's correct. So, yeah, can you introduce yourself and give that background? My background is varied, I would say. I started most of my professional type of career was in bioinformatics. It was basically using computational techniques to analyze the genetic sequences or protein sequences and so on.
1:41And that very much got me into the academic mindset of things. and I went and did a master's in bioinformatics and then a PhD in computer science. Largely focused on also quite a bit of bioinformatics, but I also worked in security applications, a lot of stuff in insider threat, for example. Worked in things with digital humanities, also like analyzing poetry for the way things are spoken and how you can visualize that type of information. So I sort of, my major interest for a lot of my life was working on visualization, data visualization. So the intersection of computer graphics, mathematics, statistics, machine learning later, I would say, and then the computer science aspect of it.
2:34And visualization opened up a lot of different avenues, in fact. So I worked in biology, I worked in security, I worked in digital humanities. I ended up doing my postdoc at CERN I spent a couple of years there then I went and worked in finance for three years taking a lot of what I learned from I would say my previous career cycles into finance also so I did a lot of work applying techniques typically used in biology or bioinformatics but using those techniques in financial time series analysis for example, are looking at different ways of correlating different signals together. And after that, I ended up working at Facebook, where I worked a lot in more machine learning and data science for all external and internal threats.
3:29So everyone trying to get into Facebook's network, preventing that happening. And then anyone trying to exfiltrate data out of Facebook from within Facebook, also stopping that as well. So more detection engineering. And then after that, I came to Proton. So I've been at Proton Life for six years. Okay. I got to ask, the digital humanities is fascinating. You're creating visualizations of language patterns in poetry. Is that essentially what you're doing? So one project, in fact, was looking at how the tongue moves. There's basically different ways. there's a representation about how you create sounds, like from the back of the throat, from the front, from the tongue, and so on.
4:17And basically we had a grid representing that, and then you were able to visualize the changes of this over time. So you could see how different poems were supposed to be read in a more smooth way versus those which are more harsh tongues where you have big dynamic changes in positions, which manifested themselves then as being very large changes in tonality, for example. So it was kind of cool. I like everything, which is interesting. And yeah, we worked on lots of stuff looking at, which is relevant also today for computing side of things, like looking how documents change over time. If you look at things like the Bible, for example, it's many manifestations of copies over many, many hundreds or thousands of years that things were added, things were removed with different languages, so you've got different translations and looking at how these are done also in other types of text as well so like the additions and deletions, which then if you think in computing terms also, it's the same, like looking at how files change over time is a complex thing or how code changes over time.
5:38And using that in security is also interesting. So we ended up using a lot of those things in security, visualization problems as well. Wow, that's fascinating. And how did Proton get started? And what was its original mission? Well, the story goes that Proton got started in the CERN cafeteria. And so Andy was doing his PhD at CERN. I think many of the founders were also still at the time doing their PhD at CERN in particle physics. And it was the time of the Snowden revelations. And, you know, Andy is, of course, he's Taiwanese-American descent. So for him, what was happening in Taiwan, for example, the threat of China, for example, but then also in America, where it's a country where it was purporting itself to be the free democratic sort of flag bearer in some way, and then to understand just how deep the surveillance operations were going.
6:46I think it scared a lot of people at the time, and they decided at that point in time to create something for meal because in the day-to-day, especially in science, I would say, it doesn't seem very scientific, but a lot of the day-to-day operations go around sending emails and all of this IP that is being sent around is suddenly globally available to players like Google or Yahoo, for example, or Microsoft. And a lot of your personal life is attached to these emails. So starting it at email seemed to be a logical step and that's when he started in 2014 and it was crowdfunded. So there was a Kickstarter project that raised, I think, around half a million dollars.
7:36And that went a long way to setting up the company, bootstrapping everything to make it have the servers and have the capacity to handle the users. And then since then, it's funded by our subscribers. We don't have investors or we don't have venture capital. It's basically just funded by the fact that we have paying users and they really believe in the product. They believe in what we're trying to do. And by paying for those subscriptions, they're able to also help bring privacy to other users who are not paying because, of course, Proton is also available for free to users as well so across all of our products.
8:20I didn't realize that. So the business model is there's a freemium offering and then you pay for premium services or something. Yeah, for instance, for VPN, if you're a free user, you get access to certain countries. You don't have very much choice. But if you're a paying user, you've got access to hundreds of countries. I think it's a lot of different gateways you can connect to. and for mail, for example, you'll have more storage. You have the ability to have custom addresses or custom domains. For Lumo, for the product eye control, or I run, basically you've got increased limits, better models and so on by paying.
9:09So we believe in the need to be able to make privacy available to all, but we also need to run a business as well, which is actually paying for all the people who are working on these things and all the infrastructure it requires to run it. And AI infrastructure in particular is quite expensive, so it's difficult to run that on a free model. Unless, of course, you're getting into things like advertising and so on, which is the way that most players are to monetize. Yeah. And we're going to talk about Born Private, this new product or project. And that fascinates me, which is that parents can register an email address for their children at birth and that follows them throughout their lives, right?
10:00But before we get to that, what are the risks, particularly in using AI? I mean, before we got on, we were talking about proprietary models. And I'm going to start paying more attention, but my understanding was that proprietary models, most of them say that they don't use your data for training. I know I've seen that notice on some of them. uh what happens with uh with your models and can you talk about is it lumo is that the name of the uh the of your model uh which is open source but how do you protect uh privacy what makes that different from from other models and if you don't have persistent memory uh doesn't that limit the capability of the model.
11:05Can you talk about that? Sure. So first, most models, so if you go to, as a consumer, let's say, if you go on to many of the big platforms like ChatGPT or Cloud and so on, you'll, by default, typically you give them permission to train on your data. That's their, I mean, If you're a free user, basically they need to make something from your free users, right? Otherwise, why give you access to the platform unless you're not paying for it? If you're a paid user, then it's a bit different. If you're an enterprise user, typically they all have it in the document in the contract saying, we do not use your data to train models.
11:50But then it's also a trust problem, right? Because they do have access to all of your data. And not only all of the conversations that are happening back and forth between the user and the model, but also all the context that is provided as well. And the business practice of these, at least at this point in time, for all these providers, the real valuable thing is data that hasn't seen before, text that has not seen before, because this is what makes the model better. It's not, you can have all these fancy derivations of algorithms or mixture of expert types of architectures and small changes to how things are being processed or attention is being given.
12:36But at the end, the only thing that really matters is compute, so how much you're wanting to throw at it, but also how much data you have. And for them, getting access to data is the most important thing. So that's why opening up to enterprise and education platforms is particularly of interest to these platforms because they want the data it hasn't seen. and within a corporate infrastructure, inside Proton, for example, we have probably hundreds of thousands of documents sitting around that have never been exposed to the public web. And education is the same. So if you have a partnership between, for example, ChatGPT and Oxford University, what do ChatGPT get in return?
13:20Because they give favorable terms to Oxford. And Oxford are presumably also giving back some data to them. But even if they didn't do it directly, all the students are still uploading all of their files into OpenAI servers that they have access to. And then you have to trust that they're not going to use your data. And I think that's a difficult thing to trust whenever you've seen that they basically threw away all sorts of regard for copyright law when they decided just to start scanning lots of books. and you see that with Claude Anthropic as well where there is$1.5 billion lawsuit against them because they basically went along and bought thousands and thousands of books, scanned each of the pages and threw away all the books.
14:12So that there'd be no paper trail because there were no digital trace, but at that time they were going to buy all these books just in the shop. At least they bought the books, I would say. But then what they did afterwards is quite bad. And then you also have the same thing with Meta, for example, where they were found to have used this big archive of pirated books as well in order to train their models. So I think trusting ChatGPT or OpenAI or Anthropic with copyright, they haven't really shown themselves to have much regard for for the law, I would say, at this point. Not to say that others are much different, but there are open...
14:59There are providers of models, such as the Allen Institute, for example, so AI2, developing the open models, which are called ALMO. There's the Swiss AI Initiative also, which has developed the Peritus, which is all based on open data too. NVIDIA are also creating their Nemetron models, which are all based on open data as well. So there are people who are creating open models, but truly open, not just open weights, but also you can see where the data comes from. You can see the code. You can see the process behind it. And I think open models has kind of been, what do you call it, the term has been degraded somewhat.
15:47it's open washing where people say something is open source or whatever, something is open, but actually it doesn't really satisfy the requirements of being open because just one small facet of it is. Yeah. That's like llama, not as llama, is that right? It's open weights, but you don't know where the data comes from. Yeah. That's right. I mean, in the end, And no one wants people to know that because we know, for example, for GPT-2, which is a long time ago now, there is like some crazy stat on the Wired article from around that time, which stated that the total data, so 0.3 % of the training data used for GPT-2 can be accounted for the entire English language section of Wikipedia.
16:39So only 0.3 % of the training data is basically everything on Wikipedia in English language. The rest, script, web pages, social media profiles, whatever else before everything got locked down by Reddit and X and so on. You know, I remember hearing and repeating in the early days of these pre-trained transformer models that they had read the entire English language internet. And I never knew whether that was really true. It sounded unlikely. But what's your view? Is that possible even? Yeah, I think so. So, I mean, with the databases that you have, you've basically got OpenCrawl, for example, that you can go and use.
17:32You can figure out which pages are interesting. You have to filter out all the garbage, of course, because there's a lot of garbage there too, and scam websites and so on. But, I mean, knowing which pages are English language is pretty easy, mostly from the domain name, I would say. And then going off and having giant crawlers getting all that content is also not a big deal. The most impressive thing at the time for companies like OpenAI, for example, where the whole transformer model thing was around for quite some time, but no one really thought it would actually give decent outputs. And the fact that they went off and created GPT-2 with, you know, there's a lot of investment behind it.
18:19There was still like hundreds of millions of investment. so spending the money to go off and build sheep crawlers to go and get the content from the internet was not a big deal so like if you know that the data is the king and the data is the oil for your machine you make sure that you go and secure the oil and that's what all these companies did and all the companies that do the best in terms of having the best models are the ones that manage to acquire the most data yeah Wow. It's so fascinating. I want to get back to Born and Private, but let me ask one more question. There's so much, you know, I do a lot of research and there's so much that's not on the end of this, that's in on paper, that's not digitized.
19:12is do you think at some point all paper archives all over the world will be digitized and then available for training i think probably yes like one one project which is kind of surprising and kind of mocked at one point was open ai's pen you know they're they had this pen thing so you could when you write it takes down your notes and so on this is an interesting concept because in some way I'd seen as being backwards. It's like, why would you create a pen? And the other side, a lot of people still write a lot of their notes on paper. I'm the same, like even to-do lists and stuff like this. I still, there's something to be said about writing something to formulate your ideas better than typing it in the computer.
19:58And sometimes I still do it via pen or paper and pencil. And I think that's interesting. thing then you have a i think a lot of these collections for example the bodden library and in oxford it's one of the biggest archives of of research material i have not confirmed it but i have no doubt that that open ai are looking at trying to get access to that because it's uh it is like literally a treasure trove of information and books and knowledge that many people haven't seen which in that case is kind of a shame as well right the fact that it is locked away is it's kind of a shame um yeah you know google books when it came out google books was amazing as a on one side is google so i have to maybe i have to say it's uh not great but on the other side uh what they did was incredible because you know there is this you know having all this knowledge but in a book that you had to go to a library and that was nice in itself but to find the information was really laborious and you had to go takes a long time to be able to find the right book and the right information then laborious to go to each shelf and take off each book and now just having this tool where you could go and search across all this knowledge and and find the references that you're looking for i remember at the time uh when i was at oxford i you know in the colleges you sit beside like everyone um at the table like different profiles and so on and there was one girl who was doing her PhD on English literature, but looking at how electricity, the advent of electricity, became evident in literature over time.
21:47And she was spending a lot of time in the library, and it was kind of Google Books was around, but it wasn't super well known. and I told her, have you looked at Google Books to type in electricity and see when the trends are but when it first is mentioned and so on. And she said, no, I never heard of it. And then I showed her it, I showed it to her, she said, wow, this is amazing. And okay, it doesn't have everything inside it but it opens up the ability to type a word and then see how that word was trending over time. it would have taken her years maybe to go through the same collective knowledge to be able to have something similar yeah and one last question are there machines that have been produced that can turn pages and scan and there must be some robotic yeah they have I mean I know that some of them are quite quick so you've got ones which are basically like a magic wand.
22:53You pass over the page. So you pass over the page, you pass over the next page and so on. And then I think there's a way then that they can control to understand the context of which page it was. But if you read it properly, of course you have the page number as well. But this is something that's been done automated by Google already for Google Books. and existing systems, I don't know exactly what they were doing at Anthropic when they did the same thing. What I understood was that they basically tore out the pages, but that seems a little bit brutalist for my liking. I would have thought that this tech company would have come up with something more interesting, but maybe Occam's Razor came to effect here and the simple solution wins.
23:45That's right. Literally a Razor, right? Yeah.
23:52So let's talk about, before we get again to Born Private, the model that you're using, that you've built, where the proton is built, if there's no persistent memory, doesn't that make it very limited in its use? so just to clarify we don't use our own model we don't build our own models the reason is primarily financial it would cost us so much money to do that instead we use open models different flavors of open I would say so like Almo was one that we were using at the start Almo 2 we're not using it anymore because it's not you know since July last year when we launched LUMO, the expectations about what these models can do has changed somewhat, I think, over that time period even.
24:53So we're basically keeping up with the frontier models by deploying the best current open models that we can deploy. So that includes things like GLM 5.1, KEMI 2.6, which is released just today, you know, QEN 3.5 we were using GPT OS S120 billion but basically any model that comes along which looks to be performant like the next frontier in performance we will use that. Nematron is another example Nematron by NVIDIA they have released two of the three flavors, so the first flavor was, I think it's I can't remember exactly if it's 20 billion parameter model, a 30 billion parameter model. They have a 120 billion parameter model, which they just released recently, and they will have a 500 billion parameter model as well.
25:55They're using fully open code, open data. It aligns very much with what we want to do as well. If that model is close enough in terms of performance, we'll use that too. then there's Apertus which is the Swiss AI one the version 1 of that was not probably not at the level where we could deploy to users yet it wasn't created for that purpose I should say also it wasn't created as a general purpose chat GPT type of competitor it was created as a base model that was going to be used for different applications so in sciences and in legal context and financial context and so on that could be fine-tuned towards those different application cases, but not as an out-of-the-box ChatGPT competitor.
26:47But their upcoming releases of their plans will also improve that as well. So if it's close enough to something that we can deploy to production that users like, we will be able to run that as well. And for the context side of things, it's more complicated, let's say. So we don't, just because we're end-to-end encrypted for the chat history, for example, doesn't mean that we cannot use the information that's already there, right? So there is no real persistent context. There's different flavors of that. So for example, memory. Memory itself is, it can be implemented in many ways. It can be like a global database that we are pushing to all the time.
27:32it can be something which is local to the user that you can just look up as you're making calls it can be something which is hybrid we keep it locally and then we sync it every so often which is basically what we do so when we sync we keep it locally it's encrypted when we sync it it's encrypted with the user's key so we can't read anything on our servers for Drive for example in projects so Lumo has this projects feature you can link your project with a dry folder and then that dry folder can have all the information that you need to be able to answer questions in that domain. So if you're an enterprise and you've got one team which is focused on some financial stuff you can have all your financial documents loaded there.
28:18What Lumo does then is it goes and looks that folder is now synced. It downloads all those files. It transforms them into text. We index them all locally. So then we know for, we can basically do two types of search. One is just basic keyword type of search. But the second is that you can, you can type in full prompts and it finds the most relevant documents based on the most relevant extracted terms. And then we inject those documents automatically into the context for the prompt and then send those back to the GPU, get the result back. And then the user has their context from their business is within the response because we provided it at a runtime so we do even if i think doing things properly in end-to-end encrypted environments is is is hard but just because it's hard doesn't mean we can't do it so we've already we've been doing it for for mail i think search is is particularly hard because you need the data on the client and not all clients are at the same level and the ability to be able to take all that information.
29:33With projects, we made the decision that if someone links their drive folder, they would link something which is more around the topic that they're looking at rather than, say, link all my drive files, which basically could be everything. It would be a lot to be able to sync onto your machine. Whereas linking a folder which is very narrow or much more narrow, it gives us the ability that we can do that quite well without overloading the user's machine. So we do have ways. And then you've got the whole search thing. Our search is pretty good. So we have web search. We have financial searches with APIs in the background.
30:14Users can enable or disable those as they see fit depending on their threat model. So if they think that they, you know, I don't want the risk of anyone seeing anything I'm typing and any query going to some third party. You turn it off, right? And if you do want it on, you can actually see what it's searched for and what results it got back and so on as well. Yeah. So born private. Describe born private to me, to listeners. It's basically, I mean, the basic premise is more that parents can go and reserve their email address for their kids. I would say the process is that you choose your email address, you donate$1 or whatever you want to the Proton Foundation to support the privacy mission.
31:03It's more of a symbolic donation in some way. You get a secure voucher back to basically a link. And if you want to unlock your account for your child immediately, you can. If you want to wait for 15 years, you can as well. That's it. It's basically, I would say, born private more than being that particular feature, which, I mean, maybe it sounds underwhelming at this point, is more about the act of thinking about your child's privacy. Where does the privacy journey start? Or where does the data collection journey start? Depending on your point of view. So some of the things that we like to talk about is that once a parent starts having emails with their gynecologist, for example, or perhaps for a fertility clinic or for something else, suddenly these systems outside are able to know, okay, these people now want to be parents.
32:04We should already start sending them advertising on maybe some hospitals that are good for birth or some advertisements for some gynecologist or whatever else or pediatricians and so on. And as soon as people start thinking more about how their profiles are being created and how the profiles of your child is being created even before they're even born, I think that part is the most powerful part of the initiative, really. It's more about thinking about your child's privacy, thinking about the privacy of not just yourself, but also the people around you. And also highlighting some, I think, important things about what tech companies and what companies in general are doing, which is not at the benefit of you at all, but only at the benefit of them in advertising.
33:01Yeah. Well, maybe you can explain how that data is collected and merged because we've all had the experience. You know, my wife is convinced that her iPhone listens to the conversations because we'll have a conversation at dinner and then she'll get served an ad. But my argument is, no, your profile is very complicated. You interact with digital systems all the time, and it's all companies, data brokers, collect all of that and create a profile. But how, I mean, it surprises me that Google, who, of course, if you have Gmail, has access to your emails, uses that data, the contents of the emails.
33:58I know they use the metadata, but can you talk about that ecosystem and where all of that data gets collected? and where all of it goes through email, I mean. Through email, specifically through email? I mean, from... Or generally, I mean, that experience that we've all had of talking to your spouse and then suddenly getting ads served related to that. Well, Google is a complex machine because not only do you have your Gmail, You have your Google Home system, perhaps, that's listening for keywords that are going to come up. Like Amazon Alexa has plenty of examples of them listening also to conversations of people as well.
34:53I think there's a lot of conspiracy theories about what's done and what's not done. And those conspiracy theories happen because there's a lot of gap in knowledge, which is also intentional. right there's no real transparency over to say how exactly these profiles are are created in the first place but let's say from from a very simplistic nature and only focusing on email you create an email address and the first thing you do is go off and and sign up for instagram right that tells me already something about your age probably yeah let's say okay then you go in and sign up for some newsletter, some political newsletter.
Read the full transcript
35:39So now I know basically your age and probably also your political leanings. Then you sign up for some book or newsletter on AI. So at this point, now I've got three data points, and then from that I can infer a whole bunch of stuff just by connecting those things. Which people are interested in AI and also have right-wing ideology, for example, and then you find a whole bunch of stuff. Which people are using Instagram or the age of Instagram users, let's say, also interested in AI and also are interested in this right-wing commentator or left-wing commentator or whatever else. And then you start off with those three data points and then you end up with a whole bunch of suggested data points.
36:30and then if you're a place like Google, you can basically say, oh, let's send an advert for this thing that's tangential, right? We think you might be interested in this and then the user clicks on it, right? And say, okay, the user was interested in that. Now you've got a new data point that helps you build a different profile. Or you might create a, send another advert and they don't click on that for things which are not super clear. For example, maybe you don't have any information about their political leanings or say religion, let's say, and you sent an advert for some Catholic church thing, whatever.
37:07Maybe you're not going to click on it anyway. I don't know. But maybe you don't click on it, but then you send another one for a Protestant one. And you say, are they clicking the Protestant one? Okay, now we know their religion too. So Google and other providers like this, which it's not just about creating the profile with what data you give them directly, but also how they change. They're able to sort of interrogate certain questions that they might be having, their own model gaps by basically proposing content and seeing what you click on or say putting a link in a Google search result higher up or lower down and seeing if you click on it or don't click on it.
37:45So changing all these visual cues and changing everything around the version of the internet which is being tailored for you also can have an impact on how you perceive the world, right? So let's say you didn't have politically right-leaning ideologies, but then there was an advertisement came up, which was super interesting and something you agree with. And all of a sudden, you're pushed down a rabbit hole of believing in this ideology, but you didn't have that inclination to start off with. And that's even a more interesting thing because you're starting to change people's behavior and change not only in changing who those people are based on what you serve them and so on.
38:29And I think there is an interesting film. I met the director, in fact, last month it was. The director is called Mark Silver. And the film is called Molly versus the Machines. I don't know if you heard the story about Molly Russell. He was 14 years old. She killed herself.
38:48things severe depression caused by basically she had suicidal thoughts and the Instagram feed kept propagating more and more adverts and more and more content let's say on suicide and in the end she killed herself and the whole thing is about you know you think your child is safe in their bedroom but all these companies are basically changing the behavior of your of your child even if you you're not in the room with them and they're not physically in the room but you know there's no real safeguards to stop uh people having a negative impact on on your on your child's psychology yeah that's that's fascinating and in the way you describe which hadn't occurred to me this profile that different companies are creating of you is kind of a cloud that's constantly morphing through time.
39:46Right. It reminds me a little of your visualization work. So how does Born Private work? Well, Born Private starts by saying, you know, a lot of people's digital identities are anchored in the email address. Right. So why you're getting information from at the start. So if the parents already are using Proton for all the communications with their gynecologists and hospitals and whatever else, basically none of that information is ever going to be used for any of to create any of those profiles. Now, that doesn't stop people from, you know, if you use Proton, but then, for example, use Google Chrome and you use Google Search, well, you've basically negated a lot of the benefits, right, from using something like Proton because on one side of things, yes, you don't have all these things that are coming into your inbox that are tracking you, making sure that you're seeing what advertisements you click on.
40:51we're not profiling at all right so you know google already knows that you're going to see your gynecologist therefore or the and maybe you've got some apartment for some fertility clinic then it already knows that you're planning to do those things so it can start creating advertising profiles if you start going and doing that also on google or on google chrome and so on then they will know that information too. So, you know, born private is more about, yeah, this is one step in the process, but it's more to highlight that we should be trying to increase the sphere, the privacy sphere for individual users.
41:33But also within that individual user, a lot of people make decisions which are also going to impact their entire family tree, right? So, you know, a classic example is like if you go to 23andMe a few years ago and got a genetic test, not only have you basically made a privacy decision on your behalf to give away your data, which can be used for something in the future, but also for all your kids, but also your family, your brothers and sisters and aunts and uncles, you've basically given away part of their privacy as well. so understanding basically if born private does anything it's about trying to increase the understanding of what privacy means where it starts reserving an email address is one part of that but it gets the conversation going in people's heads how you want to deal with your kids' privacy online so I have a daughter there are no pictures of her online every time the school asks the school does ask at least do you want the photographs to appear in or even internally I say no my family tried to my sister for example will try and take photographs and put them on Instagram or whatever I say no you can remove her from the picture we might appear in the pictures because I think our day is gone probably but for her we've made the explicit decision that we don't want anyone that we don't know seeing pictures of our daughter.
43:16And if you see what happens later on with deep fakes or the potential futures for how all that information can be misused, I think people would take a second look at their decisions and maybe not do that. Yeah. Do you think that as awareness of privacy and the negative impacts of lack of privacy spread through society, that there will develop an ecosystem around, maybe Proton will be the kernel, around an ecosystem of privacy to counter the Googles and metas of the world? so that someone with a born private email address could then participate in social media or participate in search without having somebody build this elaborate, detailed profile of them.
44:25So, I mean, at the minute, I would say maybe there are private social media systems, but I'm not sure if they're any good or if anyone's using them. You know, platforms like Facebook and Instagram and these types of tools, they're basically a TikTok or whatever. they're basically the the dominant force and and many people have made the decision that either they don't recognize the risk and they've their decision making has been basically zero they just fell into the crowd of peer pressure you know your friend does it therefore i have to do it too if you don't do it then you're left out i think a lot of people are in that boat I think maybe social media as it stands today will not be the social media that stands in a year from now or five years from now, especially with all the pressures that are coming on regulation for how people are treating people's data.
45:31And the example I give for Molly Russell, for example, is a big part of that type of drive. But people typically gloss over danger for convenience. but that's not something which is as unique to a privacy space. It's probably for almost everything, right? If something is easy, people are probably just going to do it. If something is a bit harder, then there's going to be a bit more resistance to doing it. It's easy to follow the crowd. It's harder to reject. I think with Proton, before ProtonMail came along, it was kind of difficult to use BGP I mean, for creating secure email, it wasn't that easy at all.
46:15And Proton came along and then made that whole process easy and transparent. The idea is to get to a point where privacy is not just the default, but it's also easy for people to do. and sometimes it's a long process because there will be unique cases unfortunately which show what can really go wrong it's the same what you get from data breaches it's easy to set up a website or to vibe code a project these days or whatever else it's still hard to do security to do security well and so people just ignore More security is the first stage. Then they're breached. And now suddenly all the records for all their customers or everyone that's been on their platform is now leaked online.
47:07And they say, oh, that was a mistake. I should have invested more time on the security side. People, there's drip thread, I would say, data leaks in the general population. And people are more wary of it, you know, for identity theft in the US, for example, is super easy. If you have just a small piece of information, less easy in Europe because you can't just create bank accounts so easily in Europe with someone else's name. But, you know, all these things are, if you make it easy for people to have a private ecosystem where they don't feel like they're left out because they don't use a certain tool, they're overall safer for it.
47:58They have no manipulation. You don't have cases where companies are trying to modify people's psychological profiles in real time. If we can get people to think more like this, that would be a success. Ideally, they'll come to Prouton and use our products where possible. I don't see a day where we create a social media platform. email is kind of like social media let's say you send messages, you've got connections and groups and so on but I also don't think maybe social media is not going to be super popular in a decade from now because of all the risks but on the other hand there is a positive side to this profiling because you it does make you aware of things you are served ads of things that you are interested in that you wouldn't have found before is is there a way of of managing that so that it's not abused well transparency is kind of is the important part of it so i mean they do something on this like you say why did why did I see this ad type of thing does happen on particular profile or platforms but you know the regulation call as it stands is happening because we've long treated social media as some sort of utility function rather than being a platform that was basically abusing the users, keeping them hooked the same way as you have for cigarettes or alcohol or gambling.
49:52I think gambling is the best example, and I brought this up a while back. I think now I've read more that people are using that example also in public discourse, but gambling systems have regulation in place because otherwise, what's to stop the gambling company from just taking everyone's money right and if you don't have these regulations that people abide by then you know the population would be miserable you'd have a bunch of people just taken for granted they would probably be higher divorces and mental health problems and so on but if you look at the effects of social media it's it's largely the same right you've people who are addicted to it they scroll all day because of the things that have been you know engineered to make people stay on the platform.
50:40The whole thing is about increasing active users on the platform. So increase engagement. The more engagement you have, the more time they spend on your platform, the more opportunity there is for shareholders, and they will invest more, so stocks will go up. Basically, the reward system is broken. There is no real incentive for social media companies to try and limit usage because it would damage themselves. So I think when these types of things happen where the user health and the user focus is diverging from the platform goals, then you need regulation to bring them back together. And I think that it's fair that people are talking about it now because for a long time, people just said, yeah, let them self-regulate and so on.
51:31But clearly that hasn't worked. Yeah, and also our digital lives have overtaken so much of our real lives. In the beginning, you didn't spend that much time in the digital world, but now we're, unfortunately, we're all doing what we're doing right now, speaking through a digital medium. um the uh proton so proton is kind of a prototype born private is kind of a a starting point at least you know with uh a proton uh email address that the content of your email emails are not contributing to this profiling uh although yeah uh how does somebody get, and born private is a great concept, but presumably anybody could get a Proton email address.
52:38Yeah, just go to protonmail.com. Is it protonmail.com? I'm going to check just to make sure. Yeah, protonmail.com. It's been a long time since I actually had to go to the website. Yeah, protonmail.com and create a free account, or pay also. Paying is good. It gives us jobs and so on. But yeah, protonmail.com for create your free account, and that gives you then the access into the rest of the ecosystem. You've got access then to VPN, to Proton Drive for your document storage or also photo storage. We have Proton Pass for your password manager. We have Docs and Sheets. We have Lumo for the AI assistant Meet, which we launched, I think, two weeks ago, three weeks ago, which is our video conferencing solution.
53:36So you can keep everything private there too. And more. So it's pretty easy to set up. Yeah, that's fascinating. You can build, I mean, as you were saying, you're not going to create a social media platform, but you can build a digital ecosystem through Proton that at least within that ecosystem, your privacy is protected. Yeah, and we just released Proton Workspace, which is also our plan, and more for B2B focus. But the idea is that basically companies can go along and they will have everything that they need to be able to function as a company by using Proton products. So all of your IP stays with you.
54:24You don't have to worry about confidential data getting lost somewhere. If you want to use AI assistance, you can use it within ProtonMail. We also have Scribe, which is the first AI integration we did, which is similar to what you have within Gemini now for Gmail. And you can use Lumo for everything else and connect it to your documents. You can ask questions about it and use it for your productivity gains, hopefully. And the whole idea is to create this ecosystem or workspace where you don't have to compromise on functionalities anymore to be able to use something which is fully private. And for BornPrivate, that's also it.
55:11Parents can find that on the Prouton website. Is that right? Yeah, if you search for BornPrivate in your favorite privacy-friendly browser, then you'll be able to find it. But basically, prouton.me slash mail slash bornprivate. Okay, well, why don't we leave it there? That's all fascinating. Transcription by CastingWords
From the publisher
Your child's data profile doesn't start when they get their first phone. It starts before they're born, the moment a parent emails a gynecologist or visits a fertility clinic website. That's the core argument behind Born Private, Proton's new initiative that lets parents reserve an email address for their child at birth, anchoring their digital identity in a privacy-preserving ecosystem before the profiling machine gets started. Craig Smith sits down with Eamonn Maguire, Engineering Director, Machine Learning & AI at Proton, who has spent his career at the intersection of data, security, and visualization to explore what's really happening to our data and what, if anything, we can do about it.
The conversation covers the mechanics of how just three email sign-ups can allow Google to infer your age, politics, and religion; why OpenAI and Anthropic have shown "not much regard for the law" when it comes to training data and copyright; and why social media platforms are operating like unregulated gambling companies - engineering addiction with no structural incentive to stop. It's one of the most grounded, specific, and genuinely alarming conversations about digital privacy you'll hear, and it ends with a simple, actionable proposition: privacy should be a decision you make at birth, not a problem you try to solve after the damage is done.
Subscribe to Eye on A.I. for weekly conversations with the people building and deploying the future of AI.




