In short
Snap’s experimentation data pipeline at massive scale (10+ petabytes/day) and how it accelerated Apache Spark workloads using NVIDIA Spark Rapids and GPU capacity borrowed from idle online inference GPUs, including fallbacks from GPUs to CPUs to DataProc.
Guest backgrounds
Prudhwi Vatala is Head of Engineering Platforms at Snap (about 7–8 years at the company), leading big data infrastructure, developer productivity, and enterprise AI.
Key claims
migration cut job costs ~76%, reduced required CPU cores ~62%, lowered memory footprint ~80%, and eliminated ~120 TB of disk/memory spill; production rollout from prototype to production in ~8–9 months.
Notable examples
hourly guardrail pipelines migrated first; daily pipelines as statistical authority; A/B testing methods (heterogeneous treatment effects, variance reduction, sample-size mismatch); benchmarks showed ~3x+ for join/repartition/shuffle jobs, ~2x for union, ~1.5x for aggregations; zero code changes with Spark Rapids; used NVIDIA Ether for consistent Spark tuning across GPU/CPU/DataProc environments.
Guests
Noah Kravitz (host) and Prudhwi Vatala (Snap).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOImpact of Migration on Job Costs
0:00 to 0:25
Learn how Snap reduced job costs and resource usage through migration.
“We were able to cut almost about 76 % of our job costs as a result of this migration.”
Understanding Snap's Role and Scale
0:46 to 2:18
Explore Snap's position in the tech landscape and its user engagement.
“Thanks so much for taking the time to join us.”
Accelerating Data Processing at Snap
2:18 to 4:34
Discover the challenges and strategies for handling massive data scales.
“Like, as you can imagine, with as many users as we have, and Snapchat in particular is a very complex application.”
The Importance of A-B Testing
4:34 to 7:11
Learn about A-B testing's role in decision-making for Snap's features.
“Like, is it regressing their performance?”
Integrating NVIDIA Technology
7:11 to 10:57
Find out how Snap adopted NVIDIA Spark Rapids for data processing.
“So with all of the data processing every day, what made you think that maybe some NVIDIA tech put into the stack might help things out?”
Innovative Solutions for GPU Capacity
10:57 to 14:02
Learn how Snap optimized GPU usage for their data architecture.
“So I'm into developer productivity and developer enablement.”
Building a GPU-Accelerated Platform
14:02 to 19:08
Learn about the challenges and solutions in developing a GPU-accelerated data pipeline.
“So with all of that in mind, we built out a platform ground up.”
Impact of Partnership on Job Costs
19:08 to 19:20
Discover how collaboration led to significant cost savings in job processing.
The Evolution of Snap and Social Media
19:20 to 22:36
Explore the impact of technological changes on Snap's evolution and its role in social media.
“And in terms of the roadmap, it definitely had an impact.”
Transcript
Automatic transcript. May contain errors.0:00We were able to cut almost about 76 % of our job costs as a result of this migration. 76? 76. It's phenomenal. I mean, for the engineers out there, we were able to cut down the number of cores required by like 62%. Amazing. The memory footprint, we could drop it by like 80%. So phenomenal results. The results speak for themselves.
0:25NVIDIA AI Podcast Host:Welcome to the NVIDIA AI Podcast. I'm Noah Kravitz. I'm here with Prudhwi Vatala. Prudhwi is the Head of Engineering Platforms at Snap, and we're here to talk about data processing, and in particular, how a social platform with more than 940 million active users accelerated their data pipeline. Prudhwi, welcome to the NVIDIA AI Podcast. Thanks so much for taking the time to join us. Yeah, thanks for having me here, Noah. So maybe we can start with the basics. Tell us a little bit about, well, about what Snap is now. I'm old, but I still think of it, you know, the snap glasses and everything, but Snapchat, obviously a huge social platform.
1:03NVIDIA AI Podcast Host:So maybe tell us a little bit about Snap and then your role there. Absolutely. Yeah. I mean, Snapchat at this point is pretty much a household name. You know, it's Snap as a company. It's interesting that you bring up the spectacles because Snap as a company believes that camera is at the center of, you know, improving how people communicate and improve their lives. in the digital world, so to speak. So we've been steadfast on that belief. And Snap right now is at the intersection of augmented reality, AI, and visual communication. Like you said, serving close to a billion monthly active users.
1:45I've been at Snap for a while now. And I lead a multifaceted organization. We do a little bit of it has to do with big data infrastructure, a little bit of it with developer productivity, and a little bit of it with enterprise AI and whatnot. So yeah.
2:06NVIDIA AI Podcast Host:And so when we talk about accelerating data processing, what does that mean to you? What does that mean for Snap? And thinking about the scale that you operate on, just talk a little bit about what it means to accelerate data at that level. Absolutely. That's a great question. Like, as you can imagine, with as many users as we have, and Snapchat in particular is a very complex application. So you can imagine the scale at which we operate. Especially on the data processing side, we are dealing with my team's experimentation platform is dealing with 10 plus petabytes each day. It's a massive scale, right?
2:43It's a huge scale, yeah. And then we have a strict SLA in the morning because experimentation results need to be ready for developers, product managers, data scientists to act on as early as possible so that they can take appropriate actions. So for us, accelerating data processing basically means instead of throwing more and more CPUs at the problem, figuring out a way to flatten that scale curve. So in this particular scenario, it was about figuring out how to leverage GPUs for improving our workloads, making sure they run faster, cheaper, and scale linearly or sublinearly, unlike right now.
3:27It's definitely super linear with feature areas. So that's what accelerating.
3:33NVIDIA AI Podcast Host:So you mentioned experimentations. What does that mean when you're conducting experiments at Snap? What does that look like? And then maybe how does that fit into, is that where the 10 petabytes of data each morning comes from? Or we can talk about that. Yeah, absolutely. So this 10 petabyte data is only about the experimentation platform. The big data across Snap is far wider. Sure. So experimentation, it's a little bit about Snap's product philosophy. We believe that experimentation, safety, and privacy are core pillars for our product development and iteration. When we are thinking about new product areas, when we are shipping new product features to our half a billion daily active users across the globe, we need to think about how the users are receiving it, how they're responding to it, how they're using it, whether or not this is adding value.
4:33their daily lives and also guard railing things. Like, is it regressing their performance? Is it causing their devices to slow down? We need to be very particular about protecting their experiences as well.
4:49NVIDIA AI Podcast Host:And so, Prue, along those lines with the experimentation, can you talk a little bit about the importance of A-B testing? So, A-B testing is, you know, the concept of randomized control trials has been around for a long time, you know, especially in the clinical fields and whatnot. But with the digital revolution, it has become the mode of bringing statistical rigor to decision making at scale, right? So, that's what A-B testing adds to us. When we are dealing with this massive user base that is diverse by nature, from all walks of life, across the globe, and we are trying to delight them, we are trying to bring experiences to them, we need to make sure what we are delivering is buttoned down.
5:40Like it's actually really adding value the way we think it is. And at this scale, a lot of things can happen. And that's where having the statistical rigor grounded in holdouts and well-defined controls and statistical methods comes in. Like over the years, my team has added a bunch of statistical methods to our platform. you know, heterogeneous treatment effects detection. For example, you know, you may think that feature is performing well for the global audience, but it may not perform so well for a subset. So figuring out those heterogeneous effects is one thing that we focus on. And, you know, at this scale, no matter how you slice your experiments, you're still allowing some bias to seep in, as in, you know, some power users may end up on one side of the experiment rather than the other.
6:42So how do we make sure the distributions are evened out when the experiment results are read? That's the variance reduction aspect. So that's something my team built over time. And then, you know, sometimes when we ship a feature, if people don't like it, they might even just stop showing up, you know?
7:01NVIDIA AI Podcast Host:Right, right, right. So that's the sample size mismatch problem. So we also do a bunch of that rigorously. So that's what A-B testing brings to the table. So with all of the data processing every day, what made you think that maybe some NVIDIA tech put into the stack might help things out? How did that process start? And maybe you can talk about what you've integrated and what you're using. Absolutely. So I'm really proud of this. I'm really proud of my team because over the years that have been seeing our platform, the number of users grew, like Snap, you know, ballooned, right, in terms of footprint.
7:39The number of features we shipped, like, you know, spotlight, you know, AR features, AR lenses, and all of the AI features we shipped in the recent past. So they've also been adding a lot of additional dimensions to the platform. And my team was hard at work making sure we are not,
7:59we're scaling appropriately even as all of this scale grows. And they've done a very good job of it historically for years now, maintaining the cost flat and performance predictable, meeting the SLAs and whatnot. And one thing we came across, we came across NVIDIA Spark Rapids on one of the blog posts and we saw NVIDIA is shipping this solution to speed up our PySpark workloads by anywhere from 3.6x performance versus 50 % runtime. It was phenomenal on paper.
8:39NVIDIA AI Podcast Host:So that's what drew us to it. I'm waiting to hear them. The numbers sound good. I'm waiting to hear the rest. Yeah, yeah. So we read those and we got super excited. And then our stack was, it still is entirely Google Cloud for experimentation platform. We loved working with them. The Google Cloud data proc was phenomenal. They've been a fantastic partner to us throughout the scaling journey. So when this... It's great to hear. Yeah. And then when this news came out with Spark Rapids, we wanted to try it out. We did a bunch of benchmarking. We tried, obviously, like I said, we do a lot of things.
9:19So there is a lot of complexity to the nature of the jobs we run. So we had to benchmark each kind of job as well, like taking jobs that are heavy with joins and repartitions and shuffling of data that moves data around versus jobs that are purely unioning data from various places versus jobs that are purely aggregating, like running sums and whatnot. not. So we had to benchmark across all of them. And we noticed that even on Google Dataproc with Spark Rapids, we got about, you know, I want to say 3x plus, you know, improvement for the, you know, joint jobs, and about close to 2x for, you know, the union jobs, and a little over 1.5x for aggregations.
10:13That's largely because CPUs are already good at aggregation. And then the other thing is GPUs by nature support parallelism and high bandwidth memory on the hardware itself. So that made it like a very good candidate for us to pursue.
10:30NVIDIA AI Podcast Host:And so you're running your GPU accelerated pipelines on Google Kubernetes. Is that right? Yes. Yes, that has been a very interesting journey from testing out our pipelines with Dataproc GPUs to today. And one other thing, like with Spark Rapids, I want to mention it, we didn't have to change a single thing about how we ran the jobs. That was the beauty of it. Zero code changes. Oh, that's amazing. Zero code changes. So I'm into developer productivity and developer enablement. So for me, that was music to my ears. Sure, of course. So that was very impressive. So with Dataproc, which abstracts out the Spark runtime for us, and Spark Rapids, which didn't require us to change the jobs, it was phenomenal.
11:15NVIDIA AI Podcast Host:Yeah, amazing. So it went very well. So we wanted to productionize this. We were able to, at our scale, pipelines aren't just monolithic, right? We do a bunch of sharding and then, you know, batching of work. So we were able to migrate one shard to production on Google Dataproc using 300 GPUs. The results were phenomenal. And then in the next phase, we wanted to migrate 10 shards for total 50 plus shard architecture. And then it needed about 3 ,000 GPUs, which was still doable with data proc on-demand GPUs. Because GPU capacity is on everybody's mind these days. So that was well and good. But then we didn't have a path forward after that.
12:00So we kind of hit a roadblock with on-demand GPU capacity. So we had to get creative. So we started looking around. We were like, where at Snap do we have GPU capacity that we can borrow? And that's where the real insight came for us. Snap has a global audience, and the Snapchatter's behavior is cyclical during the day. People wake up, they use Snapchat, and they go to bed, they don't. So what that meant was when some of our biggest markets went to bed, a lot of our online inference GPU capacity was sitting idle. Somewhere between 1 a.m. and 5 a.m. So that was our opening, our opportunity to go tackle.
12:47And that brought about its own set of complexity. Because online serving stack is not built for batch data processing. They were considered fundamentally different words. So all the online GPUs were tied to Kubernetes and GKE. And we were already on Google Clouds. GKE wasn't an issue for us at all. It was actually very welcome. So we had to migrate our workloads to Kubernetes-based Spark runtime and host it on GKE so that we can leverage what the online GPUs had to offer. and for that we had to actually build a data platform ground up. Okay. You know, because it's one thing for my team to just use this idle capacity, but at Snap we wanted to make sure even as the online need for GPUs increased, as our AI footprint increased, we should still have any team at Snap be able to leverage that capacity for any of their needs as available.
13:48And then we had to also acknowledge that if a user wanted to see fresh spotlight content, it supersedes GPU need for experimentation. You know, preemption had to be built in. So if we had a sudden spike in traffic, we had to give up GPU capacity. So with all of that in mind, we built out a platform ground up. And then we started migrating. And we had a lot of blockers along the way, and the team got really creative. It was a phenomenal journey. Amazing.
14:23NVIDIA AI Podcast Host:Yeah, yeah, yeah. And so you're also running an accelerated Apache Spark pipeline? Yes, yes. So a lot of our pipelines, at a high level, our pipelines are split into daily and hourly cadence. So hourly is mostly for guard railing, like I said. Like, you know, we don't want to break users' experience no matter what. And having that hourly feedback cycle goes a long way in doing that. And then we also have daily pipelines, which serve as the statistical authority for decision making. So our first migration to GKE plus NVIDIA Spark Rapids was the hourly pipeline. Because, you know, speed mattered there far more, right?
15:10So we migrated. And then we migrated and operationalized it. And during that process, we ran into a few corner cases. If the GPU capacity wasn't available at 11 a.m. when everybody was active on Snap, what do we do? So we had to figure out how to gracefully fall back from GPUs to CPUs. And then if the shared GKE resources itself was the constraint, then we had to gracefully fall back from CPUs to data pro clusters. So building all of that with operational reliability in mind was also great. Yeah. Looking back on it, what learnings would you, you know,
15:52NVIDIA AI Podcast Host:if there's a listener out there who's embarking on a similar project or trying to figure out, maybe there's a, you know, like you said, kind of a daily cycle of when the GPUs are in use for inference and when they're not, they're thinking about, you know, borrowing GPUs from other parts of the company. Learnings you would share from this whole process? Is there a big takeaway, something that surprised you? Right, right. So the direction that NVIDIA is headed in is phenomenal for these kinds of needs. You know, NVIDIA Spark Rapids, like I said, zero code written. Yeah. Zero code changed. Amazing.
16:28To enable it. We had to figure out the image building and environment difference and whatnot. The testing cycles, obviously, any production workload needs to go through that regress rollout process. So everybody needs to pay attention to it. But this is a real possibility, you know, the NVIDIA direction. The other thing that NVIDIA offered that really helped us a lot was NVIDIA Ether. It's another solution that gives us Spark tuning out of the box. Because especially when we had this fallback mechanism in place, where we had to go from GPUs to CPUs to data proc, the environments are different, the Spark parameters had to be different.
17:12So something like NVIDIA Ether giving us a starting point and making sure the tuning stayed consistent across all of these versions was also very helpful.
17:24NVIDIA AI Podcast Host:So you've mentioned, obviously, the work with NVIDIA and Google Cloud as well. Kind of from taking a step back, sort of bigger picture, what are these partnerships and working hand-in-hand so closely with Google Cloud, with NVIDIA, what is that doing to the way that you and Snap see your roadmaps for both data and AI kind of growing going forward? Yeah, it's, I mean, huge props to the NVIDIA team and the Google Cloud team, honestly. It's been a phenomenal three-way partnership like I've never seen in my career before. It was phenomenal. And the impact speaks for itself. We were able to cut almost about 76 % of our job costs as a result of this migration.
18:0976? 76. It's phenomenal. I mean, for the engineers out there, we were able to cut down the number of cores required by like 62%. Amazing. The memory footprint, we could drop it by like 80%. I mean, for the Spark nerds out there, we were able to cut out almost 120 terabytes of disk spill, disk and memory spill from our pipelines. Wow. Just vanished once we started doing all of this. Yeah. So that is one of the biggest headaches any data pipeline at scale runs into. So phenomenal results. The results speak for themselves. So without the partnership, this would not have been possible in the timescale that it was possible.
18:52Like migrating a production pipeline with 10 plus petabytes from prototyping exploration to full production in a matter of about eight to nine months is phenomenal. And without the continuous back and forth and knowledge sharing and partnership across these three companies, this wouldn't have been possible.
19:18NVIDIA AI Podcast Host:That's great. Yeah. And in terms of the roadmap, it definitely had an impact. Like I said, my team built this bottom-up data platform to enable any team at Snap to leverage the GPU capacity and what NVIDIA libraries have to offer. And we're already seeing movement with it. Even my own team started migrating other things that we haven't even tried out so far, experimenting with them, trying out. because even if we don't have ideal capacity to fit all of our workloads all the time, if we can schedule things creatively, if we can move things around, we can maximize the capacity as much as we can.
20:00And a lot of other teams are also picking this out.
20:02NVIDIA AI Podcast Host:Yeah, that's fantastic. So you've been at Snap for eight years, is that right? Seven? Close to eight, okay. And Snap's been around for about 15 years, give or take? Yes. Um, working at a social media, a huge social media platform, um, over this span of time where social media has just, you know, become such a, such a core part of the fabric of so many people's lives. Um, what's it been like to be at Snap and to see the changes both, you know, I said at the beginning, right. I remember the spectacles as my first thought of Snap and obviously now Snapchat, you know, same, same lineage, same philosophy, different product, obviously.
20:43NVIDIA AI Podcast Host:Right. But what's it like to just have seen the evolution of social media and then also so many technological changes that impact, you know, what you're able to do and how you do it, as you were just describing? What's it been like from the inside? Yeah, it's been it's been unbelievable of an experience. Noah. That's what gets me up in the morning every day. Snap, in the visual communication, AI landscape, Snap has had a massive impact on the planet, honestly. And having a direct role to play in it is a great feeling. I've seen the company grow from the camera messaging messaging, you know, picture messaging to what it is today, AR stories, which is something we invented and the whole world, including some newspapers.
21:47So the stories as a format. And then to your point about spectacles, we did it before anybody else was even thinking about it, you know. So the company is innovative. We come up with so many new things and running platforms inside means that I have to figure out a way to enable all of this, even as the company evolves. And that's been having a front row seat to that evolution and playing a big part of it has been very fulfilling.
22:17NVIDIA AI Podcast Host:Fantastic. Proof for listeners, viewers who, there are some out there who haven't used Snapchat before, for anyone who wants to get the experience, but also to learn more about Snap and maybe about some of the technical work that you're doing. Are there obviously the website, their social media, is there a research blog? Where can where can people go? Absolutely. So we have an engineering blog that's pretty active. We share a lot of phenomenal work that engineers in the company are working on. And, you know, we are also participating in events like this and sharing our knowledge with the world.
22:54So, you know, and and Snapchat, if you haven't used it, you should definitely give it a try. It's different from social media.
23:05NVIDIA AI Podcast Host:This is a true story. I got a Snap from my younger son maybe 45 minutes before we sat down to do this, and it made my day. So absolutely, if you haven't. Pruvitala, thank you so much. This has been a great conversation, and I'm sure the developers, the engineers, and the audience hopefully have taken a lot from it. But thank you so much for taking the time to join us, and all the best to you and everybody at Snap to keep changing the world for the better. Thank you so much, Noah. Thanks for having me. Appreciate it.
From the publisher
Snap processes more than 10 petabytes of experimentation data every single morning—and with NVIDIA GPU-accelerated Apache Spark on Google Cloud, Snap cut job costs by 76%, reduced memory usage by 80%, and eliminated 120 terabytes of disk spill from its pipelines.
Prudhvi Vatala, head of engineering platforms at Snap, joins the NVIDIA AI Podcast to break down how he and his team completely modernized data infrastructure for a social platform serving nearly a billion monthly active users—using NVIDIA cuDF plugin (formerly referred to as NVIDIA RAPIDS plugin) for Apache Spark on Google Kubernetes Engine, with zero application code changes.
🔬Topics covered:
How Snap runs A/B tests at planetary scale using rigorous statistical methods like heterogeneous treatment effect detection and variance reduction
Why Snap reuses idle inference GPUs between 1–5 a.m. for batch data processing—and how it built a Kubernetes-based platform to do it
How NVIDIA cuDF delivered 3x+ speedups on join-heavy Spark jobs with no code rewrites
The full business impact: 76% cost reduction, 62% fewer cores, 80% less memory, 120 TB of spill eliminated
How a three-way partnership between Snap, NVIDIA, and Google Cloud made it possible in just 8–9 months
Chapters:
0:00 Introduction and Snap overview
3:35 What is Snap’s experimentation platform?
4:05 Why experimentation, safety, and privacy are core at Snap
4:52 How A/B testing works at billion-user scale
8:14 Discovering NVIDIA cuDF plugin
9:06 Benchmarking results: join, union, and aggregation jobs
12:00 Reusing idle GPUs overnight via GKE
13:24 Building a bottom-up GPU data platform at Snap
17:48 Results: 76% cost reduction and partnership impact
20:56 Snap’s evolution and what’s next
Learn more:
NVIDIA cuDF: https://developer.nvidia.com/topics/ai/data-science/cuda-x-data-science-libraries/cudf#accel-apache




