In short
Network forecasting and data science leadership. Dr Judith Guimera-Busquets explains how air-traffic demand forecasting differs from single-route forecasting because network changes (airport pair links added/removed) cascade across the system. She outlines a multi-stage framework: (1) econometric city-pair demand generation using socioeconomic variables, (2) network evolution modeling that predicts link addition/removal using network-theory topology metrics (e.g., node degree, eigenvector centrality, clustering coefficient, weighted degree; power-law-like connectivity in the US), (3) itinerary assignment (nonstop/one-stop) and (4) segment aggregation to estimate flight operations.
Key claims
historical-data models break under structural shocks (pandemic, airspace closures, volcano); scenario planning and human-in-the-loop adjustments (e.g., removing inoperative connections, adding dummy variables) are essential.
Notable examples
COVID-19, geopolitical airspace closures, and a past volcano disruption.
Guests
Dr Judith Guimera-Busquets, Head of Data Science at DataSpark; PhD (City University London) on predicting US air-traffic network growth and connectivity cascades.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Role of Expertise in AI Predictions
0:00 to 0:18
Understanding the balance between domain expertise and predictive tools in AI.
βWe build the solutions to kind of like help us to predict the future, but I think we shouldn't be removing the expertise and domain expertise, judgment from those tools and from the math, like they complement each other.β
Introducing Dr. Judith Gamera-Busquets
1:00 to 2:26
Profile of the guest and her work in data science and forecasting.
βJudith Gamera-Busquets, Head of Data Science at Dataspark, where she oversees data science delivery across clients in aviation, logistics, supply chain, and beyond.β
The Challenge of Air Traffic Forecasting
2:26 to 2:52
Exploring the complexities of forecasting demand in aviation networks.
βI want to crack on with our sort of first question on your research, which was around air traffic forecasting in, obviously, in aviation.β
Network Effects in Forecasting
2:52 to 4:21
How changes in one part of the network affect the entire forecasting model.
βdemand forecasting, I suppose we're talking about here, in a network different and quite tricky than other problems?β
Historical Data and Aviation Networks
4:21 to 6:04
Discussing how historical data impacts future predictions in aviation.
βor a single product forecasting than the network itself.β
Analyzing Network Structures
6:04 to 8:02
Examining the topology and characteristics of aviation networks.
βand doing your research was one, I mean, there was a nice graph in it I saw, which showed the sort of distribution of connectivity.β
Modeling Connectivity Changes
8:02 to 9:36
Understanding how to model changes in aviation network connectivity.
βSo did you look at characteristics like power law distributions on connectivity?β
Dynamic Forecasting Framework
9:36 to 13:23
A look into the dynamic framework for forecasting passenger demand.
βAnd I used them for the, actually for the connectivity model.β
Data Sources in Forecasting Research
13:23 to 14:00
The types of data used for air traffic forecasting and its importance.
βknowing the granularity, the granularity of...β
Aviation Data Analysis and Economic Impact
14:00 to 17:23
Explore the use of publicly available airline data to predict flight frequency and passenger demand.
βAnd then you will move on to the next step, which will be the following year, and so on.β
Show all 17 chapters
Challenges of Predictive Models During Crises
17:24 to 21:05
Understand how unforeseen events like the pandemic can disrupt predictive models in aviation.
βObviously, there's something which was completely unanticipatable really when it did strike in 2019, 2020.β
The Importance of Human Expertise in Modeling
21:06 to 25:36
Learn why integrating human judgment is crucial for effective data modeling and forecasting.
βThe probability of that being forecasted is just like, yeah, like your model predicting that this is going to happen is just not, yeah, it's really low.β
Engagement and Support for the Podcast
25:37 to 26:06
A brief interlude encouraging listeners to engage with the podcast and provide support.
βAlgorithm, please take a moment to follow the show and leave us a review on Apple Podcasts, Spotify, or YouTube.β
Core Capabilities for Effective Data Scientists
26:07 to 28:00
Discuss the essential skills and mindset needed for data scientists to thrive in diverse projects.
βSo you're currently head of data science at DataSpark, a professional service company.β
Agile Development in Data Science
28:00 to 30:06
Learn the importance of agile methodologies and simplicity in data science projects.
βAnd it's much easier to add explainability layer into a linear regression model than neural networks and so on.β
Understanding User Needs for Data Solutions
30:06 to 31:37
Discover how to engage end users early to ensure data solutions are effective.
βIf the output is going to be just in a folder where no one can access, then what's the point of doing the solution?β
Trends in Applied Data Science and AI
31:37 to 33:16
Explore emerging trends and the shift toward specialized models in AI.
Transcript
Automatic transcript. May contain errors.0:00We build the solutions to kind of like help us to predict the future, but I think we shouldn't be removing the expertise and domain expertise, judgment from those tools and from the math, like they complement each other.
0:18Welcome to Inside the Algorithm, the show that goes beneath the surface of artificial intelligence into the scientific research, technical breakthroughs and academic thinking shaping where this technology is actually heading. I'm Jeremy Bradley, Chief AI Officer at Canberra Spark, a leader in transformational data and AI upskilling, career development and progression. Each episode, I sit down with the researchers, data scientists and technical experts working at the true frontier of this field to explore not just what's happening in AI, but also the real science and thinking driving it forward.
0:50Whether you're building sophisticated AI systems, leading teams through AI transformation, or driven to understand what's really happening beneath the headlines, this is the show for you. In today's episode, we're exploring two rather different angles when it comes to successfully delivering data and AI projects, the mathematics of forecasting in complex networks, and what it actually takes to build and lead a data science team that delivers. My guest is Dr. Judith Gamera-Busquets, Head of Data Science at Dataspark, where she oversees data science delivery across clients in aviation, logistics, supply chain, and beyond.
1:26Her PhD from City University of London focused on developing advanced methods to predict air traffic network growth, building a framework that models not just demand, but also how changes using connectivity across an entire network cascade from one node to the next. That academic foundation has shaped everything that she has done since, from working on some of the most ambitious data science projects, to thinking deeply about what separates data science teams that actually deliver from those that don't. In this conversation, we get into the technical architecture of network forecasting, the question of what she would do differently on that same problem today, with access to modern AI tools and her views on what great data science leadership really looks like.
2:10I do hope you enjoy it. Very warm welcome, Judith, to the podcast. Thank you ever so much for joining us. Thanks for having me as well, Jeremy. Pleasure to be here.
2:26I want to crack on with our sort of first question on your research, which was around air traffic forecasting in, obviously, in aviation. So before we get into the detail, technically, I'd like, you know, for listeners who maybe haven't thought about this problem before and why it's a problem? Can you just set the scene for us? Can you sort of explain what makes forecasting, demand forecasting, I suppose we're talking about here, in a network different and quite tricky than other problems? I think when people think about forecasting, they think about predicting sales for a single product or estimating passengers' demand for an isolated, standalone route.
3:15but in an air transport network the system is a complex and living system where everything is connected to everything else so I think what makes the forecasting within a network kind of like framework or setting different it comes down to for example the domino effect or the cascade effect so traditional forecasting of a product where you have a change of one of the items rarely will break the rest of your catalog, right? But in air traffic forecasting, a single change of the network, so maybe an airport pair connecting where before there wasn't a route there, or two airports that are connected now get removed that flight that was available, kind of cascades through the entire system.
4:09Right. So that kind of effect, knock-on effect on everything else, is one of the challenges and one of the differences between a single product like forecasting or a single product forecasting than the network itself. So this isn't then like a sort of standard forecasting problem at all in this respect. If I was trying to forecast how many cars of a particular model or manufacturer I was going to sell in the next six months, in the next year, and I'd been selling that car for absolutely you know for many years um then then i i would just be using you know previous sales data maybe something cyclical maybe something periodic based on the um the time of year the time of the month maybe and then some because it's quite an expensive purchase maybe some economic sort of indicators would all sort of go into the mix and go into that that forecast approach and I'd end up with something that was you know probably reasonably likely to to to to give me a decent outcome but what you'll say I think this is really fascinating in in in the world of aviation is that when you go back in time you look at the sort of previous data you're not actually looking at the same world as the one you're about about to enter you're looking at a different world because it's it's differently connected is that is that is that or is that the challenge?
5:37Your kind of like map of flight or kind of like your schedule is completely different. It kind of like evolves and changes over time. So yeah, so like you can't just say in 10 years time, my network is going to look the same as right now because actually it's not.
5:59The world you looked at when you were writing this thesis and doing your research was one, I mean, there was a nice graph in it I saw, which showed the sort of distribution of connectivity. It shows you how many links you've got of a particular distance apart. And so it looked like the world of the sort of the local hub airline, the sort of cheaper airline, which is doing shorter hops, was very much part of this world. And there weren't, proportionately anyway, many journeys in that data set, which were sort of reflected sort of 2 ,000, 3 ,000, 4 ,000 mile long haul flights. So, I mean, how do you account for that kind of network bias?
6:49Or is it a case of, well, this is the structure of the network we have. This is the model that we're trying to capture from the data. So the example application I used was the US transportation system. And it's a mature system. It's a steady state, you know, like, and it's really based on, like, having spoke, as I mentioned before. So, you know, like, the graph that you are referring to probably is, like, not degrees. So, like, you know, how many, you know, how many airports have, like, really high number of connections. And then the distribution of, like, and you could see, like, a really, really skew towards the left-hand side.
7:23So there is fewer airports that are connected to a large number of airports. And then there is less that are connected to less. Sorry, there are more that are connected to less, which basically represents the hub and spoke kind of like system and network. So I think the way that I tried to represent that was introducing the kind of like network theory metrics, which basically kind of like explain the topology of the network. and again like they look like this because it was the u.s and it was like you know like a steady state like mature system they will probably look like completely different if i was using brazil for example which is quite like a young um air transportation system and probably they are more like point to point um rather than having spoke or europe for example which kind of like yes it's many countries but within you know like it's it's kind of like operates as a you know as a one single kind of like a big aerospace let's say um so yeah so that's the way like how that's the way how i represented um the the topology and the characteristics of the network okay so i'm really keen to sort of delve a little bit more into this of the network theory element because i think this makes it quite a both quite a challenging problem and also one i don't think is particularly well tackled today so so you're looking at the structure of the network we've already talked about talked about the fact that there was a higher density of shorter hops of short-haul flights and a lower density of long-haul flights.
8:54So did you look at characteristics like power law distributions on connectivity? There's lots of interesting features that I know network theorists pick out around either power laws or... Yeah, in terms of connectivity I think the US was like a power law like follow up our law distribution. Right. For this particular case, I looked at, I think it was like four different kind of like variables that define the topology of the network. So like node degree, A game vector, centrality, cluster and coefficient, and the weighted node degree kind of thing in the end. And I used them for the, actually for the connectivity model.
9:40So the connectivity problem was kind of like a split into two. So one into kind of like two classification models. So one that will look at link addition. So like when looking at the probability of a link appearing into the network. And then the other one would be the link removal kind of like model. So like looking at the probability of like a link being removed from the network that existed before. And when I talk about links, I talk about like airport pairs. Yes. So I didn't understand that. So that's really interesting. The problem isn't just one of forecasting in the presence of a network that's previously changed considerably over time.
10:17So looking and trying to treat the data in a consistent way over those changes. It's actually looking into the future and go, what's the probability the network changes again? And new links are introduced in five years, two years, seven years, whatever it is, or taken out of the network. That's really, so how did you go about coming up with a model for that dynamic network piece? Kind of like my modeling framework that I put forward, kind of like had different kind of like models, let's say, or like steps. So the first one, we'll look at the city pair demand generation. So we'll look at the true origin, true destination passenger demand between cities, independently of like, you know, the routes that they take to go from city A to city B.
11:04And that can use socioeconomical variables like population, household income, type of destination, whether it's like leisure, like business and destination and so on. Basically, kind of like simple econometric models. The second model was the kind of like that network evolution, connectivity and an interdisciplinary assignment. So in here, what I looked at was how this connectivity changed. So again, I looked at which links were removed from the network and which links were added into the network. And then by assessing this or by predicting all those changes, then I kind of will create the new itineraries that will serve those city pairs.
11:50And then the second kind of like a step within the second stage is the itinerary assignment. So once I know how the connections have changed, now let me allocate that passenger demand between city pairs across the different itineraries that exist, either non-stop or one-stop. And then the final one, the final stage, is looking at those allocated passengers volumes aggregated at the segment level. So, like, for example, if you got, you know, like someone going from JFK to Chicago and then San Francisco, you will, you know, that volume, that passenger will be in two segments from JFK and Chicago and then from Chicago and San Francisco.
12:43So aggregating those and then translating those into like number of flights and operations, et cetera. so I think the way this work was kind of like and again the other point that I want to say is that the framework that I developed was explicitly kind of like designated for like medium and long term forecasting kind of like more serving for like policy evaluation so like what you mentioned before about like should we build an airport here or not or like a second you know terminal run away. Exactly that infrastructure that we're talking about. Yeah infrastructure kind of thing so kind of like the goal was kind of like captured that systematic kind of like shifts over a multi-year horizon rather than knowing the granularity, the granularity of...
13:26Who's going to turn up to the airport tomorrow. Yeah. Yes. So I think that's an important aspect. So basically the way the framework ran was kind of like as a dynamic kind of like feedback loop or like a simulation, let's say. So at each step, you know, you will predict how many people there will be or passenger demand across the different city pairs. Then we'll look at the connectivity, how it has changed based on previous year, and then compile those itineraries available and then compile the market share of those itineraries and then predict the flight frequency. And then you will move on to the next step, which will be the following year, and so on.
14:09Wow. Yes, complex doesn't even come close. What did the data look like for this? I mean, this is obviously publicly available data sets from the US airline market. Is that right? Yes, correct. And that's why, I mean, and that's why one of the reasons why I chose the US as an example application, because like you can get a lot of data from it. So yes, like the entirety of my thesis is based on like open source data from kind of like air traffic volumes and demand, passenger demand and like what they paid. and as well from BTS I think it's called, so this is from the Bureau of Transport Statistics but they have as well information around population and economic household income, average household income and so on.
14:59And so all the information that I took was from available social media. Wow, so capturing the propensity to want to travel, the means to do so, the economic means to do so across... I know somebody's still trying to find his flights for a summer holiday. I know it's not exactly cheap, but you really do have to engage in the network hunt to get the best thought. At the end of the day, I would say, like aviation industry is cyclical and go hand-on-hand with the economic cycles as well. so you can see you can always see like it's proven there's a lot of research out there and that's why like the you know to to model or to predict the the city pair and passenger demand usually like with a simple econometric model it's basically what is used across you know across and all the practitioners and so on partly then forecaster but partly scenario planner then in that in that you said yourself well naturally an airline or an industry will want to go what if what if we all club together and built a new hub airport in the middle of kentucky or something like that you know would it would it would it be a massive driver of traffic or what what kind of would it be worth would it be worth the investment fundamentally yeah no totally totally i think um and and like within my within my framework like i i generate the predictions in the long term in the medium long term and based as well in in based on like different kind of like socioeconomic kind of like levels for like whether like population or like income or average household income was like increasing like low medium or like high so that you could see like the differences of of that impact as well so yeah but you could you know like i mean for example like in in the case of like you know something happening like a shock event happening like the pandemic or like you know like a closed aerospace because of like geopolitical reasons and things like that and you have that framework so it's a matter of like pushing the um fiction or like pushing the boundaries of like how you want to test the model and how it reacts okay so that's really good so you mentioned the pandemic i'm really keen to come onto this.
17:24Obviously, there's something which was completely unanticipatable really when it did strike in 2019, 2020. What happens to a model like this and how could it be used or repaired, I suppose, in a scenario where you have a change that's so dramatic, that's so catastrophic really for that industry, certainly for a few months. What happens to that kind of model? What place would it have in trying to aid a recovery in that situation? I think, yeah, when something like this happens, it's almost like a structural change, right? And I think it completely kind of like breaks the core assumption of the mathematical model.
18:09And it kind of like exposes, in a way, exposes kind of like the mathematical kind of like foundations and limits of the broadcasting. I think in a way it breaks the historical correlation, right? like econometrics and machine learning models, like they are fundamentally built on the premises of like what happened in the past. Yeah. Kind of like finding those like correlations between and patterns between variables, no? So like it is proven that, you know, like a percentage increase on household income usually translate into a certain raise in passenger demand. But then when something happens like this, like a pandemic, then kind of like basically, it just doesn't drift the behavioral kind of like relationship it just like breaks everything and it snaps like you know oh and and overnight almost and again like i think in the limitation of like because we are using historical data and at the end of the day um probably this is a situation where the model never has never seen kind of like that situation so no right of yeah the boundary limits let's say um so i think that's why like going back to you know what if it's scenario planning i think that's the way we need to think in practice of like how this framework should be used um so it's it's it's kind of like it's almost like as well on like the people who are kind of like building those models and like how can quickly we can react to you know to the context changing and use those tools that we have regenerated to actually um you know make better decisions when something like this happens.
19:45So yeah, so like use the framework to generate, you know, like different outcomes based on different levels of economic variables or even like, you know, like kind of like a stress like stress testing, like your models of like, you know, like extreme situations such as like pandemic, aerospace closure and there was the volcano as well for example, like years ago that closed their aerospace as well from one day to another and and so on. And I think the other aspect as well is probably like keeping a kind of like human in the loop. So I kind of like almost like accepting that, you know, like, yeah, we build the solutions to kind of like help us to predict the future in a way.
20:30But I think we shouldn't be removing the expertise and domain expertise judgment from those tools and from the math, like they complement each other. I think ultimately it's not the limit of any forecasting model or any kind of model, etc. It's just assuming that we will do everything without the human context. So I think there is always, for me, there should be always that human expertise. Indeed. And I think you hit on two very, very important principles of modeling, data science modeling, AI modeling, whatever version you're doing, which was the first was, if the model hasn't seen historically data that it's been trained on that reflects your current scenario, the chances of it without significant aid and significant intervention, the chances of being able to produce a forecast which is in any way realistic are pretty much minimal.
21:35And I think people... The probability of that being forecasted is just like, yeah, like your model predicting that this is going to happen is just not, yeah, it's really low. Exactly. And I think that's very true of any AI model. If you're asking it to predict or perform a task that's outside of its operational exposure previously, you know, operational parameters, you're asking it to approximate wildly, hallucinate, whatever word you choose to use, really. But I mean, that's where it comes from, is when you're basically saying, we've only trained you on this bit of history, and we're asking you to predict something that's over here that's miles off of your normal understanding.
22:22Yes, totally. And I think when in the past has happened these things, you know, like the economic crisis in the past, usually you would incorporate those because you don't either you drop this data and don't use it because it's like you know it's not that because it happened you know 10 years ago it will happen 10 years you know after or you use them variables that you know like absorb that change you know and I think maybe when I was talking about like human in the loop is that maybe you know like you need to think about like when these models are in production maybe allowing for a you know like a way in to the human like adjust you know when something like big is happening yeah so let's let's let's focus on that just briefly then because i think this is this is actually quite relevant um contemporarily so you the conflict in the in the middle east which just you know has kicked off in the last couple of couple of months massive impact on on globally aviation I would imagine all those huge hub airports in um in sort of Middle Eastern countries from sort of Emirates and um and the like and and that were you know certainly very very significant connectors of global global air traffic so those all had to be routed round avoided how would a human in the loop which I think is the second great principle by the way in terms of making sure these systems work effectively how would a human in the loop then with a model such as yours maybe if you've been trained on 2025 data 2026 data as well how would they use your setup but but manipulate it and be that sort of smart arbiter to enable the model to be able to cope with a situation like that i think yeah like either like allowing to you know adjust elasticities or like add introduce like some custom dummy variables, or even kind of like, plug out some of the connections, you know, like, you know, those links don't exist anymore, because basically, those airports are inoperative, you know, like, for a period of time.
24:26And so almost like being able to remove that node, for example. And even sometimes, like, even if having that human in the loop might not been, or might not have allowed to make the change so quickly, right, because things like are a bit uncertain and like you don't really know when they're gonna you know plug the you know put the plug and pull the plug and and this is gonna explode and you know like it's gonna unfold and so on but i think um if as a as an expert or like you're you know you are monitoring what is happening in the world then then you know you could already before start running those what scenarios with those kind of like changes and that you think that might happen and then kind of like a bit you know have a bit more notice of like how it could impact yes and it's not just that it's not just the impact of the event itself it's the sort of elastic rebound afterwards both from the pandemic and something like this of people going oh thank goodness now i can do that travel that i've been postponing and i'm we're all going to do it together so suddenly got this huge bulge of demand that comes through.
25:33I hope you're finding this conversation as fascinating as I am. If you're getting value from the research and technical depth we dig into on Inside the Algorithm, please take a moment to follow the show and leave us a review on Apple Podcasts, Spotify, or YouTube. Every follow helps us reach more of the brilliant minds doing this work and keeping the rigorous substantive conversations going. You can find Inside the Algorithm on the data and AI mastery podcast feed or watch it on the Cambridge Spark YouTube channel. You will find the links to those in the show notes below. Right, let's get back into it.
26:06I'm going to switch tack a little bit, Judith. Really onto your current role. So you're currently head of data science at DataSpark, a professional service company. So you don't just have a team that solves one problem in an organization. you're fortunate enough to look at running multiple concurrent projects across multiple clients, multiple sectors in fact as well. So not just in one sector because you're running a team that works across this diverse set of projects and industries that the development process for those data scientists, those AI engineers is going to be really important. I know you have some strong views on this particularly.
26:59So what do you think then is the sort of core capabilities and qualities that you would look for in data scientists who's looking to develop themselves and deliver impactfully in this really interesting ecosystem of projects. It's not just about experimenting and doing research and building a solution just because it's interesting. It needs to, you know, you have to have that mindset of always thinking of what is going to impact and what it's going to bring to our clients. I think the other point is pragmatism over perfection. you know like in academia and you are kind of like chasing that like 0.5 percent improvement on model accuracy but actually in a business setting if you can get you know within i don't know like three or four weeks uh 80 percent model and that 80 accuracy and accuracy model um that it's already running production then it's better than you know a model that might take you know six months to develop and gives you 95 accuracy but then it's more complex or like really working production and and so on and so it's always about like thinking that like agile um and think about like what delivers quickly um quick wins and then you can you know still develop on the in parallel but at least you have something that is giving you value um already um and then and then yeah like simplicity over complexity a lot of you know like all our clients at the end of the day want to understand actually how the predictions work.
28:43And it's much easier to add explainability layer into a linear regression model than neural networks and so on. So again, simplicity. And then I guess that mindset of production first, kind of like a ready mindset. So although we might look at developing a POC or MVP, etc., and it's a bit like more research look like, research like you need to develop the code and you need to think about this is going to be in production. So think about testing, think about clean code, modularized code and so on. So it's not just like a blank canvas and like, yeah, flagging lines of code across. But yeah, it will help a lot on the subsequent phases of productionization and so on.
29:40So have you got any sort of observations from success stories that you've seen previously, which enable your data scientists, your engineers to really make that final move? Because I know it's such a challenge for many projects to get into, right, we've written this thing, we've invested in it, we're using our data brilliantly, and now how do we get the stakeholders to actually use it? I was mentioning before about the so what. So I think it's really important to understand who are the end users, how are they going to use it, how is this going to be surfaced, and so on. If the output is going to be just in a folder where no one can access, then what's the point of doing the solution?
30:23Because then no one is going to use it. and so it needs to have like that mindset of like actually thinking like you need to start talking to end users like early on even like you know at the first week of like you know when you are trying to see the and also to understand like whether do they actually need this solution or actually you know the definition of done or like what we're trying to predict even like the definition of like the outcome of the prediction you know like what is it like passenger volume so actually is it aircraft number you know like different like these aspects and then i think as well there is there is an aspect that i see sometimes which is like how much people can absorb change or like want to absorb change as well um i think um in some of the projects that you know we have the most success is also like some of the projects where we had kind of like champions within the business within the client's business and kind of like really believe on like you know these are solutions that actually could make a difference into their business so looking at everything that's happened in sort of data science and ai and we sort of alluded to some of it along the conversation um what do you think people need to be sort of really paying attention to in applied data science in AI now what do you think what do you think is going to have the most impact going going forward for you and your team for for the industry as a as a whole I think there is a kind of like a shift towards I think like a small highly specialized models for example so rather than using kind of like large language models for everything for everything yes maybe the real value actually lays into into those kind of like highly efficient kind of like language models that you know you can fine-tune for a specific kind of like from like clean specific kind of like domain data sets they are faster they're more cost effective and you know like then it's better right we win better for production as well yeah i think um the other aspect and i think this is an aspect that is usually overlooked and it's not like you know it's a bit unsexy and maybe boring but it's about like the kind of like governance you know like rigorous like evaluation and like testing frameworks and so on I think there is a yeah there is kind of like the tendency maybe of like oh yeah let's put an agent to everything and and they don't think about and and then people don't think about like the implications of using those agents with maybe enterprise data or you know like or even like personal data it's really important to keep an eye on that um so so i think and i think in the future my you know like might be where people actually need to focus more because actually we haven't taken care of this like right now julie it's been an absolute absolute pleasure talking to you thank you ever so much for for joining us today really enjoyed that yeah thanks thanks for having me thank you jeremy i really enjoy as well
33:32thank you for listening to this episode of inside the algorithm them. If today's conversation gave you something to think about, make sure you subscribe so you never miss a breakthrough and share it with someone who'd appreciate the thinking. If you're a data and AI leader looking to build deep technical capability across your organisation, Cambridge Spark is here to help. You can reach us on LinkedIn or at cambridgespark.com. And if you want to explore AI from the leadership and strategy perspective, check out our flagship show, Data and AI Mastery with Dr. Raul Gabriel Irma. Until next time, stay curious, stay rigorous, and stay ahead of the algorithm.
From the publisher
π Discover how Cambridge Spark helps organisations build the data and AI capabilities needed to turn strategy into measurable impact: cambridgespark.com
This week, Dr Judit Guimera Busquets, Head of Data Science at Datasparq, joins Dr Jeremy Bradley to trace the journey from her PhD on air traffic network forecasting through to leading data science teams delivering real-world AI projects.
Judit explains why forecasting inside a complex network is fundamentally different from standard demand prediction: when a single airport pair is removed, the cascade effect ripples across an entire system. She walks through the multi-stage modelling framework she developed, covering city pair demand generation, network evolution, itinerary assignment, and long-term scenario planning.
The conversation then turns to what actually happens when structural shocks like a pandemic break a model's core assumptions and why human-in-the-loop design is not optional. Judit also sets out what she looks for in data scientists: pragmatism over perfection, simplicity over complexity, and a production-first mindset from day one.
She closes with her view on where applied AI is heading, including the rise of small, fine-tuned specialist models and why AI governance remains the most overlooked challenge in the field.
Follow Data & AI Mastery on Apple Podcasts, Spotify, or YouTube to stay ahead of the algorithm.
If you enjoyed this episode, why not check out the Data & AI Mastery episode with Richard Masters, VP of Data and AI at Virgin Atlantic. You will learn more about how the airline leverages AI and data-driven strategies to enhance operations, optimise pricing, and deliver premium customer experiences:
Apple: https://podcasts.apple.com/gb/podcast/mastering-data-ai-insights-from-virgin-atlantics-vp/id1779783413?i=1000697801419
Spotify: https://open.spotify.com/episode/1MY3AZCvDr5eS1HlfBudHy?si=ca5ec9b2fe6d44b1
YouTube: https://www.youtube.com/watch?v=DQ3mTwTzJvA
Glossary Terms
Hub-and-spoke Model: a centralised organisational architecture where a central core connects to multiple peripheral nodes. Traffic, communication, or inventory flows through the hub rather than directly between spokes.
Network Theory: a multidisciplinary framework used to analyse complex systems by representing them as mathematical graphs
Econometrics: the application of statistical and mathematical models to economic data
Human-in-the-Loop: a collaborative AI approach where humans actively participate in an automated system's training, refinement, or operation.
Linear Regression Model: a fundamental statistical and machine learning algorithm that models the relationship between a dependent variable and one or more independent variables by fitting a straight line to the data.
Chapter Markers
(00:00) - What makes network forecasting different from standard demand prediction
(05:54) - How historical data fails when the network itself evolves
(10:05) - Modelling link addition and removal as classification problems
(13:26) - Designing for medium and long-term policy evaluation, not daily operations
(17:57) - What happens to a model when a structural shock like a pandemic hits
(22:22) - Human-in-the-loop: adjusting elasticities and running what-if scenarios
(27:20) - What great data scientists actually look like in a consulting environment
(30:03) - Getting stakeholders to use AI: champions, end users and change readiness
(32:00) - Where applied AI is heading: small specialist models and the governance gap
Useful Links
Connect with Dr Judit Guimera Busquets on LinkedIn: https://uk.linkedin.com/in/judit-guimera-busquets-696ab74a
Learn more about Juditβs PHD here: https://openaccess.city.ac.uk/id/eprint/24689/1/Busquets%2C%20Guimera.pdf
For more AI insights follow Jeremy on LinkedIn: https://uk.linkedin.com/in/jeremy-bradley
Explore Cambridge Sparkβs AI upskilling programmes at https://www.cambridgespark.com




