In short
How competing learning algorithms in markets determine “survival” (nonzero wealth share) vs “vanishing” (wealth share goes to zero), comparing Bayesian learners to no-regret learners, and proposing a “robust Bayesian update.”
Guest backgrounds
No guests are identified in the transcript; only host/interviewer dialogue appears.
Key claims
Wealth share is tied to differences in accumulated regret between agents. Low regret alone (e.g., O(log T)) may still lead to vanishing against a Bayesian with finite model support and positive probability on the true model. Survival requires regret to stay bounded within a constant additive gap versus every competitor at all times. Perfect Bayesians achieve constant expected regret; fragilities include inaccurate priors (linear regret) and “trembling hand” update noise. No-regret learners adapt better to distribution shifts. Robust Bayesian updates add decaying regularization to get constant regret in stable settings and logarithmic regret under shifts.
Notable examples
Figure 2c simulation where a Bayesian’s wealth share rises to ~1 while a no-regret learner vanishes; Figure 1b/c showing long-run vanishing but short-term presence for an epsilon-inaccurate Bayesian; Figure 2B with UCB no-regret learners fluctuating without systematic elimination.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Algorithmic Influence
0:45 to 2:53
Exploration of how algorithms dominate various market segments beyond finance.
“So, OK, if you're someone managing these incredibly complex algorithmic investment strategies and you're in these really competitive markets, a basic question kind of jumps out.”
Learning Approaches in Competitive Markets
2:53 to 4:52
Discussion on the core question of which learning approach to adopt for survival and success.
“And crucially, their investment decisions.”
Bayesian vs. No-Regret Learning Approaches
4:52 to 7:33
In-depth comparison of Bayesian learning and no-regret learning mechanisms and their implications.
“Vanishing is when your wealth share goes to zero.”
Conditions for Survival in Markets
7:33 to 8:23
Insights into the survival conditions of different learning agents in competitive markets.
“So given this tough benchmark, what's the sort of ultimate target?”
Fragility of Bayesian Learning
8:23 to 11:16
Examination of the fragility factors associated with Bayesian learners and their implications for competition.
“How does this kind of agent actually achieve this constant regret benchmark?”
Robustness of No-Regret Learners
11:16 to 14:02
Exploration of the advantages and resilience of no-regret learners in fluctuating markets.
“It depends on their inaccuracy, epsilon, and who they're competing against.”
Understanding No-Regret Learners
14:02 to 16:15
Learn about the advantages of no-regret learners and their robustness in changing environments.
“Let's pivot back to the no-regret learners.”
The Robust Bayesian Update Proposal
16:15 to 18:34
Discover the innovative robust Bayesian update method and its benefits.
“a way to potentially get the best of both worlds.”
Balancing Performance and Resilience
18:34 to 20:14
Explore how the robust Bayesian method balances high performance and resilience in markets.
“The best approach isn't one-size-fits-all.”
Impacts on Market Prices and Phenomena
20:14 to 21:14
Examine how competitive learning processes affect market prices and phenomena like momentum.
“Beyond just asking which individual agent survives, how does this whole competitive learning process, this constant battle for wealth share, actually shape the overall market?”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we're plunging headfirst into this algorithmic world that's increasingly running our markets. It really is everywhere now, isn't it? Totally. I mean, think about it. Automated trading doesn't just touch stock exchanges. It pretty much dominates them now. Yeah. From the actual stock trades themselves to the, you know, the overall strategies being deployed. It's not just finance, right? Online advertising. That's huge. Powered by algorithmic bidding. Oh, absolutely. It accounts for what? Over 1 % of the entire U.S. economy. It's massive. Wild. And then you've got retail giants, Amazon being the obvious one.
0:37Algorithms are behind. Like every price change, every recommendation you see. And deciding what's in stock, managing inventory. It's just woven into the fabric of how these platforms operate. A real fundamental shift. So, OK, if you're someone managing these incredibly complex algorithmic investment strategies and you're in these really competitive markets, a basic question kind of jumps out. Which learning approach, Which philosophy should you actually adopt to, well, survive and thrive? That's exactly the core question we're going to unpack today. Our deep dive is based on this really significant academic paper.
1:13Okay. It dissects how these different types of learning agents, these algorithms, actually behave when they compete head-to-head in markets. So what's the mission for us today? Well, our mission is to understand the performance of basically two major approaches. You've got Bayesian learning, mainly from economics. Right. And then no regret learning, which comes more from computer science. We want to really grasp what makes an agent survive or what causes it to vanish. Vanish, meaning their market share goes to zero. Exactly. And as a bonus, we'll explore this really interesting new strategy the paper proposes, one that tries to get the best of both worlds.
1:52Sounds great. So let's set the scene a bit more. What are all these different algorithmic players actually trying to do in the market? What's the goal? Well, at their core, I mean, pretty much all market participants are trying to maximize the growth rate of their wealth. OK, grow the money. Makes sense. Yeah. And often this is measured using the logarithm of their wealth. That's kind of a standard way in both economics and computer science models to think about long term growth, compounding effects, that sort of thing. Got it. So they're all aiming for wealth growth. But you mentioned these two very different philosophies for getting there.
2:26Totally different mindsets, yeah. Walk us through them. Let's start with the Bayesian approach. Okay, so Bayesian learning. This really comes out of the economics tradition. Think of these agents as starting with a set of initial beliefs, a prior, about how the market actually works. Like the underlying rules of the game. Exactly. The true probabilities, the hidden processes. And as they see more data, more market outcomes, they systematically update these beliefs. They use Bayes' rule. Right. And crucially, their investment decisions. They're based on these evolving beliefs about the process, not directly on whether they made or lost money yesterday.
3:02Ah, okay. So they're trying to understand the world first. Precisely. The idea is, if their beliefs eventually get close to the correct understanding of the market, their wealth share should grow optimally or close to it. Okay, that's the Bayesian side. Now contrast that with this no-regret learning from computer science. How's that different? It's a very different beast. No regret learners. They operate without making strong assumptions about how the market works. So no prior beliefs about the process. Pretty much none or very weak ones. In fact, they're designed to perform reasonably well, even if the market conditions are like adversarial.
3:42Adversarial, like the market's actively trying to trick them. Sort of. Yeah. Or at least the algorithm doesn't assume the market is playing nice or following some stable pattern. So success isn't measured by how accurate their beliefs are. Okay, so how is it measured? It's measured by something called regret. Regret, okay. Regret, in this context, is basically the difference between their actual wealth growth and what they could have achieved if they'd somehow known the single best fixed strategy from day one, looking back in hindsight. Ah, so comparing themselves to the perfect hindsight strategy.
4:13Exactly. And here's a key difference. Unlike the Bayesians focusing on the process, these no-regret agents explicitly adjust their strategy based on changes they see in their own wealth. Step by step, they react to performance. Okay, let's unpack this a bit. This setup sounds like a recipe for some serious competition. High stakes. Definitely. When you have these different types, these Bayjans and these no-regret learners, competing in the same market. Yeah. Does one type just, you know, drive the other one out? That's the core question of market selection, as the paper calls it. Right. You mentioned vanishing versus surviving.
4:52Vanishing is when your wealth share goes to zero. Yep. You're wiped out. And surviving means you manage to hold on to like a non-trivial chunk of the market wealth indefinitely. Exactly. Your share stays bounded away from zero. It really is a kind of zero-sum game for these wealth shares. And the paper finds a direct link between the survival and regret. Right. A surprisingly direct one. Yeah. It turns out an agent's wealth share at any time relative to any competitor is mathematically tied to the difference in their total accumulated regrets up to that point. Wow. OK. So the less regret you have compared to the competition, the better your share.
5:28Precisely. It's this difference that governs who wins and who loses in the long run. Now, this is where things get really counterintuitive, I think. You'd assume, wouldn't you, that having a pretty good strategy would be enough. What do you mean by pretty good? Like one with really low regret. Maybe regret that grows super slowly, like logarithmically with time. We hear about log regret algorithms being efficient. Right. O log T regret is often seen as a strong benchmark in learning theory. Yeah. So you think achieving that means you're safe, you'll survive in the market. But that's the surprising truth the paper reveals.
6:00It's not enough. Really? Log of regret isn't enough? Not for survival in this competitive setting. Theorem 3.1 in the paper shows that an agent can have remarkably low regret, even that OlogT kind, but still vanish completely. If it's competing against a Bayesian learner who, number one, has a finite set of initial possible models, and number two, assigns any positive probability, however small, to the correct model of the market being among them. Whoa. So even a tiny chance of the Bayesian being right lets them eventually crush, a low regret competitor. In the long run, yes. It's a pretty stark finding.
6:36So, okay, if low regret isn't the key, what is the actual condition for survival? What does it take? The paper nails it down very precisely, again, in theorem 3.1. An agent survives if and only if its regret remains bounded by a fixed additive constant relative to every single competitor at all times. Wait, bounded by a constant? Not just growing slower, but spaying within a fixed difference. Exactly. Think about it. If my regret grows at rate R1 and yours grows at rate R2, even if both R1 and R2 are tiny, like log teeth, if R1 is consistently slightly smaller than R2, that difference adds up over time.
7:12Yeah, it eventually drives the one with slightly higher regret out. Right. The difference in total regrets will grow linearly, leading to an exponential divergence in wealth shares. So to survive, your regret cannot grow relative to your competitors. It must stay within a constant bound. That's a much, much higher bar than just having low regret on your own. It absolutely is. It means constant factors matter a lot. So given this tough benchmark, what's the sort of ultimate target? Even a perfect player must have some regret against hindsight, right? That's true. Even an agent who magically knows the true market distribution and invests optimally according to that knowledge will still incur some regret compared to the absolute best hindsight strategy, which could pick the winning asset perfectly every single day.
7:58But the crucial insight shown in Theorem 5.2 is that this optimal regret, the regret of the perfect player, is expected to be constant. It doesn't grow with time. It just depends on stuff like the number of assets available. Ah, so that constant expected regret becomes the goal standard. Exactly. That's the benchmark. To survive against the best possible player, you basically need to achieve that same level of constant expected regret. Okay, let's talk about the perfect Bayesian again. How does this kind of agent actually achieve this constant regret benchmark? Well, the key is in their setup.
8:31If a Bayesian learner starts with a finite set of possible market models, their finite support prior, and if that set includes the actual true market distribution. And they have some initial belief in it. Right, some positive prior probability on it. And if they update their beliefs correctly using Bayes' rule as data comes in. Then what happens? Then, according to theorem 5.4, they achieve this constant expected regret. Their beliefs converge exponentially fast to the truth. And what does that mean for market competition? It means this perfect Bayesian survives. They hold their own against the optimal player, and they will drive out any competitor whose regret grows over time, no matter how slowly that growth is.
9:12Log t, square root t, doesn't matter. They vanish. Wow. The paper has simulations showing this, right? Yeah. Figure 2c is a good example. It shows a simulation with a Bayesian learner, let's call it Agent 1, competing against a no-regret learner, Agent 2. Both consider the correct model possible. And what happens? You can just see Agent 1's wealth share steadily climbing, heading towards one owning the whole market. Meanwhile, Agent 2's wealth share just tanks, goes to zero, vanishes. It's a stark visual. OK, so perfect Bayesians are dominant, but you use the word perfect. That sounds like a big if.
9:51It is. And this brings us squarely to what the paper calls the fragility factor of Bayesian learning. What happens if the Bayesian isn't perfect? Which seems, you know, more likely in the real world. Absolutely. So the first big fragility, inaccurate priors. Meaning? Meaning what if the true way the market works isn't actually included in the set of possibilities the Bayesian agent even considered from the start? Ah, so their worldview literally doesn't contain the truth. Exactly. Even if the truth is like really close to one of the models they do consider, if it's not exactly in their prior support, theorem 6.1 shows this leads to disaster.
10:26Disaster happens. Their regret grows linearly over time, which means they will eventually vanish when competing against basically any decent no-regret learner. So even a tiny mismatch between their assumed world and reality leads to failure. In the long run, yes. Think of an investor. Maybe they have a really sophisticated model, but it's just slightly misspecified, a tiny skew. Over enough time that tiny skew compounds, their beliefs don't converge correctly, and their performance degrades linearly relative to someone who adapts better. But do they vanish immediately, or can they hang on for a while?
11:00They can hang on. The paper talks about epsilon-inaccurate Bayesians, those whose prior beliefs are close to the true distribution, where epsilon measures that closeness. These agents might actually hold a significant market share for a noticeable period. The paper even quantifies this typical survival time. How long do they last? It depends on their inaccuracy, epsilon, and who they're competing against. Against a competitor with constant regret, their survival time is roughly proportional to 1 over epsilon squared. So, smaller error, longer survival. Makes sense. Right. But, against someone with, say, logarithmic regret, it scales like 1 over epsilon cubed.
11:37against square root t regret, 1 over epsilon 4. The better the competitor, the faster the slightly inaccurate Bayesian gets pushed out. Still, 1 over epsilon squared could be a long time if epsilon is tiny. It could be. And figure 1 in the paper shows this really well. It pits an imperfect Bayesian, agent 1, against a no-regret learner, agent 2, who can find the correct distribution. What's the picture? Well, figure 1b shows the long run, like a million steps. Agent 1, the inaccurate Bayesian, clearly vanishes. Agent 2 takes over. But then figure 1c zooms in on just the first 10 ,000 steps. And there, agent 1 still holds a very significant chunk of the wealth.
12:16It looks like a contender. It visually demonstrates this idea of short-term presence versus inevitable long-term vanishing, if you've got that fundamental flaw in your assumptions. Okay. So inaccurate priors are one big fragility. Are there others? Yes. Another really critical one is about the updating process itself, even if the prior includes the true model. What could go wrong there? The paper calls it trembling hand updates. Theorem 6.5 discusses this. It turns out that even tiny, almost imperceptible, zero mean errors in how the Bayesian balances their prior beliefs with new data at each step.
12:50Like tiny calculation mistakes or noise. Exactly. Even if the errors average out to zero over time, these tiny jitters can completely break the learning convergence. It's like that scale analogy mentioned earlier. A tiny jitter throws everything off. Precisely. These small errors prevent the agent's beliefs from properly converging to the truth, even if the truth is in their prior. And the result, again, linear regret, and eventual vanishing against better learners. It shows just how sensitive the theoretical optimality of Bayesian learning can be. Wow. Fragile indeed. But what if you have two good Bayesians competing, both with the correct model in their prior?
13:27Ah, now that's different. Figure 2a illustrates this. When two such Bayesians compete, they can actually coexist. They don't drive each other out. No, they tend to reach a kind of steady state wealth partition. They share the market. However, the initial learning phase still matters. The relative quality of their priors, like how much initial belief they assign to the true model compared to other models, influences how quickly they learn and shapes where that final wealth split ends up. So early performance still casts a long shadow. Okay, that covers the Bayesians' powerful, if perfect, but brittle.
14:02Let's pivot back to the no-regret learners. You said they're more robust. What gives them that advantage? Well, fundamentally, it's because they make fewer assumptions. They aren't trying to pinpoint the exact underlying process in the same way. Less to get wrong, essentially. Kind of, yeah. And this really helps when the environment isn't stable. For example, standard Bajan learners, as we saw, are fragile. They can suffer really badly linear regret again if the underlying market rules suddenly shift. Proposition 7.1 highlights this. Like if the economy changes or a new regulation comes in. Exactly.
14:34Those kinds of distribution shifts, a Bayesian locked into an old model can get completely blindsided. And no-regret learners handle this better. Generally, yes. Many existing no-regret algorithms, like those based on multiplicative weights or UCB, are known to adapt to these shifts more gracefully. How well do they adapt? They can often achieve regret that still grows only logarithmically in time overall, plus a penalty that's sort of linear in the number of shifts that occurred. So they take a hit when things change, but they recover and adapt. The paper shows this too. Yeah, figure 2B is interesting here.
15:08It shows two no-regret learners using a UCB strategy competing. What do we see? They tend to remain pretty comparable in the long run. Their wealth shares fluctuate, but neither drives the other out systematically. One might get an initial edge, maybe if it has a smaller set of strategies to consider, making learning faster initially. Like Agent 1 in that figure. Right. But the dynamics stay somewhat randomized. They don't settle down and forget initial conditions in quite the same way Bayesians do when they converge. They keep adapting. Okay. So summing up this part. Yeah. Bayesians are super powerful if their model is right and they execute flawlessly, but they're fragile to errors in assumptions or execution.
15:48Correct. No regret learners are more robust, especially to changes in the market. But they might accumulate regret faster initially and could get driven out by a perfect Bayesian. That's the core dilemma. It highlights this fundamental tradeoff. Speed versus robustness. Faster learning versus resilience to surprises or errors. So what does this mean for you, the listener, the decision maker, trying to choose a strategy? It seems like a tough spot. It is. And this is precisely what motivates the paper's really neat proposal towards the end. a way to potentially get the best of both worlds. Ooh, okay.
16:23What is it? They call it a robust Bayesian update. Robust Bayesian. How does that work? What's the trick? It's actually a remarkably simple modification to the standard Bayesian update rule. Simple is good. Yeah. The key idea is to introduce a small, carefully chosen regularization term. Think of it like adding a tiny bit of uniform belief across all models at every step. The paper suggests using something like epsilon-plotard decaying over time. Okay, so you add this small extra term. What does it do? It acts as a kind of safety net. It ensures that while models the data suggests are incorrect, still see their influence decay exponentially fast, just like in standard Bayesian learning.
17:04Their weight never goes completely to zero. This little regularization term effectively keeps all models alive in the background, albeit with tiny weights if they're performing poorly. Ah. So it prevents the agent from ever becoming completely certain that a model is wrong, just in case it becomes right later. Exactly. It maintains a tiny bit of adaptability, a capacity to reconsider models it had previously discarded. That sounds clever. Does it actually work? What are the benefits, according to the paper? The benefits laid out in Theorem 7.2 are pretty compelling. First, in stable stationary environments where a standard Bayesian would thrive, this robust method still achieves constant regret.
17:44Just like the perfect Bayesian? Just like the perfect Bayesian. This means it will survive against and outperform other standard no-regret learners in stable markets. Okay, so it keeps the Bayesian strength in stable times. What about when things change? That's the second crucial benefit. When those distribution shifts do happen, unlike the fragile standard Bayesian, this robust version's regret only grows logarithmically in time, plus terms related to the shifts. Logarithmically, so it adapts like a good no-regret learner. Precisely. It means it's adaptable, it can handle the unexpected shifts, and importantly, it can survive against traditional no-regret learners, even in these dynamic, changing environments.
18:24Wow. So it aims for Bayesian performance when possible, but defaults to no-regret robustness when needed. That's a great way to put it. It sort of interpolates between the two worlds. So the practical insight here for you listening seems really clear. The best approach isn't one-size-fits-all. It really depends on how much confidence you have in your understanding of the market environment. Is it stable, predictable, or prone to shifts? Absolutely. And this proposed robust Bayesian method offers a potentially powerful balance. It really underscores the paper's core message. In these hyper-competitive algorithmic markets, those seemingly small constant factors and regret rates, they aren't just theoretical curiosities.
19:04They're the difference between surviving and vanishing. Literally. They determine market life or death. OK, so let's just recap the journey we took today. We started with the sheer pervasiveness of algorithms running modern markets. Uh-huh. From stocks to ads to retail. Then we dove into this fascinating, almost brutal dynamic linking wealth share directly to differences in regret between competing strategies. And the surprising finding that just having low regret isn't enough for survival. You need bounded relative regret. We looked at the strengths of Bayesian, learning its potential for optimal constant regret, but also its critical fragilities.
19:45Vulnerability to inaccurate priors or even tiny update errors. The trembling hand. Right. Then the strengths of no regret, learning its robustness and adaptability, especially to market shifts. But its potential vulnerability to a perfect Bayesian. And finally, this intriguing step towards bridging that gap. The robust Bayesian update, aiming for that best of both worlds outcome. Constant regret in stable times, logarithmic regret when things change. A potential path to both high performance and resilience. So building on all that, the paper leaves us with a really interesting thread to pull on related to its future work section.
20:20Yeah. Beyond just asking which individual agent survives, how does this whole competitive learning process, this constant battle for wealth share, actually shape the overall market? Specifically, market prices. Ah, the aggregate effects. That's a huge question. Exactly. The paper suggests that the way these wealth shares evolve and how agents react to market feedback and competitor performance, this whole dynamic could actually induce patterns in prices, specifically positive serial correlation. What we often call momentum in finance, the idea that past winners tend to keep winning for a period.
20:56Right. So it's not just about who wins the algorithmic race, but how the race itself might fundamentally change the behavior of market prices, an emergent property of the competition. It suggests the very structure of learning in markets could be a source of phenomena like momentum. That's a really provocative thought. It really is. So for you listening, what stands out to you from today's deep dive? What further questions does this whole discussion raise for you about where these automated markets are heading?
From the publisher
This paper examines the performance of Bayesian learners and no-regret learners in competitive asset markets, identifying conditions for their survival or vanishing. It contrasts the economic focus on Bayesian learning with the computer science emphasis on no-regret learning, highlighting that low regret doesn't always guarantee market survival against a perfect Bayesian, while Bayesian learning can be fragile to slight errors. The research proposes a robust Bayesian update strategy that combines the advantages of both approaches, achieving constant regret in stable environments and logarithmic regret during distribution shifts, offering a more adaptable solution for algorithmic trading. The document also bridges the theoretical frameworks of regret minimization and Bayesian learning in asset markets, contributing to a deeper understanding of heterogeneous learning agents' dynamics and their impact on markets.




