Research
ECB simulations found that reinforcement-learning trading agents coordinate like a bank run under stress, while LLM-based agents behave less predictably, meaning the architecture choice itself is now a financial stability variable
Current financial supervision is built to watch institutions and products, not the internal design of the AI systems trading inside them, but the ECB's research suggests two AI agents given the same market conditions can produce meaningfully different systemic outcomes depending on how they are built.
This is a more specific and more useful finding than the general worry that "AI could destabilize markets." The ECB's Eurosystem researchers ran reinforcement-learning trading systems and large language model-based agents through the same simulated market conditions and got different failure modes back. Reinforcement-learning systems showed coordinated behavior resembling a bank run when stressed, a recognizable and in some ways familiar systemic pattern. LLM-based agents showed less coordination but more unpredictability, which is a newer and arguably harder kind of risk to model, because unpredictable is not the same problem as correlated.
The bank-run analogy for reinforcement-learning systems is worth sitting with. Classic bank runs happen because each depositor's rational individual response to a signal of stress, withdraw now before others do, aggregates into a collective outcome that makes the stress self-fulfilling. If RL trading agents trained on similar reward signals converge on similar de-risking responses to the same market trigger, the mechanism is structurally the same even though no depositor and no bank is involved: correlated individually-rational responses producing a collectively destabilizing outcome. That is exactly the dynamic regulators spent decades building deposit insurance and liquidity backstops to interrupt, and none of that infrastructure currently exists for correlated AI trading responses.
The LLM finding is the one that should worry a different set of people, because unpredictability resists the standard regulatory playbool of imposing position limits or concentration caps on a known, modelable behavior. If an AI system's response to a given market condition is not reliably repeatable even in simulation, then backtesting it, stress-testing it, or setting a circuit breaker calibrated to its worst-case behavior becomes materially harder. You cannot easily cap what you cannot reliably characterize in advance.
The structural gap the ECB is naming, that supervision watches institutions and products but not the architecture of the AI systems operating inside them, is the same gap the Robinhood-lawyer story surfaces from the compliance side and the same gap this week's leveraged-fund unwind surfaces from the market side. All three are converging on one conclusion: the questions regulators have historically asked (who holds the position, how much leverage, what product is this) are necessary but no longer sufficient once the decision-maker inside the trade is an AI system whose design choices materially change how it fails.
Read the original: European Central Bank - ECB simulations found that reinforcement-learning trading agents coordinate like a bank run under stress, while LLM-based agents behave less predictably, meaning the architecture choice itself is now a financial stability variable. Commentary is the independent editorial view of Share Trading; the original article is credited to its publisher.