Thought Leaders

Why the Market Is Moving Toward AI-Driven Toxic Flow Detection — and Why Nobody’s Quite There Yet

mm
Add Securities.io to your preferred sources on Google

Talk to enough brokers at industry events and one topic keeps coming up: a handful of accounts are quietly bleeding the book, and nobody’s quite sure how to stop it without also annoying the good clients. The industry calls this “toxic flow” — orders that, whatever the trader’s intent, tend to move against the broker in the seconds and minutes after execution. It’s an old problem. What’s newer is how often “AI” comes up as the answer in these conversations — even though, in practice, almost nobody has a mature, proven tool running in production yet.

For most of the last two decades, the go-to answer has been a rulebook: flag trades above a certain size, widen spreads ahead of news, cap exposure on accounts holding positions for a suspiciously short time. These rules aren’t a bad idea, and most desks still lean on them heavily. But they share one weakness — a rule only catches the pattern it was written for, and the traders worth worrying about are usually the ones quickest to notice what the rule is and adjust around it.

Why static thresholds keep losing ground

A lot of the well-known toxic flow patterns — latency arbitrage, spoofing, picking off stale quotes across venues — are well documented across the industry, and traders tend to read the same material brokers do. Flag anything held under a second, and traders start holding for 1.2. Widen spreads 30 seconds before big news, and the trade happens at 45. The rule doesn’t get smarter — the trader does. Worth noting that even the biggest platform providers are still mostly playing in this space rather than the AI one: MetaQuotes’ new MT5 matching engine ships with monitoring dashboards built into the manager terminal that let brokers see toxic order flow and A-book/B-book exposure in real time — genuinely useful, but still visibility and dashboards rather than predictive modeling.

There’s a subtler issue too: a fixed threshold treats every account the same, and that’s not really how risk works in practice. LiquidityFinder has a good breakdown of how brokers actually think about toxic flow, and one point is worth underlining — toxicity is relative to the business, not some fixed property of a trade. Something that’s a rounding error on one broker’s book can be a real problem for another, depending on hedging costs, client mix, how much runs A-book versus B-book. A single static rule can’t really hold two different definitions of “fine” at once. That’s the gap the industry keeps pointing to when it talks about AI — a model trained on a broker’s own execution data could, in theory, get a lot closer.

What “AI-driven” would actually mean here

Worth being precise about this, because “AI” gets attached to pretty much anything with a dashboard these days, and most of what’s marketed under that label is still fairly close to rules with extra steps. The version that would actually matter is a shift from static thresholds to models that learn from markouts — the price movement in the moments right after a trade closes — across far more variables than could reasonably be encoded by hand: order size relative to that account’s own history, timing relative to volatility, correlation with other accounts’ behavior, plus a number of smaller signals that mean almost nothing on their own but add up to a pattern together.

There’s early research pointing in this direction, though it’s worth being clear about how early. A 2023 paper on real-time toxicity prediction using Bayesian neural networks, tested against live FX transaction data, found a sequentially-trained model beat logistic regression, random forests, and standard maximum-likelihood estimators at predicting whether a trade would turn toxic — fast enough, in principle, to inform an internalize-or-externalize decision in real time. That’s a meaningful result in an academic setting. It’s a different thing entirely from a broker running this reliably in production against live flow, with all the false-positive tuning and data plumbing that requires — and as far as I can tell, very few firms are actually there yet.

Why this is harder than the pitch decks suggest

None of this makes the core tension go away. Every detection system, rules or models, has to sit somewhere on the line between catching genuinely toxic flow and flagging legitimate clients who just happen to trade in an unusual way. Finance Magnates has a good piece on how brokers identify toxic clients, and it makes a fair point — flag too aggressively and brokers end up widening spreads or cutting credit on people who were never a problem, and those clients don’t tend to stick around after that.

A model doesn’t remove that trade-off, it just changes its shape — in theory letting a broker draw the line with far more resolution than “over $X lot size” or “under Y milliseconds held.” But that only holds if the model is actually well-tuned, which requires a volume of clean, labeled historical data that a lot of mid-sized brokers simply don’t have yet, and ongoing recalibration that few risk teams are currently staffed to do. A poorly trained model will misclassify just as confidently as an outdated rule — it’ll just look more precise while doing it, which arguably makes it more dangerous to trust blindly.

Where this is probably heading

The underlying causes of toxic flow — informed institutional trading, HFT picking off pricing gaps, large aggressive orders around news — haven’t really changed in a decade. What has changed is how much historical execution data established desks now have sitting around, which is a large part of why model-based detection is even being discussed seriously now rather than five years ago, when most brokers simply didn’t have enough clean data to make it worthwhile.

I’d expect adoption to be slower and messier than the current wave of AI marketing implies. Rule-based systems aren’t going anywhere — they’re cheap, easy to explain to a compliance team, and they still catch the obvious cases fine. The more realistic path is a layered setup over the next few years: fast, explainable rules for the clear-cut cases, with model-based scoring gradually taking on the much larger grey area in between, once brokers and vendors alike have actually put in the work to make it reliable rather than just marketable. At Brokerpilot, that grey area is exactly what we’re watching most closely as we think about where risk tooling needs to go next.

Sergei Berezhnoy is Chief Revenue Officer at Brokerpilot, a risk management and dealing desk automation platform for FX/CFD brokers running MT4, MT5, and cTrader.