
How much control do you actually want to hand an AI agent on day one? For most compliance and fraud teams, the honest answer is: not much, yet. That instinct is correct, and it's exactly what progressive autonomy is designed for. Rather than treating autonomy as a single on/off switch, progressive autonomy breaks AI deployment into five deliberate levels, a trust ladder your team climbs one rung at a time, per queue, with evidence at every step. Here's what each rung actually looks like, and why it isn't the same question as how sophisticated the underlying agent is.
We've written before about the four levels of agentic AI in financial crime, which classify how capable an AI system is: whether it's still a copilot, genuinely agentic with human oversight, or fully autonomous end to end. Progressive autonomy is a different axis. It assumes you already have a genuinely agentic system and asks a separate question: how much of that system's output are you, today, willing to act on without a human checking it first?
A highly capable agent can sit at the lowest rung of the trust ladder in your queue for months, on purpose. That's not the agent underperforming. That's the process working as designed.
The trust ladder has five levels:
Every level above L0 keeps a human in the loop somewhere. The ladder just changes where in the process that human sits.
Two mechanisms run underneath every level of the ladder, not just the top ones:
Random sampling: pulling a sample of cases where the AI agent and analyst agreed, and a sample where they disagreed, for ongoing QA. This is what lets you catch drift before it becomes a pattern, not after.
Backtesting: running a newly configured agent against historical alerts before it ever touches a live queue, so you have evidence of how it performs before it has the chance to underperform on something real.
These two guardrails are also why L4 auto-close doesn't mean the agent is deciding on its own without a net. If you want the deeper argument for why false positive reduction has to be evidence-based rather than just faster, how to reduce false positives in AML transaction monitoring covers that ground well.
In practice, most teams follow a crawl, walk, run pattern. They start in AI assist mode, L1 or L2, where the analyst prompts the agent and validates the output against their own judgment before trusting it on anything else. Once the output is consistently trusted, usually somewhere between one and two months in, teams move the same queue to auto-review: the agent runs automatically the moment an alert lands, so by the time an analyst opens it, the summary and evidence are already sitting there.
That shift, from "let me ask the agent" to "the agent already ran," is where teams typically see time-to-investigate drop the most. It's not a bigger model. It's more trust, earned one rung at a time.
The ladder is configurable per queue, not platform-wide, and customers never start at the top. A team might run L3 on a well-understood alert type with years of historical dispositions behind it, and L1 on a newer, less-tested one, at the same time. That's the design working as intended, not an inconsistency to fix.
Human-in-the-loop is the default here, not a hedge added for regulators' sake. Your existing QA and QC processes still apply to an AI agent's output the same way they'd apply to a human analyst's. Interestingly, some teams are now running that relationship in both directions: using AI to check human analyst work for consistency and SOP adherence, not only the reverse. You own the outcome either way. A real-world example of what an L3 disposition recommendation looks like in production today is SAR narrative automation, where the agent drafts the narrative and extracts the information, and a human reviews before anything is filed. If you're building the underlying tasks that feed a specific rung of the ladder, how custom AI agents are transforming fraud and AML operations is the practical next read.
Does progressive autonomy mean the AI can't be trusted at first?
No. It means trust is built deliberately, with evidence, rather than assumed. Most agents are capable of more than L1 from day one; the ladder is about how much of that capability your team has verified against its own data.
Can autonomy levels differ by queue?
Yes. A queue with years of historical dispositions to backtest against might reach L3 quickly, while a newer alert type stays at L1 or L2 until there's enough evidence to move it up.
What stops an agent at L4 from closing something it shouldn't?
L4 auto-close is scoped to defined, low-risk case types only, backed by backtesting before go-live and ongoing random sampling afterward. It's not a general autonomy switch.
Who is accountable if an AI agent's disposition turns out to be wrong?
Your team is, in the same way it would be for a human analyst's decision. The audit trail shows exactly what the agent looked at and why, which makes that accountability easier to demonstrate, not harder.
How long does it typically take to move from AI assist to auto-review?
Most teams see one to two months of AI assist use before moving a queue to auto-review, though this depends on alert volume and how much historical data is available to backtest against.
Progressive autonomy isn't a compliance requirement bolted onto AI as an afterthought. It's the mechanism that lets a team move from skeptical to confident without ever having to take that leap on faith. See how Unit21's AI Agent supports this rollout, queue by queue.

Gal Perelman is the Product Marketing Lead at Unit21, where she spearheads go-to-market strategies for AI-driven risk and compliance solutions. With over a decade of experience in the fintech and fraud sectors, she has led high-impact launches for products like Watchlist Screening and AI Rule Recommendations.
Previously, Gal held marketing leadership roles at Design Pickle, Sightfull, and Lusha. She holds a Master’s degree from American University and a Bachelor’s from UCLA, and is dedicated to helping banks and fintechs navigate complex regulatory landscapes through innovative technology.