
Everyone says they have agentic AI. Almost no one does.
I have sat through enough vendor demos in the last year to notice the pattern. A slick chat window opens over a case. Someone types "summarize this alert," and a paragraph appears. The room nods. Somebody calls it an agent. It is not an agent. It is a summary with good manners.
This matters because the word "agentic" is now doing a lot of work in financial crime technology, and most of it is marketing. If you are going to spend budget, defend a decision to your board, and eventually explain your controls to an examiner, you need a way to cut through the language. That is what the agentic AI maturity model gives you. It is a simple way to place any system, and any vendor claim, on a spectrum from basic automation to genuine autonomy.
The framework below is grounded in independent research from Chartis, which co-produced a full report on this with us. I am going to walk through the four levels the way I would explain them to a head of financial crime who has ten vendor calls next week and no patience for hype.
The pace of AI claims has outrun the market's ability to check them. That is the honest state of things in 2026. Vendors moved fast because buyers rewarded the word, not the substance. So "agentic" got stapled onto copilots, onto rules engines with a chat box, onto retrieval systems that fetch a document and call it reasoning.
There is even a name for the two versions of this. "AI washing" is overstating how much AI is really in the product. "Agent washing" is the narrower and more current one: claiming autonomous agents while shipping automation and chatbots underneath. Both work for exactly as long as the buyer has no framework to test against.
The fix is not skepticism for its own sake. Plenty of this technology is real and materially better than what most teams run today. The fix is a shared vocabulary, so that when a vendor says "Level 4," you know what they are claiming and what to ask next.
The maturity model has four levels. The jump that matters is not level one to two. It is the jump from a system that informs the work to a system that performs the work.
Deterministic workflows triggered by predefined conditions. No AI reasoning is involved at all. Auto-routing alerts by risk score, batch screening against watchlists, threshold-based blocking, and scheduled alert generation on a 24-hour cycle. This is most of what legacy systems still do, and it is genuinely useful. It is just not AI, and no one should let it be sold as such.
AI surfaces information, writes summaries, and recommends actions. A human still performs every step and makes every decision. This is the chat window over your case data. It is a single-turn summary of alerts. It is the retrieval that pulls a document for the analyst to read. A copilot can be a real time-saver. It is also where most vendors claiming "agentic AI" actually sit today, and the distance between here and true autonomy is larger than the demo makes it look.
Now the system executes multi-step work. The agent pulls transaction histories, runs open-source research, checks watchlists, assembles the evidence package, runs anomaly detection on counterparty behavior, and drafts the narrative. The human reviews the completed work and approves the disposition. The analyst stops doing the assembly and starts doing the judgment. This is the production frontier in 2026. It is where the leading platforms operate, and where Chartis measures handle-time reductions north of 80 percent when it is done well.
The system executes end-to-end within governed constraints. It self-evaluates against defined success criteria, and it improves through feedback loops over time. Human oversight shifts from reviewing every decision to handling exceptions and sampling for quality. Clear false positives get closed automatically inside configured guardrails. Detection and investigation start to connect, so investigation outcomes feed back into the rules. This level is emerging, not everywhere, and any vendor claiming they are fully here across every use case deserves a hard follow-up question.
Here is the uncomfortable part, and it is Chartis's finding, not mine. Most vendors claiming agentic AI are operating at Level 2 or the early stages of Level 3. The gap between the pitch and the production reality is the single biggest risk in this buying cycle.
That gap is expensive in a specific way. False positive rates in AML sit above 90 percent for most institutions, which means the overwhelming majority of analyst time is spent clearing noise. A copilot makes clearing that noise slightly faster. A Level 3 agent removes most of the clearing altogether, which is the entire point of building transaction monitoring around agents rather than bolting a chat box onto it. Those are not the same purchase, and they should not carry the same price or the same expectations.
If you want to go deeper on why more analysts have never solved this and why the math only changes at Level 3, we wrote a practitioner's guide to agentic AI for AML compliance that walks through the workflow step by step. The short version is that reducing false positives is a structural problem, not a staffing one.
Three reasons this framework earns its place in your evaluation.
Procurement and budget. If you can place a vendor at a level, you can price the purchase correctly and set expectations that your team will actually hit. Paying Level 4 money for a Level 2 copilot is how programs lose credibility internally.
Regulatory defensibility. The evaluation question in agentic AI shifts from "is the model accurate?" to "what is the agent authorized to do, what triggers an escalation, and can it fail in ways I did not predict?" That is agent risk, and it is a different discipline from model risk. Knowing the level tells you how much of that discipline you need in place before go-live.
Real progress. The teams that win are not the ones that buy the biggest claim. They are the ones that start at Level 2 or 3 on a single high-volume workflow, validate accuracy against human review, and move up the autonomy ladder as the evidence comes in. Maturity is something you climb deliberately, not something you procure in one purchase order. This is also why rules and AI work better together than either alone, and why the platforms that let your own team configure the AI agents will outrun the ones that gate every change behind a support ticket.
Before your next vendor call, pick one alert type and ask yourself a plain question. On that alert, does their system do the work, or does it just help my analyst do the work faster? Everything else is detail. That one distinction is the line between a copilot and an agent, and it is the line most of this market is quietly hoping you will not draw.
If you want the full framework, including the seven dimensions to test each vendor on and the red flags that give away agent washing, download the Chartis report we co-produced, The Agentic AI Maturity Spectrum in Financial Crime. And if you want to see how Chartis assesses the vendor landscape on its own, their 2026 RiskTech Quadrant vendor spotlight covers that separately.
What is the agentic AI maturity model? It is a four-level framework for classifying AI systems in financial crime, from Level 1 rule-based automation, to Level 2 AI-assisted copilots, to Level 3 agents that execute investigations with human oversight, to Level 4 autonomous agents that run end-to-end within guardrails. It exists so buyers can test vendor claims against a common standard.
What level are most vendors actually at? According to Chartis, most vendors claiming agentic AI are at Level 2 or early Level 3, even when the marketing says Level 4. The defining test is whether the system performs the work or merely informs it.
What is the difference between a copilot and an agent? A copilot surfaces information and recommends actions, but a human performs every step. An agent executes the multi-step workflow itself and hands back completed work for review. The copilot informs. The agent performs.
Is Level 4 the goal for every team? Not immediately. Most teams should start at Level 2 or 3 on one workflow, prove accuracy against human review, then raise autonomy as confidence builds. Jumping straight to full autonomy without that evidence is how you inherit risk you cannot defend.

Tyler Allen is the CEO of Unit21 and was the company’s first hire, writing some of the first lines of code seven years ago. He previously led Unit21’s AI team as Head of AI, then served as COO, before stepping into the CEO role. He is a driving force behind Unit21’s vision as the leader in AI risk infrastructure, having led the AI team before becoming COO. A deep technical leader, Tyler recently returned to the codebase to personally build AI agent configurations, pairing his technical expertise with seven years of experience observing how compliance teams operate.