.avif)
A Chartis Research framework for evaluating what agentic AI actually means for AML, fraud, sanctions, and more, plus a practical guide for buyers navigating a market full of vendor claims.
Risk and compliance is in a major technology transition. AI, both generative and agentic, is now the primary force shaping buyer priorities. But vendor claims have outstripped the market’s ability to evaluate them.
Vendors claim “agentic” capabilities that, in Chartis Research’s assessment, are not truly agentic. The risk for buyers: procuring systems that don’t deliver, building false confidence in automation that isn’t there, and missing the window to adopt materially better approaches.
This report gives buyers a tactical framework to evaluate vendor AI maturity claims. It defines what agentic AI is and is not, applies a 4-level maturity spectrum to 5 core use cases, and gives specific, testable questions to ask any vendor.
Rather than a binary “agentic vs. not” distinction, Chartis Research developed a 4-level maturity spectrum. Use it to classify vendor claims and your own current state of innovation.
Deterministic workflows triggered by predefined conditions. No AI reasoning or capability is involved.
Where most legacy systems still operate
AI surfaces information, generates summaries, or recommends actions, but humans perform every step and make every decision.
Where most vendors claiming "agentic AI" actually are
AI agents execute multi-step workflows with tool calls and structured reasoning. Humans approve decisions but don’t perform investigation steps.
Production frontier in 2026, where leading platforms operate
AI systems execute end-to-end workflows within governed constraints, self-evaluate against success criteria, and improve via feedback loops. Human oversight shifts to exceptions and QA.
Emerging: the trajectory for the next 12–18 months
Chartis Research’s definition is precise. Knowing it helps distinguish production-grade systems from marketing claims.
Truly agentic AI systems do the following:
The key word is autonomous execution. If the system merely informs the work rather than performing it, it is not agentic AI — regardless of what it's called.
Agent washing, claiming agentic AI while delivering automation or chatbots, is widespread in the 2026 FinCrime market. These are the red flags that reveal it.
Chartis Research specifies 7 attributes that buyers should use to evaluate agentic AI vendors. For each dimension, ask vendors to demonstrate, not describe.
Each use case has a distinct maturity trajectory and different evaluation criteria. Most platforms don't serve all 5 at Level 3+, so knowing which ones they do is essential.
Fraud typologies evolve at a pace that static rule sets cannot match. A Level 3 fraud investigation agent executes the full evidence-gathering workflow — device data analysis, counterparty behavior checks, external lookups, narrative drafting in the institution's required format — without requiring the analyst to perform any step.
A critical differentiator at Level 4 is self-service agent building: compliance and fraud teams can define, modify, and extend investigation tasks themselves without an engineering dependency on the vendor. If a new fraud typology emerges, the team can add an investigation task the same day — not wait weeks for a vendor support engagement.
Protect approvals, conversions, and customer experience while blocking losses in real-time. Unit21 combines device intelligence, graph analysis, and AI-driven investigation to catch fraud before it settles, so you drive growth, not friction.



50% reduction in false positives, improvement in alert quality, & catching 2x the fraud
From rule execution and signal detection to investigation, narrative drafting, and regulatory filing, our AI agents orchestrate the entire AML lifecycle in a single place.







Modernized global AML With 57% AI-driven alert automation & 93% reduction in false positives
AML investigation is the highest-volume, highest-cost use case in financial crime compliance. Rule-based systems generate alerts when transactions match predefined patterns. Human analysts review, investigate, and disposition alerts. False positive rates of 90%+ mean the overwhelming majority of analyst time is spent on noise.Level 3 agentic AI fundamentally changes this math.
Rather than assisting the analyst, the AI agent executes the investigation: pulling transaction histories, conducting open-source research, checking watchlists, assembling evidence packages, running anomaly detection on counterparty behavior, and drafting investigation narratives. The analyst reviews the completed work and approves the disposition.
Fraud typologies evolve at a pace that static rule sets cannot match. A Level 3 fraud investigation agent executes the full evidence-gathering workflow — device data analysis, counterparty behavior checks, external lookups, narrative drafting in the institution's required format — without requiring the analyst to perform any step.
A critical differentiator at Level 4 is self-service agent building: compliance and fraud teams can define, modify, and extend investigation tasks themselves without an engineering dependency on the vendor. If a new fraud typology emerges, the team can add an investigation task the same day — not wait weeks for a vendor support engagement.
Network analysis spans two fundamentally different capabilities that are often conflated. The first is internal graph analysis: entity resolution, relationship mapping, and network visualization within a single institution's data. The second is cross-institutional network intelligence: signals derived from investigation outcomes and fraud typology classifications aggregated across multiple institutions.
Individual institutions are fighting fraud and money laundering with one hand tied behind their back — they only see their own data. The criminal sees all of them. A Level 3+ consortium provides proactive entity flagging based on cross-network intelligence: catching bad actors before they transact on your platform, because another institution has already identified them.
Sanctions screening generates more false positives than almost any other workflow in financial crime compliance — driven by name variations, transliteration differences, and the complexity of PEP relationship analysis. Level 3 agents dramatically reduce the manual review burden by autonomously evaluating match quality with corroborating evidence rather than presenting every potential match to a human reviewer.
A critical distinction for buyers: sanctions screening (binary list check) and PEP screening (relationship analysis) require fundamentally different agent architectures. A generalist agent running both in the same context window will underperform a specialized one that scopes context to the exact decision being made.
Customer onboarding involves collecting information from prospective customers or businesses, confirming validity, completing due diligence analysis, and executing a customer risk rating. The risk rating determines the period review frequency and whether enhanced due diligence is required. Level 3 agents complete the full data collection workflow — the analyst reviews and approves the completed record rather than building it.
The Level 4 frontier for onboarding extends to autonomous customer risk rating execution and CDD/EDD determination — the agent doesn't just collect data, it executes the risk rating process within the case management workflow.
The strategic question should shift from “is the model accurate?” to questions such as “what actions is the agent authorized to take, what are the escalation triggers, and can it fail unpredictably?” Here’s how to build accountability:
Compliance tasks run at or near zero temperature to minimize creative, hallucinated output. A baseline requirement, not a differentiator.
Agents write SQL or Python to query specific fields instead of letting the LLM read raw data. The model reasons, the code executes, which sharply cuts data-accuracy errors.
Test datasets with known answers validate accuracy before deployment. The gold standard is real human-reviewed decisions, benchmarked against your best analysts, not synthetic data or average performance.
Model quality degrades as the context window fills, so engineering the minimum sufficient context per task beats dumping in everything. Chartis lists it twice because it's the primary technical differentiator.
Unit21 provides AI Risk Infrastructure for a wide varity of fintechs and financial institution , with over 200 customers across 90 countries. The platform's AI agents have reviewed more than 1.5 million alerts in production.
Chartis Research identifies six trends that will shape the market over the next 12–24 months.
Download the complete Agentic Maturity Spectrum report including a Capability Matrix across 5 use cases, the vendor evaluation scorecard, and Chartis Research's calls to action for buyers and vendors.
Agentic AI in financial crime refers to AI systems that autonomously execute multi-step investigation and compliance workflows — pulling transaction histories, checking watchlists, analyzing counterparty behavior, drafting narratives, and filing reports — within governed constraints and with configurable human oversight. This is distinct from AI-assisted (copilot) systems, which surface information but require humans to perform every execution step. The defining characteristic of agentic AI is that it performs the work rather than merely informing it.
The agentic AI maturity spectrum is a Chartis Research framework defining four levels of AI system sophistication in financial crime: Level 1 (rule-based automation), Level 2 (AI-assisted copilot), Level 3 (agentic with human oversight), and Level 4 (autonomous agentic). It was developed through Chartis Research's independent analysis of the FinCrime technology market in 2026, with the goal of giving buyers a practical tool for evaluating vendor claims and setting realistic expectations for what AI can deliver today vs. on the horizon.
Agent washing is the practice of claiming agentic AI capabilities while actually delivering automation, chatbots, or basic AI-assisted workflows. It's the AI-specific version of "AI washing." Red flags include: the vendor cannot explain how context is scoped per task; there are no evaluation datasets from real human analyst decisions; the "AI agent" is a chatbot interface over a data lake with no structured workflow execution; there is no configurable autonomy ladder (it's all-or-nothing); the vendor cannot produce accuracy metrics benchmarked against human analysts; audit trails show decisions without reasoning chains or evidence citations; and the system cannot generate deterministic code (SQL/Python) for data queries. If a vendor hits three or more of these, they are likely at Level 2 regardless of their marketing.
What is context engineering, and why does Chartis call it the most important differentiator?
At Level 3 (agentic with human oversight), AI agents execute multi-step investigation workflows autonomously — gathering evidence, running analyses, drafting narratives — but require human approval at decision points. The human reviews the AI-assembled evidence package and makes the final disposition decision. At Level 4 (autonomous agentic), AI systems execute end-to-end workflows including disposition within governed constraints, self-evaluate against defined success criteria, and improve results via feedback loops. Human oversight shifts from per-decision review to exception handling and periodic QA sampling. Level 3 is the production frontier in 2026. Level 4 is emerging, with leading platforms beginning to deploy it for specific queues and use cases.
