Analyst Research Report

The Agentic Maturity Spectrum in Financial Crime

A Chartis Research framework for evaluating what agentic AI actually means for AML, fraud, sanctions, and more, plus a practical guide for buyers navigating a market full of vendor claims.

4

Maturity levels defined by Chartis Research

5

FinCrime use cases analyzed in depth

7

Evaluation dimensions buyers should test
About the Report

Why Chartis Research wrote this

Risk and compliance is in a major technology transition. AI, both generative and agentic, is now the primary force shaping buyer priorities. But vendor claims have outstripped the market’s ability to evaluate them.

Vendors claim “agentic” capabilities that, in Chartis Research’s assessment, are not truly agentic. The risk for buyers: procuring systems that don’t deliver, building false confidence in automation that isn’t there, and missing the window to adopt materially better approaches.

This report gives buyers a tactical framework to evaluate vendor AI maturity claims. It defines what agentic AI is and is not, applies a 4-level maturity spectrum to 5 core use cases, and gives specific, testable questions to ask any vendor.

What the report covers

  • A 4-level maturity spectrum from rule-based automation to fully autonomous agentic AI
  • 7 dimensions for evaluating agentic AI vendors, with specific questions for each
  • A guide to spotting “agent washing” and the red flags that reveal non-agentic claims
  • The 4 pillars of AI reliability and how to assess them in production
  • And much more

Cut through the noise. Learn how to evaluate a vendor’s AI claims.

Download the Report
The Framework

The 4-level agentic AI maturity spectrum

Rather than a binary “agentic vs. not” distinction, Chartis Research developed a 4-level maturity spectrum. Use it to classify vendor claims and your own current state of innovation.

Rule-Based
Automation

Deterministic workflows triggered by predefined conditions. No AI reasoning or capability is involved.

Where the market is →

Where most legacy systems still operate

AI-Assisted
(Copilot)

AI surfaces information, generates summaries, or recommends actions, but humans perform every step and make every decision.

Where the market is →

Where most vendors claiming "agentic AI" actually are

Agentic with
Human Oversight

AI agents execute multi-step workflows with tool calls and structured reasoning. Humans approve decisions but don’t perform investigation steps.

Where the market is →

Production frontier in 2026, where leading platforms operate

Autonomous
Agentic

AI systems execute end-to-end workflows within governed constraints, self-evaluate against success criteria, and improve via feedback loops. Human oversight shifts to exceptions and QA.

Where the market is →

Emerging: the trajectory for the next 12–18 months

Definition

What agentic AI actually means

Chartis Research’s definition is precise. Knowing it helps distinguish production-grade systems from marketing claims.

The Chartis Definition

“Agentic AI” refers to AI systems designed to pursue defined goals autonomously by planning, making contextual decisions, and executing multi-step actions within governed constraints — often orchestrating multiple models, data sources, and workflows.

Truly agentic AI systems do the following:

  • Exhibit proactive, goal-directed behavior: they initiate tasks, not just respond
  • Operate within an agent-oriented architecture that decomposes complex processes into individualized agents with scoped context
  • Adapt to changing conditions and manage end-to-end processes with limited human intervention
  • Remain aligned to policy, regulatory, and explainability requirements throughout
  • Amplify human decision-making on high-value, complex workflows

The key word is autonomous execution. If the system merely informs the work rather than performing it, it is not agentic AI — regardless of what it's called.

What Agentic is Not

  • A chatbot or copilot is not agentic.

  • A single LLM call is not agentic.

  • Rule-based workflow automation is not agentic AI.

  • RAG alone is not agentic AI.

Buyer Alert

Agent washing: how to spot it

Agent washing, claiming agentic AI while delivering automation or chatbots, is widespread in the 2026 FinCrime market. These are the red flags that reveal it.

Cannot explain context scoping 
per task

No evaluation datasets from real human decisions

The “agent” is a chatbot over a 
data lake

No configurable autonomy ladder

Audit trails 
lack reasoning chains

System can’t generate deterministic code 
for data queries

Receive a complete list of red flags & a structured vendor scorecard.

Download the Report
Evaluation Framework

7 dimensions for evaluating agentic AI

Chartis Research specifies 7 attributes that buyers should use to evaluate agentic AI vendors. For each dimension, ask vendors to demonstrate, not describe.

Most Important Differentiator
Context Engineering

Precision That Cuts Through the Noise

Context engineering is the discipline of optimizing exactly what information is provided to an AI model for a specific decision. Too little context and it can't decide correctly. Too much and performance degrades. Production-grade systems scope context per task. A sanctions-check agent receives only the entity data, watchlist matches, and relationship graph for that check — not the full transaction history used for fraud. This distinction is the primary differentiator between a production system and a demo. Chartis elevates this to first-class status in the evaluation framework.
Goal Decomposition & Modularity

Agentic tasks are broken into discrete, verifiable sub-goals.

Workflows are structured so individual components can be delegated, retried, or swapped without cascading failures. Critically: institution-controlled configuration vs. vendor-controlled is one of the clearest indicators of platform maturity. If deploying a custom agent task requires a support ticket or professional services engagement, the system will not scale.
Control & Oversight

The Autonomy Ladder

Mature platforms offer a configurable autonomy ladder — not a binary on/off. At minimum, four levels: (1) On-demand only; (2) Auto-review with human decision; (3) Auto-review with decision recommendation; (4) Auto-close within guardrails. These levels must be configurable per queue, use case, and risk tier — not applied globally. Graceful escalation paths define what happens when the agent hits ambiguity or low confidence: it escalates, it doesn't guess.
Tool & Resource Fit

Multi-Model Orchestration

Production-grade agentic systems do not rely on a single LLM. Different models excel at different tasks: document analysis, code generation, entity extraction, narrative drafting. Mature platforms dynamically select the best-performing model per task and can swap models as benchmarks change — testing against evaluation sets before production deployment. Ask: how many models, how are they selected, what happens when benchmark performance shifts?
Feedback & Iteration

Tight Feedback Loops

Analyst feedback should aggregate into decision criteria improvements over time — without retraining the underlying models. The most mature deployments implement dual-direction QA: AI checks human analysts' decisions for consistency, and humans check AI outputs. Both directions matter. Buyers should ask for specific accuracy metrics benchmarked against high-performing analysts, not average human performance.
Governance & Security

Least-Privilege Execution

Agents operate with a defined identity and minimum permissions required for each task. Scope constraints define what the agent can and cannot do. When uncertain, the default action is safe and conservative — not creative. On regulatory compliance: buyers should ask whether the vendor is in compliance with the EU AI Act in production today, not on a roadmap. Ask for documentation on demand.
Reliability

Evaluation Datasets, Determinism, and Temperature Controls

Production AI systems set model temperature to near-zero for compliance tasks to minimize hallucinations. Agents generate deterministic code (SQL, Python) to query data rather than having the LLM interpret raw data. And — critically — evaluation datasets should be derived from years of real human analyst behavior, not synthetic data. Buyers should ask: what is your eval set, how many examples, what accuracy threshold must be met before deployment, and can I see the results for my specific configuration?
Use Case Analysis

Agentic AI across 5 fincrime use cases

Each use case has a distinct maturity trajectory and different evaluation criteria. Most platforms don't serve all 5 at Level 3+, so knowing which ones they do is essential.

2026 Frontier
Level 3  →  early Level 4
Key Differentiator
Self-service agent building — fraud teams define tasks without engineering

Fraud Prevention & Investigation

Fraud typologies evolve at a pace that static rule sets cannot match. A Level 3 fraud investigation agent executes the full evidence-gathering workflow — device data analysis, counterparty behavior checks, external lookups, narrative drafting in the institution's required format — without requiring the analyst to perform any step.

A critical differentiator at Level 4 is self-service agent building: compliance and fraud teams can define, modify, and extend investigation tasks themselves without an engineering dependency on the vendor. If a new fraud typology emerges, the team can add an investigation task the same day — not wait weeks for a vendor support engagement.

What Agentic is Not

  • Can the fraud team build custom investigation tasks without a vendor engineering ticket?
  • Does the agent produce three outputs: recommended action, regulator-ready narrative, and transparent worklog?
  • Can autonomy levels be configured differently per fraud type?
  • How does the agent adapt to new fraud typologies that don't yet have rules?
  • How is context scoped differently for a device-fraud investigation vs. a payment fraud investigation?
  • Does investigation outcome data feed back into detection logic automatically?

Unit21 for Fraud 

Fraud protection built for revenue, not just risk

Protect approvals, conversions, and customer experience while blocking losses in real-time. Unit21 combines device intelligence, graph analysis, and AI-driven investigation to catch fraud before it settles, so you drive growth, not friction.

Real-Time Monitoring
Device Intelligence
Fraud Consortium

50% reduction in false positives, improvement in alert quality, & catching 2x the fraud

Read the story
Unit21 for AML 

AI That Runs Your AML Program, End to End

From rule execution and signal detection to investigation, narrative drafting, and regulatory filing, our AI agents orchestrate the entire AML lifecycle in a single place.

Transaction Monitoring
Case Management
Payment Screening
Sanction Screening
Customer Risk Rating
Regulatory Filings

Modernized global AML With 57% AI-driven alert automation & 93% reduction in false positives

Read the story
2026 Frontier
Level 3 → early Level 4
Key Metric
80%+ handle time reduction at L3

AML / BSA Investigation

AML investigation is the highest-volume, highest-cost use case in financial crime compliance. Rule-based systems generate alerts when transactions match predefined patterns. Human analysts review, investigate, and disposition alerts. False positive rates of 90%+ mean the overwhelming majority of analyst time is spent on noise.Level 3 agentic AI fundamentally changes this math.

Rather than assisting the analyst, the AI agent executes the investigation: pulling transaction histories, conducting open-source research, checking watchlists, assembling evidence packages, running anomaly detection on counterparty behavior, and drafting investigation narratives. The analyst reviews the completed work and approves the disposition.

Some questions to ask:

  • How is context scoped for the transaction history pull vs. the watchlist check? Are these the same agent or different ones?
  • What is your auto-close rate for clear false positives, and what is the accuracy rate vs. human analysts?
  • Can autonomy levels be configured independently per queue, jurisdiction, and risk tier?
  • Does the AI agent draft narratives in our institution's required format, or is it a generic template?
  • How does analyst feedback on dispositions improve the agent's decision criteria over time?
  • What does the audit trail include? Reasoning chain, evidence citations, and confidence scores — or just the decision?

2026 Frontier
Level 3 → early Level 4
Key Differentiator
Self-service agent building — fraud teams define tasks without engineering

Fraud Prevention & Investigation

Fraud typologies evolve at a pace that static rule sets cannot match. A Level 3 fraud investigation agent executes the full evidence-gathering workflow — device data analysis, counterparty behavior checks, external lookups, narrative drafting in the institution's required format — without requiring the analyst to perform any step.

A critical differentiator at Level 4 is self-service agent building: compliance and fraud teams can define, modify, and extend investigation tasks themselves without an engineering dependency on the vendor. If a new fraud typology emerges, the team can add an investigation task the same day — not wait weeks for a vendor support engagement.

What Agentic is Not

  • Can the fraud team build custom investigation tasks without a vendor engineering ticket?
  • Does the agent produce three outputs: recommended action, regulator-ready narrative, and transparent worklog?
  • Can autonomy levels be configured differently per fraud type?
  • How does the agent adapt to new fraud typologies that don't yet have rules?
  • How is context scoped differently for a device-fraud investigation vs. a payment fraud investigation?
  • Does investigation outcome data feed back into detection logic automatically?

2026 Frontier
Level 2 → Level 3
Structural Moat
Cross-institutional consortium signals compound in value over time

Network Analysis & Consortium Intelligence

Network analysis spans two fundamentally different capabilities that are often conflated. The first is internal graph analysis: entity resolution, relationship mapping, and network visualization within a single institution's data. The second is cross-institutional network intelligence: signals derived from investigation outcomes and fraud typology classifications aggregated across multiple institutions.

Individual institutions are fighting fraud and money laundering with one hand tied behind their back — they only see their own data. The criminal sees all of them. A Level 3+ consortium provides proactive entity flagging based on cross-network intelligence: catching bad actors before they transact on your platform, because another institution has already identified them.

Key Evaluation Questions for Network Analysis

  • Does your consortium include investigation outcomes and fraud typology classifications, or just binary good/bad flags?
  • Does the network AI agent see cross-institutional signals, or only your institution's data?
  • What is the diversity of the consortium? Cross-vertical and cross-typology coverage — or single-sector?
  • Can the AI agent proactively flag entities based on cross-network intelligence before they transact on our platform?
  • How are investigation outcomes from one institution used to improve signals for others?

2026 Frontier
Level 3
Core Pain
Extremely high false match rates due to name variations and transliteration

Sanctions Screening

Sanctions screening generates more false positives than almost any other workflow in financial crime compliance — driven by name variations, transliteration differences, and the complexity of PEP relationship analysis. Level 3 agents dramatically reduce the manual review burden by autonomously evaluating match quality with corroborating evidence rather than presenting every potential match to a human reviewer.

A critical distinction for buyers: sanctions screening (binary list check) and PEP screening (relationship analysis) require fundamentally different agent architectures. A generalist agent running both in the same context window will underperform a specialized one that scopes context to the exact decision being made.

What Agentic is Not

  • Can the agent distinguish between sanctions screening (binary list check) and PEP screening (relationship analysis)?
  • What is the auto-close rate for clear non-matches, and what is the accuracy rate vs. human review?
  • How quickly does the system incorporate new sanctions designations?
  • Is the agent specialized for sanctions, or a general-purpose agent running sanctions alongside unrelated tasks
  • What corroborating evidence does the agent check — and how is that context scoped?

2026 Frontier
early Level 4
Key Question
Does the agent integrate with the risk rating process — or just collect data?

Customer Onboarding (KYC/KYB)

Customer onboarding involves collecting information from prospective customers or businesses, confirming validity, completing due diligence analysis, and executing a customer risk rating. The risk rating determines the period review frequency and whether enhanced due diligence is required. Level 3 agents complete the full data collection workflow — the analyst reviews and approves the completed record rather than building it.

The Level 4 frontier for onboarding extends to autonomous customer risk rating execution and CDD/EDD determination — the agent doesn't just collect data, it executes the risk rating process within the case management workflow.

What Agentic is Not

  • Can the AI agent consume and correctly interpret customer onboarding data from multiple sources simultaneously?
  • Does the AI agent integrate with the customer risk rating process, including CDD vs. EDD determination?
  • How does the agent handle potentially conflicting data from two or more distinct sources?
  • Can the agent adapt its data collection based on account type or customer segment?
  • What is the handoff point between agent execution and human review — and is it configurable?

Discovery Call Playbook

The questions that reveal maturity

Level 2 and Level 3+ vendors give fundamentally different answers to the same questions. Use this table as a discovery call guide, or to assess your existing platform.

What to ask Level 2 response Level 3+ Response
How are investigations broken into steps? “AI summarizes the alert.” “5–8+ specialized agents per alert type, running in parallel, with distinct context windows and toolsets per task.”
How do you scope data for each task? “We send all available data to the model.” “Each task has engineered context based on what analysts actually use. We know this from years of analyst behavior data.”
How does the agent query data? “The LLM reads the data directly.” “Agents generate SQL/Python to deterministically query specific fields. The LLM reasons; code executes.”
How do you validate accuracy? “We test on sample data.” “Evaluation datasets from N years of human-reviewed data, benchmarked against top analysts specifically — not average performance.”
Can I configure automation levels? “The human always decides.” “Four levels configurable per queue, process, team, or jurisdiction — independent of each other.”
How many models do you use? “We use [single model name].” “We orchestrate N models, dynamically selected per task, updated continuously as benchmarks shift.”
How does analyst feedback improve the AI? “We update our models periodically.” “Analyst feedback aggregates into decision criteria improvements. Admins tune agents self-service — no engineering dependency.”

Cut through the noise. Learn how to evaluate a vendor’s AI claims.

Get the Report
Governance & Trust

The accountability framework for agentic AI

The strategic question should shift from “is the model accurate?” to questions such as “what actions is the agent authorized to take, what are the escalation triggers, and can it fail unpredictably?” Here’s how to build accountability:

The four pillars of AI reliability

Temperature and Determinism Controls

Compliance tasks run at or near zero temperature to minimize creative, hallucinated output. A baseline requirement, not a differentiator.

Deterministic Code Generation

Agents write SQL or Python to query specific fields instead of letting the LLM read raw data. The model reasons, the code executes, which sharply cuts data-accuracy errors.

Evaluation Datasets and Golden-Set Benchmarking

Test datasets with known answers validate accuracy before deployment. The gold standard is real human-reviewed decisions, benchmarked against your best analysts, not synthetic data or average performance.

Context Engineering

Model quality degrades as the context window fills, so engineering the minimum sufficient context per task beats dumping in everything. Chartis lists it twice because it's the primary technical differentiator.

What production accountability looks like

  • Random sampling of AI outputs for defensibility review at a level of granularity with the sample
  • AI QA of human analyst decisions for consistency and standards of practice adherence
  • Full audit trails with reasoning chains, evidence citations, and confidence scores
  • Configurable escalation triggers that pause the agent when conditions exceed defined thresholds
  • Least-privilege execution: agents operate with minimum permissions required for each specific task
  • EU AI Act compliance in production (not on a roadmap) with documentation available on demand
Chartis’ evaluation of Unit21

How Unit21 maps to the framework

Unit21 provides AI Risk Infrastructure for a wide varity of fintechs and financial institution , with over 200 customers across 90 countries. The platform's AI agents have reviewed more than 1.5 million alerts in production.

Context Engineering at Scale

Unit21's context engineering is built on years of analyst behavior data across 1.5M+ alerts. This data tells the platform exactly what information analysts actually use when making each type of decision, and that knowledge is baked into context scoping for every agent task.

Deterministic Code Generation

Agents generate SQL/Python to query data fields deterministically. The LLM reasons; code executes. This hybrid approach minimizes hallucinations and produces audit trails that are regulatorily defensible.

Multi-Model Orchestration

Multiple LLMs dynamically selected per task type. Updated as benchmarks shift. Each model change is tested against evaluation sets before production deployment. No single-model dependency.

Progressive Autonomy


Five autonomy levels configurable per queue, risk tier, and use case, independently. Start at Level 2, validate accuracy in parallel with human review, move to Level 3–4 as confidence is built.

Dual-Direction QA 


AI audits human analyst decisions for consistency. Humans review AI outputs. Feedback from both directions aggregates into decision criteria improvements — no model retraining required.

Self-Service Agent Building

Compliance and fraud teams define, modify, and extend investigation tasks without engineering dependency on Unit21. A new fraud typology doesn't require a support ticket; it requires a team that can configure the platform.

Consortium Network Intelligence

Unit21's Fraud Consortium spans banks, fintechs, and crypto platforms covering 100M+ U.S. consumers. It shares cross-typology risk signals, not binary flags. Value compounds with every member: a bad actor caught at one is flagged before they transact at the next.

Forward-Looking

What to watch in agentic AI for financial crime

Chartis Research identifies six trends that will shape the market over the next 12–24 months.

Agent-to-Agent Commerce

AI agents buying on behalf of consumers break traditional fraud signals. Platforms will need detection logic built for agent-initiated transactions.

Self-Service Agent Building

Compliance teams want to build their own agents without engineering. Vendors that enable self-service AI building will capture share, it's already a Level 4 differentiator.

Cross-Institutional Intelligence Networks

Agentic AI compounds when agents operate across a network of institutions. Consortium depth and diversity (cross-vertical and cross-typology) will become a primary evaluation criterion.

The Accountability Evolution

Watch for FinCEN guidance on task-based agents supervised by manager agents for QA/QC. Institutions building accountability frameworks now will be ready when regulatory clarity arrives.

The Interaction Model Shift

Users increasingly work through their own AI agents rather than vendor UIs. Platforms will need "receiving agents," API and MCP layers that accept instructions from buyer-side agents.

Context Engineering as Competitive Moat

As frontier models commoditize, the edge shifts to how well context is engineered per task. Vendors with historical analyst-behavior data have a structural advantage that compounds.

Cut through the noise. Learn how to evaluate a vendor’s AI claims.

Download the Report

Get the full Chartis Research report

Download the complete Agentic Maturity Spectrum report including a Capability Matrix across 5 use cases, the vendor evaluation scorecard, and Chartis Research's calls to action for buyers and vendors.

  • Full 4-level maturity spectrum
  • 5-use-case capability matrix
  • Vendor scorecard template
  • 7-dimension evaluation guide

Get the report

We’ve Got You

Frequently Asked Questions

What is agentic AI in financial crime compliance?
What is the agentic AI maturity spectrum, and how was it developed?
What is "agent washing" and how do I spot it?
What is context engineering, and why does Chartis call it the most important differentiator?
What is the difference between Level 3 and Level 4 agentic AI?
See Us In Action

Boost fraud prevention & AML compliance

Fraud can’t be guesswork. Invest in a platform that puts you back in control.
Get a Demo