AI Risk Infrastructure

What's actually holding compliance teams back from AI

Published
August 27, 2026
Read Time
7
mins
Tyler Allen
Tyler Allen
CEO, Unit21
Subscribe to stay informed
Table of contents

Around 90% of the alerts your team worked last month were nothing. That number has been roughly stable in AML for years, and every compliance leader I talk to already knows it. So when I sit down with a team that has looked at AI, run a pilot, seen it work, and then quietly let it stall, the blocker is almost never the technology.

I have had a version of this conversation dozens of times in the last year. The pattern repeats closely enough that I think it's worth naming the three things that actually stop deployments, because none of them are what people say out loud in the first ten minutes.

One: you can't prove it worked

This is the most common one, and it's the most fixable.

A team runs a successful pilot. The agents do the evidence gathering, the analysts stop assembling transaction histories by hand, and everyone can feel that the work got easier. Then someone has to go to the executive team and justify the spend, and what comes out of their mouth is "it's saving our analysts a lot of time."

That sentence dies in the room. It isn't a number, it can't be checked, and it sounds like the same thing every software purchase has promised since the invention of software.

The teams that get their second year of budget walk in with three numbers instead:

Handle time reduction. How long did an alert of this type take before, and how long does it take now. Measure it per queue, because a sanctions screening alert and a complex fraud ring investigation are not the same animal and averaging them hides everything interesting.

Cleared alert volume. How many alerts the team actually got through in a month, before and after. This is the one executives feel fastest, because it maps directly onto the backlog they already know about.

Precision and recall on decisioning. Precision tells you that nothing which should have escalated got closed. Recall tells you how much of the false positive pile you're actually clearing. Track both, because one without the other is how you end up either automating nothing or automating something you'll regret. These are also the two numbers that should govern how far up the autonomy ladder you let a queue travel.

The frame that lands is scale, not savings. You are not making the case that the team gets smaller. You are making the case that alert volume can grow without the team growing at the same rate, which is a completely different conversation and a much easier one to have with a board.

Get the baseline before you turn anything on. I have watched teams deploy first, love the results, and then have no idea what the "before" looked like. You cannot reconstruct that number afterwards.

Two: the auditor fear points the wrong way

The second blocker sounds like this: "our examiners will never go for it."

I talk to regulators regularly, and this fear is almost always aimed in the wrong direction. Examiners aren't opposed to AI. They're opposed to decisions nobody can explain. Those are very different objections, and conflating them costs teams years.

Here's the part that surprises people. The generation of technology this replaces is the one with the real explainability problem.

If your detection or risk scoring runs on a random forest or an XGBoost model, what you can hand an examiner is a score. Maybe some feature importances. What you cannot produce is the reasoning, because there wasn't any in a form a person can read. That's the black box everyone spent a decade being nervous about, and the nervousness was justified.

An LLM-based system inverts that. It produces the reasoning as a matter of course, because reasoning in language is the thing it does. You get the steps it took, the evidence it pulled, the data it accessed and when, and a narrative a human can read and disagree with. When an examiner asks how a decision was reached, you have an actual answer rather than a coefficient.

Think about the difference between a colleague who says "I don't like this transaction, trust me" and one who walks you through what they checked and why it bothered them. Both might be right. Only one of them is auditable, and only one of them can be corrected when they're wrong.

So the honest framing is that explainability got better, not worse. That's worth bringing to your examiners early rather than waiting to be asked, and in my experience the conversation goes better than teams expect. It does depend on the system actually recording the reasoning rather than just the outcome, which is not a given across the market and is worth testing before you buy.

Three: treating it as a tool you install

The third blocker is quieter than the other two and it's the one I'd most want people to sit with.

Teams treat AI as something they bought rather than something they operate. Set it up, switch it on, walk away, and check the invoice at renewal. Outsourced, in other words, the same way you'd think about a data vendor or a watchlist feed.

That mindset produces bad results, and then the bad results get read as evidence the technology doesn't work.

AI agents need configuring against your workflows, your data, your risk appetite, and your queues. They need someone who reviews what they're doing, notices the queue where agreement rates are drifting, and adjusts the tasks. They need the same quality assurance you'd apply to a new analyst, which means sampling the work, catching the misses, and feeding corrections back in.

The teams getting the most out of this treat it less like installing software and more like onboarding a colleague. You wouldn't hire an analyst, hand them a login, and never look at their casework again. The good ones get better because someone is paying attention to them.

There's a career dimension here too, and I don't think it's overstated. AI isn't going to take compliance jobs. The people who know how to direct it will take work from the people who don't, and that gap is going to open faster than most people expect. Learning to configure and interrogate these systems is quickly becoming part of the job rather than an adjacent skill.

Where to start

If your program is stuck on one of these, the smallest useful next step is usually the same: pick one queue, write down your current handle time and false positive rate before you touch anything, and run the agents against alerts your analysts have already decided so you have a comparison you trust.

Then take that to your executive team with numbers, and take it to your examiners with the reasoning trail. Both conversations get much easier when you're not asking anyone to take your word for it.

Which of the three is actually holding your program back? My guess is you already know, and it isn't the one you've been telling people it is.

If it's the first one, the fastest way to get a defensible baseline is to run agents against alerts your analysts have already decided and compare the two. Book a demo and bring your own queue rather than watching ours. If you want the independent framing first, Chartis co-produced a report with us on the agentic AI maturity spectrum, including the evaluation questions worth asking before any of this reaches your executive team.

Frequently asked questions

Why do AI pilots in compliance stall after a successful test?

Most stall on the business case rather than the technology. Teams report qualitative results like "it saved our analysts time" instead of hard metrics, and that doesn't survive an executive review. Capturing a baseline before deployment, then measuring handle time, cleared alert volume, and precision and recall, is what turns a successful pilot into a funded program.

What metrics prove ROI for AI in an AML program?

Three carry the most weight. Handle time reduction measured per queue rather than averaged across queues. Cleared alert volume before and after. Precision and recall on decisioning, where precision confirms nothing that should have escalated was closed and recall shows how much of the false positive volume is being cleared. Frame the result as scaling alert capacity without scaling the team at the same rate.

Are examiners opposed to AI in financial crime programs?

Examiners are generally not opposed to AI itself. They tend to look for decisions that can be explained, with the evidence, the reasoning, and the policy behind them. A system that records only the outcome creates the problem, not the use of AI.

Is an LLM-based system harder to explain to an examiner than machine learning?

It's usually easier. A random forest or XGBoost model produces a score and, at best, feature importances, with no reasoning a person can read. An LLM-based system produces the reasoning as part of its normal output, along with a record of what data was accessed and when, which is much closer to what an examiner is asking for.

How much ongoing work does AI in compliance require?

More than teams expect, and that expectation gap is a common reason deployments underperform. Agents need configuring against your queues and risk appetite, monitoring for drift in agreement rates, and quality assurance through sampling, in the same way a new analyst's casework would be reviewed.

Tyler Allen
Tyler Allen
CEO, Unit21

Tyler Allen is the CEO of Unit21 and was the company’s first hire, writing some of the first lines of code seven years ago. He previously led Unit21’s AI team as Head of AI, then served as COO, before stepping into the CEO role. He is a driving force behind Unit21’s vision as the leader in AI risk infrastructure, having led the AI team before becoming COO. A deep technical leader, Tyler recently returned to the codebase to personally build AI agent configurations, pairing his technical expertise with seven years of experience observing how compliance teams operate.

Learn more about Unit21
Unit21 is the leader in AI Risk Infrastructure, trusted by over 200 customers across 90 countries, including Sallie Mae, Chime, Intuit, and Green Dot. Our platform unifies fraud and AML with agentic AI that executes investigations end-to-end—gathering evidence, drafting narratives, and filing reports—so teams can scale safely without expanding headcount.
AI Tasks
|
7
min

AI Task Spotlight | Edition No. 08: Account Takeover — Before vs. After Access Change

Gal Perelman
Gal Perelman
Product Marketing Lead, Unit21
This is some text inside of a div block.
Risk and Compliance Operations
|
5
min

Progressive autonomy: the trust ladder for AI in compliance

Gal Perelman
Gal Perelman
Product Marketing Lead, Unit21
This is some text inside of a div block.
Unit21 for Crypto
|
5
min

Gate US leads in one of the fastest-growing corners of crypto, with Unit21’s flexibility and real-time detection by its side

Cassie Pallesen
Cassie Pallesen
VP, Marketing
This is some text inside of a div block.
See Us In Action

Boost fraud prevention & AML compliance

Fraud can’t be guesswork. Invest in a platform that puts you back in control.
Get a Demo