Unit21 for AML

AML quality assurance checklist: what to review and how to score it

Published
August 7, 2024
Read Time
7
mins
Subscribe to stay informed
Table of contents

Most AML quality assurance programs measure whether analysts followed the process. Far fewer measure whether the process is finding anything.

That gap is what examiners probe. A QA function that samples 5% of alerts, confirms the notes were filled in, and reports a 96% pass rate has told you almost nothing about whether your program works.

Below is a practical checklist covering what to review, how to score it, and where QA programs typically fall short, along with insights from three AML compliance leaders who have built these functions from scratch.

The 8-Step AML Compliance Checklist
A 65-point framework covering the controls, procedures, and documentation examiners look for. Use it to find the gaps before they do.
Get your copy

What is AML quality assurance?

AML quality assurance is the independent review of a sample of alerts, cases, and filings to assess whether decisions were correct, consistently applied, and adequately documented.

It is distinct from quality control, though the terms get mixed. QC is checked inside the workflow, usually a second pair of eyes before a decision is final. QA looks backward at completed work, independently, to find patterns.

Both are expectations rather than optional. Independent testing of an AML program is a supervisory requirement, and QA is how most institutions evidence it between formal audits.

The AML QA checklist

Six areas. Each one needs a documented standard, a sampling approach, and a score you can trend over time.

Review area What to check Common failure
Alert disposition Was the close or escalate decision correct given the evidence available, and consistent with how similar alerts were handled Reviewing whether the notes were completed rather than whether the decision was right
Investigation adequacy Was the depth of investigation proportionate to the risk. Were related accounts, counterparties, and prior alerts on the same entity checked Single-alert review with no check for related activity on the same customer
Documentation Does the record explain what was reviewed and why the conclusion follows, clearly enough for someone with no prior context Template language that appears identically across unrelated cases
Escalation and filing Were thresholds applied consistently. Were SAR decisions defensible and narratives sufficient. Were deadlines met Tracking deadline compliance without assessing narrative quality
Rule performance Which rules produced alerts that led to filings, and which produced consistently wrong dispositions QA and tuning run as separate exercises, so findings never reach the rule
Independence and coverage Is the reviewer independent of the reviewed team. Does the sample cover every analyst, alert type, and high-risk segment QA reporting to the owner of the alert queue

How to build the sample

Random sampling alone will under-cover your highest risk. A workable approach layers three:

  • Random, so every analyst and alert type has a chance of review. This is what supports statistical claims about overall quality.
  • Risk-weighted, oversampling high-risk customer segments, large-value activity, and alert types tied to your top risk-assessment findings.
  • Targeted, triggered by something specific: a new analyst, a newly deployed rule, a typology that just appeared, or an analyst whose disposition pattern looks unlike their peers.

Sample size should be defensible rather than convenient. If you cannot explain to an examiner why 5% is the right number for your risk profile, it is not.

How to score it

Binary pass or fail hides too much. A scored checklist per review, with weighted criteria, lets you separate a clerical miss from a wrong decision.

Two things make scores useful rather than decorative. First, weight the criteria: a missed escalation is not equivalent to a typo in a note. Second, trend the scores by analyst, alert type, and rule, because the aggregate number tells you nothing actionable. A stable 94% that hides one rule producing systematically wrong dispositions is worse than a 91% you understand.

What three AML leaders said matters most

We put these questions to Brandi Reynolds, Chris Sidler, and John Wiethorn. A few themes came up repeatedly.

Documentation quality is the tell. Comprehensive policies, procedures, standards, and quality assessments signal a program that can withstand scrutiny. Thin documentation usually means thin process underneath, whatever the pass rate says.

Independence is not optional. John stressed the need for genuinely independent testing to meet industry standards. QA that reports to the person who owns the alert queue is not independent, and examiners notice.

Staffing determines the ceiling. Experienced personnel handling compliance responsibilities matters disproportionately, and it is the most common gap at earlier-stage firms where one person carries several roles.

Integrate QA with the program, not alongside it. Brandi made the point that QA should be present in the meetings where products and services are designed, so AML considerations are built in rather than discovered later.

Measure performance, and record the measurement. Tracking outcomes on transaction monitoring alerts and KYC files is what makes a risk-based approach real. Without recorded measurement over time, "risk-based" is an assertion.

Scale the approach to the institution. How QA gets implemented varies enormously with size. A ten-person team and a thousand-person team need different mechanics for the same objective.

Where QA programs commonly fall short

Measuring process compliance instead of decision quality. The most common failure. Reviewing whether the fields were filled in is easy to operationalize and tells you almost nothing.

No feedback loop into detection. QA finds that a rule produces consistently wrong dispositions, and the finding goes into a report that nobody acts on because changing the rule requires an engineering ticket. The QA function becomes an observation exercise.

Sampling that avoids the hard cases. Reviewing straightforward closures because they are quick to assess. The value is in the ambiguous ones.

No independence. Covered above, and worth repeating because it is the finding most likely to appear in an examination report.

QA disconnected from tuning. If QA outcomes are not visible alongside rule performance, nobody can answer whether a control is reducing the risk it was built for. That is the question that matters, and most programs cannot answer it because the evidence sits in two systems.

Running QA in Unit21

Unit21 includes QA functionality built for this workflow rather than bolted alongside it.

Checklists with computable scores. Build your review criteria as a checklist that produces a numeric quality score, exportable as a report for assessment and audit.

Random sampling. QA managers generate random samples of alerts, cases, or filings, filtered to narrow the population or set as a sample percentage per agent, so risk-weighted and per-analyst sampling are both straightforward.

Classified queues. Confidential queues whose contents are invisible and inaccessible elsewhere in the application, including prior activity logs, search, and audit trails. That is what keeps a QA review genuinely independent of the team being reviewed.

The larger advantage is that QA sits on the same platform as detection. When a review finds a rule producing wrong dispositions, the person who understands the problem can change that rule directly, in a no-code interface, and test the change against historical data before it goes live. The feedback loop closes.

That matters more as investigation itself gets automated. Unit21's AI Agents work alerts end to end, gathering evidence, mapping entities, and drafting narratives, with each step recorded. QA does not disappear in that model, it changes shape: you are reviewing the agent's work with a complete audit trail rather than reconstructing an analyst's reasoning from their notes. Our practitioner's guide to agentic AI for AML covers how that works, and why configurable AI matters for compliance teams covers why the agent needs to write to your QA standards rather than a vendor's default.

If you are comparing vendors, Liminal's independent index on AML and transaction monitoring for financial services and fintechs covers how the field stacks up.

Frequently asked questions

What is the difference between AML QA and QC?

Quality control happens inside the workflow, typically a review before a decision is finalized. Quality assurance happens after the fact, independently, on a sample of completed work, to identify patterns rather than catch individual errors. Most programs need both.

What sample size should an AML QA program use?

There is no prescribed percentage. The sample needs to be defensible against your own risk profile, which usually means layering random sampling with risk-weighted oversampling of high-risk segments and targeted reviews triggered by new analysts, new rules, or unusual disposition patterns. If you cannot explain why your percentage is right for your risk, it is not.

Is AML quality assurance a regulatory requirement?

Independent testing of an AML program is a supervisory expectation. QA is how most institutions evidence that testing on an ongoing basis between formal audits. Specific obligations vary by institution type and regulator.

Who should own the QA function?

Someone independent of the team whose work is being reviewed. QA reporting to the person accountable for alert throughput creates a conflict that examiners look for specifically.

What should an AML QA review actually check?

Whether the disposition was correct, whether the investigation was adequate for the risk, whether documentation supports the decision, whether escalation and filing thresholds were applied consistently, and whether deadlines were met. Checking only the last two is the common shortfall.

How does QA connect to transaction monitoring tuning?

QA findings are the highest-quality signal you have about rule performance, because they tell you which alerts led to correct outcomes. A program where QA results are not visible alongside rule performance cannot answer whether a control is reducing the risk it was built for.

Learn more about Unit21
Unit21 is the leader in AI Risk Infrastructure, trusted by over 200 customers across 90 countries, including Sallie Mae, Chime, Intuit, and Green Dot. Our platform unifies fraud and AML with agentic AI that executes investigations end-to-end—gathering evidence, drafting narratives, and filing reports—so teams can scale safely without expanding headcount.
AI Tasks
|
6
min

AI Task Spotlight | Edition No. 09: One FinCEN Alert, Two AI Tasks

Gal Perelman
Gal Perelman
Product Marketing Lead, Unit21
This is some text inside of a div block.
AI Risk Infrastructure
|
7
min

What's actually holding compliance teams back from AI

Tyler Allen
Tyler Allen
CEO, Unit21
This is some text inside of a div block.
AI Tasks
|
7
min

AI Task Spotlight | Edition No. 08: Account Takeover — Before vs. After Access Change

Gal Perelman
Gal Perelman
Product Marketing Lead, Unit21
This is some text inside of a div block.
See Us In Action

Boost fraud prevention & AML compliance

Fraud can’t be guesswork. Invest in a platform that puts you back in control.
Get a Demo