Skip to content
Bankautomation

Anti money laundering program: structure, evidence and the honest limits of automation

Published 12 August 2026

9 min read

anti money laundering program

camt.053 · NOSTRO USD · value date 03 sep 2026

Nostro cash break, sample data

Match rate

75.0%

Breaks

4

At risk

$533.9k

Oldest

4d

Date ±1d
Amount

This console matches the first rows a side. It left out statement and ledger , so anything in them is not counted below. To run a full file, .

BRK-001

Aged 0d

$219,105.40

2026-09-03 · both sides

REF//FX/SPOT/7741 · INTERBANK FX DESK

Amount differs by 164.80 USD (0.075%)

Re-price the ledger leg on the correspondent rate source for the value date and post 164.80 USD to FX variance.

FX RATE SOURCE

→ Treasury Ops

BRK-002

Aged 4d

$187,650.00

2026-08-30 · ledger

PAY/90012/R · KESTREL LOGISTICS

Second ledger entry for 187,650.00 USD against REF//PAY/90012

Confirm the original entry REF//PAY/90012 cleared, then reverse this posting under maker-checker and note the reversal on the original item.

DUPLICATE POSTING

→ Finance

BRK-003

Aged 0d

$123,760.00

2026-09-03 · both sides

REF//TRF/554121 · HALCYON TRADING

Statement 61,880.00 vs ledger 123,760.00 (50% short)

Split the ledger posting and match the settled leg; leave the residual 61,880.00 USD open against the same reference.

PARTIAL SETTLEMENT

→ Payments Ops

BRK-004

Aged 0d

$3,410.00

2026-09-03 · statement

REF//CHG/Q3FEES · CORRESPONDENT CHARGES

On the statement, nothing in the ledger

Post 3,410.00 USD to the charges account for the period and add it to the standing accrual so it stops surfacing as a break.

FEE NOT ACCRUED

→ Finance

Everything matched under these tolerances.

Tighten the date or amount tolerance to see the breaks it was absorbing.

Correspondent statement

camt.053 · NOSTRO USD

Value date Reference Amount
2026-09-02 REF//NONREF/2026090301 1,284,500.00
2026-09-02 REF//INV/88231 96,400.00
2026-09-03 REF//TRF/554120 452,180.25
2026-09-03 REF//FX/SPOT/7741 218,940.60
2026-09-01 REF//SEPA/33421 74,220.00
2026-09-03 REF//CHG/Q3FEES 3,410.00
2026-09-03 REF//TRF/554121 61,880.00
2026-09-02 REF//PAY/90012 187,650.00
2026-09-02 REF//INV/88245 242,015.00
2026-09-03 REF//TRF/554133 18,905.50
2026-09-01 REF//PAY/90044 505,300.00
2026-09-03 REF//INV/88260 132,640.75

Internal nostro ledger

Core banking export

Value date Reference Amount
2026-09-02 NONREF/2026090301 1,284,500.00
2026-09-02 INV/88231 96,400.00
2026-09-03 TRF/554120 452,180.25
2026-09-03 FX/SPOT/7741 219,105.40
2026-08-31 SEPA/33421 74,220.00
2026-09-02 PAY/90012 187,650.00
2026-09-03 TRF/554121 123,760.00
2026-08-30 PAY/90012/R 187,650.00
2026-09-02 INV/88245 242,015.00
2026-09-03 TRF/554133 18,905.50
2026-09-01 PAY/90044 505,300.00
2026-09-03 INV/88260 132,640.75
Matched pairs are tinted on both sides. Breaks carry the brass left rule and appear in the worklist.

Sample data only. Matching runs in your browser; classification is written by the model when you press Run. Matching done in your browser. Writing classifications… Classifications written by the model on this run. Decision support, not a compliance determination. The classification model was unavailable, so the built-in rule classifier wrote these. Same matching, same numbers. This console classifies up to 12 runs a minute and this run went over, so the built-in rule classifier wrote these. Same matching, same numbers. Wait a minute for the model, or . The model wrote the first 8 classifications; the rule classifier wrote the remaining .

Open the full resolver

An AML program is a control framework, not a piece of software. Understanding which parts are judgement and which parts are logistics is what stops a bank from either automating too little or automating something it should never have automated.

The structure everyone recognizes

Programs are usually described in terms of a small number of pillars, and the labels differ by jurisdiction. In substance the components are consistent: internal policies and controls, a designated compliance officer with authority, ongoing training, independent testing, customer due diligence including beneficial ownership, and ongoing monitoring with reporting of suspicious activity.

Everything else in the program is an operational expression of those components. That is a useful lens when deciding what to buy, because software can carry an operational expression well and cannot carry a judgement at all.

What an examiner actually tests

Not whether the policy document is well written. They test whether the program operated as described, and they do it by pulling specific cases and asking specific questions.

  • Show me this alert, worked, with the rationale and the evidence the analyst had.
  • Show me who approved this disposition, and confirm that they are not the person who proposed it.
  • Show me the population of customers due for periodic review and how many are overdue.
  • Show me the version of the policy or threshold that applied on the date the decision was made.
  • Show me the independent testing that was performed, and what happened to its findings.

Each of those questions is answered instantly by a system that recorded the work as it happened, and answered over several days by a team reconstructing it from email, shared drives and memory. The difference is not diligence, it is architecture.

The risk assessment sets the cadence

The institution risk assessment is the document that justifies the shape of everything else: which customers are reviewed how often, which alerts are prioritized, what enhanced due diligence means in practice. When it is a static document produced annually, the program drifts away from it. When its outputs are configuration, the cadence stays honest, because a change to a risk rating changes the refresh population by itself.

Where automation legitimately helps

Program areaAutomate the logisticsNever automate
Customer due diligence Document requests, chasing, completeness checks, queue building The adequacy assessment and the acceptance decision
Periodic review Population building, assignment, aging, escalation The judgement on the refreshed file
Alert triage Deduplication, prioritization, evidence checklists, routing The disposition
Case narrative A draft from structured facts, for a human to edit Approving the narrative or filing on it
Independent testing Sampling, evidence gathering, finding tracking The opinion
Reporting to authorities Assembling the supporting file The determination and the filing

Swipe the table sideways to read every column.

The failure mode automation actually fixes

It is rarely a missed sophisticated typology. It is far more often a queue that grew because nobody could see it growing, a set of dispositions that were inconsistent because there was no required evidence set, or an exam response that took three weeks because the evidence had to be rebuilt. Those are operational failures, and operational failures are what workflow software is for.

A model can draft. A person decides. A second person approves. If a vendor offers to remove any of those three steps from an AML program, that is not efficiency, it is a control finding waiting to be written.

Practical starting points

  1. 01Measure the overdue periodic review population honestly, before deciding anything.
  2. 02Define required evidence per alert type, and enforce it in the tool your analysts already open.
  3. 03Enforce maker-checker separation in the system, not in a procedure document.
  4. 04Test the exam question directly: pull one case from six months ago and time how long the full file takes to produce.

Customer due diligence, and the part that is genuinely logistics

Due diligence at onboarding is a judgement supported by a lot of administration. Someone has to decide which documents are required for this customer type in this jurisdiction, request them, chase what has not arrived, check that what did arrive is complete and current, run screening, resolve the matches that screening produces, and record the reasoning behind the acceptance. Exactly one of those steps is a judgement.

The administration is where onboarding time is lost, and it is where automation is uncontroversial because nothing about it touches a decision. A requirements matrix that produces the document list, a request that goes out automatically, a chase that escalates on a schedule, a completeness check that flags an expired document before a human opens the file: all of that is logistics, and all of it is measurable.

A useful separation to hold onto: automation may assemble the file and may draft, but the assessment of adequacy and the acceptance decision stay with a named person who can be asked why.

Periodic review is a population problem before it is a review problem

Most programs describe periodic review as risk based: higher risk customers reviewed more often, lower risk less often. The description is nearly universal and the execution varies enormously, because everything depends on whether the population is calculable.

The population is a straightforward function of risk rating, last review date and any event that triggers an off cycle review. If those three live in systems that can be queried together, the population is a report and the overdue number is a fact. If they live in a rating held in one system, a review date recorded in a spreadsheet and trigger events noted in email, then the overdue number is an estimate, and an estimate is the thing an examiner will find first.

  • Can you state, today, how many customers are due for review this month and how many are overdue?
  • Is that number produced by a query, or assembled by a person?
  • When a risk rating changes, does the review population change by itself?
  • Is an event driven trigger, such as a significant change in activity, capable of moving a customer into the queue without a person noticing first?
  • Can a completed review be shown with its evidence, its author and its approver, six months later, in one action?

A program that answers all five with a system rather than with a person has removed most of the exam risk in this area, without automating a single judgement.

Alert triage: workflow, not detection

It is worth being blunt about a distinction that vendors blur. Detection is the model or rule set that decides an alert should exist. Triage is what happens to the alert afterwards. These are different products, different skills and different risks, and a workflow tool that claims to improve detection quality is describing something it cannot do.

What triage workflow can legitimately do is make the disposition consistent and provable. Deduplicating alerts that describe the same activity, grouping alerts by customer so an analyst sees the picture rather than the fragments, prioritising by a documented rule, requiring a defined evidence set per alert type before a disposition can be recorded, and enforcing that the person who proposes a closure is not the person who approves it. None of that touches the model. All of it changes the quality of the outcome, because inconsistent dispositions are usually a process failure rather than an analyst failure.

SymptomUsually diagnosed asOften actually
Backlog growing Not enough analysts Duplicate alerts and no prioritisation rule
Inconsistent dispositions Training gap No required evidence set per alert type
Slow exam responses Poor record keeping Evidence assembled after the fact rather than during
Escalations arriving late Individual judgement No aging threshold that escalates by itself
Repeat alerts on the same customer Model tuning problem Prior disposition not visible when the new alert is worked

Swipe the table sideways to read every column.

Independent testing, from the other side of the table

Independent testing exists to check that the program operated as described. The uncomfortable truth is that most findings are not about the quality of judgements. They are about the ability to demonstrate that a judgement was made, by whom, on what basis, and under which version of the policy.

That is an evidence architecture question. A program where the work is recorded as it happens produces testing samples on request. A program where the work happens in email and is written up later produces a reconstruction, and reconstructions attract findings because they are visibly reconstructions. The cheapest improvement available to most programs is not a better model, it is recording the work in the place where the work happens.

Where a language model can and cannot sit

Drafting is a legitimate use. Summarising structured facts into a first version of a narrative saves real time and does not decide anything, provided the person who edits it is accountable for the final text and the underlying facts are shown alongside the draft. Extraction from documents is also legitimate, with the same condition: the extracted value is a suggestion that a person confirms.

What a model must not do is decide. It must not dispose of an alert, accept a customer, assign a risk rating that drives a review cadence without review, or produce a filing. The reason is not caution for its own sake. It is that a determination has to be defensible by a person who can explain the reasoning, and a probabilistic draft is not a reason. Any vendor who offers to remove the human from that step is offering a control finding.

A twelve month sequence that usually works

  1. 01Measure the overdue periodic review population and the alert backlog age honestly. Do not fix anything yet.
  2. 02Define the required evidence set for the three highest volume alert types, and enforce it in the tool analysts already use.
  3. 03Make maker-checker separation a system property rather than a procedural instruction.
  4. 04Build the review population from risk rating and review date automatically, so it updates itself.
  5. 05Automate the document request and chase cycle for onboarding and refresh.
  6. 06Only then look at drafting assistance, and only where a person edits and approves the output.

The order matters. Automating drafts on top of a program whose population is an estimate produces faster work on the wrong queue.

A note on vendor claims

Two claims are worth treating with suspicion in this category. The first is any figure for false positive reduction quoted without reference to a specific institution, a specific model and a specific tuning exercise, because the number is a property of that combination and not of the software. The second is any suggestion that a tool provides regulatory assurance, since assurance comes from the institution's own testing and from its regulator, not from a supplier. A vendor can honestly describe the controls its product implements and the evidence it produces. Anything beyond that is being sold rather than described, and the gap between the two tends to appear during an examination rather than during a demonstration.

The product view of this operating layer is on AML compliance software, the customer queue on KYC automation software, and the alert side on BSA AML monitoring software.

Get started

Put your first reconciliation on rails

Create an account, and we will email you how onboarding works and what a first source connection looks like. The Reconciliation Break Resolver is open to try right now, on sample data, without an account.

Run the demo

No card required to create an account. Sample data only in the demo. Bankautomation is operations software, not a regulated service.