An AML program is a control framework, not a piece of software. Understanding which parts are judgement and which parts are logistics is what stops a bank from either automating too little or automating something it should never have automated.
The structure everyone recognizes
Programs are usually described in terms of a small number of pillars, and the labels differ by jurisdiction. In substance the components are consistent: internal policies and controls, a designated compliance officer with authority, ongoing training, independent testing, customer due diligence including beneficial ownership, and ongoing monitoring with reporting of suspicious activity.
Everything else in the program is an operational expression of those components. That is a useful lens when deciding what to buy, because software can carry an operational expression well and cannot carry a judgement at all.
What an examiner actually tests
Not whether the policy document is well written. They test whether the program operated as described, and they do it by pulling specific cases and asking specific questions.
- Show me this alert, worked, with the rationale and the evidence the analyst had.
- Show me who approved this disposition, and confirm that they are not the person who proposed it.
- Show me the population of customers due for periodic review and how many are overdue.
- Show me the version of the policy or threshold that applied on the date the decision was made.
- Show me the independent testing that was performed, and what happened to its findings.
Each of those questions is answered instantly by a system that recorded the work as it happened, and answered over several days by a team reconstructing it from email, shared drives and memory. The difference is not diligence, it is architecture.
The risk assessment sets the cadence
The institution risk assessment is the document that justifies the shape of everything else: which customers are reviewed how often, which alerts are prioritized, what enhanced due diligence means in practice. When it is a static document produced annually, the program drifts away from it. When its outputs are configuration, the cadence stays honest, because a change to a risk rating changes the refresh population by itself.
Where automation legitimately helps
| Program area | Automate the logistics | Never automate |
|---|---|---|
| Customer due diligence | Document requests, chasing, completeness checks, queue building | The adequacy assessment and the acceptance decision |
| Periodic review | Population building, assignment, aging, escalation | The judgement on the refreshed file |
| Alert triage | Deduplication, prioritization, evidence checklists, routing | The disposition |
| Case narrative | A draft from structured facts, for a human to edit | Approving the narrative or filing on it |
| Independent testing | Sampling, evidence gathering, finding tracking | The opinion |
| Reporting to authorities | Assembling the supporting file | The determination and the filing |
The failure mode automation actually fixes
It is rarely a missed sophisticated typology. It is far more often a queue that grew because nobody could see it growing, a set of dispositions that were inconsistent because there was no required evidence set, or an exam response that took three weeks because the evidence had to be rebuilt. Those are operational failures, and operational failures are what workflow software is for.
A model can draft. A person decides. A second person approves. If a vendor offers to remove any of those three steps from an AML program, that is not efficiency, it is a control finding waiting to be written.
Practical starting points
- 01Measure the overdue periodic review population honestly, before deciding anything.
- 02Define required evidence per alert type, and enforce it in the tool your analysts already open.
- 03Enforce maker-checker separation in the system, not in a procedure document.
- 04Test the exam question directly: pull one case from six months ago and time how long the full file takes to produce.
Customer due diligence, and the part that is genuinely logistics
Due diligence at onboarding is a judgement supported by a lot of administration. Someone has to decide which documents are required for this customer type in this jurisdiction, request them, chase what has not arrived, check that what did arrive is complete and current, run screening, resolve the matches that screening produces, and record the reasoning behind the acceptance. Exactly one of those steps is a judgement.
The administration is where onboarding time is lost, and it is where automation is uncontroversial because nothing about it touches a decision. A requirements matrix that produces the document list, a request that goes out automatically, a chase that escalates on a schedule, a completeness check that flags an expired document before a human opens the file: all of that is logistics, and all of it is measurable.
A useful separation to hold onto: automation may assemble the file and may draft, but the assessment of adequacy and the acceptance decision stay with a named person who can be asked why.
Periodic review is a population problem before it is a review problem
Most programs describe periodic review as risk based: higher risk customers reviewed more often, lower risk less often. The description is nearly universal and the execution varies enormously, because everything depends on whether the population is calculable.
The population is a straightforward function of risk rating, last review date and any event that triggers an off cycle review. If those three live in systems that can be queried together, the population is a report and the overdue number is a fact. If they live in a rating held in one system, a review date recorded in a spreadsheet and trigger events noted in email, then the overdue number is an estimate, and an estimate is the thing an examiner will find first.
- Can you state, today, how many customers are due for review this month and how many are overdue?
- Is that number produced by a query, or assembled by a person?
- When a risk rating changes, does the review population change by itself?
- Is an event driven trigger, such as a significant change in activity, capable of moving a customer into the queue without a person noticing first?
- Can a completed review be shown with its evidence, its author and its approver, six months later, in one action?
A program that answers all five with a system rather than with a person has removed most of the exam risk in this area, without automating a single judgement.
Alert triage: workflow, not detection
It is worth being blunt about a distinction that vendors blur. Detection is the model or rule set that decides an alert should exist. Triage is what happens to the alert afterwards. These are different products, different skills and different risks, and a workflow tool that claims to improve detection quality is describing something it cannot do.
What triage workflow can legitimately do is make the disposition consistent and provable. Deduplicating alerts that describe the same activity, grouping alerts by customer so an analyst sees the picture rather than the fragments, prioritising by a documented rule, requiring a defined evidence set per alert type before a disposition can be recorded, and enforcing that the person who proposes a closure is not the person who approves it. None of that touches the model. All of it changes the quality of the outcome, because inconsistent dispositions are usually a process failure rather than an analyst failure.
| Symptom | Usually diagnosed as | Often actually |
|---|---|---|
| Backlog growing | Not enough analysts | Duplicate alerts and no prioritisation rule |
| Inconsistent dispositions | Training gap | No required evidence set per alert type |
| Slow exam responses | Poor record keeping | Evidence assembled after the fact rather than during |
| Escalations arriving late | Individual judgement | No aging threshold that escalates by itself |
| Repeat alerts on the same customer | Model tuning problem | Prior disposition not visible when the new alert is worked |
Independent testing, from the other side of the table
Independent testing exists to check that the program operated as described. The uncomfortable truth is that most findings are not about the quality of judgements. They are about the ability to demonstrate that a judgement was made, by whom, on what basis, and under which version of the policy.
That is an evidence architecture question. A program where the work is recorded as it happens produces testing samples on request. A program where the work happens in email and is written up later produces a reconstruction, and reconstructions attract findings because they are visibly reconstructions. The cheapest improvement available to most programs is not a better model, it is recording the work in the place where the work happens.
Where a language model can and cannot sit
Drafting is a legitimate use. Summarising structured facts into a first version of a narrative saves real time and does not decide anything, provided the person who edits it is accountable for the final text and the underlying facts are shown alongside the draft. Extraction from documents is also legitimate, with the same condition: the extracted value is a suggestion that a person confirms.
What a model must not do is decide. It must not dispose of an alert, accept a customer, assign a risk rating that drives a review cadence without review, or produce a filing. The reason is not caution for its own sake. It is that a determination has to be defensible by a person who can explain the reasoning, and a probabilistic draft is not a reason. Any vendor who offers to remove the human from that step is offering a control finding.
A twelve month sequence that usually works
- 01Measure the overdue periodic review population and the alert backlog age honestly. Do not fix anything yet.
- 02Define the required evidence set for the three highest volume alert types, and enforce it in the tool analysts already use.
- 03Make maker-checker separation a system property rather than a procedural instruction.
- 04Build the review population from risk rating and review date automatically, so it updates itself.
- 05Automate the document request and chase cycle for onboarding and refresh.
- 06Only then look at drafting assistance, and only where a person edits and approves the output.
The order matters. Automating drafts on top of a program whose population is an estimate produces faster work on the wrong queue.
A note on vendor claims
Two claims are worth treating with suspicion in this category. The first is any figure for false positive reduction quoted without reference to a specific institution, a specific model and a specific tuning exercise, because the number is a property of that combination and not of the software. The second is any suggestion that a tool provides regulatory assurance, since assurance comes from the institution's own testing and from its regulator, not from a supplier. A vendor can honestly describe the controls its product implements and the evidence it produces. Anything beyond that is being sold rather than described, and the gap between the two tends to appear during an examination rather than during a demonstration.
The product view of this operating layer is on AML compliance software, the customer queue on KYC automation software, and the alert side on BSA AML monitoring software.