Back to Articles
Thought Leadership5 min read

AI Can Rank Your KYC Alerts. Who Checks What It Leaves at the Bottom?

AI models that score and rank KYC screening alerts can remove a large share of false positives. For management companies and TCSPs, the harder questions are who reviews the alerts ranked lowest, what historical decisions the model learned from and whether AI-written summaries match the evidence.

Fredrik Gröndahl
Dark bronze trays of index cards on a stone desk, with one stack held apart under a thick glass block

Moody’s describes its AI screening model as producing an alert confidence score that mimics a human analyst, trained on millions of analyst decisions, with firms setting the threshold that suits their own risk appetite. It says the approach can reduce false positives by as much as 70% while prioritising true alerts. Figures like that come from the vendor, not from independent testing, but the direction is clear: ranking alerts with machine learning is now a mainstream option.

For management companies, fund administrators and trust and company service providers, a ranked queue is attractive. Screening hits on common names and generic company terms consume hours. The useful question is not whether a model can put the likely matches first. It is what happens to the alerts it puts last.

A low score is still a decision

Consider a hypothetical firm that screens 4,000 names a month and sets a threshold so that the lowest-scoring 60% of alerts are closed with light review. Nobody has chosen to ignore those alerts. Yet in practice the threshold has made a risk decision for 2,400 of them every month.

That decision belongs to the firm, not the model. The threshold should be approved in the same way as any other risk appetite setting, with a named owner, a documented rationale and a date for review. If it is tuned only to clear a backlog, it will drift towards whatever keeps the queue short.

The record for each closed alert should also show that it was closed because of its score and the threshold in force at the time. When the threshold changes, the firm should be able to say which alerts were handled under which setting.

Sample below the line

The only way to know what the bottom of the queue contains is to look at it. A proportionate control is a regular sample of low-ranked alerts reviewed in full by an analyst who does not see the score first. The sample should include the alerts just below the threshold, where a miss is most likely, as well as a random draw from further down.

Track how often the sample finds an alert that should have been escalated. Record why the model scored it low: a transliterated name, a missing date of birth, an unfamiliar jurisdiction or a new type of structure. A single meaningful miss is a reason to investigate the pattern, not a statistic to average away.

Cards pulled from the back of a sorted tray and fanned under a magnifying glass, representing a sample of low-ranked alerts
Cards pulled from the back of a sorted tray and fanned under a magnifying glass, representing a sample of low-ranked alerts

Challenge the history the model learned from

A model trained on past analyst decisions learns those decisions, including their weaknesses. Alerts closed quickly during a backlog, inconsistent judgements between reviewers and closures made on incomplete information can all be learned as normal.

Before relying on a ranking, ask what the historical data contains. Which customer types and jurisdictions are well represented, and which are rare? Were past closures checked by a second reviewer? A firm whose client book is heavy in trusts, foundations and multi-layer holding structures should be cautious about a model whose history is dominated by retail banking names.

The same question applies after deployment. If analysts start to agree with the score because it is shown first, their decisions become new training data that confirms the model rather than testing it.

Check the summary against the source

Many tools now pair a score with an AI-written summary of the match: who the person is, why the name appeared and what the adverse media says. Summaries save time, and they can also omit the one sentence that mattered or state a connection the source does not support.

A practical standard is that every factual statement in a summary should point to the document or record it came from, and the reviewer should confirm the key facts against that source before relying on it. Moody’s own guidance on AI in investigations describes human oversight as central to investigative conclusions, escalation decisions and regulatory reporting. That principle only works if the reviewer checks the evidence, not just the prose.

A brass ruler aligning a short summary card with a line in a stack of source documents, representing verification of an AI summary
A brass ruler aligning a short summary card with a line in a stack of source documents, representing verification of an AI summary

What to settle before switching it on

A ranking model can be a sound control. It becomes a weak one when nobody owns what it filters out. Before going live, agree the following:

  • Who approves the threshold, how often it is reviewed and how changes are recorded.
  • How large the below-threshold sample is, who reviews it and what result triggers a review of the model.
  • What the historical training data contains, and which customer segments it covers poorly.
  • How AI-written summaries are checked against source evidence before a decision is recorded.

Fidify’s earlier analysis of AI agents in KYC argued that the value of automation depends on the data and the reasoning it leaves behind. Alert ranking is a clear test of that idea. The alerts at the top of the queue will get attention anyway. The quality of the control is decided by what happens at the bottom.

Sources and further reading

Moody’s, AI Review: intelligent screening (product brochure). Performance figures are vendor claims and are attributed to Moody’s.

Moody’s, AI augmentation for compliance investigators, 21 April 2026.

Related Fidify analysis: AI Agents Are Coming to KYC. Here’s What They Will Need from the Platforms They Run On.