Artificial Intelligence has the potential to transform how fraud investigations are conducted - cutting through large volumes of complex, unstructured data in seconds rather than weeks. But there is a fundamental challenge standing between that potential and real-world use in criminal proceedings: the fact that AI does not always give the same answer twice.
That tension was at the heart of a discussion during a recent webinar on AI and fraud, led by Rob Savage, VP for Public Sector at NUIX - a technology company specialising in software for fraud investigation and the handling of large unstructured datasets.
What AI Could Do for Investigations
Rob set out what becomes possible if you temporarily set aside governance constraints and look purely at technical capability. He described a typical fraud investigation as involving a sprawling mix of data types: emails, chat messages, Word documents, PDFs, photographs, videos, handwritten notebooks, financial transactions, and call records - both structured and unstructured, spread across multiple systems.
Historically, Rob said, Investigators would need to navigate all of that data manually, relying heavily on keyword searching once documents had been made machine-readable through optical character recognition. The challenge was always understanding what the data actually meant in the context of the investigation as a whole.
With the right AI system, Rob explained, it becomes possible to extract entities from all of that material - people, organisations, locations, bank account numbers, telephone numbers, vehicles, dates, and transactions - and build them into what he described as an ontology or knowledge graph. Rather than storing information as individual documents, the system understands the relationships between them.
That shift, he said, is where AI becomes transformational. Instead of searching for keywords, Investigators can ask questions in natural language. Rob gave examples such as: "Show me communications discussing payments within 30 days of a contract," or "Identify individuals who appear in handwritten notes but not in official procurement records." Questions that might have taken manual reviewers hundreds of hours could potentially be answered in seconds.
Rob added that this also reduces the need for extensive tool training. Investigators would not need a background in data science to query the data effectively.
The Repeatability Problem
However, Rob identified a significant challenge in using generative AI within a forensic context: it is not deterministic. Unlike traditional algorithms, which will always produce the same output from the same input, large language models are probabilistic - and their outputs can vary over time even when nothing about the input has changed.
To illustrate this, Rob described a test he conducted using the script of The Mousetrap, Agatha Christie's long-running murder mystery. He fed the script into a large language model and asked it - using only the information contained in the script - to identify the murderer and explain its reasoning. The model gave a detailed, well-structured, and apparently well-evidenced answer.
Rob repeated the same test daily. For approximately a week, he received the same answer each time. Then, without any change to the input, the model changed its conclusion - identifying a different suspect and providing an equally compelling justification for that conclusion.
"That sort of uncertainty is not necessarily a problem when generating something like a meeting summary," Rob said. "But if we're using it to support investigations, then obviously that's a very different matter."
What Forensic Standards Require
In the UK, forensic science is governed by the Forensic Science Regulator's Codes of Practice, which place specific emphasis on validation, repeatability, and reproducibility. The expectation is that two competent practitioners using the same technology and the same data, following the same process, will arrive at the same answer.
Generative AI systems do not behave in that way - and that creates a real challenge for their use in proceedings where evidence must be capable of being tested and challenged.
Rob was clear, however, that the answer is not simply to avoid using AI in investigations. He offered two reasons for that. First, data volumes are growing and cases are becoming more complex. If the tools used in investigations do not develop alongside that complexity, investigations will become progressively less effective over time.
Second, humans are not a reliable benchmark. Rob noted that studies have shown human document reviewers operate at around 80% recall in the best-case scenario. An AI system operating at 86% recall would outperform that - yet the AI system would likely be rejected on the grounds that it misses 14% of documents. "There's nothing to say that AI systems are less effective than humans in doing certain tasks," he said.
The solution, Rob argued, lies in understanding how these models work well enough to build in the right safeguards - so that the output of a probabilistic system can, with appropriate controls, be used reliably to support findings in a formal investigation.
VP Sales - UK&I Public Sector, Nuix
This post is based on a webinar on "AI and fraud prevention in the public sector", featuring speakers from the Public Sector Fraud Authority, the Hertfordshire Shared Anti-Fraud Service, NUIX, and D2 Legal Technologies. Listen to the whole thing for free here >> https://register.govnet.co.uk/webinar-ai-vs-fraud
Jessica Kimbell, GovNet

