The Model That Cried Wolf


Suppose you were screening for a rare medical condition that occurs in just 1 in 1,000 people. The test you're using is highly accurate, with a true positive rate of 99% and a true negative rate of 99%.

A person tests positive. What is the probability they actually have the condition?

The answer, which can be verified using Bayes' theorem, is just 9.02%. That means it's more than 10X more likely that a person who tests positive has received a false positive result than that they actually have the condition.

The medical profession is now well aware of this and has developed work-arounds, such as performing multiple tests before diagnosing patients with rare conditions. However, there was a time when this wasn't the case, resulting in patients receiving life-changing diagnoses that weren't correct, sometimes with tragic consequences.

Back when I taught Bayesian statistics, each semester I presented this example. But as a data scientist, rather than a medical professional, I never thought it would be directly relevant to my own work.

Recently, while working with a client on a financial risk monitoring tool, I realised I was wrong.

Rare events aren't confined to medicine. They occur across all kinds of applications, including fraud detection, equipment failure, financial distress. And yet, unlike the medical profession, most data scientists building monitoring tools don't think explicitly about the probability of their model crying wolf.

The consequences of not thinking about it can be as devastating as those in medical screening.

If a monitoring tool repeatedly flags warnings that turn out to be false alarms, the people watching it recalibrate. They start ignoring the alerts. And when a true positive result finally arrives, it inevitably gets ignored too. The tool becomes useless at precisely the moment it matters most.

Optimising for accuracy alone isn't enough. If you're building a monitoring tool for rare events, you need a strategy for investigating false positives before you deploy it.

The medical profession learned this the hard way. Data scientists don't have to.

Talk again soon,

Dr Genevieve Hayes

Data Science Impact Algorithm

Twice weekly, I share proven strategies to help data scientists get noticed, promoted, and valued. No theory — just practical steps to transform your technical expertise into business impact and the freedom to call your own shots.

Read more from Data Science Impact Algorithm

If your stakeholders could take two rocks, bang them together three times, spin around, and make a better decision, they would. That’s not a criticism. It’s just the truth about what stakeholders actually want. Your stakeholders do not wake up in the morning hoping that a new predictive model will be waiting in their inbox when they arrive at work. And they do not lie awake at night, wishing that the following day their data scientists will present them with more accurate forecasts. What they...

If you discovered AI had drafted large chunks of a report you'd paid a Big 4 consultancy $435k to deliver, how would you react? The mayor of Wellington recently did just that. But the internet is outraged about the wrong thing. Last November, Deloitte recommended Wellington City Council cut 20% of its staff, with ChatGPT to pick up the slack. However, the analysis subsequently turned out to be flawed. And now it turns out AI wrote the report, too. Based on the way it's being reported, it...

In the first four months of 2026, Uber burned through its entire AI budget for the year. It did so by giving its engineers access to AI tools like Claude Code and then encouraging them to use them “as much as possible”. To make sure their staff got the message, they even set up internal leaderboards to rank staff based on usage. In retrospect, it’s a pretty clear example of what an AI strategy shouldn’t look like. Most organisations today have AI strategies a lot like Uber’s. Their biggest...