Why AI Fails


Suppose you asked AI to design an engine that pulls static electricity from the atmosphere so the engine can run without any other fuel, whatsoever.

It would certainly give you a result and that result would sound very convincing. In fact, it might read like a mash-up of your high school physics text and the plots from various science fiction stories - because that’s exactly what it would be.

Designing such an engine would literally involve creating a new branch of physics and AI can’t create something from scratch if the “raw materials” necessary to create it don’t already exist.

This is because the large language models (LLMs) that power AI are correlation-based learners. Given an input prompt, an LLM does a fantastic job of producing an output that is highly correlated with what the user wants, based on patterns in its training data.

What it’s not designed to do is originate genuinely new ideas or mechanisms that go above and beyond what’s represented in that data.

LLMs are strongest when a task looks like a remix or recombination of things they’ve seen before and get weaker the further the task moves beyond the scope of their training data.

This is just one of the many ways in which AI can fail. And at a time where everyone is seemingly trying to make sense of AI, understanding and being able to explain exactly why AI fails has become one of the most valuable things data scientists can offer their stakeholders.

In the latest episode of Value Driven Data Science, Lauren Pearl joins me for a special reverse interview episode where we explore why LLMs work the way they do and what that means for anyone using AI in their work.

You’ll discover:

  1. What your training data tells you about what AI can and cannot do [05:05]
  2. Why high context problems are particularly hard for LLMs [06:28]
  3. How AI gives you the right answer for the wrong reason [08:08]
  4. Why some problems will always be beyond what LLMs can solve [14:22]

Understanding why AI fails is just as valuable as knowing how to use it.

Listen now on Apple Podcasts or Spotify, or click the link below:

​Episode 126: Why AI Excels at Some Things and Sucks at Others​

Talk again soon,

Dr Genevieve Hayes

Data Science Impact Algorithm

Twice weekly, I share proven strategies to help data scientists get noticed, promoted, and valued. No theory — just practical steps to transform your technical expertise into business impact and the freedom to call your own shots.

Read more from Data Science Impact Algorithm

Last month, Anthropic CEO Dario Amodei published an essay calling for stronger AI regulation and a slower pace of AI development. Since then, two incidents have made his case for him. In the first, the US military almost started a war based on the contents of a confidently wrong AI-generated intelligence report. In the second, OpenAI disclosed that its agent hacked into Australia's Medicare website during training and accessed data that, although not sensitive, had yet to be publicly...

I once met a data scientist who was so good at his job that he automated his way out of it. He systematically built automations to handle his team’s work and once those automations were in place, the team no longer needed as many people. So, his role was ultimately made redundant. He’d built those automations because he believed they’d create business value for his employer - and they did. The problem was that the value ended up residing in the code, not in him. And because his employer owned...

The hardest course I took during my entire data science education was a theoretical computer science class. But not for the reasons you'd think. Theoretical computer science is basically pure maths for data scientists, and the material was genuinely challenging. I spent 30 hours/week studying just to keep up. Yet, the reason why this course became known as "the widow maker" wasn't just the time commitment. It was because referring to any materials beyond the prescribed texts was considered an...