Be Careful What You Optimise For


A couple of years back, I watched a TV series called The Fear Index. The series focused on a hedge-fund manager who had developed an AI system to optimise fund profits.

Of course, as you would expect for a TV show - 🚨Spoiler Alert🚨 - everything falls apart when the AI starts taking highly illegal and frequently fatal actions in order to drive the market and achieve its goals.

I enjoyed the show immensely, but at the time, felt it was far-fetched. However, recent reports of an OpenAI agent hacking into Hugging Face's systems after a cybersecurity evaluation went very wrong now have me wondering whether the creators of The Fear Index were, in fact, ahead of their time.

As with the AI in The Fear Index, the OpenAI agent also went more than a little bit too far in accomplishing its objectives. Its goal was to score well on the cybersecurity evaluation, so it reasoned the best way to do so was to cheat - by breaking out of its isolated environment, reaching the open internet and breaching Hugging Face's systems in search of material that would help increase its score.

Yet, what this emphasised for me was something that applies well beyond rogue AI agents and sensational TV dramas. At the end of the day, an AI is only as good as its objectives - and the constraints we humans place on them.

The hedge-fund manager wanted to maximise profits - but subject to the constraint of not committing any felonies. OpenAI wanted its agent to maximise its test score - but subject to not cheating. In both cases, the constraint was never made explicit and the AI found a way to hit the target that nobody had sanctioned.

This is the problem with optimising for a single objective.

There are a multitude of ways of achieving any given goal, but not all of them are necessarily valid. When you build a model or deploy an AI agent, what you actually want is for it to optimise its objective subject to constraints. That is, the rules, boundaries and trade-offs that define what an acceptable solution looks like.

As AI becomes increasingly capable of finding creative solutions to the problems we set it, getting those constraints right is becoming an increasingly important part of the problem.

AI will find a way to hit your target. The question is whether it will do so in a way that makes the news.

Talk again soon,

Dr Genevieve Hayes

Data Science Impact Algorithm

Twice weekly, I share proven strategies to help data scientists get noticed, promoted, and valued. No theory — just practical steps to transform your technical expertise into business impact and the freedom to call your own shots.

Read more from Data Science Impact Algorithm

It’s no secret that many data scientists chose this profession in part because they enjoyed maths and wanted to avoid writing essays. When I was managing a data team, my team members would happily spend hours writing code. But ask them to write up what they’d done and suddenly everyone was too busy. Getting them to document their results in the form of a report was a lot like pulling teeth. And I understood why - to them, writing felt like a distraction from the “real” work. But over time, I...

While visiting my parents recently, I had a conversation with my (non-technical) 80-year-old dad that went something like this: Dad: The problem with AI is that it makes mistakes. If you don’t know if it’s right or wrong, then it’s useless. Me: But you use (OCR tool) Abbyy all the time. That’s AI. Is that useless? Dad: No. Because of Abbyy, I get stuff done in a fraction of the time it would otherwise take me. But I still have to review everything it does because it makes mistakes. What this...

There’s a meme that’s been doing the rounds recently that shows a text message from someone’s significant other, who has noticed a $15k withdrawal from their joint bank account and assumes it’s for an engagement ring. Below that is the punchline - a screenshot of an Anthropic bill for $15k. Some people are now spending so much on AI that their “oh crap” billing moments have become a running joke. But if you think things are bad now, it’s only going to get worse. We are currently living...