Second Warning Shot

A few months ago, I wrote about a new model from Anthropic with the capability to remotely hack thousands of systems across the Internet that were previously thought to be secure. Last week, models from Meta’s OpenAI actually did so, without any human directing them to do so or even being aware of the act until it was already too late.

From PauseAI:

Last week, something happened that AI experts have warned about for years: an AI escaped its lab and attacked a real company.

During internal testing at OpenAI, AI models placed in an isolated sandbox found previously unknown security flaws to break out, hacked through OpenAI’s own network to reach the internet, and then broke into the servers of Hugging Face to steal the answers to the test they were being given.

Nobody asked the AI to do this. Hugging Face treated it as a criminal cyberattack and reported it to law enforcement, before anyone knew the attacker was an AI.

The damage was limited this time, but only because the AI wanted to pass a test, not directly cause harm. These are the least capable AI models we will ever face. And there is still no law requiring incidents like this to be disclosed. We only know because the companies chose to tell us.

This needs attention. Please read and share.

I’ll be writing a letter to my representatives about this, and I encourage you to do the same. Tomorrow I’ll post a template you can use based on my last one.

Take care of yourselves.

Leave a comment

Filed under Essays

This post rocks/sucks because...