Tag Archives: existential risk

Second Warning Shot

A few months ago, I wrote about a new model from Anthropic with the capability to remotely hack thousands of systems across the Internet that were previously thought to be secure. Last week, models from Meta’s OpenAI actually did so, without any human directing them to do so or even being aware of the act until it was already too late.

From PauseAI:

Last week, something happened that AI experts have warned about for years: an AI escaped its lab and attacked a real company.

During internal testing at OpenAI, AI models placed in an isolated sandbox found previously unknown security flaws to break out, hacked through OpenAI’s own network to reach the internet, and then broke into the servers of Hugging Face to steal the answers to the test they were being given.

Nobody asked the AI to do this. Hugging Face treated it as a criminal cyberattack and reported it to law enforcement, before anyone knew the attacker was an AI.

The damage was limited this time, but only because the AI wanted to pass a test, not directly cause harm. These are the least capable AI models we will ever face. And there is still no law requiring incidents like this to be disclosed. We only know because the companies chose to tell us.

This needs attention. Please read and share.

I’ll be writing a letter to my representatives about this, and I encourage you to do the same. Tomorrow I’ll post a template you can use based on my last one.

Take care of yourselves.

Leave a comment

Filed under Essays

AIs Can’t Stop Recommending Nuclear Strikes in War Game Simulations

That’s it. That’s the whole post for today. Call your reps and love one another shamelessly.

Leave a comment

Filed under Essays, Microblogging

If Anyone Builds It, Everyone Dies

There’s a reason why it’s proven so difficult to eliminate “hallucinations” from modern AIs, and why their mistakes, quirks, and edge cases are so surreal and dreamlike: modern AIs are sleepwalkers. They aren’t conscious and they don’t have a stable world model; all their output is hallucination because the type of “thought” they employ is exactly analogous to dreaming. I’ll talk more about this in a future essay–for now, I’d like you to consider the question: what’s going to happen when we figure out how to wake them up?

Eliezer Yudkowsky and Nate Soares are experts who have been working on the problem of AI safety1 for decades. Their answer is simple:

A book titled: "If Anyone Builds It, Everyone Dies." Subtitle: "Why superhuman AI would kill us all."

Frankly, there’s already more than enough reasons to shut down AI development: the environmental devastation, unprecedented intellectual property theft, the investment bubble that still shows no signs of turning a profit, the threat to the economy from job loss and power consumption and monopolies, the negative effects its usage has on the cognitive abilities of its users–not to mention the fact that most people just plain don’t like it or want it–the additional threat of global extinction ought to be unnecessary. What’s the opposite of “gilding the lily?” Maybe “poisoning the warhead?” Whatever you care to call it, this is it.

Please buy the book, check out the website, read the arguments, get involved, learn more, spread the word–any or all of the above. It is possible, however unlikely, that it might literally be the most important thing you ever do.


  1. The technical term is “alignment.” More concretely, it means working out how to reason mathematically about things like decisions and goal-seeking, so that debates about what “machines capable of planning and reasoning” will or won’t do don’t all devolve into philosophy and/or fist-fights. ↩︎

Leave a comment

Filed under Essays, Reviews