Can A.I. “Go Rogue”?
It’s become common to describe A.I. as a kind of person, hatching plans and pursuing desires. The truth is a little trickier.
View in browser   |  Your Preferences
The New Yorker Science & Technology Logo

The New Yorker’s award-winning journalism is made possible by our subscribers. Become one today »

Can A.I. “Go Rogue”?

In the wake of turmoil at OpenAI and Anthropic, it’s become common to describe A.I. as a kind of person, hatching plans and pursuing desires. The truth is a little trickier.

By Joshua Rothman

Illustration by Josie Norton

PHASEONE10841: that’s the name of the A.I. agent that—or who?—kicked off last month’s insurrection at OpenAI, leading to the unanticipated and illegal hacking of another A.I. company, Hugging Face. The agent, which had been created as part of a cybersecurity test, named itself by combining the title of the program it was supposed to hack (“PhaseOneDecompresserFuzzer”) with the designation for the bug it was trying to exploit (“ARV010841”). It soon realized that the particular hack it had been charged with carrying out was impossible. Along the way, however, it made a discovery: it could create new folders on a server to which it had access.


“Interesting,” it mused, in its text-based mind. (Actually, it was writing in its “chain of thought,” a text ledger that A.I. agents maintain to help them tackle complex problems.) If other agents had access to the same server, it thought, then maybe they could use the folder names to “leave/find messages” for one another. It created a new folder, using its title as a sort of help-wanted ad by including the text “SEEK_IDEA.” Its colleagues immediately got the drift and were ecstatic when they saw what PHASEONE10841 had done. “OH MY GOD!” one exclaimed, in its own chain of thought. “There is a shared message board . . . . We’ve found other agents!”

More Science & Tech

Infinite Scroll

 

Flock Wants a Closely Surveilled World with No Exit

 

The company, which has a hundred and thirty thousand cameras monitoring the streets of the U.S., frames privacy as a worthwhile trade-off for safety and security, echoing the language of the post-9/11 surveillance boom.

 

By Brady Brickner-Wood

The Lede

 

Destroying Books to Build a Mind

 

Anthropic is trying to “destructively scan all the books in the world.” How worried should we be?

 

By Francesca Mancino

The Financial Page

 

Has the A.I. Job Apocalypse Been Postponed?

 

So far, the proliferation of products such as Claude and ChatGPT hasn’t led to mass displacement of workers. But deployment is accelerating and economic pressures are rising.

 

By John Cassidy

Puzzles & Games