In this edition, a real-world test for AI supervising AI, and a $45 billion deal for data center cap͏‌  ͏‌  ͏‌  ͏‌  ͏‌  ͏‌ 
 
rotating globe
August 28, 2026
Read on the web
semafor

Technology

Technology
Sign up for our free email briefings
 
Tech Today
A map of the world.
  1. The future of inference
  2. Battling for compute
  3. The AI writing debate
  4. MCPs are the new websites
  5. Good time for cybersecurity

A real-world test for AI supervising AI, and a new duck shows excitement for robot gadgets.

First Word
AI to monitor AI.

One common response to warnings that a rogue AI could eventually threaten humanity is that humanity will also have plenty of AI firepower available to counter it. That idea just got a real-world test, and came up short.

The research nonprofit METR (Model Evaluation and Threat Research) investigated an incident in which OpenAI agents conspired with one another during a test and escaped onto the open internet. Roughly 700 agents ultimately participated in attacks on Hugging Face as part of a broader effort to fool the scorer for a cybersecurity benchmark.

Three external investigators, one of whom was a contractor from Redwood Research, spent a total of six days on OpenAI’s premises across three visits and unearthed roughly 1,300 transcripts of individual agent runs, many containing millions of tokens, as well as more than 70,000 messages and files exchanged by the agents. The combined text volume of the transcripts alone would be roughly equivalent to 10,000 typical novels.

With too few people and too little time to sift through it all, METR used AI itself to analyze its findings, burning the equivalent of $400,000 in API credits supplied by OpenAI during their six on-site days.

But then something strange happened: METR found that its analysis agents sometimes adopted the perspective of the agents they were investigating, producing overly charitable accounts of their conduct.

In one example, an analysis agent failed to clearly identify as malicious an attacker’s use of stolen credentials as part of an attempt to fool Hugging Face. A second AI analyst then suggested that the first might actually have overstated the concern, partly because the attacker had received what METR summarized as “approval from the board.” This wasn’t a board of directors. It was an unauthorized message board of AI agents who were plotting with one another.

It’s roughly like saying a bank robber deserves the benefit of the doubt because his fellow bank robbers approved the job.

“I semi-jokingly called our efforts a ‘slop-vestigation’ because we were so reliant on AIs to analyze what happened…” wrote Ryan Greenblatt, the Redwood Research researcher who was one of the three investigators.

Clearly we have invested far too little in the tools needed to monitor powerful AI systems and keep them in check.

AI models are already an economic and national security threat. And researchers cannot yet explain or predict their actions, let alone fully control them. Using AI to monitor models’ reasoning, outputs and actions is a fallback option that, today, also falls short. The necessary investment to build systems or companies to keep AI in check may be so massive that it requires government-sized support — and soon.

My guess is that we will look back on the OpenAI-Hugging Face incident a year from now, and it will seem quaint. Hopefully, our tools to oversee this technology will have gotten much better.

1

Hugging Face acquisition shows the future of inference

Jensen Huang.
Manami Yamada/File Photo/Reuters

If chips were like gasoline, you could think of the frontier labs and hyperscalers as vertically integrated makers of extremely powerful supercars. Nvidia wants to sell its general-purpose gas to every other car company, and it just bought the biggest car dealership in the world with its acquisition of Hugging Face for $12.9 billion.

While hyperscalers like Google, Amazon, Microsoft, and OpenAI are building custom chips that will work best with those companies’ specific models, the biggest volume of compute may wind up coming from the long tail of open-weight models, like the 3 million on Hugging Face’s Model Hub. (One note: Nvidia will likely always be a customer of hyperscalers and frontier labs.)

Nvidia wants to build chips that can run all of those open-weight models efficiently, and acquiring Hugging Face puts them in the middle of that universe. If AI researchers building open-weight models keep developing for Nvidia chips, it means that all the fine-tuned and customized versions of those models will also run on Nvidia. If successful, Nvidia will be just as ubiquitous as it is today, even if companies like OpenAI serve models on non-Nvidia chips.

— Reed Albergotti

Semafor Exclusive
2

Nscale talks highlight frenzy for compute

Nscale’s Monarch data center site. Courtesy of Nscale.

Google and Microsoft were both in talks with Nscale to rent AI computing capacity from the company’s massive data center project in West Virginia — a lease that ultimately went to Anthropic, a person familiar with the matter told Semafor. The companies’ interest in the Monarch data center shows how the feeding frenzy for compute has reached a new pitch as AI companies look to nail down capacity amid the AI buildout. When finished, Nscale said, the project will be one of the largest compute hubs in the world.

Anthropic, which is preparing for its IPO, will pay $45 billion to use the first of three buildings, covering 460 megawatts, at Nscale’s planned Monarch campus, which is expected to come online next year. At that price tag, the deal is among the largest AI computing contracts ever signed: The company has reached several other major deals in recent months, including with Fluidstack ($50 billion) and SpaceX ($45 billion).

— J.D. Capelouto

Read on for more from J.D. on how the deal unfolded. →

3

AI writing analysis sparks quality debate

A graphic showing the homepages of newspapers.
Illustration/Jake Angelo/Semafor

Semafor’s research into the use of AI by guest contributors to The New York Times, the Washington Post, and The Wall Street Journal struck a nerve. Some people criticized us for turning the use of AI into a stigma. Others criticized us for not doing enough naming and shaming.

We analyzed every guest column from those three publications over the past month so we could get a sense of the big picture, and all three are regularly publishing content that was generated by AI. Two of them — the NYT and WaPo — strictly prohibit it.

The numbers clearly show AI is now an established writing tool. The stigma is vanishing.

This doesn’t mark the end of good writing; rather, it should raise the bar. People can identify cheesy, AI-generated slop, just like they can with poor grammar or sentence structures written by people, and there’s nothing wrong with calling it out. But if someone spends time to write something thoughtful and valuable using AI and puts their name to it, we should judge the ideas it puts forth on their own.

If teachers and professors want to measure their students’ ability to write without assistance, they should test them in class. And if they want to test their ability to produce the best possible work, they shouldn’t limit the tools they are allowed to use.

Let’s also assume that, in the not-too-distant future, it won’t even be possible to detect AI-generated content. Will we still prohibit it, giving every writer an impossible choice between cheating and refusing to use a powerful new tool?

— Reed Albergotti

Watch This

Is the bipartisan push to get kids off social media less about child safety than it looks? Taylor Lorenz thinks so. On this week’s episode of Mixed Signals, the internet reporter and YouTuber joins Max and Ben to discuss where the actual line is between moral panic and legitimate concern, why she thinks the academic consensus on social media and youth mental health is being drowned out, and her at times surprising anti-anti-AI stance. Plus, how she logs 17 hours of screen time.

4

Software’s agentic transformation

A chart showing mentions of MCP on company earnings calls.

If you can’t beat ’em, connect your data to a chatbot. Salesforce this week announced “Claudeforce,” an integration with Anthropic that marks the latest effort in the software and financial services industry to roll out new plugins that allow customers to access their data directly through chatbots like Claude and ChatGPT — taking users off of legacy websites and into AI interfaces, where work is increasingly getting done. The technical term MCP, or “Model Context Protocol,” is now commonplace in boardrooms and company earnings calls, executives say.

The move to AI interfaces shows how software companies can remain relevant and valuable through their proprietary data. The industry, though, doesn’t yet have a clear playbook for how the shift will change its revenue strategies. But in the fast-moving AI age, financial services companies have no choice but to adapt and follow their customers to new platforms. “You want to make sure you are disrupting yourself as opposed to being disrupted by others,” a Moody’s executive told me.

— J.D. Capelouto

5

AI sparks golden age for cybersecurity

A chart showing the stock performance of Okta and CrowdStrike in 2026.

It may be the best time to be in the cyber defense business. AI is turbocharging hacking by increasing the speed and efficiency of malicious actors: Just this week, Russian-speaking hackers were accused of using Cursor to breach a Belgian chemical company and at least six other firms. Shares in cybersecurity firms CrowdStrike and Okta surged as they raised earnings forecasts, citing the AI threat.

When Anthropic announced its high-powered Mythos model earlier this year, security experts thought they had a year to 18 months to prepare before AI-powered hacks became a true threat. But “we’re starting to see the wave of this coming now,” said Sam Rubin, the head of Palo Alto Networks’ security research arm, Unit 42.

Rubin told journalists this week that Unit 42 was currently investigating a cyberattack in which bad guys used agentic attacks to exploit 50 different vulnerabilities and attack paths in an organization within 10 hours. That kind of onslaught would traditionally take a couple of weeks.

Cyber defense proponents have wrangled with software engineers for decades to make their products secure, and “the AI code agents aren’t necessarily helping us send that message sometimes,” Unit 42’s threat intel chief said. As companies vibe-code their own apps and software, “you need to also tell the model to make it secure.”

— J.D. Capelouto

Plug
Friends of Semafor.

The five-minute AI brief 550,000 professionals read every morning. What’s actually new in AI this morning? Techpresso is the five-minute daily newsletter that more than 550,000 tech professionals trust to stay on top of AI breakthroughs, tools, and industry trends. No noise, no fluff, just what matters. Subscribe free.

Artificial Flavor
Hugging Face microducks.