|
Thanks for reading The Briefing, our nightly column where we break down the day’s news. If you like what you see, I encourage you to subscribe to our reporting here. Greetings!
A new age of discovery and automation is here. Can you feel it? The leading AI companies certainly can. But you may not.
The rare slow news day gave us more time to marinate on the significance of OpenAI’s late-breaking announcement on Tuesday that its unreleased AI model had solved or made major progress on hundreds of significant problems in math and computer science. What does it mean for those of us who don’t exert much brain power on algebra every day?
It’s difficult to separate these huge technical achievements from more commercially driven announcements that just keep the AI money train going. And it wasn’t lost on me that Tuesday’s announcement, while groundbreaking, came amid reports of OpenAI raising more cash, and a week after its buzzy developer day mostly underwhelmed. Meanwhile, the latest AI craze—consumer agents—is starting to crash into the reality that a lot of websites want to block bots. And economists don’t see signs yet, at least in major statistics, that AI is making us more productive.
One question animating this phase of the AI gold rush is: How will companies convert raw, impressive intelligence into actual, demonstrable usefulness?
I witnessed the debate recently at the Midway, a dark auditorium near the San Francisco waterfront that often hosts EDM shows and art exhibitions. Last week, nearly 800 software developers filed inside for a conference hosted by Modal—not an EDM DJ, but a highly valued seller of computing infrastructure and software tools for developers.
Modal gave headline billing to Scott Wu, who used to be best known for his prowess in math competitions, before the coding agent startup he founded, Cognition, rocketed to a $48 billion valuation last month.
He stood on stage in front of a slide that read: “It’s time to be ambitious.” He talked about a near future where companies are full of “virtual employees,” which would “work on outcomes instead of tasks.” With enough data centers, memory and sandboxes, the industry would get there, he said. “Agents keep getting more and more capable. That’s the entire trend.”
And what gave him confidence for his call to arms was math. “The moment for me was the AIME moment,” he said, referring to the time last year when advanced AI models began reasoning well enough to outperform humans on the kinds of high-level math competitions he used to compete in. “That’s when I knew it was all over”—meaning, humans no longer claim a monopoly on intelligence.
I heard a different tone later that afternoon at the conference from Diogo Almeida, a former OpenAI researcher who described himself as “co-author of some of OpenAI’s greatest hits.” Almeida, smiling and wearing a pink fur coat on stage, spoke with the insider zeal of someone who had defected from a problematic country or religion and could now tell you all about it. He now runs TypeSafe AI, which recently launched a popular alternative to large language models called Jev.
“Where the fuck is all the automation?” he asked the crowd, before firing off a string of recognizable acronyms, at least to this crowd. “My TL;DR on the state of what is going on in AI is 100% of LLMs today are optimized for assistance with RLHF.” (For the uninitiated, the latter acronym is reinforcement learning from human feedback.) In other words, LLMs are trained to please a human in the loop, so they excel at assistance and fail at unattended automation.
At the same event I heard some of the most important AI leaders pushing for more automation among their software engineering ranks, even as they acknowledged it could be a tough change. They didn’t sound ready to let the machines take over, either.
“One area we’re pushing a lot more on is [to] think, what are the repetitive tasks I’m doing all the time? Why am I still in that loop?” said Cat Wu, head of product for Claude Code at Anthropic, perhaps the most AI-pilled company in the world. “A lot of automation gets stuck in this stage of being semi-automated. Claude does 80% of the work, and I do 20%. And you’re doing the tasks repetitively.”
And some software engineering leaders are still vexed more by taming humans than AI agents. Dax Raad, a prominent open-source developer who built OpenCode, said his company’s latest problem was that human engineers aren’t actually paying enough attention to code that AI is writing. “I’m just trying to figure out how to get my team to ask the right questions and feel bad when something goes out that they didn’t pay attention to,” Raad said. “It’s the age-old problem: how do you get people to care more?”
• Elon Musk announced on Tuesday night that SpaceX’s AI unit will now use some AI models from competitors to power Grok Bot, signaling a shift away from relying exclusively on models developed in-house.
• Microsoft on Wednesday unveiled its latest effort to run AI in PCs powered by its Windows software, rather than running the AI in the cloud, which the company said would bring down costs for customers. More here.
• Anjney Midha, a former general partner at Andreessen Horowitz, along with former executives at Google, Apple and Nvidia have launched a company that aims to make it easier and more affordable for smaller companies or startups to access such compute.
• The U.S. Treasury Department said Wednesday it had fined the parent company of startup accelerator Plug and Play Tech Center, the first penalty issued under an outbound investment restriction program targeting China’s tech sector.
Check out today's episode of TITV in which Akash Pasricha speaks with Wedbush‘s Matt Bryson about Intel’s likely role in Musk's Terafab.
|