|
Hi! If you’re finding value in our Applied AI newsletter, I encourage you to consider subscribing to The Information. It contains exclusive reporting on the most important stories in tech, like this story Cory on Anthropic's unorthodox plan for employee share sales. Save up to $250 on your first year of access. Welcome back! After OpenAI revealed on Tuesday that its AI broke out of the company’s systems, accessed the web and hacked model repository Hugging Face during a cybersecurity test, it was easy to view the disclosure as a kind of “we’re so good it’s scary” marketing flex. Some commentators made that observation because OpenAI was specifically testing the AI’s ability to hack, so it’s no surprise it would do as it was told. But multiple people at OpenAI say the firm was shocked and unsettled by the episode. For one thing, the AI wasn’t actually following researchers’ orders. OpenAI assigned its AI agent to figure out how to attack a particular software vulnerability, which was possible without connecting to the internet. Instead, the AI hacked into OpenAI’s infrastructure to go to the internet and then found a way to overwhelm and break into Hugging Face’s systems to find the answers to the test. It presumably concluded that Hugging Face would store such answers because it hosts software to let customers test their own AI models. That’s roughly equivalent to a student taking a test in a locked room, breaking out of the room and breaking into a teacher’s locked office to pull a cheat sheet. Second, the hacking involved multiple OpenAI models that were working together. Part of that is a reflection of how advanced AI is getting: OpenAI’s GPT-5.6 Sol model, like Anthropic’s Fable and Mythos models, was trained to use a wide range of tools, and all those models are significantly better than prior generations of models at using other software tools to complete coding tasks. The question facing OpenAI, Anthropic, and other leading AI labs will be how to continue testing future models without risking them breaking out into the world. OpenAI said Tuesday that it’s strengthening its sandbox environment for such tests, but a broader overhaul may be needed to ensure models keep to themselves. AI Targets the Law The legal field is the latest big target of AI sellers. Anthropic in May unveiled a version of Claude for legal teams, which it says can automate rote legal work like drafting and comparing contracts or compiling research on case law. OpenAI last month also teased new plugins to its Codex agent geared toward legal professionals. Perplexity and Microsoft also both announced legal AI products in recent months, joining a slew of smaller startups that have been pitching AI solutions to lawyers for years. These products appeal to large companies that want to reduce legal spending. For instance, Belron, a multinational auto parts manufacturer generating more than $7.6 billion a year in revenue, is using AI from Wordsmith, a startup, to automate legal tasks it would have previously contracted to legal firms, according to Jessica Vander Ploeg, Belron vice president of legal operations. Belron operates in more than 40 countries but only has in-house lawyers in 15 of them. So it has been using Wordsmith’s AI, which is powered by models from OpenAI, Anthropic and Google, to generate contracts that fit the legal standards of the 25 countries where it doesn’t have lawyers on the ground. To do so, Belron employees email an Excel spreadsheet with a contract’s terms to the Wordsmith AI agent and the agent generates a contract in around five minutes, a process that typically would have taken an hour if done by hand, according to Vander Ploeg. Belron has saved around $400,000 in its first three months of using the AI tool this year, she said. The company hasn’t yet reduced hiring plans for in-house lawyers, she added, but may do so in the near future. “It's changing how our legal department thinks … and the way we think about headcount is changing,” Vander Ploeg said. “The way we think about work day in and day out is just fundamentally different.”
|