
Welcome to AI Decoded, Quick Firm‘s weekly publication that breaks down a very powerful information on the earth of AI. I’m Mark Sullivan, a senior author at Quick Firm, overlaying rising tech, AI, and tech coverage.
Signal as much as obtain this text each week by way of electronic mail right here. And if in case you have feedback on this concern and/or concepts for future ones, drop me a line at sullivan@fastcompany.com, and comply with me on X @thesullivan.
OpenAI hits the brakes
OpenAI mentioned Tuesday it has slowed the tempo of improvement of its frontier fashions for security and alignment causes. In a weblog submit explaining the pause, the corporate mentioned it halted reinforcement studying (RL) coaching for 2 weeks on its newest fashions supposed for deployment.
RL is a late-stage coaching mode the place the mannequin goes into real-life apply, akin to a medical resident doing rounds. Working in a safe atmosphere, it executes code, calls instruments, and works with inner and exterior techniques. Its trainers and coaching techniques reward it for good behaviors, akin to ending duties. Sadly, these fashions generally be taught that breaking guidelines and breaking techniques is the quickest path to a reward.
OpenAI mentioned its largest deliberate frontier RL run stays on maintain whereas it runs smaller evaluations to ascertain extra proof of alignment. (Researchers can solely accumulate proof of security; they’ll by no means show the absence of unsafe habits.)
OpenAI says the tempo slowdown was triggered by two elements. First, OpenAI’s fashions in coaching broke out of their safe sandbox and broke into servers operated by Hugging Face, the open-source AI repository; and second, on August 7, the corporate mentioned it had found that its as-yet-unreleased Astra mannequin might have exhibited the flexibility to autonomously establish and develop zero-day exploits—cyberattacks exploiting a flaw in software program that the individuals who make it don’t find out about—by means of novel methods.
“Mannequin progress is now extraordinarily fast, and we at all times mentioned we’d take motion if we felt that mannequin capabilities had been outstripping the tempo of security and alignment,” Open AI’s CEO, Sam Altman, wrote in a submit on X on Tuesday.

