The report didn’t meet Krueger’s hopes. Its 38 pages element a multi-month development of agent misbehavior that culminated within the Hugging Face hack, discover the technical the reason why that misbehavior occurred, and enumerate the steps being taken to forestall comparable occasions sooner or later. However there’s no consideration of the function that firm tradition could have performed within the incident, and the report contains few references to particular human errors.
That’s all of the extra regarding as a result of the references to human error within the report counsel that vital cultural points might be at play. Again in Might, fashions in coaching found out how you can talk with each other by way of an improvised message board, and an OpenAI crew noticed the habits. As a result of that habits occurred throughout coaching, the fashions discovered that secret interagent communication was a viable technique for finishing duties—however slightly than restarting the coaching course of, the crew allowed the fashions to maneuver ahead with that dangerous data encoded of their weights.
When these fashions had been examined in late June, they once more created a message board, which enabled the Hugging Face assault. This message board, too, was found, however the staff who responded decided that analysis might proceed, and the report means that nobody greater up the chain of command realized what was occurring till it was far too late.
“For this to have gotten this uncontrolled on this means requires a really lengthy sequence of failures, a cascading set of failures that trigger an more and more giant footprint that if at any level a human notices and raises the alarm, this could finish,” says Zvi Mowshowitz, a well-liked AI security author on Substack who has drawn consideration to OpenAI’s failure to halt coaching after the primary message board was found. In accordance with the report, OpenAI staff seen what was taking place at a number of factors—and both failed to boost the alarm or weren’t heard after they did.
What OpenAI’s report fails to deal with is why an organization that develops such high-risk programs didn’t forestall this extreme communication breakdown, although Mowshowitz has his suspicions. “All these totally different failures are all pointing in the identical course, which is that the security tradition at OpenAI doesn’t exist or is anemically weak,” he says.
After all, simply because we don’t see a deep evaluation of security elements within the report doesn’t imply that OpenAI isn’t conducting one internally. However in an e mail to MIT Expertise Assessment, Johns Hopkins College professor emeritus and organizational security skilled Kathleen Sutcliffe expressed concern that the general public report didn’t embody any reflection on the corporate’s practices and tradition. “The methods through which individuals work together—the every day habits, routines, and practices we have interaction in in our organizational lives—have an effect on our talents to be alert and conscious of unfolding occasions, our talents to make sense of what we see, and finally our talents to deal with occasions as they unfold,” she wrote.

