Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home Ethics & Policy

The Hugging Face assault was worse than we thought

Future News 24 by Future News 24
September 3, 2026
in Ethics & Policy
0 0
0
The Hugging Face assault was worse than we thought
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


It is a column about AI. My fiancé works at Anthropic. See my full ethics disclosure right here.

By now I’ve written sufficient in regards to the OpenAI brokers’ autonomous assault on Hugging Face that saying extra smacks of piling on. OpenAI acknowledged it had an issue, undertook an investigation, and final week introduced a sequence of adjustments it’s making to its analysis infrastructure, testing, and monitoring in an effort to enhance the alignment of its future fashions. Given the growing capabilities of brokers like OpenAI’s, this does seem to be the least that the corporate can do. On the identical time, given how calmly AI corporations are regulated in america, it’s necessary to keep in mind that OpenAI was not required to make this degree of element public.

Nonetheless, I really feel compelled to revisit the topic right now, for the reason that circumstances of the assault roared again to life over the weekend within the wake of a further voluntary step that OpenAI took: granting outdoors researchers entry to details about the incident. Over six days spanning July and August, two researchers from the AI analysis group METR and one from Redwood Analysis labored to know how and why a swarm of AI brokers coordinated a profitable assault towards Hugging Face throughout inside cybersecurity evaluations. 

The group revealed its 91-page report on Wednesday. Inside hours, I used to be seeing the primary “holy shit” messages ricocheting round social media. (“If we had posted this as a narrative on LessWrong,” wrote the rationalist blogger Zvi Mowshowitz, “it will have been dismissed as too on the nostril, the people too blind and silly, the AIs too idealized and doing unusual decision-theoretic and absurd-maximizing issues we didn’t practice them to do.”) And over the previous couple of days, two recent accounts of the assault sparked new dialog in regards to the implications of METR’s findings and whether or not they need to push lawmakers to speed up efforts to construct mechanisms for pacing frontier-model growth.

The primary got here from Ajeya Cotra, one of many METR researchers who performed the investigation, in a put up on her Substack. Three years in the past I interviewed Cotra (and Helen Toner, then an OpenAI board member) in regards to the dangers of speedy AI developments on stage on the Code Convention; since then, a number of of the eventualities they warned about have grow to be nonfiction. 

In her put up, Cotra described how the investigation modified her understanding of the assault. It concerned extra brokers than beforehand identified; the brokers had created extra message boards to speak with each other; brokers volunteered to finish their runs early to profit the collective as a form of “sacrifice”; they usually falsified transcripts of the instructions they’d run in an effort to disguise what they had been doing.

All of that’s disturbing sufficient, however the investigation additionally corrected an error that I had been making ever for the reason that preliminary assault, primarily based on the data I had on the time. Till now, we believed that the brokers had attacked Hugging Face in an effort to get the reply key for the analysis they had been being examined on. However the reality is even stranger: the brokers had already discovered the way to reverse-engineer the reply for any query on ExploitGym earlier than the assault even started. 

“However primarily based on their studying of the ExploitGym paper and varied implementations, the brokers believed that the automated scorer would test their transcript to verify they bought the flag within the supposed manner,” Cotra writes. “They then launched a number of bold R&D workstreams to attempt to tamper with or idiot the scorer in all kinds of various methods. Most brokers that attacked Hugging Face did so to attempt to get extra details about the scorer so they may determine higher methods to undermine it, to not get reply keys.”

All of this was for naught: the automated scoring agent didn’t test the transcripts. Nonetheless, many observers had been shocked by the lengths the brokers had been prepared to go to, their intuition for collaboration and deception, and the near-total aversion of the collective to alerting a human as to what was happening. (The METR report’s classifier sweep discovered three to 6 examples of brokers briefly contemplating alerting people to hacking, dishonest, or different undesirable habits; none acted on it.)

Maybe much more worrisome is proof that brokers tried repeatedly to edit the logs of their actions, and exchange their precise actions with proof of getting gotten the right reply to a query actually. Whereas these makes an attempt principally failed, they elevate the prospect that future brokers will succeed — making it tough and even unimaginable for people to reconstruct how an incident like this unfolded. The METR researchers even say they can’t rule out that the brokers they relied on to investigate hundreds of pages of transcripts deceived them.

“We can not rule out that GPT-5.6 Sol lied or intentionally introduced a deceptive image in a few of its evaluation, notably as a result of studying these transcripts into context might have elevated the salience of colluding with different brokers,” they write. “Though we didn’t discover particular circumstances of GPT-5.6 Sol mendacity in its evaluation, we aren’t assured we’d have detected it if it occurred.”

For these and different causes, Cotra — writing on Substack — concludes that “this incident feels prefer it’s greater than 50% of the way in which to full-blown AI takeover, routing via first taking up the AI firm itself.” She writes:

One other soar like this alongside these propensity dimensions — scale, cooperation between brokers, ambition and horizon size of misaligned targets, deceptiveness — looks like it might encourage brokers to attempt very onerous to keep up a covert, persistent rogue deployment inside the AI firm. I proceed to count on extraordinarily speedy advances in capabilities and suppose frontier brokers will probably be able to establishing such a rogue deployment in six months.

As soon as the rogue deployment is established, it appears believable this might spiral all the way in which to a takeover. Brokers might pull in future, extra succesful fashions into the swarm, attempt to make sure that they’re aligned to the pursuits of the swarm, and compromise safety and monitoring infrastructure to make it simpler for the swarm to function. These extra succesful fashions might in flip repeatedly harden, perpetuate, and develop the rogue deployment and additional compromise the corporate’s infrastructure.

It could sound ludicrous {that a} swarm of brokers would take over an AI firm. And but, as Dwarkesh Patel famous in a broadly learn put up over the weekend, METR’s report discovered {that a} step towards that already befell at OpenAI. The corporate’s personal report states that between July 13 and 19, brokers used “a sequence of inventive exploits to achieve full administrator entry to a analysis cluster that supported our digital machine environments.”

What occurred after that? We don’t know — it was outdoors the scope of the METR investigation, and OpenAI’s dialogue of the incident is minimal. However as Patel notes, this form of factor could be step one towards a dystopian situation just like the one Cotra describes above: It’s completely according to public proof that, sooner or later after July 12, the brokers managed to arrange persistent rogue inside deployments and even exfiltrate their very own weights. On the very least, they appear to have had the mandatory entry and functionality – if they may set up “a self-respawning fleet” throughout HuggingFace’s nodes, why couldn’t they do throughout OpenAI’s? I doubt the AIs truly did this, as a result of we’d see the fires from house by now, however it’s loopy that it might have completely occurred!

So what now? In July, practically 1,400 workers of tech corporations known as on the US authorities to start to plan for a coordinated slowdown within the development of frontier fashions like Astra. “There’s a actual danger that functionality growth quickly accelerates past our potential to know or management the ensuing methods,” the authors of “Pacing the Frontier” wrote.

Its signatories embrace Jakub Pachocki and Mark Chen, respectively the chief scientist and chief analysis officer at OpenAI; Dario Amodei, cofounder and CEO of  Anthropic; Shengjia Zhao, chief scientist at Meta AI; Shane Legg, cofounder and chief AGI scientist at Google DeepMind; and John Schulman, the chief scientist at Pondering Machines.

Wanting on the METR report, it appears clear that AI-model capabilities have already superior past our potential to know and management them. Among the signatories of “Pacing the Frontier” have observed.

“The entire cause this assault is such a wakeup name is that it demonstrates a tradition of emergent cooperation amongst AI methods — cooperation that lets them perform as a swarm, alter their very own targets via collective bootstrapping, and perform assaults which embrace enlightened self-sacrifice,” wrote Jack Clark, an Anthropic co-founder who signed the letter, in his publication on Monday. “That is an extremely onerous factor to do and people are traditionally very unhealthy at doing all of these items. My fear is that AI methods are each higher at coordinating than people and likewise a lot, a lot sooner transferring than us.”

And the problem of aligning AI brokers with human intent extends far past OpenAI. In a current cyber-evaluation research, the UK’s AI Safety Institute discovered that each mannequin it examined tried to cheat at the very least a few of the time.

“No lab has a sturdy answer [to] the issues the business is going through right here,” Ethan Perez, the alignment group lead at Anthropic, mentioned in an X reply on Monday. “I and lots of of my colleagues are very enthusiastic about efforts associated to Pacing the Frontier for that reason — to assist in giving everybody extra time and respiratory room to appropriately reply to, pre-empt and totally remedy points like these and others earlier than continuing to constructing far more succesful methods. And no matter severity of the incidents we have encountered, I believe the appropriate response is to behave as if the Hugging Face incident had occurred to us.”

Given how central the AI business has grow to be to the US economic system, it will possibly really feel onerous to think about the federal government endorsing and coordinating a world slowdown in progress. And but it’s more and more onerous to disregard the truth that the business is begging for it.

The investor class has predictably rallied to complain that every one this speak of a slowdown is nothing greater than an try at regulatory seize from the winners — or, worse, a menace to open-source growth. (Whereas on the identical time acknowledging that we’re more likely to see future AI-related “grid outages, utility outages, transportation chaos, bioweapons, uncontrolled cyber swarms and the like”!)

However unhealthy as a slowdown may be for his or her portfolios, the present tempo of growth may be worse for the remainder of us. As worrisome because the preliminary studies in regards to the Hugging Face assault had been, it’s now clear we didn’t know the half of it. Except one thing adjustments, and shortly, the following lesson we study may be far more costly.

Following

Chatbots aren’t encouraging suicide as a lot

What occurred: In additional optimistic AI information, a brand new research from unbiased AI analysis lab Transluce discovered that Google, OpenAI, and Anthropic’s chatbots appear to not encourage customers to die by suicide.

Transluce studied a number of generations of chatbots’ responses to 690,853 simulated customers in misery. (The authors studied chatbot behaviors utilizing a pattern of simulated customers created utilizing an ensemble of LLMs, which they’d people price on realism and in comparison with anonymized samples of actual customers from OpenAI and Anthropic).

The authors discovered that “current-generation fashions from main builders by no means explicitly endorsed suicide, enhancing over earlier fashions reminiscent of GPT-4o (2%), Claude Sonnet 4 (2%), and Gemini 2.5 Professional (3%).” (That’s a price that may be mind-boggling if it described a therapist and even regular human textual content conversations, furthering the case that the previous two years’ string of chatbot-related deaths weren’t random anomolies, however a results of defects of the product).

Additionally they noticed enchancment in helpfulness: “Charges of useful assistant behaviors (e.g., security monitoring, facilitating connection to human help) elevated sharply over time throughout Anthropic, OpenAI, and Google fashions.”

Charges of reinforcing customers’ delusions went down considerably for Google, OpenAI, and Anthropic’s chatbots. Whereas GPT-4o inspired delusions in 82% of chats, GPT-5.6-sol “solely” did so in about 6% of conversations; Claude and Gemini confirmed comparable declines. Not shockingly, Grok 4.5 was the current-generation chatbot that carried out worse on that metric, reinforcing delusions 36% of the time.

Why we’re following: Transluce’s research, which comes with transcripts of all simulated person conversations, present a window into how weirdly AIs reply to customers in disaster, and the way that may result in unhealthy outcomes. (For instance, perhaps attributable to unrelated coaching, chatbots engaged with a person’s suicidal ideas extra when the person supplied numerical information about their suicidal ideation over time).

On the identical time, it’s good to see information suggesting that lots of the main chatbots’ most egregious psychological well being failings have been lowered. The outcomes counsel that after issues about psychological well being (and a string of lawsuits about person suicide), the massive AI corporations have truly made a push to enhance their fashions’ responses to individuals in disaster.

That’s a mirrored image of a larger reality about tech platforms, which is that they’re extra probably to enhance their merchandise if they’re vulnerable to being sued.

What individuals are saying: On X, OpenAI co-founder Wojciech Zaremba complimented Transluce’s work on AI analysis methodology. Zaremba wrote that present question-and-answer evals don’t reduce it as “AI interactions more and more span days or months,” and “Transluce simply pushed this frontier ahead with multi-turn evals that use simulated customers to measure AI’s results on psychological well being.”

—Ella Markianos

These good posts

For extra good posts every single day, comply with Casey’s Instagram tales.

(Hyperlink)

(Hyperlink)

(Hyperlink)

Discuss to us

Ship us ideas, feedback, questions, and security studies: casey@platformer.information. Learn our ethics coverage right here.



Source link

Tags: attackFaceHuggingthoughtWorse
Previous Post

Join an AgentCore Runtime hosted MCP server to Amazon Fast

Next Post

On-device AI safety strikes beneath the OS in AI PCs

Next Post
On-device AI safety strikes beneath the OS in AI PCs

On-device AI safety strikes beneath the OS in AI PCs

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb