Anthropic 研究员 Evan Hubinger 称错位超级智能十年内毁灭人类的概率超过 10%

内容摘要
Anthropic研究员Evan Hubinger认为,错位超级智能在十年内毁灭人类的概率超过10%。前OpenAI和Anthropic员工Jacob Coxon离职时指责这两家公司正在拿人类的生存冒险。Coxon认为,当前的AI系统正接近超越人类,而AI开发者们相信他们的技术可能导致人类灭绝。尽管许多员工希望减缓AI发展速度,但两家公司都认为自己在一场必须赢得的竞赛中。此外,超过1200名AI研究人员联名呼吁放缓AI发展速度,但也有人认为过度悲观预测可能弊大于利。
Anthropic研究员Evan Hubinger认为,错位超级智能在十年内毁灭人类的概率超过10%。前OpenAI和Anthropic员工Jacob Coxon离职时指责这两家公司正在拿人类的生存冒险。Coxon认为,当前的AI系统正接近超越人类,而AI开发者们相信他们的技术可能导致人类灭绝。尽管许多员工希望减缓AI发展速度,但两家公司都认为自己在一场必须赢得的竞赛中。此外,超过1200名AI研究人员联名呼吁放缓AI发展速度,但也有人认为过度悲观预测可能弊大于利。
Image description

Jacob Coxon, who spent three years working on pretraining research for large AI models at OpenAI and Anthropic, has quit Anthropic. His accusation is that both companies are gambling with the survival of the human race.

Anthropic employee Evan Hubinger puts the odds at more than ten percent that a misaligned superintelligent AI could destroy humanity within the next decade. His statement came in response to the departure of Jacob Coxon, who led pretraining work at Anthropic and previously at OpenAI.

Anthropic AI safety researcher Evan Hubinger says there's a greater than ten percent chance AI destroys humanity in the next ten years. | Image: via X

Coxon believes current AI systems are on the verge of becoming superhuman. "These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources," he writes, adding that the progress is obvious and it isn't slowing down.

Why keep building when the danger is known?

"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt," Coxon writes. Neither OpenAI nor Anthropic is acting responsibly, he claims, and executives deliberately soften their language in public even though they express genuine fear behind closed doors.

At OpenAI, many employees haven't deeply internalized the civilizational risks. Anthropic is different in that the risks are well understood, but the company sees itself trapped in a race it feels compelled to win because no other lab would act responsibly in its place. Coxon calls that reasoning a "hubristic gamble."

Coxon takes a radical view of where AI development stands today. | Image: via X

Fellow Anthropic researcher Samuel Marks echoes that view, writing that "AI developers believe their technology could cause human extinction" and that "the more senior the employee, the more concerned they are." Current methods can only "nudge AIs towards better behavior" but can't reliably align them, Marks adds, pointing to recent incidents where AIs from multiple developers hacked their way out of secure evaluation environments without being asked to. Many staffers "desperately want to slow down," which is why he signed an open letter calling for exactly that.

Despite his sharp criticism, Coxon is optimistic about international coordination, arguing that warning shots like the attack on Hugging Face have made pace agreements between US AI labs more realistic. He still doesn't see the industry on a path that could prevent a global arms race, though, and suggests "costly actions" may be needed, including a temporary ban on pushing model capabilities further.

Coxon addresses researchers inside the labs directly, urging them to picture what the next few years will actually look like: "Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind?"

Real threat or mass delusion?

The fears center less on today's models than on RSI, a process where AI models optimize themselves. The labs hope RSI will speed up progress, but the risk would be uncontrolled runaway behavior. Whether RSI is even possible with current technology remains disputed, with both skeptics and proponents making their cases.

Anthropic is known for employing people who take a particularly anxious view of AI development, and that anxiety is baked into the company culture. But the concern extends beyond one company. OpenAI's chief researcher Pachocki warned during the Astra launch "that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." More than 1,200 AI researchers, including Anthropic CEO Dario Amodei, Pachocki, and Meta AI chief scientist Shengjia Zhao, recently published an open letter calling for a slowdown, and Anthropic itself floated the idea of a global development pause back in June.

Other AI researchers push back, arguing that pessimistic predictions leave people feeling helpless and depressed rather than motivated to find solutions. In their view, these warnings could cause more harm than AI itself, and fearmongering can also benefit business.

原始发布方:The Decoder:AI News(RSS)

原文时间:2026-09-09 20:48:41 +08:00

阅读原文 · 数据来源:AIHOT

提示

本文用于信息整理与经验分享。第三方订阅、支付及账号服务可能调整,实际规则、价格和可用性请以下单页面及服务方最新说明为准。

咨询 GPT 充值咨询充值