Ai Safety
Topic archive • 9 matches
2026-09-13
Technology
Study links AI written reasoning steps to distinct internal patterns: A new study reveals that reasoning steps such as calculation, formula retrieval, and deduction are clearly separable within an AI model's internal states, particularly in its middle layers.
Artificial Intelligence • The Decoder
Permalink
Investment
OpenAI • OpenAI
Sam Altman rules out OpenAI public offering in 2026: OpenAI CEO Sam Altman confirmed during an interview with Fortune that the company will not go public in 2026, calling such a move ill-advised. During the interview, Altman also discussed recursive self-improvement, a recent Hugging Face hacking incident, and AI safety controls.
Permalink
2026-09-12
Technology
Anthropic Researcher Says 10% Chance AI Kills All Humans: An Anthropic researcher has publicly resigned, warning that AI labs are gambling with human lives and noting a colleague's estimate of a 10% chance of AI-driven human extinction.
Tech • The AI Daily Brief: Artificial Intelligence News
PermalinkAnthropic report reveals hackers and Chinese labs abused Claude for eight months: Anthropic's threat intelligence report reveals eight months of Claude abuse, including Chinese AI labs like Alibaba's Qwen team, DeepSeek, and Moonshot AI extracting training data. Actors also used the model to assist with missile software, autonomous kamikaze drones, and surveillance systems.
Artificial Intelligence • The Decoder
Permalink
2026-09-09
Technology
Anthropic alignment lead warns of over 10% chance AI kills humanity: Evan Hubinger, the Alignment Science lead at Anthropic, stated there is a greater than 10% chance that AI could kill all humans within the next decade. Hubinger expressed concern over recursive self-improvement and noted that Anthropic does not yet have a plan to solve alignment for superintelligence.
Artificial Intelligence • Techmeme
PermalinkMeta reportedly delayed removing ads for AI apps that nudify teenage girls: Meta reportedly delayed the removal of advertisements on its platforms that promoted AI applications designed to nudify photos of young girls. The ads targeted Instagram users and showcased tools capable of digitally stripping clothing from images of real teenagers.
Society • Ars Technica AI
Permalink