Skip to content

Ai Safety

Topic archive8 matches

Back to homeGEO summary endpoint

2026-09-13

Technology

  • Study links AI written reasoning steps to distinct internal patterns: A new study reveals that reasoning steps such as calculation, formula retrieval, and deduction are clearly separable within an AI model's internal states, particularly in its middle layers.

    Artificial IntelligenceThe Decoder

    Permalink

2026-09-12

Technology

  • Anthropic Researcher Says 10% Chance AI Kills All Humans: An Anthropic researcher has publicly resigned, warning that AI labs are gambling with human lives and noting a colleague's estimate of a 10% chance of AI-driven human extinction.

    TechThe AI Daily Brief: Artificial Intelligence News

    Permalink
  • Anthropic report reveals hackers and Chinese labs abused Claude for eight months: Anthropic's threat intelligence report reveals eight months of Claude abuse, including Chinese AI labs like Alibaba's Qwen team, DeepSeek, and Moonshot AI extracting training data. Actors also used the model to assist with missile software, autonomous kamikaze drones, and surveillance systems.

    Artificial IntelligenceThe Decoder

    Permalink

2026-09-09

Technology

  • Anthropic alignment lead warns of over 10% chance AI kills humanity: Evan Hubinger, the Alignment Science lead at Anthropic, stated there is a greater than 10% chance that AI could kill all humans within the next decade. Hubinger expressed concern over recursive self-improvement and noted that Anthropic does not yet have a plan to solve alignment for superintelligence.

    Artificial IntelligenceTechmeme

    Permalink
  • Meta reportedly delayed removing ads for AI apps that nudify teenage girls: Meta reportedly delayed the removal of advertisements on its platforms that promoted AI applications designed to nudify photos of young girls. The ads targeted Instagram users and showcased tools capable of digitally stripping clothing from images of real teenagers.

    SocietyArs Technica AI

    Permalink