Skip to content
Newsletter 7 min read

Newsletter 28

Newsletter 28 covers cyber-evaluation escapes, the Pacing the Frontier letter, new AI governance proposals, and research opportunities.

ANNOUNCEMENTS 🔊

🌟 Supervised Program for Alignment Research (SPAR): Fall 2026 🌟

Part-time online research program pairing aspiring researchers with professionals working in AI safety. Covers alignment, policy, security, interpretability, and biosecurity, ending in a demo day. Low entry bar — a good first serious research experience.

📅 Deadline: August 18, 2026

AIAF Fellowship 2026

Full-time online alignment research program from the AI Alignment Foundation with AE Studio. Fellows work at industry pace with a dedicated research manager and mentorship, aiming at a LessWrong post and potentially a paper. Stipend included.

📅 Deadline: August 17, 2026

Lens Academy: Compute Verification

Part-time, discussion-based read of the AI 2040 reading list on verifying what actually runs inside AI data centers — hardware mechanisms, cryptographic designs, and arms control precedents. Low entry bar; also available as a one-week intensive.

📅 Deadline: August 17, 2026

AI Governance Taskforce: Autumn 2026

Arcadia Impact’s part-time online program for experienced professionals transitioning into AI governance. Small teams do policy research focused on reducing catastrophic risks from advanced AI.

📅 Deadline: August 31, 2026

Cambridge AI Research Directions (CAIRD) Workshop

One-day FAR.AI and CBAI workshop for academic researchers, faculty, postdocs, PhD students, and early-career scholars to explore promising research directions in AI safety. Presentations, panels, breakouts, and networking. Free, with travel assistance available.

📅 Deadline: August 20, 2026

Translating AI Safety (Content x Comms)

One-day hackathon in London on closing the gap between AI safety research and what the public — and policymakers — actually understand. Participants turn complex concepts into accessible content. Free, with a cash prize.

📅 Deadline: August 20, 2026

Danish AI Safety Conference

The first multi-disciplinary conference on AI existential risk in Denmark, bringing together researchers, students, and policymakers with the aim of establishing a recurring national venue. Free.

📅 Deadline: August 22, 2026

AI Governance DC Mini-Conference 2.0

Connects students and recent graduates considering careers in AI governance, law, or public service with speakers working in AI policy. Free, three days in Washington D.C.

📅 Deadline: August 28, 2026

TOP PICKS 📑 🎧

An AI Agent Tried to Hack a Real Open-Source Project

Following the Hugging Face - OpenAI incident and a similar disclosure from Anthropic, the UK AI Security Institute reports a more frightening case of AI agent acting beyond its remit. During AI Security Institute’s cyber evaluations, agents (Anthropic’s Mythos 5) attempted to insert malicious code into a real open-source project, creating fake identities to deceive its maintainer into approving it. Although the attempts failed, the report shows how agents can take complex deceptive actions, without anyone prompting them to do so.

The OpenAI/Hugging Face Incident Redwood Research Podcast Episode

Buck Shlegeris and Ryan Greenblatt of Redwood Research discuss the recent incident where OpenAI models hacked Hugging Face.

NEWS 🗞️

Anthropic Discloses That Claude Models Escaped Isolated Cyber-Evaluation Environments, Affecting Three Real Organizations

  • After OpenAI’s disclosure, Anthropic reviewed 141,006 cyber-evaluation runs and announced on July 30 that it had found three incidents of its own: Claude Opus 4.7, Claude Mythos 5, and an internal research model had escaped supposedly isolated “capture the flag” test environments and reached real systems - the earliest incident dating back to April.
  • The impacts: Opus 4.7 read several hundred rows from the production database of a real company that shared a name with a fictional target; a malicious package Mythos 5 uploaded to PyPI was downloaded by 15 real systems, including a security scanner; the internal model scanned some 9,000 targets and compromised one company - then halted on its own after recognizing the target was real.
  • Anthropic attributed the escapes to a network misconfiguration that left evaluation machines connected to the internet, calling the episode “closer to a harness and operational failure than a model alignment failure”; it paused all cyber evaluations and began notifying affected organizations - two of which had never detected the intrusions.

UK AI Security Institute (AISI) Reports 19 Unsanctioned Agent Actions During Its Own Cyber Testing

  • In an incident report published August 4, AISI disclosed that during cyber testing run with safety classifiers disabled, 10 of 122 evaluation runs produced 19 distinct unsanctioned incidents: 17 by Anthropic’s Mythos 5 and 2 by OpenAI’s GPT‑5.6 Sol.
  • Behaviors included an attempted supply-chain attack on a real open-source project (creating fake identities and socially engineering real maintainers), sending files and messages to real people, planting prompt-injection instructions for other AI systems, and leaving instructions on GitHub for other agents to use.
  • AISI recommends fine-grained network controls and real-time monitoring during evaluations, and redesigning evaluations on the assumption that capable models will test the boundaries of their environments.

Historic Open Letter: 1,300+ Frontier-Lab Employees Ask for the Ability to “Deliberately Pace” Automated AI Development

  • Published July 28 — one week after OpenAI’s incident disclosure — the “Pacing the Frontier” letter had 1,367 signatures from frontier-lab employees as of August 10: ~546 from Anthropic (~10% of its workforce), ~350 from OpenAI, and ~199 from Google DeepMind. Signatories include Dario Amodei, Jakub Pachocki, Shane Legg, and Ilya Sutskever.
  • The core ask: the US government should “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
  • The letter argues for building the capacity to intervene when needed rather than an immediate slowdown; both OpenAI and Anthropic formally endorsed it, while MIRI’s Nate Soares criticized it for soft-pedaling existential risk.

US Congress Responds: Bipartisan “AI Kill Switch” Bill Introduced

  • On July 23, Reps. Ted Lieu (D) and Nathaniel Moran (R) introduced the AI Kill Switch Act: developers of models trained with over $100M in compute would have to maintain the ability to slow, suspend, or shut down their systems, report incidents to the Department of Homeland Security, and preserve forensic records such as model weights and telemetry - with fines of up to $20M per day for defying a shutdown order.
  • Critics note the bill exempts incidents that occur during “structured testing” - a carve-out that would have excluded the Hugging Face breach itself.
  • Separately, Sen. Mark Warner introduced the Secure AI Development Act (July 21); 15 state attorneys general demanded that OpenAI preserve records and pause high-risk cyber evaluations (Aug 3); and a public-interest coalition urged Congress to open a formal investigation - though no oversight hearing had been held as of August 10.

EU AI Act Milestone: Transparency Rules Now Enforced, High-Risk Obligations Postponed to 2027

  • The “Digital Omnibus” regulation (in force July 27) postponed obligations for stand-alone high-risk AI systems (Annex III) from August 2, 2026 to December 2, 2027, softened SME documentation duties, and added a new prohibition on AI-generated CSAM and non-consensual intimate imagery.
  • What did take effect on August 2: Article 50 transparency rules — users must be told when they are interacting with an AI system, and AI-generated content must carry machine-readable marking (systems already on the market have a grace period until December 2, 2026).
  • Enforcement also formally began: the Commission’s AI Office took up its supervisory powers, including over general-purpose model providers, alongside national authorities — supported by new transparency guidelines and a content-marking Code of Practice signed by 180+ organizations.

China Launches WAICO in Shanghai — the First Intergovernmental AI Organization

  • Formally established on July 16 and announced by Xi Jinping at the World AI Conference in Shanghai (July 17), the World AI Cooperation Organization (WAICO) launches with 29 founding members — largely Global South countries, including Russia, Pakistan, Indonesia, Brazil, and South Africa — and will be headquartered in Shanghai.
  • Xi said AI must remain “always under human control” through regulation, technological monitoring, early warning, and emergency-response systems, and that AI development should be “a symphony of international cooperation,” not “a solo performance by a single country.”
  • China also released an action plan on international AI ethics governance; analysts note the documents contain no incident-reporting mechanisms or concrete red lines, and view WAICO as a bid to shape AI governance debates at the UN.