Skip to content
Newsletter 4 min read

Newsletter 26

Newsletter 26 covers frontier model oversight, METR’s risk report, biosecurity safeguards, and emerging cyber risks from AI systems.

ANNOUNCEMENTS 🔊

🌟 Constellation Visiting Fellowship: Fall 2026 🌟

Lightly-structured 3-6 month funded fellowship for full-time AI safety researchers to work from Constellation’s Berkeley research center. Fellows connect with other researchers and have their housing, travel, meals, and workspace covered.

📅 Deadline: June 12, 2026

🌟 CAMBRIA: August 2026 🌟

Full-time ML upskilling program run by Cambridge Boston Alignment Initiative (CBAI) for aspiring technical AI safety researchers. Based on the ARENA curriculum; covers transformers, mechanistic interpretability, and alignment science. Meals and travel support provided.

📅 Deadline: June 14, 2026

AI Policy Leaders Programme: Autumn 2026

Fellowship by Talos Network which trains aspiring policy professionals for careers shaping AI governance. Provides mentorship, training, and hands-on policy experience - including potential placement within a think tank or policy organization.

📅 Deadline: June 14, 2026

Global South AIS Hackathon

Hackathon by Apart Research bringing together researchers, engineers, and policy professionals from Latin America, Africa, and Asia to produce AI safety tools, evaluations, and policy research. Includes prizes and potential invitations to the Apart Fellowship.

📅 Deadline: June 19, 2026

TOP PICKS 📑 🎧

A “Voluntary” Government Review for Frontier AI Models

Many AI safety experts have long argued that powerful AI models should be tested by the government before release. Yet even today’s most capable models (such as Mythos or GPT-5.5) with very strong cyber capabilities had no such requirement.

Now they do: in a reversal of his administration’s earlier hands-off stance, Trump has signed an order requiring frontier models to be tested for cyber risks before release. Zvi Mowshowitz explains what this will look like in practice, the “drama” behind it, why it matters, and some concerns.

A third-party audit reveal how AI agents could run unsupervised

For the first time, a third party conducted a direct audit of internal AI agents at major labs, showing these agents could plausibly launch unsupervised operations. The assessment warns that as capabilities advance, this vulnerability will only grow—making third-party oversight increasingly critical.

NEWS 🗞️

Anthropic calls for coordinated frontier AI slowdown plan if risks rise

  • Reuters reported on June 4 that Anthropic is calling for AI labs to develop a coordinated plan to slow or halt frontier AI development if risks become too high.
  • Anthropic argues that a unilateral pause would be insufficient unless competing frontier labs can verify that one another are also slowing down.
  • The company says AI systems’ ability to complete autonomous tasks is doubling roughly every four months, increasing the urgency of stronger oversight.

White House executive order creates voluntary pre-release review path for advanced AI models

  • The White House issued an executive order on June 2 titled “Promoting Advanced Artificial Intelligence Innovation and Security.”
  • The order creates a voluntary process through which developers of the most capable AI models can provide models to the federal government for pre-release review.
  • The review process is intended to help assess national-security, cybersecurity, and public-safety risks before broad deployment.

AI leaders back mandatory synthetic DNA and RNA screening to reduce biosecurity risks

  • On June 4, leaders from OpenAI, Anthropic, Google DeepMind, Microsoft AI, and other organizations signed a letter calling for stronger screening of synthetic DNA and RNA orders.
  • The letter argues that advanced AI could lower the knowledge barriers for designing or obtaining dangerous biological materials.
  • The signatories call for mandatory screening of synthetic nucleic-acid orders and relevant equipment to reduce the risk of AI-enabled biological misuse.

METR publishes Frontier Risk Report on misalignment and rogue-deployment risk inside frontier AI companies

  • METR published its Frontier Risk Report on May 19, focusing on misalignment and rogue-deployment risks inside frontier AI companies.
  • The pilot involved Anthropic, Google, Meta, and OpenAI, and examined how frontier AI agents may be used internally by leading AI developers.
  • The report is especially concerned with risks that arise when AI systems are used to accelerate AI research, software engineering, evaluation, and deployment workflows.

Anthropic says Claude Mythos found over 10,000 high and critical vulnerabilities through Project Glasswing

  • Anthropic published an initial update on Project Glasswing on May 22, reporting results from using Claude Mythos for defensive cybersecurity work.
  • The company says Claude Mythos helped identify more than 10,000 high and critical vulnerabilities through work with partners.
  • Anthropic says the bottleneck is shifting from finding vulnerabilities to verifying, disclosing, and patching them responsibly.

Researchers demonstrate AI-driven adaptive computer worms

  • Researchers at the University of Toronto released a paper in early June describing a proof-of-concept AI-driven adaptive computer worm.
  • The system uses an AI agent to generate attack logic dynamically at runtime rather than relying only on fixed, pre-written exploit behavior.
  • The authors intentionally withhold operational details, reflecting the dual-use nature of the research.