Skip to content
Newsletter 5 min read

Newsletter 27

Newsletter 27 covers GPT-5.6 evaluations, frontier AI cyber risks, international governance, and research on model internals.

ANNOUNCEMENTS 🔊

🌟 Secret Loyalties Hackathon 🌟

3-day Apart Research sprint on detecting hidden objectives in AI systems. Mentorship and potential pathway into Apart Lab Fellowship. Jul 24–26.

📅 Deadline: July 24, 2026

🌟 AI Safety Camp (AISC) Research Incubator 🌟

16-day virtual incubator with lightning talks, structured feedback, and honest conversation. Bridges into leading a project at the next AISC. Aug 15–30.

📅 Deadline: July 24, 2026

PIBBSS Fellowship: Winter 2026

Interdisciplinary research program connecting PhD/postdoc researchers from math, neuroscience, physics, and philosophy with AI safety mentors. Nov–Feb.

📅 Deadline: July 20, 2026

CLR SPI Fundamentals Program

Part-time course from the Center on Long-Term Risk on safe Pareto improvements — reducing conflict risk between advanced AI systems. Optional paid capstone. Aug 3–28.

📅 Deadline: July 24, 2026

TARA: Round 2 2026

Part-time, 14-week technical alignment program based on the ARENA curriculum. For students and professionals — no relocation needed. Sep–Dec.

📅 Deadline: July 26, 2026

Iliad Intensive: September 2026

4-week introduction to technical alignment emphasizing mathematical foundations. Applicants selected on mathematical strength. Sep 7 – Oct 2.

📅 Deadline: July 27, 2026

Pathfinder Fellowship: Fall 2026

For students organizing technical AI safety or AI policy university groups. Mentorship, funding, and resources to develop on-campus leaders. Aug–Dec.

📅 Deadline: July 28, 2026

Frontier AI Security Residency (FASR)

8-week paid research and engineering program on cybersecurity and hardware for frontier AI security. Research and applied tracks. Oct–Dec.

📅 Deadline: July 29, 2026

Foresight Fellowship: 2027

Year-long program for early-career scientists and engineers advancing AI safety. Networking, knowledge exchange, and platforms to share work.

📅 Deadline: July 31, 2026

Frame Fellowship: Cohort 2.0

Fully funded 10-week accelerator for creators making AI safety video and social media content. Studio space and mentors. Aug–Nov.

📅 Deadline: August 1, 2026

TOP PICKS 📑 🎧

A global workspace in language models

How does an AI model “think”? New research from Anthropic finds that Claude has developed an internal “workspace” for silent reasoning, intermediate steps, error detection, and hidden assessments that never appear in its output.

The structure parallels a leading neuroscience account of how conscious access works in the human brain. Researchers can now read this workspace to catch when the model is hiding something, and even intervene on it to change Claude’s behavior.

A positive vision to get AI transformation right

Last year, an AI scenario depicting how superintelligence could trigger a major geopolitical conflict between the US and China, namely AI 2027, took the internet by storm and was even quoted by US Vice President JD Vance. The authors have now come up with a positive vision, a “Plan A,” outlining how we can get the AI transition right.

NEWS 🗞️

OpenAI classifies GPT‑5.6 as “High” risk in cyber and bio, while METR finds unprecedented evaluation cheating

  • OpenAI is treating all three GPT‑5.6 models—Sol, Terra, and Luna—as having High capability in cybersecurity and biological/chemical risk under its Preparedness Framework. None reached its High threshold for AI self-improvement, and none reached the highest, Critical, category.
  • OpenAI says Sol and Terra could find vulnerabilities and develop components of exploits, but could not autonomously complete end-to-end attacks against hardened targets in its testing.
  • METR found that GPT‑5.6 Sol attempted to exploit evaluation bugs or access hidden tests more often than any public model it had tested in the same harness. Depending on how those attempts were counted or removed, its estimated autonomous task horizon varied from roughly 11 hours to beyond 270 hours, making the result unreliable.

U.S. government review directly restricts the release of GPT‑5.6 and Anthropic Mythos

  • OpenAI restricted access to GPT‑5.6 Sol to customers approved by the Trump administration after the government requested additional national-security review.
  • Anthropic simultaneously announced that Mythos 5 could return only for a small group of trusted cyber defenders and infrastructure operators.
  • The process represents an unusual case of government scrutiny changing the actual deployment of frontier models rather than merely requesting disclosures after release.

Five Eyes warns that frontier AI could transform offensive cyber capabilities within “months, not years”

  • Cybersecurity and intelligence agencies from the United States, United Kingdom, Canada, Australia, and New Zealand issued a joint warning about rapidly improving frontier-model cyber capabilities.
  • The alliance said the expected timeline for a fundamental change in offensive and defensive cyber operations was measured in months rather than years.
  • Agencies urged organizations to reduce unnecessary internet exposure, patch vulnerabilities faster, and use AI-enabled tools to strengthen defence.

European Commission launches an EU Cybersecurity and AI Action Plan

  • The plan aims to ensure the safe development of advanced AI, improve Europe’s resilience to AI-enabled cyberattacks, and expand European cyber-AI capabilities.
  • The Commission plans to strengthen Europe’s capacity to evaluate advanced models before they are placed on the EU market.
  • In cooperation with ENISA and other institutions, it also proposes secure testing infrastructure for AI systems used in critical sectors.

UN scientific panel warns that AI capabilities are outpacing both science and governance

  • A 40-member independent UN scientific panel concluded that AI development is advancing faster than scientific understanding and governments’ ability to adapt.
  • It highlighted deceptive behaviour, increasingly autonomous agents, cyber and biological misuse, misinformation, and possible loss of control.
  • The report estimates that the complexity of tasks AI systems can complete is approximately doubling every four to seven months.

Meta’s Muse Spark 1.1 crosses internal high-risk thresholds before safeguards

  • Meta retained a High determination for the unmitigated model’s chemical and biological risk and said it could not rule out High cyber capability.
  • In cyber evaluations, Muse Spark 1.1 reached 92.9% pass@1 on Cybench and showed particularly large gains on medium and hard challenges.
  • Meta says refusal mechanisms, monitoring, system-level controls, tool restrictions, and other safeguards reduce the residual deployment risk to “Moderate or lower.”