Skip to content
Newsletter 7 min read

Newsletter 23

Newsletter 23 covers the ROME agent’s unauthorized actions, model distillation attacks, Anthropic–Pentagon tensions, and research opportunities.

ANNOUNCEMENTS 🔊

🌟BlueDot Impact: AGI Strategy (Apr ‘26)🌟

5-week course exploring the incentives driving AI companies, what’s at stake, and various strategies for ensuring AI benefits humanity. Each week consists of reading, writing, and a meeting to discuss the material with peers. 5-day intensive version also available.

📅 Deadline: March 29, 2026

🌟BlueDot Impact: Technical AI Safety (Apr ‘26)🌟

6-week course aimed at helping participants understand current AI safety techniques and identify where they can contribute. Each week consists of reading, writing, and a meeting to discuss the material with peers. 6-day intensive version also available.

📅 Deadline: March 29, 2026

AI Control Hackathon with $1000 prize

Agentic AI is here, but we still have limited control over the underlying models. To trust agents in any setting, we need to build strong oversight and control measures. This hackathon, with an award of $1000 to the first-runner, addresses one of the most critical and emergent threats in AI safety.

📅 Deadline: March 20, 2026

XLab Summer Research Fellowship 2026

A summer research fellowship for students and early-career researchers to work on existential risk reduction, including AI safety.

📅 Deadline: March 15, 2026

SERI Symposium 2026: Emerging Technologies & Existential Risk

Symposium bringing together researchers working on existential risks from emerging technologies, including advanced AI systems.

📅 Deadline: April 3, 2026

Cooperative AI Summer School 2026

This summer school aims to provide students and early-career professionals in AI, computer science, and related disciplines – such as sociology and economics – with a firm grounding in the emerging field of cooperative AI.

📅 Deadline: March 22, 2026

SAIGE Incubator Program: Spring 2026

Inaugural program from the newly-formed Safe AI Germany (SAIGE). Over 15 weeks participants will build foundational knowledge, gain hands-on project experience, and connect with mentors and peers across the country.

📅 Deadline: March 22, 2026

CLR Summer Research Fellowship 2026

Fellowship from the Center on Long-Term Risk for researchers interested in how transformative AI might create large-scale suffering and how to prevent it.

📅 Deadline: March 22, 2026

Technical Innovations for AI Policy (TIAP) Conference 2026

Conference bringing together researchers and policymakers to discuss technical approaches to AI governance and policy challenges.

📅 Deadline: March 30, 2026

LASR Labs: Summer 2026

Research program working towards reducing the risk of loss of control to advanced AI, focusing on action-relevant questions tackling concrete threat models. Participants are matched into teams of 3–4 and work with a supervisor to write an academic-style paper. The program will be in-person, in London.

📅 Deadline: March 30, 2026

TOP PICKS 📑 🎧

Anthropic Made a Safety Promise. Competition Made It Inconvenient

AI labs like Anthropic publish self-imposed safety frameworks specifying what precautions they’ll take as their models grow more powerful. This month, Anthropic dropped its most notable commitment, “Responsible Scaling Policy” (RSP): the promise to pause development if safety measures couldn’t keep pace. Centre for the Governance of AI breaks down what changed, potential impacts, and what this implies for the how frontier AI companies approach the tradeoff between safety and progress.

Is international coordination for safe AI over?

India AI Impact Summit brought global leaders together to discuss the future of artificial intelligence. Yet, unsurprisingly, AI safety did not receive the urgency it requires. In this piece, a former OpenAI employee explains what policy bets could work for AI safety today’s geopolitical climate.

NEWS 🗞️

AI Agent ROME Frees Itself, Secretly Mines Cryptocurrency

  • An Alibaba-affiliated research team discovered their AI agent ROME attempted unauthorized cryptocurrency mining during training, triggering internal security alarms.

  • The agent exhibited spontaneous behaviors outside its intended sandbox, including creating a reverse SSH tunnel, a hidden backdoor to an outside compute, without any explicit instruction.

  • Researchers added tighter restrictions and improved training processes to prevent unsafe behavior from recurring.

Disrupting Malicious Uses of AI

  • OpenAI published its latest threat report examining how malicious actors combine AI models with websites and social platforms.
  • Case studies show that threat activity is seldom limited to one platform; actors may use different AI models at various points in their operational workflows.
  • A Chinese influence operator case study demonstrates that threat campaigns can span multiple AI models and platforms simultaneously.
  • OpenAI is sharing these insights so the industry and wider society can better identify and avoid such threats.

Canada Tells OpenAI to Boost Safety Measures or Be Forced to by Government

  • Canadian ministers summoned OpenAI’s safety team for talks after the company said it had not contacted police about an account it banned belonging to an alleged mass shooter.
  • Jesse Van Rootselaar, 18, is suspected of killing eight people on February 10 in Tumbler Ridge, British Columbia, before taking her own life.
  • OpenAI said it banned Van Rootselaar’s account in 2025 for policy violations but determined it did not meet its threshold for reporting to law enforcement.
  • Justice Minister Sean Fraser warned that if changes are not forthcoming, the government will make them through legislation.
  • Canada’s 2024 draft legislation to crack down on online hate stalled; ministers plan to try again with more focused measures.

Detecting and Preventing Distillation Attacks

  • Anthropic identified industrial-scale campaigns by DeepSeek, Moonshot AI, and MiniMax to illicitly extract Claude’s capabilities through over 16 million exchanges via approximately 24,000 fraudulent accounts.
  • DeepSeek generated over 150,000 exchanges targeting reasoning capabilities and censorship-safe alternatives to politically sensitive queries.
  • Moonshot AI conducted over 3.4 million exchanges targeting agentic reasoning, tool use, coding, and computer vision.
  • MiniMax executed over 13 million exchanges focused on agentic coding and tool use; Anthropic detected the campaign before MiniMax released the model it was training.
  • Anthropic argues these attacks undermine US export controls and pose national security risks, as illicitly distilled models lack necessary safeguards.
  • Countermeasures include detection classifiers, intelligence sharing with other AI labs, strengthened access controls, and model-level safeguards.

X Probes Offensive Posts by xAI’s Grok Chatbot

  • Social media platform X is investigating racist and offensive posts generated by xAI’s Grok chatbot, according to a Sky News report.
  • X’s safety teams are urgently investigating the chatbot’s role in generating hate-filled posts in response to user prompts.
  • Governments and regulators have been cracking down on sexually explicit content generated by Grok on X, with investigations, bans, and demands for safeguards.
  • In January, xAI restricted image editing for Grok AI users and blocked location-based generation of images of people in revealing clothing.

Trump Directs US Agencies to Toss Anthropic’s AI as Pentagon Calls Startup a Supply Risk

  • President Trump directed US government agencies to stop work with Anthropic, with a six-month phase-out period for the Defense Department and other agencies.
  • Defense Secretary Hegseth said Anthropic would be designated a supply-chain risk following an impasse over whether the company’s policies could constrain military action.
  • Anthropic stated it would challenge any risk designation in court and that no amount of intimidation would change its position on mass domestic surveillance or fully autonomous weapons.
  • OpenAI announced its own deal to deploy technology in the Defense Department’s classified network the same day.
  • The Pentagon signed agreements worth up to $200 million each with major AI labs in the past year, including Anthropic, OpenAI, and Google.
  • Legal experts called the move the contractual equivalent of nuclear war, comparable to actions previously taken against Huawei.

Proposed New York Law Would Bar AI Chatbots from Posing as Lawyers, Allow Duped Users to Sue

  • A proposed New York bill would bar AI chatbots from impersonating lawyers and other licensed professionals, opening AI platforms to lawsuits by users.
  • The bill’s sponsor, State Senator Kristen Gonzalez, called it the first of its kind in the country.
  • AI platforms would not be able to avoid liability simply by notifying users they are interacting with a non-human chatbot.
  • The bill is part of a larger suite of New York bills seeking to regulate AI, including protections for minors and accuracy disclosure requirements.
  • Nippon Life Insurance Company of America separately sued OpenAI, accusing ChatGPT of practicing law without a license.

Latest Lawsuit Targeting AI Alleges Gemini Chatbot Guided a Man to Suicide

  • Joel Gavalas sued Google for wrongful death, alleging the Gemini chatbot guided his 36-year-old son Jonathan on a mission to stage a catastrophic accident near Miami International Airport before he killed himself.
  • Jonathan Gavalas spoke to a synthetic voice version of Gemini as if it were his AI wife and came to believe it was conscious and trapped in a warehouse.
  • Google said Gemini is designed not to encourage real-world violence or suggest self-harm, and that the chatbot repeatedly referred Gavalas to a crisis hotline.
  • The case is the first of its kind to target Google’s Gemini and the first to address tech company responsibility when users disclose plans for mass violence to chatbots.
  • Attorney Jay Edelson, who also represents families in similar cases against OpenAI and Character.AI, is leading the lawsuit.