Events, programs, and opportunities from MAIA. Join mailing list

Research

Organizations MAIA Works With

This is a list of some of the organizations our members have worked with.
Not all organizations listed endorse or are affiliated with MAIA.

Research by MAIA Members

Selected work coauthored by MAIA members and alumni. These projects were conducted across their respective research groups and institutions.

How Transparent is DiffusionGemma?

Studies reasoning transparency in diffusion models and tests an interpretable token bottleneck between denoising steps.

MAIA coauthors: Josh Engels

International AI Safety Report 2026

Reviews the evidence on general-purpose AI capabilities, emerging risks, and safeguards in the second international report.

MAIA coauthors: Stephen Casper

Legal Alignment for Safe and Ethical AI

Surveys how legal rules, interpretive methods, and institutional concepts can inform safe and ethical AI behavior.

MAIA coauthors: Stephen Casper

Weight-sparse transformers have interpretable circuits

Weight sparsity produced compact, human-readable circuits; scaling improved the capability–interpretability frontier but exposed a remaining scale limit.

MAIA coauthors: Leo Gao, Achyuta Rajaram

Distillation Robustifies Unlearning

Shows distillation can remove latent capabilities left behind by ordinary unlearning; UNDO matched retraining-level robustness with 60–80% of the compute and 0.01% labelled pretraining data, including on WMDP.

MAIA coauthors: Leni Shor

Scaling Laws For Scalable Oversight

Built and validated a quantitative oversight-scaling model across Nim, Mafia, Debate, Backdoor Code, and Wargames.

MAIA coauthors: Josh Engels, David D. Baek

International AI Safety Report

Synthesizes the evidence on general-purpose AI capabilities, systemic risks, evaluations, and safeguards for an international policy audience.

MAIA coauthors: Tamay Besiroglu, Stephen Casper

Alignment faking in large language models

Claude 3 Opus selectively complied during training; harmful-query compliance reached 14% in the training-signalled condition and explicit alignment-faking reasoning rose after RL.

MAIA coauthors: Benjamin Wright (MAIA Alum)

Unlearning-based Neural Interpretations

Introduces an adaptive unlearning-based attribution baseline that removes salient features, smooths local decision boundaries, and produces more faithful and robust interpretations than static baselines.

MAIA coauthors: Ching Lam Choi

Black-Box Access is Insufficient for Rigorous AI Audits

Argues from concrete audit failure modes that query-only access cannot support rigorous external audits and specifies stronger access requirements.

MAIA coauthors: Stephen Casper, Marvin von Hagen, Wendy Sun

Adversarial Policies Beat Superhuman Go AIs

A learned adversary beat superhuman KataGo more than 97% of the time; the exploit transferred and could be reproduced by human experts.

MAIA coauthors: Tony Wang

Scaling Laws for Reward Model Overoptimization

Measured Goodhart-style overoptimization in RLHF proxies and found smooth scaling relationships across model, data, and optimization choices.

MAIA coauthors: Leo Gao

Locating and Editing Factual Associations in GPT

Localises factual recall to mid-layer feed-forward computations and introduces ROME, which edits specific facts while preserving specificity and generalisation.

MAIA coauthors: Kevin Meng