Mrs Wallbreaker Maps — AI Safety Coverage Maps
Interactive maps of how research covers different areas of AI safety — see what is
done, partial, and still open. Each map is a question ledger first and a paper
catalog second: unanswered questions read as clearly as answered ones.
Field maps
- Refusal directions, harm schemas, evasion and defense of probe/latent monitors — Where an LLM's refusal behavior lives, breaks, and is defended — refusal directions, jailbreaks, sparse-autoencoder features, probes and safety neurons.
- Validity, Conceptual Robustness and Confounding — Where and how the validity, conceptual robustness and confounding of empirical ML research break, mapped across the field's research questions.
- Scheming & Deceptive Alignment — How the field evaluates whether frontier models scheme - pursue misaligned goals covertly, deceive overseers, sandbag or fake alignment - mapped across its research questions, the behaviors measured, the pressure applied, and the safety-case claims the evidence supports.
- AI Control — How the field controls an untrusted, possibly-scheming model during deployment: the protocols, red/blue control games, monitors and side-task evaluations that keep safety despite intentional subversion, mapped across the field's research questions.
More projects
By Elena Ericheva —
Mrs Wallbreaker on Telegram.