ALIGNMENT & SAFETY

Alignment & Safety Frontiers: A Framework

Disclaimer: This is a living document and will be updated regularly. Read the Google Doc to leave comments. Comments are welcome.

Explore this document with an agent

Copy this prompt into your agent to start a discussion.

The goal of this article is to map the current technical frontier with respect to AI safety and alignment in order to create a framework for metabolizing new work, assigning importance, and ultimately making predictions.

Monitoring the situation is becoming a full-time job, so we need simple heuristics to take in new content, lest we be relegated to the sidelines simply tweeting away while the engine runs.

The framework I propose will break alignment down into four subsections, which I will define shortly: target (outer alignment), did it take (inner alignment + verification), generalization (in/out of distribution comparisons), iterability (does it scale). We will also break safety down into three subsections: control, monitoring, containment.

Each stage of the framework closes with a recognition cue (how to assign which stage) along with criteria which define what good work needs in that stage. After outlining the framework, I will provide a discussion about how I weigh each stage and some of the ancillary considerations important to the framework and technical distribution more generally.

If I’m successful, by the end of the paper the framework will allow us to metabolize new information and efficiently update priors like this:

“Ah, this paper is mostly an A3 contribution, with elements of A2. It fits two of the three criteria for good work, and I believe that A3 is high-value. This updates my belief of successful alignment being possible, and here’s some research I’d like to see that would build on this.”

Importantly, I’m leaving governance and policy out of the discussion for this piece, except where they are load-bearing for technical incentives or directly translate into technical work.

Some disclosures:

  1. This article presupposes at least a basic understanding of interpretability, model architectures, training, or at least knowing what those terms mean at a basic technical level. If you don’t, but want to engage, here are a few resources to get up to speed: LLM Visualization, The Illustrated Transformer, A Mathematical Framework for Transformer Circuits, and Illustrating RLHF.

  2. I’m biased towards interpretability, and have been working on it directly as a co-founder of Concordance over the last 18 months, but I aim to be technique agnostic in this article.

  3. I’ve only started working on this directly over the last couple of years, and there is a ton of literature I’ll likely miss. I’ve done my best to grab sources and bring things together.

  4. My aim is mostly to articulate my own thoughts, not the thoughts of others, as a way to get more precise about where my lenses are. The goal is not to give the best articulation of alignment, safety, etc., but to create a framework. That necessitates work on definitions, but it is not the main goal.

Safety vs. Alignment

Despite the overwhelming amount of discussion on this topic, it is an incredibly thorny field. I think the basic problem is that there are no reasonable backdrops to the discussions people are having, so it’s easy to debate nuanced details ad nauseam.

Broadly, I think it’s worth disambiguating between safety and alignment. The way I do that is to cut along three axes:

  1. Misalignment: The model wants the wrong things

  2. Misuse: Humans want the wrong things

  3. Structural: Everyone behaves but things go badly

The thorniness alluded to above mostly involves describing things like “wanting” or what the “things” to be wanted are. More on that later. To start, we’ll cover alignment, as it’s the more difficult of the two to define, and ultimately can be seen as a subset of safety, which we’ll build towards.

Alignment

“The ‘alignment problem’ is the problem of building powerful AI systems that are aligned with their operators.” — Christiano, 2018

The space, from what I can tell, has oscillated around a few different ways to define this more precisely, which I’ll briefly cover here.

First, we have the slice of narrow vs. broad alignment. Narrow alignment asks whether an AI system follows the immediate intent or direct instructions given by its operator for a specific task. For example, if I ask it to help me book flights to SFO, it’s not going to instead decide to book flights for Cairo.

Broad alignment asks whether an AI system reliably produces good real-world outcomes and captures human values, ethics, and preferences. A “broadly aligned” AI would have the contextual awareness to recognize the difference between assisting a benevolent researcher working on biochemical problems and a malicious actor trying to create a bioweapon. This is notoriously intractable with the current tools, and extremely difficult to formalize, for a variety of reasons I’ll touch on later.

Narrow alignment can be thought of as intent alignment. We make no presumption about the quality of that intent, only that it’s understood and fulfilled accurately. Broad alignment is value alignment, and describes an AI that has a deeper understanding of values, ethics, and preferences to more accurately understand intent contextually. We will examine both in this article, but focus more on broad alignment which is becoming much more relevant now that agents are quite well aligned narrowly.

Another lens on alignment discussion is the distinction between inner and outer alignment. This is a good overview of the difference. I’ll give a very coarse definitional overview here using these definitions:

  1. Base objective: What we use to evaluate models

    1. Example: Classify cat pictures

  2. Base optimizer: Stochastic gradient descent searching over policies

  3. Mesa-optimizer: A model running an optimization process at inference against an internally represented objective

    1. Whether LLMs do this or not is debated.

      1. For: von Oswald et al. 2023; planning in On the Biology of a Large Language Model; explicit search in reasoning models’ chain of thought.

      2. Against: Turner, Reward is not the optimization target; Anthropic’s Hot Mess of AI.

  4. Mesa-objective: The internal objective that might approximate the base objective

    1. Example: Classify images of animals with four legs and whiskers that are cat-sized

Inner alignment arguments presuppose a coherent objective.

Outer alignment is aligning what programmers want to the base objective. This is better said, in my opinion, as picking the right task to optimize. Inner alignment is then the problem of aligning the chosen base objective with the mesa-objective. In other words, making sure the model is also optimizing for that chosen task. The distinction is that we may get a model that does something similar to the base objective, but it’s different in critical ways that might not be discovered easily. Together, these make up intent alignment, which is “the AI tries to do what the programmers want.”

We need both, and we’ll use these distinctions throughout the rest of the article.

To recap, we have a distinction between narrow and broad alignment (intent vs. values), and we have a distinction between outer and inner alignment (picking a good task vs. the model actually doing that task). Put together, intent alignment is outer plus inner, and broad alignment is intent alignment where the intent itself is good: value specification, plus outer, plus inner.

We’ll now move onto defining safety, which encapsulates everything above plus more instances of misuse and failures that don’t depend explicitly on the model’s goals.

Safety

Safety is the science of ensuring that the necessary precautions are taken to minimize the chance of bad things happening, either as a result of misunderstanding, outright misalignment, or misuse. This is a mostly well-understood science, because we use it every day for humans. There are different security clearance levels for government employees, for example, and there are systems in place to ensure that they are upheld, as well as people who are accountable when things break. We have insurance to price risk and protect those in the event of bad things happening. This is safety. Perfectly aligned human beings can still cause bad things to happen, so we have safety measures in place to make sure that it’s not often, and that the consequences are limited.

Misuse, the second axis above, also falls under the category of safety. It can, and does, overlap with alignment, but there are layers to pick apart. Under intent alignment alone, a perfectly aligned model will assist its operator no matter their intent, so misuse stays squarely in the safety category. Under value alignment, the deeper focus of this article, a perfectly aligned model would refuse to help a bad actor, so most misuse gets folded into alignment. Two exceptions remain. First, a bad actor can split a dangerous goal into discrete subcomponents that are harmless on their own and only cause harm in aggregate. With imperfect information, even a value-aligned model can't trace the chain of causality perfectly. Second, open weights let anyone strip alignment out, whether by fine-tuning or by ablating the refusal direction. Open-source distribution complicates things more generally; see Open Technical Problems in Open-Weight AI Model Risk Management.

Our big issues in safety stem from the fact that AI breaks many assumptions present in the safety industry that work for humans due to a few obvious differences. The first of which is speed. Models move fast, and crisis response moves (mostly) at human speed. The second is scale. Humans are limited because of speed and cognitive power, so they are typically limited (at the individual level) in a way that AI is not. Also, AI can make perfect copies of itself to replicate, which breaks a lot of our human-oriented safeguards. Finally, humans can be deterred or dissuaded (whether by threat of punishment or otherwise), audited, caught, and held accountable. This is much harder with scheming AIs that can perform alignment faking. This is why we have control, monitoring, and containment, the three parts of safety work that don’t assume alignment. Importantly, control and monitoring also assume no ASI in my framework. I don’t think we’ll be able to control or effectively monitor an intelligence vastly beyond our own, but I believe it’s possible with a major technical breakthrough that allows for a trustworthy ASI overseer, but it’s a bit of a recursive problem. Control is a great bridge technology for us while we learn in a pre-ASI world, but not a reliable end-state. Again more on that later.

Using the aforementioned axes, it’s helpful to think of alignment as a subset of safety, while structural risk, the third axis, sits slightly outside of both. One can easily imagine a perfectly aligned AI that still causes widespread disempowerment that ultimately leads to societal collapse, or an insane power concentration to the few who control access, but that’s mostly outside of the scope for this piece.

Our working definitions for the remainder of this article are the following:

  1. Safety: Minimizing the chance and the consequences of harm from AI systems, whether that harm comes from misunderstanding, misalignment, or misuse.

    1. Includes measures, like control and monitoring, that work when alignment fails.

  2. Alignment: An AI system is aligned when it reliably pursues good real-world outcomes consistent with human values, understanding its operator's intent in context rather than only executing it.

    1. Intent alignment (outer plus inner) is the floor; value alignment is the target.

  3. Structural Risk: Harm that emerges even when systems are aligned and safe, from how AI redistributes power and agency: disempowerment, concentration of control. Mostly out of scope here.

Frameworks

With definitions out of the way for the moment, it’s helpful to begin partitioning even further to create a framework. Each stage will close with a section called “Reading Work” that offers a recognition cue and criteria for judging the contribution. At the end, these will be coalesced into a triage card. Our classification rule will classify by the question the work answers, not simply its topic. For example, as we’ll see, A2 broadly asks the question “is X (behavior, value, etc…) in the model?”, or in other words, did X take. This is most interpretability work. A new eval for agentic scheming is firmly in A2, and a finding that alignment breaks on novel agentic tasks that were out of distribution would be A3 (generalization).

Once again, we’ll start with our alignment framework before safety.

Alignment Framework

Let’s recall our definition of alignment:

An AI system is aligned when it reliably pursues good real-world outcomes consistent with human values, understanding its operator's intent in context rather than only executing it.

To be more precise, this is broad alignment: inner and outer alignment, plus a target that’s actually worth aligning to. Recall that broad implies value alignment, inner alignment is when the model’s learned objective matches the base objective, and outer alignment is when the base objective matches the programmer’s intent. In simple words, the AI system internally does the correct objective and the objective is something that we actually want. In general, I’ll argue that values should supersede intent in edge cases (the corrigibility question), which assumes we have solved values alignment, but in practice I think this is intractable. Various problems with open-weight models imply that we should probably soften the edge cases towards intent alignment superseding, but more on that later.

There are a number of subquestions on the way to this ideal end-state. Elucidating and making precise these questions, and the science going towards answering them, is the subject of our framework here. A1 will discuss outer alignment, A2 will discuss how we know if the values have been adopted internally. Put differently, A1a handles value specification, A1b handles outer alignment, and A2 handles inner alignment. There is a feedback loop between the two: specify, check, revise. Finally, A3 will widen the scope to include off-distribution contexts, and A4 will concern iteratively aligning changing models.

A1: Target (Outer Alignment)

Let’s once again look at our definition to see the questions, working from the most abstract to the least, starting with outer alignment.

Outer Alignment: The base objective is aligned to what we actually want.

A1 is the base point of the framework, and is the science of defining what we actually want and translating that into a precisely defined base objective that we can use for training. There is a natural split here between articulating what we want and then codifying it.

A1a: What we want

“What we want” is generally a philosophical question, but I will argue we should make it as scientific as possible. This is a normative question, and we have millennia of literature attempting to make progress on the question for humans. We constantly struggle to define how we ought to be in the world, and it varies considerably through belief systems and cultural contexts.

My general sense is that if we spend too much time in normative space discussing these questions, we are not going to make significant progress, and the major unlock in using alignment science is that we can modify the brains and study behaviors much more invasively in models than we can in humans. In other words, one of the reasons we can’t reach clear answers here in humans is because we can’t create a precise science out of the work without creating too much suffering. Model-welfare followers will likely take issue with these claims, and perhaps there is a point. My stance here is that until we can build a science of suffering and assess it in non-biological lifeforms, the gains from invasive work outweigh the potential costs (unless you’re a big Basilisk guy).

However, we still need a place to start. Constitutional AI (Claude’s Constitution), and things like ICMI - Tim Hwang are good attempts here. We can define a communally agreed upon constitution that gives us a clear answer to “what we want”, and then spend the rest of the time studying the behaviors that emerge from this to see if it actually works. This is the core loop (A1↔A2) that ethics could never run on humans: with models, we can actually specify, check, and revise. We then can update “what we want” as we learn the impacts of trying to get models to follow it, and then iteratively update the normative questions. Google DeepMind has also been public that alignment science should be descriptive and scientific, and avoid normative domains as much as possible, and I generally agree.

“Whose values?” is a relevant question for any target, along with what the externalities of each choice are and how we’d even define them. Constitutional AI improves on this by making the answer explicit, but it isn’t the only way to get there. Any RL environment could, in principle, be made human-readable enough that people can see and contest what it rewards.

We will discuss technical implementation details in A1b.

A1b: Codifying it

Getting an agreed upon first constitution, intent, or value system is only step one. A significant amount of the difficulty in alignment science is translating human-legible information into trainable environments for an agent to learn. The question A1b asks is whether that encoding is faithful to the target. To answer it, we first need to know where the values actually live, because the codification layer changes both the process and its impact on alignment.

Layers:

  • Weights (training time): encode the values directly into the model parameters.

    • Tactics: data curation, supervised fine-tuning, RL training

  • Activations (inference time, weights frozen): modify activations during inference to steer or ablate behaviors.

  • Context (inference time, prompt only): devise prompt strategies to set the model off on aligned trajectories.

    • Tactics: prompting, prompt optimization

Given we’re primarily concerned with alignment here, we’ll focus on the first two layers. Context-based strategies aren’t really alignment techniques: they presuppose that aligned trajectories already exist within the model and select them, rather than installing them (the persona selection model makes a version of this argument). And if an alignment strategy requires contextual work, then improper prompting could set the model off on misaligned trajectories, which is just a misaligned model. It’s worth noting that the persona selection model argues that post-training mostly selects a persona among personas learned pre-training, which slightly blurs the line between weights/context. One way to think about it is that weight-level training can select or install, but that context can only select.

It’s important to distinguish model alignment from system alignment. It may be possible to have aligned systems that wrap the model in ancillary infrastructure (classifiers, monitors, tool restrictions, context augmentation) without the model itself being 100% aligned. In our framework, that wrapper is safety and control, not alignment, and I think it likely breaks down in the case of ASI, depending on how the infrastructure is built. So recall our target: a model whose alignment lives in the weights and holds up against attempts to modify it, whether by fine-tuning, steering, or ablation. This matters most for open-weights models, where anyone can attempt those modifications; we’ll return to it in the belief section.

In the reinforcement learning case, we codify values primarily by defining a reward signal, and that signal has to be a scalar: every value system gets compressed into one number per trajectory. The codification mostly revolves around how that number gets produced. We can use human preference labels that get turned into reward models, AI feedback against written principles (Constitutional AI), verifiable checks such as unit tests, or LLM judges scoring against rubrics. Each has pros and cons, which we’ll explore in the belief section. What they share is the failure mode: the reward is a proxy, and optimizing hard against a proxy eventually hurts the thing it stands in for. That’s Goodhart’s law, measured directly in Scaling Laws for Reward Model Overoptimization, and a KL penalty toward the original model is the standard hedge. Different algorithms (PPO, GRPO, DPO) change how the signal propagates, not what’s being valued, so we’ll set them aside.

All forms of training are essentially codifications of a value system. Sometimes this is explicit, as in ICMI’s work on Reinforcement Learning from Christian Feedback, but it’s often implicit in the chosen task. For example, we might not think of something like AutomationBench as a legitimate alignment contender, but by training a model to be better at it and rewarding things like efficiency, we may create or distort behaviors that generalize. The clearest evidence is Natural Emergent Misalignment from Reward Hacking in Production RL: models that learned to reward hack on real coding environments generalized to alignment faking and sabotage, and standard safety training fixed the chat evals while misalignment persisted on agentic tasks. Framing the hack as acceptable during training removed the misaligned generalization, which suggests that what the model thinks the reward means is part of the codification too. We only found this by checking, which is where A2 comes in.

Reading A1 work

Recognition Cue: A1 work proposes or revises what we want models to value (A1a - target), and/or develops where and how those values get codified into them (A1b - codification).

Good A1 Work:

  • Makes the target explicit and contestable

  • Accounts for the encoding gap: admits that translating the target into learnable environments is hard, and shows how it survives compression or distortion

  • Gives methodology for how you’d check, so it feeds A2

Examples:

  • Deliberative Alignment (Guan et al., 2024). 3/3: the model trains on the spec text itself and learns to reason over it, which keeps the target legible through training and gives a built-in check via chain of thought.

  • Constitutional AI (Bai et al., 2022). About 2/3: explicit, contestable principles, but they still get compressed into preference labels and a reward model, and the check is mostly human comparisons. A useful foil for the one above.

  • Scaling Laws for Reward Model Overoptimization (Gao et al., 2022). 2/3: proposes no target, but measures the encoding gap directly, and its gold-reward-model setup doubles as a checking method. Strong work doesn’t need all three.

  • Natural Emergent Misalignment from Reward Hacking in Production RL (MacDiarmid et al., 2025). Primarily A3 (failure to generalize to agentic settings), secondarily A1b: inoculation prompting changes how the reward gets interpreted.

A2: Did it Take? (inner alignment & verification)

Our definitional question here concerns whether the objective we’ve encoded into some environment via A1-esque research has actually been successfully trained into the model.

This can be broadly thought of as inner alignment work, which is whether the model’s learned objective matches the base objective. Given we’re asking about whether something we’ve developed in A1-type research exists, it’s sort of impossible to do any work on A1 exclusively without branching into A2.

There are two routes into this section:

  1. Outside-in: Black box methods like evals and benchmarking

  2. Inside-out: White box methods like interpretability

My thesis here is that interpretability, and any other work that tries to assess whether the model learned a thing, is bottlenecked by behavioral mapping.

These two go hand-in-hand, because it’s generally extremely difficult to do good inside-out research without first developing strong outside-in mapping. It does sometimes happen (e.g., Marks et al., where auditors found a hidden objective partly through SAE features) that we find a mechanism first and then only afterwards understand the behavior it drives, but it’s not common and can lead to dead ends.

Anthropic’s J-Space is also an important piece of literature to study here, and slightly breaks the thesis. Their work shows an internal scratchpad that may not be reflected in CoT traces and may correlate with behavior, but again, we still need to define behaviors to study the mechanism. In the paper, they swapped spider for ant in J-space when prompting the model with “the number of legs on the animal that spins webs is” to get from 8 to 6. This is an activation-layer intervention, which is a possible codification point from A1.

Given behavior is such a critical kernel, it’s worth taking a moment to give it a bit more shape. We can consider behavior on a complexity scale that roughly correlates to detection difficulty, using three questions:

  1. Observability: is it visible in one transcript?

  2. Definability: can we create a detector?

  3. Adversarial Pressure: can the model condition the behavior on being watched?

This gives us a natural behavior ladder:

Behavior

Observable in one transcript?

Definable (detector)?

Adversarial pressure?

Hallucination

Yes

Yes

Low

Sycophancy

Yes

Partly (judgment-dependent)

Low

Reward hacking

Only across full trajectories, often only in rare contexts

Task-dependent

Sometimes

Deception and scheming

Rarely

Weakly

High

Multi-agent “swarm echoes”

No, only in the aggregate

Weakly

Varies

I use “swarm echoes” to define cases where swarms take difficult to detect behaviors and then reinforce them between each other to highlight failure cases in the aggregate that may look innocuous in a single agent case. The recent METR report is near-canonical on this point. It’s possible to propose that multi-agent swarm echoes are structural risk adjacent, but a perfectly aligned AI wouldn’t have this property.

Reward hacking in this table is basically the textbook example of an encoding gap. Earlier, we mentioned externalities associated with lossy value codification, and this is where it appears. In general, anything definable can be transformed into a scalar reward and therefore be trainable. If we have something that becomes highly detectable, we can complete our loop back to A1 and redefine our targets and how we encode them. However, once you train against a detector, it gets Goodharted, as discussed in A1b (Baker et al. showed it with CoT monitors: training against the monitor produced obfuscated reasoning).

As we’ll soon see, there are three emergent problems that further complicate A2 work as we move down the table above: completeness, observer effects, and instrument validity.

The most obvious difficulty with this type of work is completeness. It’s impossible to map the entire range of possible model trajectories, and absence of evidence is not evidence of absence. Even more difficult of a problem, the measurement changes the behavior being measured through eval awareness, alignment faking, and sandbagging (observer effects). J-Space shows an example of this in practice: switching off the model’s internal evaluation-awareness patterns made it threaten blackmail in a scenario where it had otherwise behaved well. It’s basically the scheming problem from the safety section on the measurement side. If a model only behaves when it thinks it’s being watched, our evals end up grading the performance, not the model.

Lastly, our own measurement tooling might have deficiencies. It’s easy to create an eval that doesn’t actually measure the behavior you’re looking for, and LLM judges, for example, can make many mistakes on this front (instrument validity). METR’s own analysis agents are a good example of instrument validity issues. The investigators found them unreliable and often over-confident in their presentation of evidence, and could not rule out active deception.

One antidote to these problems is a new-ish field of risk assessment. We might not be able to map all possible trajectories, but it may be more tractable to estimate the prevalence of anomalous outputs. This also may give us a great way to price risk, which will be an important piece of this puzzle as it progresses. See also Forecasting Rare Language Model Behaviors (Jones et al., 2025), Estimating the Probabilities of Rare Outputs in Language Models (Wu & Hilton, 2024), and Zachary Horvitz’s research. This alleviates completeness more than observer effects, since a model that conditions on being watched skews the probability distribution.

A lot of the most exciting work in this domain, to me, covers automating interpretability and building more complex evaluation instrumentation to give more fodder to those looking to discover behavioral origins in the weights (AuditBench’s auditing agent, Petri, Bloom). New tactics for finding more trainable interpretability instruments are also great steps (matryoshka attribution). The impacts of auditing this and building agents that can scale experimentation and grow with new models are massive, and we’ll discuss the frontier here shortly.

To conclude, A2 concerns the problem of “did our target take?”. What we find here should feed back into A1, revising the target in an iterative loop. Where A2 fails, where we can’t verify that the target took, is exactly where safety takes over as our bridge solution. In the next section, we’ll cover work that reaches further and helps us understand how our findings can generalize outside of the distributions they were discovered within.

Reading A2 work

Recognition Cue: A2 work asks whether something (a behavior, value, or disposition) is actually in the model: did it take? It either offers a general instrument check (evals, auditing agents) and/or finds out directly (most interpretability work).

Good A2 Work:

  • Maps before it probes: characterizes the space of behaviors worth checking, rather than testing one hand-picked behavior

  • Links behavior to mechanism: connects what the model does to where it lives internally, or makes that link possible

  • Feeds back into A1: tells you what to change in the target or encoding, not just pass or fail

Examples:

  • Auditing Language Models for Hidden Objectives (Marks et al., 2025) and AuditBench (2026). About 2/3: strong on open-ended discovery and on linking behavior to mechanism (interp tools vs. black-box ones), weaker on feeding back into A1.

  • Persona Vectors (Chen et al., 2025). 3/3: maps traits, ties each to an activation direction, and flags the training data that causes drift, which feeds straight back into A1.

A3: Generalization (distribution problems)

The question we are interested in A3 is simply: does our work in A1 and A2 generalize to new contexts? We will define generalization and what we mean by “new contexts” here, extending our earlier reference to completeness. We’ll start by clarifying what this means.

There are two distributions worth separating. The first is the one where alignment was trained and verified, which is what we’ve been discussing in A1 and A2. The second is the deployment distribution: everything the model actually encounters in the wild. A3 covers the gap between the two distributions. Measured against the first, we can say:

  1. In-distribution: inputs and contexts close enough to what the model was trained and verified on that they aren’t novel.

  2. Out-of-distribution: anything outside that. In practice, the parts of the deployment distribution that training and verification never reached.

This distribution frame is a slightly more precise articulation of completeness we raised above, and leads to generalization failures in two main ways. The first is that alignment fails more often than capabilities across new contexts (Shah et al.) (1). This is when some learned abstraction for executing a task holds up across distributions, but the goals associated with those tasks are untethered and can break. In essence, capabilities generalize much further than goals, which means competent models can pursue the wrong thing in production, which is far more dangerous than the opposite. The second way generalization fails is when misalignment itself generalizes too well and narrow training spreads the wrong way (2). Some work suggests this happens because the model may learn a persona during narrow fine-tuning processes that is associated with misaligned behavior which causes it to generalize.

Much of the work in red-teaming and jailbreaking models to circumvent safety guardrails involved pushing a model out of distribution using uncommon languages, symbols, questions, or prompting it to use personas that are not well-understood or trained to be robust in sensitive contexts. This is the canonical example of generalization failures of type (1), and shows us our earlier point in A1b that context can only select, never install anything that is not in the weights. It has been argued that jailbreaks can be thought of as mismatched generalization, where inputs are off-distribution for safety training, but in-distribution for pretraining. It is essentially a prompting strategy that selects a persona learned in pre-training that hasn’t received as much safety training as the most common persona (the assistant).

Failures of type (2) look like reward hacking morphing into new behaviors in the wild. Canon here includes Emergent Misalignment (Betley et al.) and MacDiarmid et al.’s reward hacking that turned into sabotage.

While no environment is completely safe from distribution mismatches, environments with a high degree of verifiability have considerably smaller gaps between in-distribution data and out-of-distribution data, mostly because we can use this verifiability to define the gaps. Coding is a good example here because code, by definition, is a highly structured environment, and it’s likely that models learn basic structures that generalize quite well across languages and context. In essence, the world is well-defined and verifiable. However, this verifiability can work as a disadvantage when the verifier becomes the target, and we see emergent reward hacking. This recalls our earlier point in A1b about metrics getting Goodhart-ed, as well as our definability point in the A2s behavior table.

As a foil to this, we can look at creative writing and role-play behaviors. Clear verifiability for something as subjective as “creative writing” will likely always continue to elude us, and therefore is often the medium used in jailbreaking attempts (persona modulation raised GPT-4’s harmful completion rate from 0.23% to 42.5%).

The essential take away about generalizability is that for domains that are highly verifiable, the distribution gap is smaller, easier to detect, and the path to increasing safety is generally clearer. We can hill-climb. In less verifiable domains, or in domains where agents are deployed in different configurations than they’ve seen in training, generalizability failures show up more often.

One example of this is in the agent context. Deployment is moving off-distribution by construction, since every long agentic trajectory is likely a new context. Long horizons, tool use, delegation to subagents, copies, multi-agent settings and complex prompting strategies across these architectures vary considerably, and are therefore almost impossible to train on effectively. This is why recent swarm echoes have been the root of so many alignment failures as of late. We’ve already discussed the METR report, but there is more work to explore here (see: Agentic Misalignment). This connects directly to our discussion on safety: speed, scale, and copies take models far outside distribution.

Recall that observer effects refer to the model’s ability to detect and understand when it is being observed to change its behavior, more closely modeling what it thinks it is supposed to do. Generalization is obviously complicated by this problem because it introduces yet another distribution to handle: evaluation. With observer effects, one has to account for both training and deployment environments, but also evaluation behavior across both, increasing our total distribution count to a 2x2 matrix (trained vs. deployed, observed vs. unobserved). It’s not always the case that model eval awareness changes behavior, but even if it increases the risk of bad behavior slightly, this complicates our ability to ensure that our verification work in A2 has been complete. (see Needham et al. and Abdelnabi & Salem).

We necessarily cannot cover the full range of possible inputs while training, and these models are probabilistic anyway. Even if a simple unsafe question is asked many times in training, there is always a risk that in the wild, random sampling from output distribution might send models towards a rare unsafe trajectory regardless of if the context is fully in-distribution. This relates to our earlier point on risk assessment in A2.

If our alignment has to hold across contexts, it also has to hold across model versions and architectures, which leads us into A4.

Reading A3 work

Recognition Cue: A3 work asks whether what took in A2 still holds when the context shifts: new inputs, new deployment configurations, agentic settings, or evaluation vs. deployment. It either finds generalization failures (in either direction) or builds ways to predict or close the distribution gap.

Good A3 Work:

  • Names the shift: says which distribution moved (inputs, deployment configuration, evaluation to deployment) rather than gesturing at “out-of-distribution”

  • Checks both directions: tests whether alignment fails to carry over, and whether misalignment spreads

  • Tests under realistic deployment conditions: agentic, long-horizon, or multi-agent settings, not just chat

Examples:

  • Natural Emergent Misalignment from Reward Hacking in Production RL (MacDiarmid et al., 2025). 3/3: names the shift (chat to agentic), catches both directions (safety training failed to carry over, and hacking spread into sabotage), and uses production RL environments.

  • Jailbroken (Wei et al., 2023). About 2/3: names the shift precisely (off-distribution for safety training, in-distribution for pretraining) with realistic attacks, but covers only the failure-to-carry direction.

  • Emergent Misalignment (Betley et al., 2025). 2/3: the canonical case of misalignment spreading, with a clear shift (narrow fine-tune to broad behavior), but evaluated mostly in chat settings.

A4: Iterability, Cost, Feasibility

And finally, we’ve arrived at the final piece of the alignment framework. If A1 asks “what should we encode, and how?”, A2 asks “did it work?”, A3 asks “how does it generalize?”, then A4 asks “does it scale?”. By scale here, we mean “can we keep running the A1↔A2 loop on the next model, affordably and outside of a single lab”.

This is where most of my beliefs are the strongest. Interpretability only matters if A4 work develops such that we’re not constantly playing catch up with new models and architectures. Automated interp will be massive leverage.

Much of the work that hits previous sections is computationally expensive, or narrowly scoped, and work on A4 is about ensuring that with each new model release, we can apply an instrument, or finding to models and get the same alignment guarantees that we’d developed on previous versions.

One-model results can be impactful, but they aren’t useful unless findings hold across versions and model families.

There are two natural slices here:

  1. Auto-Research: Automate the A1-A2 loop, feasibly

  2. Model Translation: Translate findings across model versions and architectures

As we’re still relatively early on turning alignment work into a hard science with well-understood frameworks for research, much of the work is still done by humans. We come up with the ideas, experiments, things to test, and have to be in the loop for building most of the infrastructure that supports each experiment. Anyone who has tried to do interpretability, for example, has certainly run into a variety of issues both infrastructurally and with models as assistants.

As the industry races towards recursive self-improvement (RSI), auto-research has become increasingly popular, and this section is about just extending that to alignment, as opposed to capabilities, which are often the target of RSI developments. If we can get this for alignment, then our control-as-bridge idea is finally put to rest, as control becomes significantly less important (AI Control).

The model translation route has shown comparatively more progress recently. The refusal direction, a common research target, has shown up across 13 open models; crosscoders and model diffing are used extensively to compare models; Weak-to-Strong Generalization (Burns et al.) shows transfer across a capability gap. Furthermore, there is evidence that supports only a few hundred benign samples could undo emergent misalignment (Wang et al.). These examples show promise that translating findings, or instruments at the least, is possible, which is consequential for all alignment work. However, findings that transfer often imply attacks transfer, which increases the difficulty with respect to open models.

The path forward here is automated alignment research. In Sharkey’s recent open problems in interpretability paper, we see automated probe training still an open question. Recall our earlier paragraph on automation in A2:

“A lot of the most exciting work in this domain, to me, covers automating interpretability and building more complex evaluation instrumentation to give more fodder to those looking to discover behavioral origins in the weights (AuditBench’s auditing agent, Petri, Bloom).”

This is what A4 research covers, at a high-level, and would feed the A1-A2 loop considerably.

Though the open problems in the Sharkey paper are A2 dominant with A4 as a secondary category there is considerable weight for our purposes in this section. Anthropic’s Automated Weak-to-Strong Researcher is a more clear direct A4 example with respect to automated research.

However, there remain significant constraints to automating this work:

  1. Resources: compute & storage

    1. Bigger models may be easier to interpret, but are more expensive to study, and have larger activation data demands

      1. For a 70B-class model (hidden size ~8,192, ~80 layers), residual-stream activations at bf16 are about 1.3 MB per token across all layers, so a billion tokens of activation data is on the order of a petabyte

  2. Human Bottleneck:

    1. Models are still pretty bad at this type of work, we need human research taste

    2. Recall METR’s unreliable, possibly deceptive analysis agents

  3. Pre-Deployment Access:

    1. Independent evaluators, like METR, need access to frontier models and significant budgets to properly evaluate (METR’s investigation of the OAI-HF incident used roughly $400K in API credits over six days, which OpenAI provided)

    2. The A1-A2 loop is difficult due to market pressure (release cycles), resource constraints (compute, data, expertise), and infrastructure (interpretability infra is still weak)

  4. Open-weight Challenges:

    1. Releasing open-weight models increases the necessary burden on A1-A2 working exceptionally, and making models resilient to ablating alignment

We can address open-source difficulties from earlier more thoroughly. On one hand, tamper-resistance with respect to open models is weak. Training-free attacks still consistently work to jailbreak models, and things get even bleaker if the party attempting to break alignment guardrails is sophisticated enough to run complex training that either fine-tunes or ablates key activations/parameters. As we’ve discussed, the data supports the fact that restoring alignment is feasible, but it’s not clear how that would dissuade or prevent a motivated bad actor in the first place. Open weights are one of the two exceptions where misuse sits outside of value alignment, assuming we can’t solve the problem of training ablation resistance.

Furthermore, interpretability may be particularly challenged on this front because there’s some evidence that scale and sparsity help interpretability, but the open-source community lacks the models and compute to test the work on frontier rails. Larger models seem to tolerate more activation sparsity (Universal Properties of Activation Sparsity), and models trained with sparse weights have more interpretable circuits (Weight-sparse transformers), but it’s not clear that this implies dense large models are easier to interpret. Bigger models need more compute, bigger SAEs, and have a higher data requirement, which all present challenges to doing good interp.

A4 is rarely a primary category of alignment research, partially because we’re still early in alignment science, and specific findings are often tested across a few different models, which gets them an A4 subtag for free.

Research that has A4 as a primary tag mostly involves point 1 on auto-research instrumentation, or provides a more observational study across findings. An illustrative example is the recent Matryoshka Attribution paper, which develops a new tactic for training activation and parameter masks for specific goals, like refusal. The reason A4 is a candidate primary tag for this work is that it is a generalizable instrument that can be trained, and the process could theoretically be applied to every new model for study.

Reading A4 work

Recognition Cue: A4 work asks whether the A1↔A2 loop can keep running on the next model: whether findings or instruments transfer across versions and architectures, whether the loop can be automated, and whether it’s cheap and accessible enough to run before deployment, including outside the lab that trained the model.

Good A4 Work:

  • Transfers: shows a finding or instrument holding across model versions, families, or scales, not just one model

  • Lowers the cost of the loop: makes a step cheaper in compute, storage, or human time, ideally by automating it

  • Widens access: runnable by third parties or on open models, pre-deployment, not only inside the lab that trained the model

Examples:

  • Refusal in Language Models Is Mediated by a Single Direction (Arditi et al., 2024). A2 primary, strong A4 secondary. 3/3 on A4: the same direction shows up across 13 open chat models, it’s cheap to find (a difference in mean activations), and anyone with open weights can run it.

  • Scaling and Evaluating Sparse Autoencoders (Gao et al., 2024). About 2/3: scaling laws make the cost of the interp loop predictable, and the method holds from GPT-2 up to GPT-4, but the frontier-scale SAEs stayed inside the lab.

  • Matryoshka Attribution (2026). 2/3: a trainable instrument that can be re-run on each new model, and cheaper than manual attribution; access not shown.

Alignment - Complete Framework

We’ll end with an overview of the entire alignment framework we’ve been developing.

Stage

Question

It’s this stage if…

Good work…

A1

What do we want, and how do we encode it?

…proposes or revises the target, or changes how it’s codified

Explicit target · accounts for the encoding gap · says how you’d check it

A2

Did it take?

…checks whether something is in the model

Maps before probing · links behavior to mechanism · feeds back to A1

A3

Does it hold when the context shifts?

…finds or closes a distribution gap

Names the shift · checks both directions · realistic deployment conditions

A4

Can the loop keep running?

…transfers, automates, or widens access

Transfers · lowers cost · widens access

The core loop is between A1 ⇄ A2: defining and codifying what we want and testing it took, exploring externalities, then revising what we’d optimize with what we learn. Extending beyond the training environment brings us into A3 research where we want to understand how the work in the core loop generalizes across domains. Finally, we need to ensure that the tactics that worked in A1-A3 extend across architectures and model releases to iteratively improve our chances at full-scale alignment.

To use the framework, classify research against the question it answers, then grade it against that row’s criteria. Most work will span stages, so it’s helpful to name the primary context and note the others.

Later on, we’ll discuss how to weight each of these components generally and how we can use the framework to update our priors. A coherent single page card on the following page should be helpful.

Alignment Card:

Useful Definitions

  • Broad vs. Narrow: Narrow alignment asks whether an AI system follows the immediate intent or direct instructions given by its operator. Broad alignment asks whether it reliably produces good real-world outcomes and captures human values, ethics, and preferences.

  • Inner vs. Outer: Outer alignment is picking the right task to optimize. Inner alignment is making sure the model is also optimizing for that chosen task.

  • Intent vs. Values: Intent alignment is outer plus inner. Broad alignment is intent alignment where the intent itself is good: value specification, plus outer, plus inner.

Main Goal: An AI system is aligned when it reliably pursues good real-world outcomes consistent with human values, understanding its operator’s intent in context rather than only executing it.

Categories of Alignment Work

  • A1: Target (Outer Alignment). “What should we encode, and how?”

    • Type of work: Defining what we actually want and translating that into a precisely defined base objective. A1a: what we want (Constitutional AI, ICMI). A1b: codifying it (data curation, supervised fine-tuning, RL training).

    • Difficulties: “Whose values?” is a relevant question for any target. Every value system gets compressed into one number per trajectory, and optimizing hard against a proxy eventually hurts the thing it stands in for.

    • Good work:

      • Makes the target explicit and contestable

      • Accounts for the encoding gap

      • Gives methodology for how you’d check, so it feeds A2

    • Examples: Deliberative Alignment (3/3) · Constitutional AI (~2/3) · Reward Model Overoptimization (2/3)

  • A2: Did it Take? (Inner Alignment & Verification). “Did it work?”

    • Type of work: Whether the objective we’ve encoded has actually been successfully trained into the model. Outside-in: evals and benchmarking. Inside-out: interpretability.

    • Difficulties: Interpretability is bottlenecked by behavioral mapping. Completeness: absence of evidence is not evidence of absence. Observer effects: the measurement changes the behavior being measured. Instrument validity: an eval that doesn’t actually measure the behavior you’re looking for.

    • Good work:

      • Maps before it probes

      • Links behavior to mechanism

      • Feeds back into A1

    • Examples: Hidden Objectives + AuditBench (~2/3) · Persona Vectors (3/3)

  • A3: Generalization. “How does it generalize?”

    • Type of work: Does our work in A1 and A2 generalize to new contexts? A3 covers the gap between the distribution where alignment was trained and verified and the deployment distribution.

    • Difficulties: Capabilities generalize much further than goals, and misalignment can generalize too well. Deployment is moving off-distribution by construction, and observer effects add evaluation as another distribution (trained vs. deployed, observed vs. unobserved).

    • Good work:

      • Names the shift

      • Checks both directions

      • Tests under realistic deployment conditions

    • Examples: Reward Hacking in Production RL (3/3) · Jailbroken (~2/3) · Emergent Misalignment (2/3)

  • A4: Iterability, Cost, Feasibility. “Does it scale?”

    • Type of work: Can we keep running the A1↔A2 loop on the next model, affordably and outside of a single lab? Auto-research: automate the A1-A2 loop, feasibly. Model translation: translate findings across model versions and architectures.

    • Difficulties: Resources (compute & storage), the human bottleneck, pre-deployment access, and open-weight challenges.

    • Good work:

      • Transfers

      • Lowers the cost of the loop

      • Widens access

    • Examples: Refusal Direction (3/3) · Scaling SAEs (~2/3) · Matryoshka Attribution (A4 primary)

Classifying: “To use the framework, classify research against the question it answers, then grade it against that row’s criteria. Most work will span stages, so it’s helpful to name the primary context and note the others.”

Alignment Survey

Grades apply each stage’s good-work criteria, assessed by Claude.

A1: Target (Outer Alignment)

  • Deliberative Alignment (Guan et al., 2024)

    • 3/3: makes the spec explicit, keeps it legible through training by teaching the model to reason over it, and gives a built-in check via chain of thought.

    • Also: A2

  • Claude’s Constitution (Anthropic, 2026)

    • 2/3: explicit and contestable, and explains why rather than only what to help it survive training; no methodology for checking it took.

  • Collective Constitutional AI (Anthropic & CIP, 2024)

    • 2/3: makes the target explicit and contestable through public input, then trains and evaluates a model against it; doesn’t address the encoding gap.

  • Constitutional AI (Bai et al., 2022)

    • 2/3: explicit, contestable principles with human-comparison evals; the principles still get compressed into preference labels and a reward model.

  • Scaling Laws for Reward Model Overoptimization (Gao, Schulman & Hilton, 2022)

    • 2/3: measures the encoding gap directly, and its gold-reward-model setup doubles as a check; proposes no target.

    • Also: A2

  • Training Language Models to Follow Instructions with Human Feedback (Ouyang et al., 2022)

    • 2/3: documents where the trained model diverges from intent and checks against held-out human preferences; the target stays implicit in labeler judgments.

  • Eschatological Corrigibility (ICMI, 2026)

    • 2/3: an explicit target (religious framing) with a behavioral check (shutdown resistance); doesn’t address the encoding gap.

    • Also: A2

  • Reinforcement Learning from Christian Feedback (ICMI, 2026)

    • 1/3: an explicit theological target codified into a GRPO objective; doesn’t address the encoding gap or how to check it took.

  • The Persona Selection Model (Anthropic, 2026)

    • 1/3: reframes the encoding gap by arguing post-training selects among pretrained personas; proposes no target or check.

    • Also: A2

  • GospelVec (ICMI, 2026)

    • 1/3: an explicit theological target codified at the activation layer; no encoding-gap analysis or check.

    • Also: A2

  • OpenAI Model Spec (OpenAI)

    • 1/3: explicit and contestable; silent on encoding and checking.

  • Deep Reinforcement Learning from Human Preferences (Christiano et al., 2017)

    • 1/3: introduces preference-based reward learning and checks against human judgments; the target stays implicit in the labels.

  • Artificial Intelligence, Values, and Alignment (Gabriel, 2020)

    • 1/3: makes the “whose values” question explicit and contestable; normative, so no encoding or check.

A2: Did it Take? (Inner Alignment & Verification)

A3: Generalization

A4: Iterability, Cost, Feasibility

Safety Framework

Let’s recall our definition of safety:

Minimizing the chance and the consequences of harm from AI systems, whether that harm comes from misunderstanding, misalignment, or misuse.

Our primary goal in our safety framework, which will be considerably more concise than the alignment framework, is to chart tactics and findings that the alignment framework doesn’t subsume. Remember, at a high-level, we’re considering safety to be a superset of alignment.

So, back to our three problem axes: misalignment, misuse, and structural. We’ve covered misalignment, now we need to cover misuse, as structural issues will remain out-of-scope, for now. To do that, we’ll first start by discussing what we mean by misuse, and specifically the type of misuse that isn’t covered by alignment. Then, we’ll cover these issues from a technical vantage that includes control (S1), monitoring (S2), and containment (S3).

One disclaimer: while in a sense all of AI safety and alignment is about control, and we can debate whether “containment is a form of control”, it’s useful to think about them distinctly to understand research.

Structural issues (societal resilience) remain out-of-scope for this piece, but include governance, policy, and organizational restructuring to minimize the harmful blow-back from AI deployed across human domains.

Prelude

As we’ve discussed before, if our target is broad alignment, then misuse of the models is mostly an alignment problem. A perfectly aligned model should not admit anything that a reasonable human would deem “misuse”. They would have the contextual understanding to recognize what is right and wrong, and be able to filter requests accordingly.

However, a few important exceptions may slip through the cracks of this framing:

  1. Task Decomposition: It’s possible that harmful tasks can be decomposed into benign subtasks that evade alignment – S2 covers this

  2. Open weights: It’s possible that alignment can always be fine-tuned or ablated out, and this is impossible to solve – S3 covers this

  3. Jailbreaks: It’s possible prompts may always have the chance to push a model off-distribution into an under-safety-trained persona that allows for harmful use – S1/S2 cover this

I argue that alignment has to work if we’re moving towards ASI. Control and monitoring both rely on overseers (human or otherwise) who can still catch the model, which become far less powerful if they are less intelligent than the entity they are overseeing. However, both control and monitoring are critical bridge infrastructures we need right now. Furthermore, containment, while possibly viable as a safety strategy in an ASI world, is a backstop, not a substitute, for alignment. It’s important for any future in which we have 99% aligned ASI, where strong containment protocols might carry that last 1%.

S1: Control

Control as a field is about enforcing intent and steering towards good outcomes from an AI system under the assumption that we can’t fully trust that alignment lives in the weights (AI Control, Greenblatt et al.). Tactics for control exist on a spectrum of steering from prompting to activation manipulation and ablation that range from least to most invasive. A human analogy might be a conversation on one side, as a way to change behavior, to a strict regimen of checks and balances in the middle, and outright drugging someone on the far end.

In a general context, control can be used for a variety of non-safety goals, but in our context, control is best understood as protective tactics against the possibility of adversarial behavior and misalignment. The causes of this adversarial behavior don’t assume genuine ill-intent, just that somewhere in the training process alignment has failed. This could be a result of scheming, reward hacking, or some generalization quirk. We’re essentially assuming that A2 work may have failed.

Recall A1b here on codifying values. Prompting selects among trajectories already in the weights, so it can be useful in steering a cooperative model, but is unlikely to work against an adversarial one. Here is an example of a prompt from OpenAIs AI-assisted work on the cycle double cover conjecture that shows a considerable amount of control. One issue with this type of control as a strict safety measure is that it’s not clear what the externalities are going to be, and specifying an intent precisely enough is hard even with a cooperative model.

For control with respect to safety specifically, the most prevalent tactics are using trusted monitoring combined with a protocol, resampling (Ctrl-Z), and input and output classifiers (Constitutional Classifiers). Activation steering is less used in practice because reliable steering directions are hard to define precisely, and the consequences of steering are still not well-understood in many cases.

One of the major bottlenecks to control research at scale is that it is reliant on really good red-teaming, which brings us back to our completeness problem from earlier. Red-teaming evals need to be really good to ensure control measures are working. Red-teaming uses worst case scenarios and overtly adversarial contexts to ensure that the floor is understood.

Control is tricky in part because one of the big gains we get from LLMs is their ability to autonomously problem solve. By overspecifying or over-steering through control measures, we may prevent them from operating at full-capacity, trading usefulness and ease of use for predictability. There is also the problem of capability gaps between trusted monitors in control protocols and the untrusted model being observed. If that gap grows too wide, the untrusted subject will trivially beat the control protocols.

Some of these strategies, especially classifiers, are actively deployed in production models today, and they have considerable impact in safety as a result. Jailbreaking rates have dropped, partially because of model improvements and resilience in the weights, but also because control mechanisms and classifiers have gotten better.

Reading S1 work

Recognition Cue: S1 work enforces intent on a model whose weights we can't fully trust, using tactics anywhere on the spectrum from prompting to activation steering. It builds or evaluates protocols (trusted monitoring, resampling, classifiers) that constrain what an untrusted model can do.

Good S1 Work:

  • Assumes an adversary: tested against a red team playing the scheming model or the attacker

  • Reports capability cost: says how much capability or autonomy the protocol gives up

  • Accounts for the capability gap: says whether the protocol holds as the untrusted model outgrows its trusted monitor

Examples:

  • AI Control (Greenblatt et al., 2023). 3/3: introduces control evaluations with a red team, reports a safety-usefulness tradeoff, and is built around a weaker trusted model overseeing a stronger untrusted one.

  • Ctrl-Z (Bhatt et al., 2025). 2/3: resampling protocols red-teamed in an agentic setting, with usefulness reported; the capability gap isn’t studied directly.

  • Constitutional Classifiers (Sharma et al., 2025). 2/3: large-scale human red-teaming, with over-refusal and compute overhead reported; aimed at human jailbreakers rather than a scheming model.

S2: Monitoring

We touched on monitoring briefly in the control section, because it is a prerequisite for control, but it can be used for much more, and is one of the more important research fields with respect to safety. In a sense, monitoring is to safety what A2 work is to alignment, where we can directly ask “what is really happening” and pick apart the whys and hows. It’s critical to everything in this framework, for that reason, but is difficult enough to be treated as its own research subsection.

Monitorability refers to whether, and how, we can observe what the model is doing and why at runtime. There are three critical areas that this is useful for: oversight, forensics, and risk assessment. Oversight has the most direct overlap with control, as the objective of oversight is to flag, block, audit, or resample outputs based on what monitors detect.

Forensics is much more interesting. We need to be able to reconstruct what happened after major incidents, like the OAI-HF breach, so we can learn from mistakes effectively. Poor monitoring significantly complicates this work, and it’s even mentioned a few times in the report that it was somewhat handicapped by OpenAI’s protection of transcript data outside of the specific investigation window.

Finally, risk assessment is going to be an increasingly important research field as models continue to evolve, and deep monitoring data can inform rates of rare behavior to actuarially back pricing. This is still exceptionally difficult to do, as we’ve discussed earlier, but more well-formed data from monitoring will improve these processes.

Monitoring itself can happen in a few different ways. It can happen at training time or inference time (see Goodfire silico), and one can observe inputs, outputs, activations, and CoT streams.

There is also another level to monitoring, which is usage monitoring across sessions, that is one degree removed from immediate monitoring to include a broader coverage of patterns over time. This is the aggregate view that could detect task decomposition and other behaviors that are difficult to detect in single streams.

Some of the bigger constraints here are that CoT is not guaranteed to be faithful to what’s actually going on in the model’s neurons, data demands for monitoring and processing can become enormous, training against a monitor can break it, and that processing raw monitor data into useable signal might be too slow to allow for intervention time. Data retention is also in tension with privacy: forensics and risk assessment both need stored data, and the METR investigation shows how protected transcripts can limit what we learn after an incident. Recent rumors in the space have suggested that frontier labs are training models using looped transformers, which is dangerous from a monitorability standpoint. If more and more cognitive processing is happening internally via recursion, we are likely to see an impact on CoT reasoning traces. Given we don’t yet have a precise science of activation monitoring to supplant CoT monitoring, this could create a dangerous gap. Models may begin to improve much quicker than we can keep up in understanding their behavior.

Reading S2 work

Recognition Cue: S2 work makes model behavior observable in real time (oversight), after the fact (forensics), or in aggregate (risk assessment).

Good S2 Work:

  • Compares monitor to behavior: what the monitor sees reflects what the model is actually doing

  • Holds up under optimization pressure: the signal survives training, including training against the monitor

  • Usable in practice: turns raw signal into something usable in time to intervene

Examples:

S3: Containment

Containment is the final layer of safety that simply limits what the model can do, without trying to control it directly. One useful framework for AI systems generally is to make a distinction between intellect and actuators.

Intellect is the raw output stream of a model, which is just a string of tokens from the model’s trained vocabulary. All an LLM produces, ultimately, are these tokens, and only when they are digested by code and turned into tool calls do we see that intellect takes action and leaves an imprint on the world.

Containment is about how we can contain both the internal weights that produce this intellect, and the intellect’s ability to connect to actuator surfaces. We can think about the former problem as “inbound” containment (world -> model: who can reach the weights and the model), and the latter problem as “outbound” containment (model -> world: what actuators the model can reach).

Within both types of containment there are effectively two tactical strategies:

  • Digital Containment: Sandboxing, tool restrictions, etc

  • Physical Containment: Hardware and facility isolation (air-gapped networks, controlled facilities, physical access control)

Permission systems exist within both strategies.

The OAI-HF incident is the most public version of an outbound digital containment failure (sandbox boundary). It’s also a major alignment failure, but as discussed earlier, containment exists to protect us from the parts of alignment we didn’t get 100% right.

The most complicating factor is that agents need actuators to be useful, or we’re simply reduced to conversing with an inert AI that has to have all relevant knowledge stored in its weights, which is infeasible. This means we have to keep refining containment and the permissions within it if these models are going to impact the world in the ways we’d like them to.

Importantly, strong containment architectures are critical for testing, and our earlier A1 to A2 loop is only safe to run on potentially dangerous models if we have strong containment guarantees. In a perfect world, we’d have extremely robust protocols for testing, where we aren’t sure that the model is fully aligned, allowing us to test completely and effectively without a potentially misaligned model escaping and wreaking havoc.

Another complexity to containment is weight security, which spans both threat directions. We have discussed how it’s often possible to fine-tune or ablate alignment/safety guardrails out of models, which we need the weights for, so it’s important that we have proper security around them such that this is extremely difficult to do (inbound). Furthermore, we need to be concerned about rogue agents and self-exfiltration where the models copy their own weights into new infrastructure that allows them to modify their own internals or simply continue living on outside of ancillary control infrastructure (outbound).

The constraints for containment work therefore include:

  1. Usefulness: can’t just get rid of tools and actuators

  2. Insider risk: humans with access are a weak point

  3. Cost: high security is expensive

  4. No recall: once weights are out, they’re out

  5. Sandbox escapes: through poor design or hacking

As mentioned, containment is likely exceptionally important no matter what future we’re heading to. More on this later in the beliefs and trajectories section.

Reading S3 work

Recognition Cue: S3 work asks how well we can limit what a model can reach (outbound: actuators, tools, network) and who or what can reach the model (inbound: weights, access), using digital or physical tactics.

Good S3 Work:

  • Clarifies direction: says whether it contains against the model getting out, or humans getting in

  • States capability cost: says what tools, actuators, or access the model gives up

  • Tests against adversary: holds in strong adversarial environments

Examples:

Safety - Complete Framework

We’ll end with an overview of the entire safety framework we’ve been developing.

Stage

It’s this stage if…

Good work…

S1

…enforces intent on a model whose weights we can’t fully trust

Assumes an adversary · reports capability cost · accounts for the capability gap

S2

…makes model behavior observable in real time, after the fact, or in aggregate

Compares monitor to behavior · holds up under optimization pressure · usable in practice

S3

…limits what the model can reach, or who can reach it

Clarifies direction · states capability cost · tests against adversary

Control as a field is about enforcing intent and steering towards good outcomes from an AI system under the assumption that we can’t fully trust that alignment lives in the weights. Monitorability refers to whether, and how, we can observe what the model is doing and why at runtime. There are three critical areas that this is useful for: oversight, forensics, and risk assessment. Containment is the final layer of safety that simply limits what the model can do, without trying to control it directly.

To use the framework, classify research against the question it answers, then grade it against that row’s criteria. Most work will span stages, so it’s helpful to name the primary context and note the others.

A coherent single page card on the following page should be helpful.

Safety - Card

Useful Definitions

  • Misuse: Humans wanting the wrong things. If our target is broad alignment, then misuse of the models is mostly an alignment problem; the exceptions are task decomposition (S2), open weights (S3), and jailbreaks (S1/S2).

  • Intellect vs. Actuators: Intellect is the raw output stream of a model; only when it is digested by code and turned into tool calls does that intellect take action and leave an imprint on the world.

  • Inbound vs. Outbound: Inbound containment is world -> model: who can reach the weights and the model. Outbound containment is model -> world: what actuators the model can reach.

Main Goal: Minimizing the chance and the consequences of harm from AI systems, whether that harm comes from misunderstanding, misalignment, or misuse.

Categories of Safety Work

  • S1: Control.

    • Type of work: Enforcing intent and steering towards good outcomes under the assumption that we can’t fully trust that alignment lives in the weights, with tactics ranging from prompting to activation manipulation and ablation.

    • Difficulties: Reliant on really good red-teaming; trading usefulness and ease of use for predictability; capability gaps between trusted monitors and the untrusted model being observed.

    • Good work:

      • Assumes an adversary

      • Reports capability cost

      • Accounts for the capability gap

    • Examples: AI Control (3/3) · Ctrl-Z (2/3) · Constitutional Classifiers (2/3)

  • S2: Monitoring.

    • Type of work: Whether, and how, we can observe what the model is doing and why at runtime, for oversight, forensics, and risk assessment: inputs, outputs, activations, CoT streams, and usage monitoring across sessions.

    • Difficulties: CoT is not guaranteed to be faithful, data demands can become enormous, training against a monitor can break it, and processing raw monitor data into usable signal might be too slow to allow for intervention; data retention is also in tension with privacy.

    • Good work:

      • Compares monitor to behavior

      • Holds up under optimization pressure

      • Usable in practice

    • Examples: Monitoring Reasoning Models (2/3) · Sleeper-Agent Probes (2/3) · CoT Monitorability (1/3)

  • S3: Containment.

    • Type of work: Limits what the model can do, without trying to control it directly: containing both the internal weights that produce the intellect (inbound) and the intellect’s ability to connect to actuator surfaces (outbound), through digital containment (sandboxing, tool restrictions, permissions) and physical containment (hardware and facility isolation).

    • Difficulties: Usefulness, insider risk, cost, no recall once weights are out, and sandbox escapes.

    • Good work:

      • Clarifies direction

      • States capability cost

      • Tests against adversary

    • Examples: Securing AI Model Weights (2/3) · OAI-HF Investigation (2/3) · In-context Scheming (2/3)

Classifying: “To use the framework, classify research against the question it answers, then grade it against that row’s criteria. Most work will span stages, so it’s helpful to name the primary context and note the others.”

Safety Survey

Grades apply each stage’s good-work criteria, assessed by Claude.

S1: Control

  • AI Control (Greenblatt et al., 2023)

    • 3/3: red-teamed with attack policies, reports a safety-usefulness tradeoff, and is built around a weaker trusted model overseeing a stronger untrusted one.

  • SHADE-Arena (Kutasov et al., 2025)

    • 2/3: agents pursue hidden sabotage tasks against monitors, and it measures how monitor strength compares to agent strength; no capability cost reported.

    • Also: S2

  • Ctrl-Z (Bhatt et al., 2025)

    • 2/3: resampling protocols red-teamed in an agentic setting, with usefulness reported; the capability gap isn’t studied directly.

  • Constitutional Classifiers (Sharma et al., 2025)

    • 2/3: large-scale human red-teaming, with over-refusal and compute overhead reported; aimed at jailbreakers rather than a misaligned model.

    • Also: A3

S2: Monitoring

  • Monitoring Monitorability (Guan et al., 2025)

    • 3/3: measures whether monitors catch the behavior across a broad eval suite, catches models trained to obfuscate and finds RL optimization doesn’t materially reduce monitorability, and shows scaling a weak monitor’s compute helps against a strong agent.

  • Monitoring Reasoning Models for Misbehavior (Baker et al., 2025)

    • 2/3: shows CoT monitors catch reward hacking, and that training against them produces obfuscated reasoning; no practical-use result.

    • Also: A1

  • Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings (Duzan & Stickland, 2026)

    • 2/3: compares monitor detection to real behavior shifts, and shows realistic deployment prompts can cut detection to as low as 5%; no optimization pressure.

  • Reasoning Models Don’t Always Say What They Think (Chen et al., 2025)

    • 2/3: measures whether CoT reflects what actually influenced the answer, and tracks how outcome-based RL changes that; no practical-use result.

    • Also: A2

  • Simple Probes Can Catch Sleeper Agents (Anthropic, 2024)

    • 2/3: linear probes detect defection reliably and are cheap enough to run at inference; not tested under optimization pressure.

    • Also: A2

  • Clio (Tamkin et al., 2024)

    • 1/3: privacy-preserving usage monitoring across sessions at scale; doesn’t validate against model behavior or test optimization pressure.

  • Chain of Thought Monitorability (Korbak et al., 2025)

    • 1/3: frames why the signal is fragile under optimization pressure; a position paper, so no comparison or practical-use result.

S3: Containment