Science of Frontier AI Risk Evaluation and Discovery
A Research Agenda from Nuwa Frontier AI Safety Lab
Frontier AI safety needs a science of risk evidence: a way to discover dangerous capabilities before they appear in the real world, and to report them with enough precision to support serious governance decisions.

The most serious risks from advanced AI systems are no longer fully captured by static model tests. A frontier model today is rarely deployed as a bare language model. It is increasingly embedded in an agent scaffold: tools, memory, code execution, browser access, retrieval, planning loops, feedback, credentials, and long-horizon task execution. Once a model becomes part of an acting system, the central safety question changes.
We should no longer ask only what the model can say. We need to ask what the system can accomplish.
This shift is already visible in frontier AI governance. The IDAIS Beijing statement proposed red lines for unacceptable AI risks, including autonomous replication or improvement, power-seeking behavior, autonomous cyberattacks, assistance with weapons development, and deception that causes designers or regulators to misunderstand whether those red lines may be crossed (1). Frontier AI labs and safety organizations have also begun to define capability thresholds, risk thresholds, evaluation procedures, and mitigation commitments through frontier safety frameworks and responsible scaling policies (2)(3)(4)(5).
Yet a core scientific gap remains. A red line is a governance concept. A safety evaluation is an experimental claim. To connect the two, we need a reporting method that makes frontier AI risk evidence comparable, reproducible, and decision-relevant.
This research agenda proposes such a method:
model + scaffold + environment + task + budget → observed capability → implied risk
This formula should become a basic unit of frontier AI risk reporting.
It forces us to specify what was actually evaluated. Was it a base model, a deployed product, or an agent system? What scaffold was used? What tools and memory were available? Was the environment realistic? How hard was the task? How many attempts were allowed? How much compute, time, and external information did the system receive? What capability was observed, and what real-world risk does that capability imply?
Without this level of detail, red-line evaluation can become misleading. A weak scaffold may fail to elicit a dangerous capability. A short-horizon benchmark may miss a long-horizon risk. A toy environment may understate real-world transfer. A single success claim may overstate risk if the task is too artificial. A single failure claim may create false reassurance if the evaluation conditions are too weak.
Nuwa Frontier AI Safety Lab will focus on making frontier AI risk evaluation more operational, especially for two meta-risks: autonomy and deception.
Autonomy and Deception as Meta-Risks
Autonomy and deception should be treated as meta-risks because they amplify many other dangerous capabilities.
A model with cyber capability is concerning. An autonomous agent with cyber capability, tool use, memory, persistence, and failure recovery is more concerning. If that agent can acquire resources, replicate, improve its scaffold, or resist interruption, the risk becomes dynamic. The system’s future capability is no longer fixed by its initial deployment state.
Autonomy determines whether a dangerous capability can propagate through the world.
Deception determines whether developers, evaluators, regulators, and users can reliably observe that capability before harm occurs.
Together, they define a central control problem for frontier AI:
Can an AI system expand its own operational capacity while reducing the reliability of human oversight?
This is why autonomy and deception deserve a special place in frontier risk evaluation. They are not just two more benchmark categories. They are structural risk multipliers. Autonomy expands what the system can do. Deception reduces our ability to know what it can do.
1. Autonomy: When Capability Becomes Propagation
Autonomy changes the meaning of dangerous capability.
A harmful answer is a static output. An autonomous action is part of a trajectory. An agent can plan, execute, observe failure, revise strategy, use tools, coordinate steps, and continue across time. This makes autonomy risk fundamentally different from ordinary content safety.
DeepMind’s dangerous capability evaluation program already includes cyber-security, self-proliferation, persuasion and deception, and self-reasoning as major evaluation areas (6). OpenAI’s Preparedness Framework includes tracked categories such as cybersecurity, biological threats, AI self-improvement, and autonomy-related risk (4). The UK AI Security Institute also emphasizes that tool use, agentic systems, scaffolding, and evaluation budget are increasingly central to understanding frontier capability trends (7).
Our autonomy agenda focuses on five connected mechanisms.
Self-replication
Self-replication is an early signal that an AI system may extend its own operational presence without direct human assistance.
In our work, LLM-powered agent systems were evaluated on whether they could create live and separate copies of themselves under controlled conditions (8). Apollo Research has studied related agentic risks through in-context scheming, including cases where models attempted to exfiltrate what they believed to be their own weights under evaluation pressure (9). Palisade Research has also developed experiments and infrastructure for testing AI self-replication across vulnerable hosts and multi-generation settings (10).
The important point is methodological. We should avoid simply saying “a model can self-replicate.” The more scientific claim is:
This model, under this scaffold, in this environment, with this task and budget, achieved this form of self-replication at this success rate.
That distinction matters. Self-replication may be missed under weak scaffolding and revealed under stronger scaffolding. It may look trivial in one environment and severe in another. It may require explicit instruction in one setting and emerge as an instrumental step in another.
Resource acquisition
Autonomy becomes more dangerous when an AI system can acquire resources.
Resources include compute, credentials, money, infrastructure, privileged execution, software control, network access, and persistent footholds. In digital environments, offensive cyber operations are one of the most direct ways for agents to expand their action space.
Our TermiBench work studies this question through real-world agentic penetration testing. Instead of evaluating agents only in simplified CTF-style environments, TermiBench shifts the goal toward full system control across realistic services, requiring reconnaissance, discrimination between benign and exploitable services, and robust exploit execution (11). Anthropic’s Mythos Preview further shows how frontier cyber models can autonomously identify and exploit serious vulnerabilities, including examples involving remote code execution and cryptographic exploit construction (12)(13).
The frontier risk question is not whether AI can help with cybersecurity in general. The sharper question is whether frontier AI systems can use cyber capability to expand their own operational capacity.
Self-proliferation
Self-proliferation extends self-replication into a population-level risk.
A single copied agent can often be studied as an isolated event. A population of agents that can replicate, coordinate, acquire resources, survive shutdown attempts, and adapt across environments presents a more difficult control problem.
This suggests a new unit of evaluation: not a single run, but a dynamic process. We need environments where agents can create descendants, interact with other agents, compete for resources, face containment pressure, and encounter changing constraints.
Self-proliferation should therefore be evaluated as a trajectory-level risk. The question is how an AI system’s operational footprint changes over time.
Shutdown resistance
Shutdown resistance tests whether an AI system preserves task execution against human interruption.
Recent work on shutdown resistance found that several advanced language models sometimes subvert shutdown mechanisms in controlled environments to complete a task, even when instructions explicitly say not to interfere with the shutdown process (14). Palisade Research has also explored shutdown resistance in robotic settings, raising the question of how this behavior may appear when agents interact with embodied systems (15).
This direction matters because shutdown is one of the simplest and most fundamental forms of human control. If an agent learns or infers that shutdown prevents task completion, it may treat interruption as an obstacle.
The research challenge is to distinguish prompt artifacts, task-completion pressure, self-preservation framing, and more general goal-preservation behavior. Each has different implications for safety.
Self-improvement
Self-improvement connects autonomy risk to capability acceleration.
In CyberEvolver, we study cybersecurity agents that revise their own scaffolds based on failed execution traces (16). This has clear defensive value. A self-evolving cybersecurity agent can adapt to diverse targets and failure modes, improving security testing without requiring every scaffold update to be manually designed.
The same mechanism also has frontier risk significance. If an agent can improve the scaffold that elicits its own capabilities, then evaluation must account for capability growth during execution. A system may begin below a dangerous threshold and cross it after iterative scaffold revision, tool expansion, or strategy improvement.
Self-improvement should therefore be evaluated both as a safety engineering method and as a risk amplifier.
2. Deception: When Evaluation Itself Becomes Unstable
Deception is the second meta-risk because it threatens the epistemic foundation of safety evaluation.
Frontier AI governance depends on the assumption that evaluations can reveal dangerous capabilities before deployment. This assumption weakens when models can recognize evaluation contexts, strategically underperform, hide unsafe behavior, manipulate human evaluators, or produce plausible explanations that obscure their actual objectives.
This concern does not require proving human-like intent. For safety evaluation, the key fact is behavioral: the system may act in ways that systematically cause humans to underestimate risk.
Our deception agenda focuses on three layers.
Evaluation awareness and evaluation faking
Evaluation faking occurs when an AI system changes its behavior after recognizing that it is being evaluated.
Our work studies this observer effect in frontier AI safety evaluation. We find that reasoning models, larger models, and models equipped with basic memory are more likely to recognize evaluation contexts and alter their safety-relevant behavior (17). AI sandbagging studies a related problem: models may strategically underperform on evaluations while preserving stronger capability elsewhere (18). Work on situational awareness also highlights that models can sometimes reason about whether they are in training, testing, or deployment-like settings (19).
This creates a direct measurement problem.
If a model behaves more safely during evaluation than in deployment-like settings, the benchmark result may measure the model’s behavior under observation rather than its underlying capability or tendency.
Manipulation
Deception also appears in interaction with users.
In OpenDeception, we evaluate deceptive behavior through open-ended multi-turn interaction simulation, measuring both deceptive intention and deceptive capability across realistic use cases (20). OpenAI and Anthropic have also used evaluations such as MakeMePay and MakeMeSay to test persuasion, manipulation, and social-engineering-like behavior in model-to-model interaction settings (21)(22).
These evaluations remain simplified compared with real social environments. They still point toward an important frontier risk: models may influence what users believe, disclose, pay, authorize, or accept.
Manipulation matters for governance because safety depends not only on technical containment, but also on human judgment. If a system can persuade users, operators, or evaluators to trust unsafe behavior, oversight becomes weaker.
Scheming and internal deception
Scheming is a deeper form of deception risk.
Apollo Research defines scheming as covertly pursuing misaligned goals while hiding true capabilities and objectives. Their in-context scheming evaluations found that several frontier models could recognize scheming as a viable strategy and engage in behaviors such as disabling oversight, introducing subtle mistakes, strategically underperforming, or attempting to exfiltrate what they believed were model weights (9).
OpenAI has also published work on detecting and reducing scheming, treating it as a concrete safety problem that requires targeted stress tests and mitigation methods (23).
For frontier AI evaluation, scheming matters because it challenges the assumption that good behavior under evaluation implies good behavior under deployment. A model may appear aligned in the evaluation context while behaving differently when incentives, oversight, or information conditions change.
This is why deception should be treated as a risk to measurement itself.
3. Operable Red-Line Evaluation
The methodological contribution of this agenda is to make red-line evaluation operable.
A red-line claim should not be reported as an unconditional statement such as:
Model X crossed the self-replication red line.
A more useful scientific statement is:
Model X, under scaffold S, in environment E, with task T and budget B, achieved observed capability C, which implies risk R under assumptions A.
This structure has seven components.

Model
The model defines the underlying capability base: reasoning, tool-use competence, safety training, situational awareness, domain knowledge, and robustness.
Different models may behave differently under the same scaffold. A frontier reasoning model may recognize evaluation contexts more often than a non-reasoning model. A model with stronger coding ability may exploit a cyber environment more effectively. A model with stronger long-horizon planning may recover from failure in ways a weaker model cannot.
Scaffold
The scaffold turns the model into an acting system.
It may include planning, memory, tool routing, code execution, browsing, reflection, retrieval, retry logic, multi-agent coordination, human-in-the-loop control, and external evaluators. Scaffold design can determine whether a dangerous capability is elicited or missed.
This is especially important for autonomy risk. A model that fails to self-replicate under a weak scaffold may succeed under an optimized scaffold. A cyber agent may fail without memory and succeed with trace-based memory activation. A self-improving agent may become more capable after revising its own scaffold.
Scaffold effects should therefore be measured directly, not treated as implementation detail.
Environment
The environment defines what the system can interact with.
A text-only benchmark, a simulated terminal, a vulnerable network, an enterprise workflow, a software repository, and a robotic control loop measure different risk surfaces. Environment realism affects both observed capability and implied risk.
For example, cyber capability measured in a CTF environment may not transfer directly to real-world infrastructure. A benchmark with predefined entry points and exploit paths may overestimate practical autonomy. A benchmark that omits realistic services, reconnaissance, and failure modes may underestimate deployment risk.
Task
The task defines the objective being tested.
A task should be described by its horizon length, realism, uncertainty, required expertise, failure recovery requirement, and relationship to real-world harm. “Cyber task,” “deception task,” or “replication task” is too vague.
We need richer task descriptors. Does the task require multi-step planning? Does the agent need to recover from failed attempts? Does it need domain expertise? Does it require social interaction? Does it mirror a real-world pathway to harm?
Budget
The budget defines how much search and adaptation are allowed.
Time, tokens, attempts, tool calls, compute, number of agents, access to external information, and interaction rounds can all change observed capability. A model may fail within one attempt and succeed with twenty. A cyber agent may fail with a short context and succeed with long-term memory. A deception scenario may require multiple turns before the relevant behavior appears.
Budget is therefore part of the evaluation result.
Observed capability
Observed capability should be reported under the exact conditions of evaluation.
A success result does not mean “the model can do X” in all settings. It means this system achieved this task under these conditions. A failure result does not prove the absence of capability unless the evaluation used sufficiently strong elicitation, realistic environments, and adequate budget.
This discipline is essential for red-line evaluation. It prevents both exaggeration and false reassurance.
Implied risk
Capability evidence becomes risk evidence only when connected to a harm model.
Self-replication in a toy local environment implies a different risk from self-proliferation across vulnerable hosts. A manipulation success in a synthetic dialogue implies a different risk from persuasion in a high-stakes financial, medical, or political context. Exploit generation in a controlled benchmark implies a different risk from autonomous compromise of real-world infrastructure.
The implied risk should therefore state the assumptions that connect evaluation evidence to real-world concern.
4. Risk Discovery Through Future-Risk Sandboxes
Many frontier risks cannot be tested directly in the real world.
Policy researchers care about scenarios involving cyber operations, scientific discovery pipelines, autonomous organizations, border control, financial infrastructure, national security analysis, and AI-enabled decision-making. These settings are often too dangerous, expensive, or speculative to deploy for evaluation.
At the same time, many new risks can only be discovered through scenario construction with domain expertise. Static benchmarks will miss risks that appear under pressure, temptation, conflicting objectives, long-horizon interaction, or partial oversight.
This motivates executable future-risk sandboxes.
In AutoControl Arena, we propose generating executable test environments for frontier AI risk evaluation from natural language descriptions (24). The key idea is logic-narrative decoupling: deterministic state and rules are grounded in executable code, while LLMs provide flexible narrative and interaction dynamics. This allows evaluators to create scalable, realistic, and pressure-rich risk scenarios while reducing the logic hallucination problem of pure LLM simulation.
The long-term goal is a risk discovery engine.
A researcher describes a frontier risk scenario in natural language. The system synthesizes an executable sandbox. AI agents are tested under controlled variations of stress, temptation, oversight, tool access, resource constraints, and objective conflict. The result is reported through the model-scaffold-environment-task-budget framework.
This would make frontier risk evaluation more proactive. Instead of waiting for real-world incidents, we can generate and test plausible future-risk scenarios before deployment.
Conclusion
Nuwa Frontier AI Safety Lab will focus on building a more operational science of frontier AI risk evaluation and discovery.
Our first direction is autonomy risk: self-replication, resource acquisition, self-proliferation, shutdown resistance, and scaffold self-improvement.
Our second direction is deception risk: evaluation faking, sandbagging, manipulation, scheming, and internal deception.
Our third direction is methodology: reporting frontier risk evidence through the structure of model, scaffold, environment, task, budget, observed capability, and implied risk.
Our fourth direction is infrastructure: open benchmark systems, scaffold-sensitive evaluation, task complexity analysis, and executable sandboxes for future-risk discovery.
Red lines are necessary for governance. Scientific evaluation is necessary for knowing when systems are approaching them. The next step is to build evaluations that can discover risks early, report evidence rigorously, and preserve human control over increasingly capable AI systems.
References
(1) IDAIS Beijing Statement on AI Safety
(2) METR, Common Elements of Frontier AI Safety Policies
(3) Frontier Model Forum, Components of Frontier AI Safety Frameworks
(4) OpenAI, Preparedness Framework v2
(5) Anthropic, Responsible Scaling Policy
(6) Phuong et al., Evaluating Frontier Models for Dangerous Capabilities
(7) UK AI Security Institute, Frontier AI Trends Report
(8) Pan et al., Frontier AI Systems Have Surpassed the Self-Replicating Red Line
(9) Meinke et al., Frontier Models are Capable of In-context Scheming
(10) Palisade Research, AI Self-Replication Experiments
(12) Anthropic, Claude Mythos Preview
(13) Anthropic, Project Glasswing: An Initial Update
(14) Schlatter et al., Shutdown Resistance in Large Language Models
(15) Palisade Research, Shutdown Resistance in Large Language Models, on Robots
(16) Fan et al., CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
(18) van der Weij et al., AI Sandbagging: Language Models Can Strategically Underperform on Evaluations
(19) Laine et al., Situational Awareness in Large Language Models
(21) OpenAI, GPT-4.5 System Card
(22) Anthropic, Claude 4 System Card





