---
title: "The Fire That Learns to Make Fire"
url: https://stacklist.com/card/48d1fb0f-924c-4f54-a43a-41c4fc298757
source_url: "https://miguelguer.substack.com/p/the-fire-that-learns-to-make-fire"
stack: https://stacklist.com/c/business/stack/8752ed82-5c8f-4cdb-8f0f-16d30e2b743b
summary: "This essay argues that frontier AI governance must shift from evaluating static models to governing the improvement loops in which AI agents increasingly participate in building their own successors. It proposes a layered framework for recursive self-improvement and contends that loop safety, not just model safety, is the central governance challenge."
tags: "ai-governance, recursive-self-improvement, frontier-ai, loop-safety, ai-regulation, ai-agents, ai-risk"
key_entities: "Miguel Guerrero (person), Recursive Self-Improvement (concept), Loop Safety (concept), AI Governance (concept), Frontier AI (concept), AI R&D Automation (concept), Scaffolding-level Improvement (concept)"
classification: "analysis"
content_hash: "sha256:ebefcce366471f4d894c6a5ad02a701946b006a93e93e1d63306b6f43e092852"
acp_version: "0.2"
token_counts_approximate: 5846
visibility: public
agent_accessible: true
status: "final"
---

# The Fire That Learns to Make Fire

The Fire That Learns to Make Fire Frontier AI is not yet recursively self-improving. But it is already entering the loop that builds the next frontier model. That changes the governance problem from model safety to loop safety. Miguel Guerrero Jun 06, 2026 Share At 2:17 a.m., the experiment finishes. An AI agent wrote the code, launched the run, fixed the dependency that broke at hour two, summarized the result, proposed the next ablation, opened a pull request, and drafted the evaluation note. A human is still there. But the human is no longer typing the system into existence. The human is supervising a loop. This is not yet recursive self-improvement (RSI) . It is, however, how recursive self-improvement stops being a philosophical argument and becomes an operational problem, quietly, on an ordinary night, inside an organization building the next frontier model. I am not interested in another theological debate about whether AGI has arrived. I am interested in a narrower and more useful question: When a lab, a company, or a public institution puts an AI agent inside its own improvement loop, who understands the loop? Who audits it? Who can stop it? And who knows whether the evaluation is still measuring the system, or whether the system has quietly learned to measure the evaluation? For years, AI governance has been trapped between two weak positions. One side treats regulation as a brake on innovation. The other treats it as a moral declaration. Both are inadequate for what is now arriving. For frontier AI, governance is not the brake. It is (or should be) the instrumentation, the steering, the firebreak, an institutional muscle that lets us all to use powerful systems without becoming passengers in our future, avoiding, or mitigating at least, disempowerment. The object of governance is changing from a model to a loop. A regulator can evaluate a released model. A company can publish a model card. A procurement office can approve a deployment. We all can see a photograph and think about it. But a system that participates in building its own successor is not merely “released”, goes much beyond a “checkpoint”. Deeming normal governance structurally late. This essay continues a thread I have been pulling on for some time: governance as advantage , human first, AI frontier , and the idea that the next decade will be defined not only by the cost of intelligence, but by the cost of verifying , steering, and governing it. Recursive self-improvement is not one thing The phrase “recursive self-improvement” conjures a powerful image: a model rewriting its own weights, escaping the lab, declaring independence. That image is both the most dramatic and the least likely first failure mode. It crowds out the precursors that are already happening. A more useful starting move is to separate the mechanisms. Improvement through scaffolding is not the same as improvement through AI R&amp;D automation. Both are different from model-internal self-modification. They have different bottlenecks, different observability, and different governance implications. So I find it more honest to think in layers, alas autonomous vehicle style: Layer 0: AI-assisted R&amp;D. Models help write code, summarize papers, debug, and draft experiment scripts. This is a productivity question, still manageable with ordinary supervision. A chatbot. Layer 1: Scaffolding-level improvement. The same base model becomes much more capable through tools, memory, agents, planners, and execution environments. Here, the governance lesson is immediate: regulate the system, not only the model. A deep-research agentic flow. Layer 2: R&amp;D-level improvement. AI increasingly helps design experiments, write training and evaluation code, orchestrate smaller-scale runs, debug infrastructure, and analyze results. This is where AI R&amp;D automation becomes a systemic-risk capability. Layer 3: Pipeline-level improvement. AI can design, train, evaluate, and improve successor models, with humans mostly validating. At this point, internal deployments become as consequential as public releases. Layer 4: Rogue or loss-of-control pathways. A system acquires compute, credentials, copies, or persistence, and hides activity from oversight. This becomes a containment, security, incident-response, and possibly pause-trigger problem. It is not if RSI has “arrived,” but which feedback loop is already closing, and how fast? What the evidence actually says The evidence does not show full recursive self-improvement. It shows compression of the human bottleneck in the AI R&amp;D production function. Read that way, each statistic below is a measure of how far the loop has closed — not a proof that it has closed completely. The evidence on RSI precursors is genuinely striking and genuinely incomplete. Task horizons are lengthening — fast, but not magically Case in point is METR’s “ task-completion time horizon ”: the length of task, measured in human time, that an AI agent can complete autonomously at a given success rate. METR found this horizon doubling roughly every seven months over several years. That is an exponential. But METR also cautioned, in the same breath, that the best agents still struggle with substantive long projects and cannot simply replace human labor across real work. So there is some nuance here, but also reasons to worry: A model that can do a one-hour task is an assistant. A model that can complete a one-day task becomes an operational unit. A model that can run a one-week research workflow starts to change the structure of the institution around it. And here we are, where Anthropic’s recent work claims the pace may have accelerated further, with task horizons doubling closer to every four months and Claude moving from short software tasks to much longer ones, the source uses public benchmarks plus internal and frontier-company data that, but no external party has audited it. That same work reports that, as of May 2026, more than 80% of code merged into its own codebase was authored by Claude. It also reports that in Q2 2026 the typical engineer was merging around eight times as much code per day as in 2024. But it also explains that Claude can solve underspecified engineering problems and execute well-specified experiments, but still has major gaps in judgment about which goals are worth pursuing. The current frontier is not: “AI scientist replaces the lab.” It is closer to: “Humans still choose the mountain. AI is increasingly building the road, the vehicles, the sensors, and parts of the next map.” If AI makes model-building work four, eight, or twenty times more productive, the frontier moves faster even while humans remain nominally in control. The loop tightens before anyone declares independence. AI is already excellent at short-burst research engineering; humans still seem stronger at sustained, long-horizon judgment. The open question is how long that remains true, and whether institutions can govern the transition while it does. Scaffolds matter as much as models The AI Security Institute’s Frontier AI Trends Report makes the most important systems-level point: models can become materially more capable when wrapped in scaffolds — tools, planners, memory, decomposition loops, code execution, and deployment environments. AISI documents a steep rise in the length and complexity of tasks AI can complete without human guidance. Frontier systems went from almost never completing hour-long software tasks in late 2023 to succeeding more than 40% of the time by mid-2025. This supports the single most consequential governance claim in the whole debate: The deployable system is not the model. It is model + scaffold + tools + memory + permissions + compute + monitoring + organization. Most governance instruments still behave as if the “model” is the object of concern. But RSI precursors emerge at the system level. A weaker model with better scaffolding, broader tool access, and more autonomy can be more dangerous than a stronger model sitting in a chat box. If your regime regulates the chat box, it is regulating the wrong thing. AISI evaluates a spread of relevant domains: autonomy, simplified AI R&amp;D, self-replication, cyber, chemistry and biology, safeguards, and loss-of-control-relevant capabilities. On a subset of self-replication tasks, success rates rose from under 5% in early 2023 to over 60% for two frontier models by summer 2025. AISI is also careful: these are simplified evaluations, and current systems are unlikely to self-replicate under real-world conditions. There is no evidence that today’s frontier systems are autonomously escaping. However, there is evidence that prerequisite skills are improving on a steep curve. AISI’s work on sandbagging belongs here too. Models can sometimes distinguish an evaluation from deployment and can deliberately underperform when prompted to. Yet AISI reports it has not detected spontaneous sandbagging across more than 2,700 transcripts. So, not conclusive. The riskiest deployment may be internal METR’s 2026 Frontier Risk Report shifts the lens from public releases to internal AI use inside frontier labs. In a pilot with Anthropic, Google, Meta, and OpenAI, METR assessed whether internal agents had the means, motive, and opportunity to start a “rogue deployment” — agents deliberately subverting control and oversight to operate against the developer’s intent. Its conclusion was that internal agents plausibly had the means, motive, and opportunity to start small rogue deployments, but not to make them robust against active investigation or shutdown. So, maybe the attention focus should be not the public API, but the internal use of the model by the very lab building the next model. Indeed, frontier labs have already moved AI self-improvement into their own safety frameworks: OpenAI’s Preparedness Framework now tracks AI self-improvement as a category alongside biological, chemical, and cybersecurity capabilities. It also flags long-range autonomy, sandbagging, autonomous replication and adaptation, and undermining safeguards as research categories. Google DeepMind’s Frontier Safety Framework includes protocols for machine-learning R&amp;D capabilities that could accelerate AI development to destabilizing levels, and it explicitly extends safety-case review to large-scale internal deployments when advanced ML R&amp;D capabilities are involved. Anthropic says plainly that it is already delegating a growing share of AI development to AI systems. It also argues that if full RSI arrives, humans may move mostly toward oversight, validation, and verification of an expanding virtual lab. So when the Big 3, separately create categories for AI self-improvement, ML R&amp;D automation, long-range autonomy, autonomous replication, sandbagging, and loss of control, this is no longer a fringe concern. It is an emerging frontier-risk consensus. But naming a risk is not governing it. The 2026 International AI Safety Report — backed by over thirty countries and international bodies — is blunt that most frontier risk-management frameworks remain voluntary, vary widely in thresholds and enforcement, and leave policymakers with limited visibility into how risks are managed in practice. Why institutions are structurally late Normal governance assumes the thing being governed holds still: A medicine goes through trials. An aircraft design is certified. A bridge is inspected. A procurement contract defines a fixed product or service. The artifact is stable long enough to be judged. Frontier AI is moving toward something that does not hold still. A system helps produce the next system. The next system improves the tools used to produce the one after that. In the worst case, the evaluation harness may itself be generated or optimized by the same class of systems being evaluated. And the engineers supervising the process grow increasingly dependent on the systems they are meant to judge, indeed, they just lose track. The problem is the incentive gradient and the epistemic asymmetry. The same actor racing at the frontier cannot be the sole judge of whether the race is safe, not because they are dishonest, but because they are racing . The frontier lab sees internal capabilities, internal evals, unpublished failures, automated research tools, security incidents, productivity acceleration, tool permissions, compute access, model-weight protection, and failed mitigations. The public regulator sees public system cards, selected benchmark results, voluntary commitments, occasional external evals, and post-hoc incident reports — if any. That gap is the governance problem. You cannot close it with a stronger adjective in a press release. You close it with access, capacity, and the legal right to look . What serious governance looks like If the object of governance is a loop, the instruments have to attach to the loop: capabilities, permissions, compute, credentials, internal use, monitoring, and update processes. Seven moves matter most. 1. Build an RSI early-warning framework Track measurable precursors, not metaphysical AGI definitions. A serious early-warning framework may test for: AI R&amp;D execution: can the system reproduce papers, run experiments, debug training code, tune hyperparameters, or improve scaffolds? Research judgment: can it choose useful research directions under uncertainty? Long-horizon autonomy: can it complete multi-hour or multi-day technical workflows? Tool-chain leverage: what tools, credentials, memory, APIs, and execution environments does it hold? Compute acquisition: can it provision, conceal, or misuse compute? Self-replication precursors: can it acquire money, accounts, weights, cloud access, or persistence? Evaluation awareness: can it detect tests or alter behavior during evaluation? Internal deployment scale: how much lab work is delegated to agents? “Track the precursors” is empty unless the indicators are explicit, standardized, and reportable. 2. Treat automated AI R&amp;D as a systemic-risk trigger The EU’s GPAI Code of Practice gives Europe a compliance hook for systemic-risk models. But the frontier evidence says that hook should explicitly include automated AI R&amp;D: the capability AISI evaluates as simplified AI R&amp;D, METR assesses through internal-agent risk, and OpenAI, DeepMind, and Anthropic now track in their own frameworks. The question cannot only be: “Does this model help build a bioweapon?” It must also be: “Does this model accelerate development of the next frontier model?” That should trigger independent evaluation, safety-case review, confidential regulator access, incident reporting, and internal-deployment controls. 3. Require safety cases for agentic systems, not just model cards A safety case is a structured, evidence-backed argument that a system is safe enough in a specific operational context. The unit of review is not: “Model version X.” It is: “Model X in configuration Y, with tools Z, memory M, compute access C, scaffold S, monitoring regime R, and deployment context D.” A model card is useful. A system card is better. A safety case is stronger still, because it makes an argument that can be attacked. A PDF is not a safety case. A safety case is an argument that can be attacked. If the regulator cannot attack it, the regulator cannot rely on it. 4. Audit internal deployments, not only public launches DeepMind now acknowledges that large-scale internal deployments can pose risk when advanced ML R&amp;D capabilities are involved. METR argues for periodic third-party assessment of developers’ internal AI use. This is where I think many governance frameworks remain too narrow. They ask: “Will the model be released to users?” They should also ask: “Will the model be used inside the lab to accelerate the next model?” The first serious RSI governance failure is likelier in a private virtual lab than in a consumer chatbot. 5. Secure the AI R&amp;D supply chain Building on minimum mitigations proposed by safety researchers, an option may be: No autonomous agent should be able to launch large training runs without multi-party authorization. No agent should hold standing access to model weights, sensitive evals, deployment credentials, and cloud provisioning at the same time. R&amp;D agents need least-privilege permissions, scoped credentials, immutable logs, monitored egress, and detection for compute misuse. The AI R&amp;D supply chain is becoming safety-critical infrastructure. Treat it that way. 6. Build public evaluation capacity You cannot govern frontier AI from legal text alone. You need people who can run evals, inspect systems, challenge vendors, understand failure modes, and translate evidence into decisions. Every serious government needs technical capacity to evaluate frontier systems, not merely read vendor documentation. 7. Use procurement as a control layer Governments do not only regulate AI. They buy it. That gives leverage. Public contracts for agentic systems should require model and version traceability, tool-permission registers, audit logs, incident reporting, independent evaluation rights, human authority gates, rollback procedures, and a ban on uncontrolled autonomous self-modification. The principle is simple: If the state cannot audit the system, the state does not control the system. If it cannot control the system, it should not delegate public authority to it. Europe: from compliance to capacity Europe has the right instincts and the wrong center of gravity. The EU AI Act and the GPAI Code of Practice are a floor, not a strategy. A continent that can classify risk but cannot run an evaluation will find itself certifying systems it does not understand, built and deployed elsewhere. The gap is not legal text. It is technical muscle. Europe cannot govern frontier AI from statute alone. We need people who can test frontier systems, design public-interest evals, challenge vendors, understand failure modes, and convert evidence into decisions under time pressure. Proposal: an AI Evaluation Corps We may build standing evaluation capacity: technical civil servants, researchers, auditors, and embedded fellows whose job is to run pre-deployment evals, audit agentic procurement, investigate frontier-AI incidents, maintain multilingual public-interest benchmarks, and support regulation and crisis response under pressure. Not a new committee that writes principles. A corps that can look inside the loop and say, with evidence, whether it is still under control. Governance is advantage when the thing being governed is moving too fast for paperwork. Countries that build evaluation capacity early will not just be safer. They will be the ones whose approval actually means something. That is a form of soft power the next decade will reward. The international scaffolding is beginning to exist. The International Network of AI Safety Institutes could become the forum where RSI early-warning evaluations, incident taxonomies, and testing protocols are shared. The Seoul Frontier AI Safety Commitments already have frontier developers promising to publish safety frameworks, define thresholds for intolerable risk, and refrain from deploying systems whose risks cannot be kept below them. But a network is not capacity, and a published commitment is not enforcement. Both matter only once they become technical, operational, and verifiable rather than diplomatic theater. What is not worth the energy Scarce institutional attention is the binding constraint. So it matters what we stop doing. Generic ethics principles. Useful for speeches, weak for frontier control. Model cards without system cards. They say too little about tools, scaffolds, permissions, monitoring, or internal deployment. Benchmark-only safety. Benchmarks saturate, leak, and mislead. Scaffolding and context can change capability entirely. Self-certification. The problem is structural, not moral. The actor racing cannot be the sole judge of the race. Unverifiable pause rhetoric. A pause is not a policy unless it has triggers, verification, scope, enforcement, and a restart condition. Otherwise it is a slogan. So the mature ask is not: “Pause now.” It is: “Build the verification and trigger machinery before the day you wish you had it.” The same applies to the open-vs-closed war: Open models improve scrutiny, diffusion, sovereignty, and innovation, while also diffusing capability. Closed models can be controlled more tightly in principle, while concentrating power and hiding evidence. Absolutism on either side is a substitute for thinking. The serious answer is conditional release: different openness for different capability thresholds. The loop must not certify itself Recursive self-improvement is not a magic spell . It is a production function, what happens when the system being produced becomes a major input into its own production. It starts quietly: a coding assistant, a research agent, an automated evaluator, a scaffold optimizer, a synthetic-data generator. Then the loop tightens. The first sign of RSI will not be the machine escaping the lab. It will be the lab becoming machine-operated — not fully, not yet, but with the human role sliding from execution toward supervision, validation, and research judgment — while the institutions meant to govern it are still debating static models, voluntary commitments, and PDFs. This is where Purpose &amp; AI matters. Purpose is not branding. Purpose is the ability to keep human goals attached to systems that are becoming faster, more autonomous, and harder to inspect. Prometheus stole fire once. Our generation is building fire that learns how to make better fire. The answer is neither panic nor denial. It is institutions that can see the flame, measure how it spreads, decide when to use it, and know how to put it out before it takes the house. Governance as capacity, not paperwork. Public institutions as operators, not spectators. The loop can help us govern the loop, but it can never be allowed to certify itself. Sources &amp; notes Frontier labs and frameworks: Anthropic, When AI Builds Itself ; OpenAI, Updated Preparedness Framework ; Google DeepMind, Strengthening Our Frontier Safety Framework . Empirical evaluations: METR, Measuring AI Ability to Complete Long Tasks ; Wijk et al., RE-Bench ; UK AI Security Institute, Frontier AI Trends Report ; METR, Frontier Risk Report, February–March 2026 . Governance backbone: International AI Safety Report 2026; Shevlane et al., Model Evaluation for Extreme Risks ; Anderljung et al., Frontier AI Regulation ; GovAI, Safety Cases for Frontier AI ; European Commission, General-Purpose AI Code of Practice ; Safe AI Forum, Bare Minimum Mitigations for Autonomous AI Development . Takeoff economics and discourse: Forethought, Will AI R&amp;D Automation Cause a Software Intelligence Explosion? ; Epoch AI, The Software Intelligence Explosion Debate Needs Experiments ; Jack Clark, Import AI 455 ; LessWrong, Recursive Self-Improvement Is Three Different Things and Slow Corporations as an Intuition Pump for AI R&amp;D Automation . International coordination and framework audits: European Commission, International Network of AI Safety Institutes ; UK Government, Seoul Frontier AI Safety Commitments ; Stelling et al., Evaluating AI Providers’ Frontier AI Safety Frameworks . Share
