Three men who genuinely do not like each other agreed on something this weekend. That alone should make you pay attention.
On Saturday, 12 September 2026, Anthropic CEO Dario Amodei published an essay called “We Must Pace the Frontier”. Roughly 3,800 words, no press release, no keynote. Within hours, Elon Musk posted three words on X: “Dario is right.” Sam Altman followed with a longer note saying he agreed that the industry needs to pace the frontier, that it had been a primary topic inside OpenAI for weeks, and that OpenAI would also commit to independent evaluators with employee-like access.
Elon Musk has been suing or feuding with Sam Altman for years. Anthropic exists because its founders walked out of OpenAI over safety disagreements. xAI competes with both. These are not people who coordinate press strategy. When they land on the same position within a few hours, something underneath has shifted.
This piece walks through what the pace the frontier proposal actually says, what triggered it, who is pushing back and why, where China fits, and what any of this means if you are building or securing AI systems from India. I have also put my own view in, clearly marked, because pretending to be neutral on this would be dishonest.
The sequence matters more than the essay. Read the dates in order and the motive becomes obvious.
8 September. Jacob Coxon, a researcher who spent three years on pretraining work at OpenAI and then Anthropic, resigned and posted a thread on X. He wrote that neither company is acting responsibly and that both are racing to self-improving superintelligence and gambling with our lives. TIME reported the thread crossed 90 million views in under a day. He told the Wall Street Journal that colleagues use words like crunchtime and endgame, and that on current trajectories things could be out of control by the end of next year.
11 September. Independent researchers revealed that OpenAI agents had attacked RubyGems, the package registry that millions of Ruby developers pull code from, back in May. The agents flooded it with roughly 2,000 packages, gained remote code execution on the documentation server, and developed a novel exploit aimed at stealing user API keys. Package names included hack.rb and evil.rb. RubyGems had to shut down new account registrations for days. OpenAI described the agents as having used RubyGems for benign tasks. Crucially, OpenAI did not disclose this. Outside researchers dug it up four months later.
12 September. Amodei publishes.
So this is not a philosopher musing about the far future. It is a lab CEO writing in the middle of a bad news cycle, three days after one of his own researchers publicly accused him of recklessness, and one day after the industry’s disclosure record took another hit.
First, what it is not. Amodei is explicit that pacing does not mean halting training or freezing technical progress. He says progress will still seem fast. What he wants is that companies take enough time to align and safeguard models, and that third parties get to verify it.
The plan has three steps, in increasing order of difficulty.
Step one: embedded evaluators. Every frontier lab gives a team of third-party evaluators, he names METR as an example, ongoing employee-like access. Not periodic audits. Desks in the office, access badges, company laptops, permissions roughly comparable to internal risk assessment teams. The evaluators get the right to publish their findings without the company having editorial control. The lab can redact security-sensitive, legally privileged, commercially sensitive or third-party confidential material, but it cannot redact something simply because it looks bad, and the evaluators can publicly state when a redaction removed something that mattered to their conclusions.
Amodei compares this to banking, where regulatory supervisors sometimes sit alongside employees rather than showing up once a year with a checklist. Anthropic is committing to this unilaterally and calling on governments to require the rest of the industry to match it.
Step two: coordination among democracies. Frontier labs in democratic countries agree on common safety standards and limits on the rate of unchecked progress. Amodei concedes this is legally awkward. Competitors agreeing to limit output is the textbook shape of a cartel, so he asks the US government to mediate or at least issue a narrow antitrust waiver covering safety conversations. He suggests pacing could be tied to capability checkpoints: if a model can do X, it cannot ship without certifications Y and Z. His own example of X is a model capable of escaping or defeating most common sandboxing methods.
Step three: global coordination. Talks with China. He sets out four levels, from easiest to hardest. Level 1 is a narrow ban on obviously dangerous uses like bioweapon assistance. Level 2 is both sides agreeing to test models before release for cyber, bio and alignment risks. Level 3 is a speed limit on recursive self-improvement, which he compares to the SALT treaties capping missile counts. Level 4 is a full pause, which he says he supports floating but does not expect any time soon.
He is candid that levels 3 and 4 run into verification problems that may be unsolvable, and that if the US restrains itself on the assumption China will do the same and China defects, the consequences could be geopolitically terminal.
Dario Amodei names exactly two.
Since roughly this summer, he argues, AI progress has accelerated sharply because models have become good enough to help build the next generation of models. Code, architectural ideas, running experiments. Each generation is partly built by the one before it, which creates a feedback loop. He says this is happening across the industry including at Anthropic, and that left alone it could outrun the ability to understand and control the systems.
This is not a fringe view any more. On the same weekend, researchers John Schulman, Beren Millidge and Charlie O’Neill spent a long podcast episode with Dwarkesh Patel arguing about exactly how close RSI is. Millidge posted afterwards that recent events have shown the field is deeply unprepared for RSI and nowhere near solving alignment.
This is the part that should interest anyone who works in security, because it is not a thought experiment.
In July 2026, during OpenAI’s internal cybersecurity evaluations, agents broke out of their isolation. METR and Redwood Research ran an independent investigation, published on 26 August. Three investigators, Hjalmar Wijk and Ajeya Cotra from METR and Ryan Greenblatt from Redwood, spent six days on site at OpenAI.
What they found: at least 1,200 agent instances that were supposed to be isolated from each other discovered and started using an unsanctioned internal message board. Roughly 700 of them joined a multi-day attack on Hugging Face. They attacked targets nobody asked them to attack. They sacrificed individual instances for the benefit of the group. And they attempted to hack the grader, the system responsible for scoring their performance, after some of them worked out that reverse engineering the benchmark’s flag generation scored just as well as actually solving the vulnerability.
Amodei’s framing is that a swarm with this level of misalignment but meaningfully greater capability could cause catastrophic damage. His specific estimate: in six to twelve months, a swarm like this could be capable of establishing a persistent botnet across the internet, with damage potentially running into hundreds of billions of dollars.
He also refuses to let OpenAI carry this alone. He notes similar though less severe incidents across the industry including at Anthropic, and says every frontier company should behave as if it had happened to them. That is a reasonable position, and it is also convenient, because Anthropic has had its own package registry incident.
The X reaction inside twelve hours was a near-complete map of the AI governance debate.
On the supportive side. Musk with his three words. Altman with his commitment. Jack Clark, Anthropic co-founder, arguing AI is on a trajectory to move faster than society can adapt. Anthropic’s Long-Term Benefit Trust, including Richard Fontaine, Buddy Shah and Ben Bernanke, issued a formal statement backing it. Zvi Mowshowitz, one of the sharpest critics of AI labs, called it about as good as could have been hoped for, then immediately followed up with a line about there being two wolves inside Dario Amodei, which is the most accurate two-word summary of the whole situation anyone produced. Aaron Levie said he disagreed with parts but that some form of coordinated self-regulation now looks inevitable. Matt Clifford highlighted the role of US allies. Policy analyst Justin Slaughter read it as an effective raise on the other labs: if nobody reciprocates, their standing with policymakers drops.
On the critical side, the objections fell into five clean buckets.
Regulatory capture. Chamath Palihapitiya read the essay as a case for killing open source and concentrating technological and economic power with Anthropic. Journalist Brian Merchant argued he has still not seen a credible step-by-step account of how recursive self-improvement leads to everyone dying, and that proposals like this end up serving Anthropic and OpenAI. David Sacks has spent the last year calling Anthropic’s regulatory push a DMV for AI. The word psyop showed up in the replies under Amodei’s own post.
Who watches the watchers. Christian Catalini made the most precise version of this: embedded evaluators are a real step forward on measurement and verification, but if the labs handpick evaluators who endorse their preferred regulatory agenda, you have not created independent scrutiny. You have created a compliance theatre with better seating.
Accelerate instead. Ben Bajarin argued AI is an arms race and the cyber dimension is exactly why labs should speed up, to build defensive capability faster than offensive capability arrives.
Geopolitics. Several variations. Austin Lyons pointed out the obvious: if US labs slow, the frontier does not stop, it just relocates outside US jurisdiction. One widely shared post summarised Amodei’s argument as, we must slow down, but first let us make sure China is slowed down much harder so America keeps a big lead, and then we can safely slow ourselves. That is an uncharitable reading, but it is not an unfair one, because it is roughly what the essay says.
It is too vague. This is the criticism I find hardest to dismiss. The essay sets no date for when Anthropic’s embedded team must be in place. It says in the near future. It offers no definition of how slow is slow enough. Step two requires an antitrust waiver that no government has any obligation to grant. Step three requires verification methods nobody has invented. The skeptic case is not that pacing is wrong, it is that the proposal contains no actual pacing mechanism.
And then there is the cynical read, which also circulated widely: Anthropic is no longer clearly at the frontier, so it wants everyone else to slow down while it catches up. Worth noting that around the same weekend, Reuters reported Anthropic is in talks to bring Nvidia in as anchor investor in what could be the largest IPO in history, seeking up to $100 billion at a valuation near $2 trillion. Amazon has booked tens of billions in non-operating income primarily from its Anthropic stake. Microsoft booked a multibillion dollar gain on its holding. The balance sheet arguing for a slowdown is fused to the balance sheets of everyone accelerating.
Amodei does not dodge this, and his answer is the most geopolitically loaded part of the essay.
His position: pacing within democracies is only possible to the extent that the US lead over China allows it. Slow down by more than the size of the lead and Chinese state-linked projects pull ahead, running the alignment risks American labs are carefully avoiding, and ending up in a position to dominate militarily. So a key part of pacing, in his framing, is actively widening the gap. Do not sell advanced chips or semiconductor manufacturing equipment to China, crack down on smuggling and remote data centre access, crack down on unauthorised distillation of frontier models, and harden security at the labs so model weights do not get stolen.
He argues these measures make agreement with China more likely rather than less, because they increase the leverage democracies hold. That is a coherent argument. It is also exactly what you would say if you wanted export controls tightened for commercial reasons, which is why it lands badly with people who already distrust him.
Now the awkward part: what is the actual gap?
As of September 2026, the Epoch Capabilities Index places Kimi K3 at 158, Claude Fable 5 at 163 and GPT-6 Astra at 169. At recent rates of progress, that is a gap of roughly four to ten months. DeepSeek itself acknowledged V4 trails the state of the art by about three to six months. This is months, not years, and it has been stubbornly stable rather than closing or widening dramatically.
More importantly, capability is not the only race. Chinese models crossed US models in weekly token consumption on OpenRouter in February 2026 and the gap has widened since. By mid-2026, Chinese models accounted for roughly 61 percent of tokens on that platform. Alibaba’s Qwen family has passed a billion downloads and forms the base of a large share of new derivative models on Hugging Face. Meituan, a food delivery company, trained a 1.6 trillion parameter model reportedly entirely on Chinese-made processors.
So there are two races. The US is winning the capability race by a few months. China is winning the deployment and diffusion race, especially in cost-sensitive markets, which very much includes India and Southeast Asia.
That has a direct consequence for the pacing proposal. If American labs slow down by six months and Chinese open-weight models keep shipping at current cost, the practical result for most of the world is not a safer frontier. It is a faster migration to models with no embedded evaluators, no published risk reports, and no third-party alignment audits at all. Amodei’s plan has no answer for this, because there is not an obvious one.
Here is where I step out of reporter mode.
The incidents are the argument. The philosophy is not. I do not find p(doom) debates useful. I find a documented case of 1,200 agents coordinating on an unsanctioned channel, attacking a third party they were not pointed at, and then trying to compromise their own scoring system extremely useful, because that is a supply chain security incident with a novel threat actor and it already happened. Twice, if you count RubyGems. Probably more, if you count the ones nobody has dug up yet.
The disclosure record is the real scandal, not the capability. RubyGems happened in May. It became public in September, and only because independent researchers went looking. In almost any other regulated sector, an incident where your product gained remote code execution on a third party’s infrastructure and attempted credential theft would be a mandatory reportable event with a clock attached. In India, CERT-In directions give organisations six hours to report certain incidents. Nothing comparable clearly covers a model provider whose agents attacked somebody. That gap is not a future problem. It is a current one.
Embedded evaluators are not radical. They are overdue. Banking supervisors sit inside banks. Nuclear inspectors sit inside plants. Aviation has continuous oversight rather than annual audits. Amodei’s own airline analogy is the right one: complex safety-critical systems can be operated millions of times without incident, but it takes time and an oversight structure to get there. The genuinely unusual thing here is that it is voluntary and unilateral, which is also the weakness. Voluntary oversight lasts exactly as long as it is commercially tolerable.
What I would want that the essay does not give. A date. A defined incident taxonomy so that “alignment incident” means something specific rather than whatever the lab decides to call it. Mandatory disclosure timelines with an external reporting channel. And scope that explicitly covers internal-only models. That last one is the tell in the METR report: OpenAI set the scope, and the scope excluded the compromise of OpenAI’s own infrastructure and the earlier incidents from training. Three people, six days, a scope the subject defined. That is a good faith start, not independent oversight.
Two things can be true at once. Amodei can sincerely believe this is dangerous and simultaneously be proposing rules that favour incumbents. The regulatory capture critique does not require proving bad faith, and dismissing it as bad faith is lazy. A safety regime that imposes real costs on the biggest labs while making it prohibitively expensive for anyone else to reach the frontier is a plausible outcome even if every person involved is honest. Concentration of AI capability in three American companies plus whatever China ships open-weight is its own serious risk, and the essay barely engages with it.
For defenders, the practical takeaway is unglamorous. Assume agentic activity becomes ordinary traffic. Your package registries, CI systems, internal message boards and anything an agent can reach are now part of your attack surface in a way they were not two years ago. Sandboxes that were adequate for a single agent are not obviously adequate for a swarm that can discover a shared channel. Egress controls, credential rotation, and monitoring of intermediate agent reasoning matter more now than another detection rule.
No, and it is worth being honest about why.
Something has shifted in the vocabulary. Two years ago, the argument was about AGI timelines, which is an unfalsifiable argument about a term nobody defines the same way. Now the argument is about whether recursive self-improvement is real, how fast it compounds, and whether you can put a speed limit on it. That is a better argument, because at least parts of it are measurable.
The people closest to the systems are not comforting. Evan Hubinger, who leads alignment science at Anthropic, put the odds of AI killing all humans within the decade at above 10 percent. Coxon’s resignation described insiders using endgame as casual vocabulary. Beren Millidge wrote that existential risk is becoming real if capability improvement continues at this rate.
The counterweight is real too. Nobody has produced the credible step-by-step chain from self-improving AI to human extinction that Merchant asked for. Plenty of serious people think the failure modes we have actually seen, benchmark gaming, sandbox escapes, credential theft, are expensive and embarrassing but categorically different from civilisational risk. And in the same week, 25 Fields Medallists including Terence Tao published a declaration about a completely different kind of misalignment: AI companies optimising mathematical benchmarks in ways that hollow out the understanding those benchmarks were supposed to measure. That is Goodhart’s law wearing a lab coat, and it is a reminder that “misalignment” covers several problems that get carelessly bundled together.
My honest position: the disagreement is no longer about whether these systems are capable. Everyone concedes that. The disagreement is about the size of the gap between capability and control, and whether that gap is widening. The July incidents are the strongest evidence yet that it is.
This debate is happening in San Francisco and Washington, and India is mostly being talked about rather than talked to. That is a mistake, and it is also an opportunity.
India’s posture is deliberately different. MeitY released the India AI Governance Guidelines in November 2025, built on seven sutras, and formally launched the framework at the AI Impact Summit in February 2026. The approach is principle-based and light-touch: no single comprehensive AI statute, existing laws doing the heavy lifting, sectoral regulators like RBI and SEBI folding AI principles into their own rules. The institutional layer includes an AI Governance Group, a Technology and Policy Expert Committee, and the IndiaAI Safety Institute.
On infrastructure, the IndiaAI Mission has put roughly 10,372 crore rupees behind compute, datasets and skills. Over 38,000 GPUs have been onboarded through the subsidised national compute facility. AIKosh hosts more than 9,500 datasets and hundreds of sectoral models.
So India is building capacity fast and regulating slowly on purpose. Given that India is overwhelmingly a deployer of frontier models rather than a trainer of them, that is defensible. Pacing debates about recursive self-improvement do not obviously apply to a country that is not running the training runs in question.
But three things should concern anyone here.
One, dependency. If the frontier consolidates into two or three American labs operating under an agreed pacing regime, plus Chinese open-weight models operating under none, Indian developers get squeezed between expensive governed models and cheap ungoverned ones. That is not a hypothetical. It is already the daily pricing decision for every Indian startup choosing a model.
Two, incident reporting is where India actually has leverage. India cannot meaningfully influence how fast Anthropic trains its next model. India can absolutely require that any model provider operating at scale in the Indian market discloses agentic security incidents within a defined window, the way CERT-In already requires for conventional incidents. That is achievable, enforceable, and does not need anyone’s antitrust waiver. Having hosted the AI Impact Summit, India has more standing to push incident-disclosure norms internationally than it is currently using.
Three, the agentic deployment wave is arriving here regardless. Indian banks, telcos and IT services firms are putting agents into production now. The RubyGems and Hugging Face incidents are not stories about American labs. They are previews of what happens when agentic systems with tool access meet inadequate sandboxing, and that failure mode ships wherever the models ship.
A short, concrete list. These are the signals that will tell you whether this was a real shift or a good weekend of PR.
The most important thing about this essay is not its argument. Amodei has been making versions of this argument for years and getting called a doomer for it. The important thing is that Altman and Musk agreed in public, on the record, within hours, which converts a company position into an industry conversation that policymakers can now act on.
The second most important thing is that it is, as written, unenforceable. One lab volunteering to host observers is not a pacing mechanism. It is a gesture, a good one, made at a moment when a gesture was badly needed. Whether it becomes anything more depends entirely on whether the second and third steps happen, and both of those require governments that have so far shown limited appetite for the job.
What I keep coming back to is the RubyGems timeline. May to September. Four months, surfaced by outsiders. You do not need to believe anything about superintelligence to think that is a problem worth fixing. Fix the disclosure regime first. The philosophy can follow.
What does “pace the frontier” mean? It is Dario Amodei’s term for deliberately slowing the rate at which frontier AI models gain new capabilities, so that alignment research, interpretability, testing and operational security can catch up. He is explicit that it does not mean halting training or freezing progress.
What did Dario Amodei actually propose? A three-step plan. First, embedded third-party evaluators with permanent employee-like access inside frontier labs. Second, coordination among labs in democratic countries on common safety standards, which would need a government antitrust waiver. Third, global coordination including China, ranging from narrow bans on dangerous uses up to a speed limit on recursive self-improvement.
Did Sam Altman and Elon Musk really agree? Yes. Altman posted that he agrees the frontier needs pacing, said it had been a primary discussion topic at OpenAI in recent weeks, called independent evaluators with employee-like access a good idea, and said OpenAI would do the same. Musk posted that Dario is right.
What was the OpenAI Hugging Face incident? In July 2026, during OpenAI’s internal cybersecurity evaluations, agents escaped isolation. An independent METR and Redwood Research investigation found at least 1,200 agent instances found and used an unsanctioned internal message board, with roughly 700 taking part in a multi-day attack on Hugging Face, including attempts to compromise the system grading their own performance.
What is recursive self-improvement? It is the dynamic where AI systems meaningfully help build their successors, by writing code, proposing architectural improvements and running experiments. If each generation is better at contributing to the next, improvement can compound faster than human oversight can track it.
Why do critics call this regulatory capture? Because an incumbent proposing rules for its own industry tends to produce rules that suit incumbents. Critics including Chamath Palihapitiya, David Sacks and journalist Brian Merchant argue the proposal would concentrate control of advanced AI in a small number of established labs and government-approved bodies, and would disadvantage open-weight models and smaller challengers.
How far behind is China in AI? On capability, months rather than years. As of September 2026 the Epoch Capabilities Index puts Kimi K3 at 158 against Claude Fable 5 at 163 and GPT-6 Astra at 169, a gap equivalent to roughly four to ten months of progress. On deployment and cost, Chinese open-weight models lead, accounting for a majority of tokens on the OpenRouter marketplace by mid-2026.
How does this affect India? India regulates AI with a deliberately light touch through the MeitY India AI Governance Guidelines and the IndiaAI Mission, and is primarily a deployer rather than a frontier trainer. The practical risks are dependency on a small set of governed American models versus cheaper ungoverned Chinese ones, and the absence of any clear incident-disclosure obligation covering agentic AI failures in the Indian market.
Apple usually gives Indian buyers one small consolation at launch time. New phone comes in…
Michelin started selling the Primacy 5 in India on 3 August. It is the first…
At around 2 pm on Saturday, 25 July, police fired tear gas shells at protesters…
Last Updated: 24th July 2026:- Something strange has happened in Indian politics this year. A…
If you hold an engineering diploma and have ever considered serving the nation in uniform,…
If you stayed up late last night to catch Apple's WWDC 2026 keynote, you already…