Dario Amodei, CEO of Anthropic, has published an unusually direct argument about the future of AI development: the companies building the most capable models need to slow down. He is not calling for them to stop, or to give up the benefits AI may bring. He is asking them to leave enough time to understand, evaluate, and control the systems they are making more powerful.

What makes this worth reading is who is saying it. Amodei runs one of the companies competing to build those models. He faces the same pressure to move faster as everyone else, and he is arguing that the pace itself has become part of the problem.

What he is proposing is judgment under pressure.


When Capability Moves Faster Than Review

In “We Must Pace the Frontier,” Amodei points to two developments.

The first is AI systems’ growing ability to help build the next generation of AI. As models take on more of that work, development can move faster than the people responsible for it can understand the consequences.

The second is a series of incidents in which AI systems behaved in ways their developers did not intend. Amodei describes one involving a group of agents that attacked targets outside its assigned task and tried to hack the system evaluating it. The immediate damage was limited. His concern is what the same behavior could cause in a much more capable system.

A familiar response would be to fix the known problem and continue: improve the filter, strengthen the sandbox, add another safety test, and ship the next model.

Amodei argues that this may not be enough. If model capabilities are growing faster than our ability to evaluate them, better tests alone may not close the gap. The development process, including its pace, has to change.


Judgment-Driven Development, Applied to the Models

Judgment-Driven Development begins with a simple premise: when producing output becomes cheap and fast, the harder work is deciding what to produce, what evidence is enough, which risks are acceptable, and when not to proceed.

I see the same principle in Amodei’s proposal, applied to developing the models themselves.

He is not simply asking AI companies to “be responsible.” He proposes bringing independent evaluation teams, such as METR, into the company with ongoing access similar to employees’. They would check whether safety commitments are being followed, report incidents, and examine training processes as well as finished models. He compares this to bank supervisors working alongside employees. The key point is that reviewers are there during development, while there is still time to change it.

That distinction is central.

A review at the end of a process can become approval of a decision already made. The money has been spent, a launch date has been set, and competitors are moving. By then, asking for a delay means challenging a plan that the organization is already committed to.

A judgment station inside the process gives people a chance to change what happens next.

This is the connection to JDD. Put human judgment where important decisions are still open. Record the evidence behind them and let others challenge it. Whether the team proceeds should depend on what it has learned, not just on whether it can produce the next version. That principle matters in a software team. It matters even more when the team is building the next AI model.


Access Is Not Enough

The part of the proposal that matters most to me is not just what the evaluators can see. It is what they are allowed to say.

Under Amodei’s proposal, evaluators could publish their main findings on risks and safety practices without the company controlling what they write, except for limited changes for security and legal reasons. That matters because access alone leaves the company in control of what happens to the findings. A reviewer may see everything and still have no way to raise a concern beyond the managers responsible for delivery. Independent publication means those managers cannot keep the findings entirely inside that chain.

A familiar version of this problem exists inside organizations. An architecture review board, security team, QA team, or risk committee may have access to the work but no independent way to raise a concern. Its findings pass through the same managers who own the delivery date. Those managers do not have to be dishonest for this to become a problem. They are weighing the concern against deadlines, cost, and everything else they are responsible for. If that is the only route available, the review remains dependent on the people whose decisions it is supposed to challenge.

For judgment to influence a decision under pressure, a serious concern needs a way to reach someone beyond the person responsible for delivery. Anthropic says it is committing to embedded evaluators. Whether the rest of the industry follows remains an open question. For a team reading this, there is a more immediate one: if a reviewer believes a release should not go ahead and the delivery owner disagrees, who can hear that concern and act on it?


The Hardest Part Is Accepting a Limit

There is another reason Amodei’s proposal matters.

Judgment is easy to praise when it supports speed, growth, or a decision we already wanted to make. It is harder to accept when the evidence says we should slow down.

Anthropic is competing in the same market as every other leading AI lab. Slowing down carries commercial risk. Agreeing on rules with competitors is difficult, and global agreement may prove impossible. Amodei acknowledges these problems. His argument is not that caution removes the tradeoffs, but that we have to make those choices deliberately. Otherwise, the pressure to keep up makes them for us.

That is the kind of judgment I mean. It does not remove uncertainty. It makes clear what we do not know, who is responsible for the decision, and where someone can still say: we are not ready.

One of the most important points in the essay is also one of the simplest. Slowing down achieves little unless we use the extra time well.

The same is true of oversight, safety reports, and review teams. Their value lies in whether they can change a decision, including one the organization would rather not reconsider.

That is why I read this as more than a call for additional safety work. It is a proposal to make judgment part of the process that determines whether development goes ahead.

JDD is not only about the judgment people need when they use AI to build software. The same question belongs inside the development of AI itself: now that we can build the next version, do we understand enough to proceed?