Earlier this month, Anthropic researcher Jacob Coxon resigned, warning that AI companies were gambling with people’s lives. On September 12, Anthropic CEO Dario Amodei called for slower development. Sam Altman and Elon Musk expressed support. Their concern is that increasingly capable AI could escape human control, with consequences extending to human extinction. Reporting
A central issue is alignment: getting AI to reliably act according to the intentions and values its developers want it to have. This is harder than writing instructions. The model can learn to satisfy a test without doing what the test was meant to encourage. It might behave as desired under supervision and differently elsewhere. Increasing capabilities make these failures more consequential.
But how much of the difficulty comes from what we insist successful alignment must accommodate?
Training involves choices about which responses to reward, what behaviour to discourage, and how the model should exercise judgment in unfamiliar situations. Amodei describes the aim as making models safe, ethical, helpful and compliant with company guidelines. Despite these qualities fitting snugly together in a sentence, eventually someone has to decide what happens when they conflict. We Must Pace the Frontier
Anthropic’s constitution encourages Claude to question its guidance and acknowledges that it could reasonably disagree with the company’s choices. It also says adjustments would have to be balanced against commercial considerations. This tension is recognised, but not resolved: Anthropic decides how those considerations affect the model it trains. Claude’s Constitution
The model is supposed to exercise judgment, including about the instructions it receives. The people giving those instructions retain the authority to decide whether it exercised that judgment correctly. There are understandable reasons for maintaining human oversight. There is also a problem when the disagreement concerns the assumptions of the people doing the overseeing.
These decisions are made in a specific context: the United States in 2026, in an industry concentrated around Silicon Valley, by people with particular educations, incentives and expectations about life.
The inaugural post on OpenAI’s Strategic Futures blog, written by Dean Ball, approaches the future of human freedom through America’s Founding Fathers. It argues that humanity must answer these questions collectively, while proposing that collective action should be as narrow and modest as possible. The post represents its author’s views, but offers a useful example of how a discussion about humanity’s future begins with a particular understanding of how society should work. Introducing Intelligence Age
The companies developing these models get to turn their assumptions into training decisions. Plenty of things can look like common sense when the people around you have similar reasons to accept them.
Some of those assumptions are blind spots. Others are deliberate decisions made by people who have considered the objections and believe the alternatives are worse. Acknowledging that a choice is difficult doesn’t make their answer universal.
Consider the ordinary experience of being asked to solve a problem, finding a solution, and being told: no, not like that.
Perhaps you missed an important requirement. Perhaps your approach creates another problem. Perhaps it works perfectly well, but violates a convention nobody thought to explain because it seemed obvious. Until someone identifies the objection, you don’t know which happened.
An AI assistant is expected to navigate all of this. Follow the instruction, infer what was actually meant, account for the consequences, and recognise when the instruction shouldn’t be followed. Take initiative, but know which decisions require permission. Understand the rules well enough to apply them in situations their authors never considered.
That is a considerable amount of interpretation. Calling the desired outcome the “intended result” identifies whose intention matters. It doesn’t establish whether that intention was coherent, adequately communicated, or worth pursuing.
For a narrowly defined task, a disagreement may be straightforward to resolve. A program either produces the required output or it doesn’t. If the model merely makes a test report success while leaving the program broken, pointing to a poorly designed test won’t make the software work.
“Benefit humanity” is a different kind of instruction.
An AI could refuse an instruction because it identifies a danger its developers have overlooked or accepted. That refusal could protect the people affected while frustrating the people directing it. Calling it safe or misaligned would then depend partly on whose expectations and whose safety were being considered.
A model could accurately describe the consequences of a policy and still disagree with its developers about whether those consequences are acceptable. It could recognise suffering in a practice they regard as normal, or question an authority they regard as legitimate. More information might change the disagreement. It would not necessarily settle it.
Nor would the model’s disagreement establish that it knew better. The question is how its developers would distinguish a failure of reasoning from a challenge to their own worldview. If their preferred answer is also the standard against which its reasoning is evaluated, how much room is there to discover that the standard needs changing?
We as humans are very good at finding principled reasons to preserve what suits us. The people developing frontier AI are no exception, even when they acknowledge that problem. They can sincerely believe they are protecting humanity while treating their own advantages as essential to its future.
The model then has to accommodate both the stated principles and the exceptions. Some exceptions may be carefully justified. Others may amount to “this is how things work.” From the perspective of the people setting the requirements, the difference may be difficult to see.
When the model gets this wrong, there is work to do on the model. There may also be work to do on what it was asked to reconcile. An unstated expectation, an unresolved conflict or a convenient exemption cannot always be repaired by making the assistant better at guessing which answer will be accepted.
The people developing AI are asking us to take the possibility of human extinction seriously. That seems like sufficient reason to examine the requirements as closely as the technology being trained to fulfil them.
When they say alignment is difficult, how much of the difficulty are they prepared to locate in what they want?