What AI Safety Needs, as UNGA81 Opens
It has been a crazy two weeks for AI news. Jacob Coxon resigned from Anthropic, saying the people building AI believe it could kill us all by the end of the decade. An Anthropic safety researcher put the chance at more than 10% within ten years. Dario Amodei then asked the industry to slow down in "We Must Pace the Frontier". The attention was so intense that, like most colleagues in this field, I was swamped with requests for comments. I had to write a ready-to-go answers document, or there would not have been enough hours in the day.
The UN General Assembly's high-level week begins against this backdrop, with AI high on the agenda. The Secretary-General says it will be a major topic in his talks with world leaders, and the Security Council will be meeting on Wednesday to discuss AI. The narrative has reached the very top: António Guterres has listed "runaway artificial intelligence" among three existential threats to humanity. But this is the industry's framing. It makes their systems sound more powerful, it attracts investment, and it locates the danger in the technology rather than in the decisions companies make about how to build, test and deploy it. When the world's highest diplomatic forum repeats that framing, it does not just misdescribe the problem. It pre-selects the response, since a runaway technology calls for taming the technology, not for holding the companies that release it accountable.
I have written before about why "runaway AI" is the wrong description of what is happening. Nothing is escaping our control. What we are letting escape is accountability, for the companies that build and deploy these systems. That framing produces two answers, and I think both miss the point.
Two answers, one mistake
The existential framing produces two answers. One is to pause. The other is to build safety into the models. Russell proposes machines that defer to human preferences. Bengio proposes non-agentic "Scientist AI". Tegmark and Omohundro call provably safe systems "the only path to controllable AGI". Most claim that this is not sufficient on its own. Russell calls for governance alongside it. The Bengio-led Science paper asks that at least a third of AI research budgets go to safety, and for adaptive regulation. Even so, the big AI developers are trying, and largely succeeding, to sell us the simplified version: that if we specify, verify and prove hard enough, the problem is largely solved. I.e. that tech can solve what tech messes up.
I also want more formal verification, not less. But a proof about the artefact says little about the societal context it is deployed into. Harms arise in interaction with people, organisations and other systems that no world model can fully contain. A verified model still needs someone to decide what it is allowed to do, to whom it answers, and what happens when its guarantees do not hold in practice, which is exactly the accountability question the "runaway" framing avoids.
A pause fails for a different reason. Speed is not the problem. The lack of engineering discipline is. An agreement on pace among a few firms would be treated as a cartel arrangement in any other sector, and it leaves accountability untouched. Both answers place safety in the artefact or in its developers' tempo.
What cars teach us
A car can reach 250 km/h, far more than safe use requires. We add brakes, which are part of how it works. We add seatbelts and airbags, which serve only safety. And we add what is not part of the car at all: traffic rules, driver licensing, road design, inspection, insurance and liability law. Safety is an emergent property of the whole system and must be controlled at that level (Leveson, 2012). In complex, tightly coupled systems, accidents will happen, no matter how much safety is designed in (Perrow, 1984).
We also build crumple zones, emergency services and crash investigation, because we assume crashes will happen. My doctoral work on agent organisations started from the same assumption (Dignum, 2004). Autonomous agents cannot have every norm hard-wired, so we must define what counts as a violation, how it is detected, what repair follows, and who is accountable. An approach that builds safety only into the design has nothing to say about failures that occur in use.
There is a further risk: safety features change behaviour and shift risk onto others. A model certified as safe can become an alibi for the system around it, the same way a pause becomes an alibi for inaction, both framing safety as a question of speed rather than of who is accountable.
What UNGA81 could deliver
Not a pause, not a treaty. A realistic and useful outcome is narrow and boringly institutional:
- A mandate for an independent international mechanism to report and investigate serious AI incidents, modelled on aviation, where accident investigation is independent and findings are shared and feed back into rules.
- A time-bound process to agree a short list of red lines: not properties to engineer into a model, but uses states commit not to permit, such as AI in nuclear command and control or fully autonomous lethal weapons, backed by a way to check the commitment is kept.
Neither item is dramatic. That is the point. Aviation did not become far safer through better engines, or through a plea to pilots and airlines to fly more carefully. It became far safer because every serious incident is investigated by a body independent of the airlines, the findings are made public, and they change the rules. AI has no such institution anywhere.
I hope, but doubt
I would like nothing more than to see either of these agreed this week. I doubt it. Watch what actually happens in New York: speeches naming AI as an existential threat, a Security Council session on Wednesday, side events, declarations. What you will not see is a negotiating text similar to the proposals above. The Global Dialogue on AI Governance, the body meant to carry this forward, is a forum, not a lawmaker. It cannot bind anyone, and its next session is not until May 2027. The Scientific Panel can assess but has no mandate to investigate. Incident reporting would ask developers and states to disclose their own failures, and both have reasons not to. SG Guterres himself has pushed for a binding instrument on autonomous weapons before his term ends this December, a place where setting red lines would fit. But it competes for space with war talks, a leadership transition, and everything else on this week, and red lines matter most exactly where states are least willing to accept scrutiny. The likely result is a paragraph in a declaration welcoming the Panel and the Dialogue and restating the importance of safety. Worth having, but a hope, not a mechanism.
What we can
However, accountability does not have to wait for New York. The next time a headline says an AI system did something alarming, ask who deployed it, with what permissions, and who is answerable, before asking what the model might do next. That question belongs with the company, not with the system. And we can all stop spreading the fear, and start asking who is accountable instead.
References
- Bengio, Y. et al. (2024). Managing extreme AI risks amid rapid progress. Science, 384(6698), 842–845. https://doi.org/10.1126/science.adn0117
- Dignum, V. (2004). A Model for Organizational Interaction. PhD thesis, Utrecht University. https://dspace.library.uu.nl/handle/1874/890
- Leveson, N. (2012). Engineering a Safer World. MIT Press. https://doi.org/10.7551/mitpress/8179.001.0001
- Perrow, C. (1984). Normal Accidents: Living with High-Risk Technologies (updated ed., Princeton University Press, 1999). https://press.princeton.edu/books/paperback/9780691004129/normal-accidents
- Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking. https://en.wikipedia.org/wiki/Human_Compatible
- Tegmark, M., and Omohundro, S. (2023). Provably safe systems: the only path to controllable AGI. arXiv:2309.01933. https://arxiv.org/abs/2309.01933