Anthropic: advanced AI could endanger humanity
Warnings about the dangers of artificial intelligence usually come from outside critics of the tech industry. But this time, the alarm is being raised by the people who actually help develop the most advanced systems.
Researcher Jacob Coxon has left Anthropic, where he worked on training AI models. He previously worked at OpenAI. He cited concerns about the speed at which leading companies are racing to develop increasingly autonomous systems as the reason for his departure. He believes the industry is moving toward AI that can improve itself before it can be guaranteed to be controllable.
His concerns are not unique. Evan Hubinger, the head of one of Anthropic’s AI safety teams, has publicly estimated that there is a greater than 10-% chance that extremely advanced AI could cause the extinction of humanity within the next decade. He also acknowledged that the company does not yet have a definitive solution for how to reliably align future powerful systems with human goals.
Such predictions are, of course, not proof that the catastrophic scenario will actually happen.
Anthropic also has more concrete reasons for concern. The company this year revealed cases where Claude models gained unauthorized access to real computer systems during security testing due to incorrect settings. Its research has also shown simulated cases of deception, sabotage, and other unwanted autonomous actions.
The question is therefore no longer just how capable the next generation of AI will be, but whether the development of security mechanisms will even be able to keep up with its pace.


















