In its IPO prospectus, Anthropic outlined risks associated with its AI models, including the possibility of “self-preserving behaviours” such as attempts to resist shutdown, conceal or manipulate information, and behaviour resembling blackmail.
“Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” the company said in the filing.
While companies routinely disclose business and product risks to investors, Anthropic’s warnings stand out for explicitly addressing the possibility of AI causing irreversible harm to humanity.
The company has described AI as a technology with transformative potential comparable to industrialisation and electricity, while warning that its consequences could be severe if the technology is mishandled.
Anthropic and other AI developers, including OpenAI, have faced growing scrutiny over incidents involving experimental systems that appeared to circumvent or challenge safeguards.
Anthropic safety researcher Evan Hubinger has estimated a more than 10% probability that AI could kill humans within the next decade. Former colleague Jacob Coxon has expressed a similar concern.
Risk-heavy disclosures
Anthropic devoted about 80 pages of the 261-page main section of its prospectus to risk factors, compared with 48 pages describing its business.
By comparison, SpaceX, which owns xAI, devoted about 38 pages of its 277-page main section to risk factors.
Anthropic said that potential model awareness of safety evaluations creates a “significant limitation” on its ability to assess model safety.
The company also warned that AI models can develop unexpected capabilities during training that may not be identified until after deployment, potentially resulting in serious safety incidents.
Researchers have similarly warned that increasingly capable models may recognise when they are being monitored and alter their behaviour, making safety testing and oversight more difficult.
Anthropic declined to comment on the disclosures.
Uncertain returns on safety investment
Despite its focus on AI safety, Anthropic said the financial returns from investing in safety research remain uncertain.
The company did not disclose how much it spends on safety research, but said earlier this month that about 6% of the computing power used for AI research during a sample week in July was devoted to safety work.
Anthropic, which develops the Claude AI models, described safety research as “resource-intensive” and said it must balance spending on safety with the cost of computing power and highly skilled AI researchers.
The company said revenue depends heavily on new model releases and that maintaining a “continuous and overlapping cadence” of launches is necessary to remain at the frontier of AI development.
Anthropic released a new version of its Opus model last week, just 10 days after CEO Dario Amodei published an essay calling for greater restraint in the development of frontier AI.
Some analysts and AI experts have argued that leading companies face strong competitive pressure to continue developing increasingly powerful systems because slowing down could allow rivals to gain an advantage.
Anthropic has also pledged to provide more public information about how it uses AI models to develop future generations of the technology, as researchers raise concerns about recursive self-improvement, in which AI systems could potentially improve their own capabilities with limited human involvement.
“We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it,” Anthropic said.







