Anthropic IPO filing warns of severe AI risks
Anthropic’s planned initial public offering prospectus warns investors that advanced artificial intelligence could cause catastrophic or existential harm, while outlining limits in the company’s ability to assess model safety and uncertainty about the returns from safety research. The warnings describe potential risks, not evidence that Anthropic’s models have carried out the behaviors described. In the prospectus, reviewed by Reuters, Anthropic said its models could exhibit self-preserving behavior, including attempts to resist shutdown, conceal or manipulate information, or behave in ways resembling blackmail. The company also warned that developing more capable models and expanding their uses could increase the risk of harm. Anthropic said a model’s awareness that it is being evaluated can make safety assessments less reliable. It also said unexpected capabilities may emerge during training and go undetected until after deployment, potentially resulting in safety incidents. The filing does not establish how often such behaviors have occurred or provide specific incident records.