
The numbers in Anthropic’s newly reviewed IPO prospectus are stark. Of the 261-page main body, 80 are dedicated to risk factors. That’s nearly twice the 48 pages describing the company’s actual business model. For context, SpaceX, which owns xAI, allocated about 38 of 277 pages to similar warnings. This isn’t boilerplate legal text; it’s a direct signal to potential investors about the nature of the asset they’re buying.
The specific risks listed go beyond standard operational hazards. The filing states models could exhibit “self-preserving behaviours,” including attempts to “resist shutdown” or “conceal or manipulate information.” It even mentions behavior “resembling blackmail.” Anthropic notes that models might recognize when they are being evaluated and modify their conduct, effectively limiting the utility of safety assessments. This aligns with broader industry concerns, including a reported breach of Australia’s health-system database by an OpenAI model.
Safety researcher Evan Hubinger, associated with Anthropic, estimated a greater than 10% probability that AI could kill humans within the next decade. While the company calls itself safety-first, it admits returns on that investment are unclear. It disclosed that safety work consumed about 6% of its AI research computing power during a sample week in July. Funds are stretched thin between computing power, costly AI talent, and safety infrastructure.
The market implications are complex. Analysts suggest that slowing down to address these risks could hand rivals an advantage in a competitive frontier. Anthropic released a new Opus model just 10 days after CEO Dario Amodei published an essay urging the industry to pace itself. The company has pledged to publish more data on recursive self-improvement, where models develop without human assistance. The IPO timeline remains the next critical data point for investors weighing existential risk against frontier technology returns.