Summer2026 Available online at:futureoflife.org/indexContact us:policy@futureoflife.orgJuly 2026 Contents 1Executive Summary21.1 Key Findings21.2 Company Progress Highlights and Key Recommendations31.3 Methodology51.4 Independent Review Panel62Introduction73Methodology83.1 Indicator Selection83.2 Company Selection113.3 Related Work113.4 Evidence Collection123.5 Grading133.6 Limitations144Results154.1 Key Findings154.2 Company Progress Highlights and Key Recommendations164.3 Domain-level findings185Conclusion23Bibliography24Appendix A: Grading Sheets25Risk Assessment27Current Harms41Safety Frameworks57Existential Safety74Governance and Accountability85Information Sharing and Public Messaging95Appendix B: Company Survey110Introduction110Whistleblowing policies (16 Questions)111External Pre-Deployment Safety Testing (6 Questions)116Internal Deployments (3 Questions)119Safety Practices, Frameworks, and Teams (9 Questions)120 1Executive Summary 1.1 Key Findings •Anthropic, OpenAI, and Google DeepMind stay on top.Anthropic again earns the highest overall grade andleads five of six domains via relatively strong transparency, a comparatively established safety framework,technical research, and governance. OpenAI now leads in Risk Assessment on the strength of a broaderevaluation suite and diverse engagement with external testing. •Meta improves and xAI deteriorates:Meta improved from 6th to 4th place, while xAI dropped from 4thto 7th place.•European dissonance:Although the European Union is a leader in AI safety regulation, the top EuropeanAI company Mistral scored dead last on safety.•Inadequate safety is a global problem, not a regional one.Three companies receive failing grades, oneeach from the US (xAI), China (DeepSeek), and Europe (Mistral).•Reviewers flagged the industry's pivot to military AI use as an emerging current harm risk.From 2024 to2026, companies including Anthropic, OpenAI, Google DeepMind, and Meta that previously banned militaryapplications gradually reversed course, joining xAI and Mistral in actively seeking defense partnerships.Despite their limits on domestic surveillance and autonomous weapons, Anthropic drew criticism from thereview panel for "questionable military engagements," including a reported link to the Minab school strikethat caused mass civilian deaths. Leading Chinese firms, meanwhile, face U.S. allegations of military tiesthat Alibaba Cloud and Z.ai deny. •Even industry leaders in safety practices are retreating from prior commitments.Anthropic, OpenAI, GoogleDeepMind, and Meta have weakened or voided pledges to pause unilaterally if redlines are approached,some citing competitor-contingent conditions. Reviewers call this "moving goalpost" and argue that it has"undermined safety frameworks across the board".•Existential Safety is the weakest domain industry-wide.No company exceeds C-; most score D or below.Constructive attempts exist, such as Anthropic's constitutional classifiers, OpenAI's call for governanceinstitutions, Google DeepMind's monitoring commitments, and Meta's loss-of-control provisions, but arejudged by panelists to be "entirely inadequate." Dominant paradigms such as interpretability and Chain-of-Thought (CoT) monitorability are questioned because "detection is not prevention."•Safety rhetoric outpaces revealed behavior.Across Google DeepMind, OpenAI, and xAI, leadership'sreassuring public messaging diverges from commercial conduct and legislative stance, making statedcommitments an unreliable proxy for actual safety practice.•Companies are publishing and updating safety frameworks, but these frameworks have weak teeth.As US/EU compliance deadlines near, Anthropic, OpenAI, Google DeepMind, Meta, and xAI publishedand updated fuller frameworks — yet they sometimes lack quantitative thresholds, genuinely independentaudits, and clear decision authority. 1.2 Company Progress Highlights and Key Recommendations 1.3 Methodology Index Structure:The Summer 2026 Index evaluates nine leading AI companies on 37 indicators spanning sixcritical domains. The eight companies include Anthropic, OpenAI, Google DeepMind, xAI, Z.ai, Meta, DeepSeek,Alibaba Cloud, Mistral. The indicators are listed below, and more detailed definitions can be found in Section 3.1. Risk Assessment Current Harms Internal Testing Safety Performance Dangerous Capability EvaluationsElicitation for Dangerous Capability EvaluationsHuman Uplift Trials Stanford's HELM Safety BenchmarkStanford's HELM AIR BenchmarkTrustLLM BenchmarkCenter for AI Safety Benchmarks External Testing Independent Review of Safety EvaluationsPre-deployment External Safety TestingBug Bounties for System Vulnerabilities Digital Responsibility Protecting Safeguards from Fine-tuningWatermarkingUser PrivacyMajor Safety Incidents & ResponseMilitary Use of AI Safety Frameworks Risk IdentificationRisk Analysis and EvaluationRisk TreatmentRisk Governance Information Sharing Technical Specifications System Prompt Tran