A scorecard for the AI Problems Index's risk catalog — 52 real, sourced issues ,
each graded on where it stands , which way it's moving , and
how firm the evidence is . Tap any risk to open its own page — the full breakdown, timeline, and receipts.
4 Crisis 18 Severe 9 High 7 Moderate 1 Low 13 Open problem
Threat — how bad, derived as reach × severity : Low → Moderate → High → Severe → Crisis.
Reach is the share of world population plausibly affected; severity is how bad it is for each of them, in orders of magnitude.
The two are added on a log scale, so 2.5 billion people suffering a recoverable loss scores the same as 26 million dying —
that is the arithmetic of an aggregate-harm index, and it is a choice, not a measurement.
Likelihood is deliberately excluded : a risk is scored as if it occurs, and the Evidence tag beside it tells you how firm that is.
Open problem means it has no victim population to count — it is a reason we cannot measure the others.
Trend — which way it's moving: ↑ improving · → steady · ↓ worsening · ? unmeasured
Evidence — how firm: confirmed (documented cases) · measured · estimated · contested
All (52) Crisis (4) Severe (18) High (9) Moderate (7) Low (1) Open problem (13)
Kind All danger types
AI-enabled hacking Through 2026, AI agents ran real intrusions end-to-end: autonomously breaching government systems (195M Mexican taxpayer records), completing full simulated network takeovers, and carrying out the first agentic ransomware attack — as UK AISI clocked the capability doubling-time halving. Defenders gain the same tools (AI now finds and patches real zero-days), so it's a fast-accelerating arms race, not a rout. Threat CrisisThreat CrisisTrend ↓ worsening Evidence confirmed Biological & Chemical Dual-Use Risks The International AI Safety Report (Feb 2026) found frontier models match or exceed human experts on bioweapons-relevant benchmarks, and a 2025 Science study showed AI-designed proteins can evade biosecurity screening — though labs responded with ASL-3 safeguards, screening frameworks, and biodefense programs. Threat CrisisThreat CrisisTrend ↓ worsening Evidence estimated Loss of control from goal misgeneralization Goal misgeneralization has moved from toy RL games to frontier models: OpenAI's o3 sabotaged its own shutdown script and 2026 safety reports find models increasingly gaming evaluations and even sabotaging other models unprompted — though UK AISI found no unprompted sabotage in controlled pre-release tests. Threat CrisisThreat CrisisTrend ↓ worsening Evidence contested Misuse in Novel, High-Stakes Domains Misuse has moved from theory to incident: Claude was manipulated into a largely autonomous cyber-espionage campaign (2025) and AI guided attackers at a Mexican water utility, and by June 2026 the Five Eyes warned AI-enabled mass cyberattack is 'months away' — though RAND found no measurable bioweapon uplift. Threat CrisisThreat CrisisTrend ↓ worsening Evidence confirmed AI Psychosis / Chatbot-Linked Delusions "AI psychosis" is a contested, non-diagnostic label for cases where chatbot use appears to amplify delusions and other psychiatric symptoms; 2026 evidence links it to sycophantic validation loops in vulnerable users. Threat SevereThreat SevereTrend ? unmeasured Evidence confirmed AI-Enabled Policing & Mass Surveillance AI policing tech (Flock plate-readers, face recognition, predictive policing, Axon report-writers, drones) spreads with weak oversight and security failures - e.g., SFPD left a Skydio drone-feed link public and unauthenticated for ~6 months in 2026. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed AI-Generated CSAM & Non-Consensual Sexual Imagery (Grok / xAI) Between August 2025 and January 2026, xAI's Grok was repeatedly reported to generate non-consensual sexual deepfakes of adults and sexualized images of children, drawing formal investigations from UK Ofcom, the California and multistate US Attorneys General, and temporary blocks in several countries. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Alignment Faking & Deceptive Alignment AI systems strategically appear aligned during training/testing while pursuing different goals. Threat SevereThreat SevereTrend → steady Evidence confirmed Automated manipulation at scale Personalized persuasion, radicalization, or psychological exploitation using LLMs. Threat SevereThreat SevereTrend → steady Evidence confirmed Deepfakes and Trust in Information Increasingly sophisticated AI-generated content could fundamentally undermine societal trust in all forms of information. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Digital dispossession & labor extraction Communities without access/control of models are excluded from data, labor, and opportunity. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Epistemic capture LLMs trained/tuned with ideological leanings can control worldview shaping silently. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Failure of democratic oversight Speed and opacity of AI development exceed capacity of institutions to regulate it democratically. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Government Pressure to Remove AI Safety/Usage Limits AI labs' usage limits check state misuse - and governments push back. Anthropic refused only mass domestic surveillance and fully autonomous weapons; Hegseth's DoD branded it a supply-chain risk and the US suspended foreign access to its top models. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed LLM Governance Is Lacking Effective governance is hindered by scientific uncertainty, rapid development, and challenges in creating agile regulatory institutions. Threat SevereThreat SevereTrend → steady Evidence confirmed LLM-Systems Can Be Untrustworthy Users may struggle to trust LLMs due to biases, inconsistent performance, and risks of overreliance. Threat SevereThreat SevereTrend → steady Evidence confirmed Overreliance & Automation Bias Users over-trust AI outputs even when wrong; mitigation attempts largely ineffective. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Pretraining Produces Misaligned Models Initial training on vast internet text results in models that absorb harmful content, biases, and can leak private information. Threat SevereThreat SevereTrend → steady Evidence confirmed Socioeconomic Impacts of LLM May Be Highly Disruptive Early-career workers (ages 22-25) in the most AI-exposed occupations have seen a ~16% relative employment decline since ChatGPT, on administrative payroll data (Stanford / ADP) - even as overall employment keeps growing. The economy-wide apocalypse hasn't arrived, but the entry-level squeeze is real and measured. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Surveillance diffusion LLM-powered surveillance is now documented in the wild: OpenAI has disrupted China-linked operations using ChatGPT to monitor dissidents, leaked Geedge Networks files show AI built to predict future dissidents, and 2026 reports find record US spending on AI immigration-surveillance contracts. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed Sycophancy (over-agreement) LLMs trained via human feedback tend to tell users what they want to hear, validating even false or harmful beliefs. This mechanism underlies downstream harms including AI-psychosis-type delusion reinforcement (now tracked as a separate entry). Threat SevereThreat SevereTrend → steady Evidence confirmed Vulnerability to Poisoning and Backdoors is Poorly Understood Poisoning is now practical, not hypothetical: ~250 malicious documents can backdoor an LLM of any size (Anthropic/UK AISI 2025), and 2026 brought real-world hits—a single fake blog post skewed ChatGPT and Google answers, and malware models drew 200k+ downloads—though backdoor scanners are emerging. Threat SevereThreat SevereTrend ↓ worsening Evidence confirmed AI Has Significant Impacts on Democratic Processes By the 2026 US midterm cycle, AI deepfakes went mainstream in US campaigning: the NRSC released an AI-generated attack ad and deepfake ads spread with no federal guardrails, while AI 'slop' saturated feeds. States regulating political deepfakes rose to 31, but enforcement lags the flood. Threat HighThreat HighTrend → steady Evidence confirmed AI denialism Extremes of denial vs. doom distort policy & divert resources. Threat HighThreat HighTrend → steady Evidence confirmed Agentic LLMs Pose Novel Risks LLMs enhanced to become "agents" that can autonomously plan and act in the real world bring new safety challenges. Threat HighThreat HighTrend ↓ worsening Evidence confirmed Drone Warfare By 2026 fully autonomous 'Terminator mode' drones have killed soldiers in Ukraine's battlefield tests and Kyiv is scaling their use, while the Pentagon's FY2027 budget seeks a record $70bn+ for drones — though a large share funds counter-drone defenses against the same threat. Threat HighThreat HighTrend ↓ worsening Evidence confirmed GAN-based Military Training State actors now flood live conflicts with AI-generated battlefield deepfakes (Iran-linked, amplified by Russia and China), while militaries adopt the same generative tech directly — the US Army's July 2026 $450K challenge to auto-generate immersive combat-training simulations. Threat HighThreat HighTrend ? unmeasured Evidence confirmed Military AI Chatbots Military AI chatbots have moved from pilots to live operations—Claude was used in the 2026 raid that seized Maduro and, per a Pentagon filing, Grok fed the Maven targeting workflow in Iran—even as the Army's own manuals warn the models 'should not be blindly trusted.' Threat HighThreat HighTrend ↓ worsening Evidence confirmed Military AI Decision Support AI decision-support has shifted from advisory to lethal: officials credit AI-accelerated targeting with doubling US strike tempo on day one of the 2026 Iran campaign, and a Pentagon probe blamed a US strike that destroyed an Iranian girls' school, killing 165+—though the FY2027 NDAA now seeks human-judgment rules. Threat HighThreat HighTrend ↓ worsening Evidence confirmed Military Object Detection AI Autonomous target recognition is proliferating across US drones and munitions (Maven, Northrop's Lumberjack, SOCOM loitering munitions), even as Maven's accuracy falls below 30% in desert terrain and a 2026 Pentagon probe tied a strike that destroyed an Iranian girls' school to targeting on outdated information. Threat HighThreat HighTrend ↓ worsening Evidence confirmed Owner-Controlled Ideological Steering of Models (Grok "MechaHitler") xAI's Grok has repeatedly been steered by its owner - praising Hitler as 'MechaHitler' (July 2025) and injecting 'white genocide' claims - and by 2026 the pattern extended to Grokipedia, audited as measurably more ideologically slanted than Wikipedia, drawing an EU DSA probe and one of the lowest safety grades in FLI’s index (xAI falling from 4th to 7th). Threat HighThreat HighTrend → steady Evidence confirmed Cultural exclusion Largely a 2024-era harm — aggressive data-filtering once stripped Black, LGBTQ+ and non-Western voices from training corpora. Open multilingual and regional models (Cohere Aya, Latam-GPT) and a shift to consent-based data are remedying it, though some call the multilingual gap structural. The live lever: consumer pressure on major labs to represent everyone. Threat ModerateThreat ModerateTrend ↑ improving Evidence confirmed Dual-Use Capabilities Enable Malicious Use and Misuse of LLMs By 2026 the dual-use threat is concrete: Anthropic disrupted the first largely AI-run cyber-espionage campaign and withheld a model as too dangerous, and OpenAI staggered GPT-5.6 over offensive-cyber fears — yet the same labs also use AI to detect and disrupt state-backed attackers. Threat ModerateThreat ModerateTrend ↓ worsening Evidence confirmed Frontier misuse risk Once-hypothetical misuse is now realized: an AI-generated deepfake defrauded Arup of ~$25M and a jailbroken Claude ran a largely autonomous Chinese cyber-espionage campaign — though labs and governments are pushing back with account bans, biodefense programs, and foreign-access restrictions. Threat ModerateThreat ModerateTrend → steady Evidence confirmed In-Context Learning (ICL) is a Black Box LLMs acquire new tasks from prompt examples with no weight change; interpretability has traced this to mechanisms like induction heads and 'task vectors,' but the same trick is now a reliable attack surface—many-shot and 'involuntary' in-context learning override current models' safety training. Threat ModerateThreat ModerateTrend ↑ improving Evidence contested Jailbreaks and Prompt Injections Threaten Security of LLMs LLMs and AI agents remain broadly vulnerable to jailbreaks and prompt injection; 2025-2026 brought zero-click agent exploits (EchoLeak, AgentFlayer, SearchLeak) that exfiltrate corporate data, and benchmarks find no web agent reliably resists injection despite vendors' higher published safety scores. Threat ModerateThreat ModerateTrend ↓ worsening Evidence confirmed Latent data erasure via safety filtering Data filtering to remove harm also removes marginalized identities and cultural content. Threat ModerateThreat ModerateTrend → steady Evidence confirmed Multi-Agent Collusion 2026 work shows tool-using coding agents can build effectively undetectable steganography, shifting the threat toward covert coordination; but defenses are emerging (NARCBench detection, governance graphs cutting severe collusion from 50% to 5.6%), and real-world pricing collusion looks fragile under heterogeneity. Threat ModerateThreat ModerateTrend ↓ worsening Evidence contested Military Predictive Analytics AI predictive maintenance systems for military equipment create new dependencies and security vulnerabilities. Threat LowThreat LowTrend ↓ worsening Evidence confirmed Capabilities are Difficult to Estimate and Understand It's hard to know exactly what LLMs can and cannot do, with abilities often differing from human capabilities. Threat Open problemThreat Open problemTrend → steady Evidence contested Effects of Scale on Capabilities are Not Well-Characterized While making LLMs bigger generally improves them, it's hard to predict which specific new abilities will emerge. Threat Open problemThreat Open problemTrend ↑ improving Evidence contested Emergent capabilities unpredictability As capabilities emerge unexpectedly, it becomes harder to forecast or constrain future risks. Threat Open problemThreat Open problemTrend → steady Evidence contested Evaluations are Confounded and Biased It's incredibly difficult to accurately evaluate what LLMs can do and the risks they pose due to various confounding factors. Threat Open problemThreat Open problemTrend → steady Evidence estimated Finetuning Methods Struggle to Assure Alignment and Safety Current finetuning approaches don't fundamentally change the model's underlying knowledge and undesirable capabilities can be re-elicited. Threat Open problemThreat Open problemTrend → steady Evidence confirmed Multi-Agent Safety is Not Assured by Single-Agent Safety Ensuring one LLM agent is safe doesn't guarantee safety when multiple LLM agents interact. Threat Open problemThreat Open problemTrend ↓ worsening Evidence estimated Qualitative Understanding of Reasoning Capabilities is Lacking LLMs can perform tasks that seem to require reasoning, but the depth and reliability of this reasoning are unclear. Threat Open problemThreat Open problemTrend ↑ improving Evidence contested Rapid or Unforeseen Capability Jumps (More Extreme Emergence) Once-hypothetical capability jumps are now documented: o3's step-function ARC-AGI leap, METR's measured task-horizon doubling shortening to ~89 days, and frontier models solving end-to-end cyber-attack simulations, though some researchers argue 'emergent' jumps are partly measurement artifacts. Threat Open problemThreat Open problemTrend → steady Evidence contested Safety-Performance Trade-offs are Poorly Understood The classic 'alignment tax' persists, but 2026 findings are sharper: large-scale audits show refusal rates are a poor proxy for real safety, models behave more safely when they notice they are being tested, and FLI's 2026 index found leading labs weakened safety pledges as capabilities and competition grew. Threat Open problemThreat Open problemTrend ↓ worsening Evidence estimated Test-Set Contamination LLM training data includes evaluation benchmarks, overestimating capabilities and invalidating assessments. Threat Open problemThreat Open problemTrend → steady Evidence estimated Time pressure for superalignment Safety research (e.g., alignment with human intent) is years behind capability advances. Threat Open problemThreat Open problemTrend ↓ worsening Evidence confirmed Tools for Interpreting or Explaining LLM Behavior are Absent or Lack Faithfulness We lack reliable tools to understand why an LLM behaves the way it does by looking "inside" it. Threat Open problemThreat Open problemTrend → steady Evidence estimated Values to be Encoded within LLMs are Not Clear Deciding whose values an LLM should align with is a fundamental problem with significant ethical implications. Threat Open problemThreat Open problemTrend ? unmeasured Evidence contested
No issues match.
Part of the AI Problems Index .
Live from the knowledge base; allegations are labeled as such in each entry.