climate-fever

**Climate-FEVER** is the **specialized fact-checking benchmark that applies the FEVER methodology to real-world climate change claims** — testing whether NLP models can verify scientific statements against evidence while navigating the complexity of nuanced, contested, and cherry-picked environmental science claims. **What Is Climate-FEVER?** - **Origin**: Extends FEVER (Fact Extraction and VERification), adapting its pipeline to climate science. - **Claims Source**: Claims scraped from real web articles, climate-skeptic sites, scientific publications, and news media — not artificially mutated sentences. - **Label Set**: SUPPORTED, REFUTED, NOT_ENOUGH_INFO, and DISPUTED — the last label acknowledging that some claims are genuinely contested even among experts. - **Evidence**: Wikipedia articles serve as the evidence corpus, requiring retrieval-then-verify pipelines. - **Scale**: ~1,535 climate claims, each requiring multi-sentence evidence retrieval and label prediction. **Why Climate-FEVER Is Harder Than FEVER** - **Scientific Nuance**: A claim can be "technically true but misleading." For example, "Arctic sea ice has recovered since 2012" is factually accurate for one year but ignores the long-term declining trend. FEVER binary labels cannot handle this. - **Cherry-Picking**: Climate misinformation often involves citing real data out of context. The model must understand statistical trends, not just sentence-level facts. - **Consensus vs. Outlier**: Some claims cite genuine scientific papers that are outliers. The model must distinguish mainstream consensus from fringe positions. - **Temporal Sensitivity**: Climate data changes. A claim verified true in 2010 may be refuted by 2023 measurements. **Why Climate-FEVER Matters** - **Misinformation Combat**: Climate change disinformation is among the most consequential misinformation categories. Automated fact-checking can assist human moderators at scale. - **Scientific Literacy AI**: Forces models to develop genuine understanding of scientific evidence, not just surface pattern matching. - **Benchmark Gap**: General fact-checking benchmarks (FEVER, VitaminC) use balanced Wikipedia-style claims. Climate-FEVER exposes domain-specific failure modes. - **Policy Impact**: Erroneous AI fact-checking of climate claims can directly influence public discourse and policy narratives. - **Retrieval Pressure**: Requires retrieving domain-specific scientific evidence from a large corpus — stress-testing retrieval-augmented verification systems. **The Pipeline** **Step 1 — Document Retrieval**: - Given claim: "The Greenland ice sheet is gaining ice." - Retrieve top-k Wikipedia documents about Greenland glaciology. **Step 2 — Sentence Selection**: - Extract the 5 most relevant sentences from retrieved documents. - These become the evidence set for classification. **Step 3 — Label Prediction**: - Classify: SUPPORTED / REFUTED / NOT_ENOUGH_INFO / DISPUTED. - State-of-the-art models (2024): ~65-72% label accuracy — well below human performance (~85%). **Comparison to Related Benchmarks** | Feature | FEVER | Climate-FEVER | MultiFC | |---------|-------|---------------|---------| | Domain | General Wikipedia | Climate Science | Multi-domain web | | Labels | 3 | 4 (incl. DISPUTED) | Diverse site-specific | | Claims Source | Mutated sentences | Real web articles | Professional fact-checks | | Difficulty | Moderate | High (scientific nuance) | High (label normalization) | | Real-world misinformation | Low | High | Very High | **Key Challenges for Models** - **Long-chain Reasoning**: Climate claims often require synthesizing data from multiple evidence sentences. - **Quantitative Reasoning**: "Temperatures rose 1.1°C since pre-industrial levels" requires numeric comparison. - **Counter-intuitive Labels**: Factually accurate statements can still be REFUTED if they imply a false overall conclusion. **Tools & Repositories** - **Dataset**: Available via Hugging Face Datasets (`climate_fever`). - **Baseline Models**: TF-IDF + BERT-based verifiers; DPR-based dense retrievers. - **Evaluation**: Standard FEVER scorer — label accuracy + evidence F1. Climate-FEVER is **the stress test for scientific reasoning in AI** — moving fact-checking from simple Wikipedia trivia into the contested, consequential territory of environmental science where model failures have real-world impact.

Go deeper with CFSGPT

Get AI-powered deep-dives, save terms, and run advanced simulations — free account.

Create Free Account