Home Knowledge Base SciBERT

SciBERT is a BERT language model pre-trained from scratch on 1.14 million scientific papers from Semantic Scholar, with a custom scientific vocabulary that efficiently tokenizes domain-specific terminology — outperforming general-purpose BERT on scientific NLP tasks including paper classification, citation intent prediction, Named Entity Recognition (NER) for chemicals and proteins, and relation extraction from biomedical and computer science literature.

What Is SciBERT?

Performance on Scientific NLP Tasks

TaskSciBERTBERT-baseImprovement
Paper Classification (SciCite)85.5%83.1%+2.4%
NER - Chemicals (BC5CDR)90.1% F187.2% F1+2.9%
NER - Proteins (JNLPBA)77.3% F174.8% F1+2.5%
Relation Extraction (ChemProt)76.8% F173.4% F1+3.4%
Citation Intent (SciCite)84.0%82.1%+1.9%

Why SciBERT Matters

SciBERT vs. Domain BERT Models

ModelDomainTraining DataVocabularyKey Strength
SciBERTScience (CS + Bio)1.14M papersCustom scientificBroad scientific coverage
BioBERTBiomedical onlyPubMed abstractsBERT vocabBiomedical NER
ClinicalBERTClinical notesMIMIC-IIIBERT vocabEHR understanding
MatSciBERTMaterials scienceMaterials papersCustomMaterials NLP

SciBERT is the foundational domain-adapted language model for scientific text processing — proving that pre-training on domain-specific data with a custom vocabulary produces substantial improvements on scientific NLP tasks, and establishing the methodology that spawned dozens of subsequent domain-adapted BERT variants across medicine, law, finance, and materials science.

scibertscientificpapers

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.