Home Knowledge Base Med-PaLM

Med-PaLM is a medical domain language model developed by Google Research that was the first AI system to achieve "expert" level performance on the US Medical Licensing Exam (USMLE) — with Med-PaLM 2 reaching 86.5% accuracy on MedQA (surpassing the ~60% passing threshold by a wide margin), built by fine-tuning Google's PaLM foundation model using instruction tuning on curated medical question-answering datasets (MultiMedQA) and rigorously evaluated for clinical safety, accuracy, and potential harm.

What Is Med-PaLM?

Performance Evolution

ModelMedQA (USMLE)MedMCQAPubMedQANotes
Med-PaLM 1 (2022)67.6%57.6%79.0%First AI to pass USMLE
Med-PaLM 2 (2023)86.5%72.3%81.8%Expert physician level
GPT-4 (2023)~86%~70%~80%Comparable to Med-PaLM 2
ChatGPT (GPT-3.5)~60%~55%~75%Near passing threshold

Safety and Evaluation

Deployment and Access

Med-PaLM is the benchmark-setting medical AI that proved language models can reach physician-level accuracy on structured medical examinations — while simultaneously demonstrating the critical importance of rigorous safety evaluation, harm assessment, and controlled deployment for AI systems operating in high-stakes clinical domains.

med palmgooglemedical

Explore 500+ Semiconductor & AI Topics

From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.