Nature | Medicine
Follow
Why high scores do not mean application readiness for health AI
Large language models achieve high scores on health application benchmarks, yet adversarial stress tests now reveal prevalent brittleness — shortcut reliance, fragile visual grounding and fabricated reasoning traces — which exposes substantial gaps between benchmark performance and the robustness evidence needed to support claims of readiness for medical decision-support and patient-facing applications.