Nature | Medicine
Follow
Lessons from deploying the ChatEHR system at Stanford Medicine
In piloting and deploying a large language model within a large medical center, we learned that benchmark-based evaluations are insufficient for monitoring and evaluating interactions driven by clinicians, and that this requires new methods for monitoring performance.