Skip to content
AI情报2026年8月19日AI情报
文章

LLMs for Medical Consultation Are Evaluated Too Late

The Preformulation Gap: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24...

Frontier 编辑部来源: arXiv
01

来源简报

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap: Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24...