Chrome Extension
WeChat Mini Program
Use on ChatGLM

Medical Large Language Model Benchmarks Should Prioritize Construct Validity

Ahmed Alaa,Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, Inioluwa Deborah Raji,Travis Zack

CoRR(2025)

Cited 0|Views2
AI Read Science
Must-Reading Tree
Example
Generate MRT to find the research sequence of this paper
Chat Paper
Summary is being generated by the instructions you defined