Hugging Face Blog·· 2024-04-19AI 评分35
Open Medical-LLM Leaderboard:医疗大语言模型基准测试
The Open Medical-LLM Leaderboard: Benchmarking Large Language Models in Healthcare
AI 导读
Hugging Face 发布 Open Medical-LLM Leaderboard,旨在通过标准化平台评估和比较大型语言模型在医疗任务中的表现。该基准测试涵盖 MedQA、MedMCQA、PubMedQA 及 MMLU 医疗子集,以准确率作为主要评估指标。测试结果显示 GPT-4-base 和 Med-PaLM-2 等商业模型在多个医疗数据集上表现优异。
来源:Hugging Face Blog · huggingface.co