Google DeepMind·· 2025-12-09精选AI 评分68
Google DeepMind 发布 FACTS Benchmark Suite 评测大模型事实准确性
FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
AI 导读
Google DeepMind 联合 Kaggle 发布 FACTS Benchmark Suite,包含参数知识、搜索、多模态及 Grounding v2 共 3,513 个公开测试用例。评测显示 Gemini 3 Pro 以 68.8% 的 FACTS Score 领先,所有模型准确率均低于 70%,多模态任务得分最低。
推荐理由
提供涵盖参数知识、搜索与多模态的4项事实性评测基准,揭示前沿模型在事实准确性上的具体短板。
来源:Google DeepMind · deepmind.google