跳到正文
原文
Google DeepMind·· 2025-12-09精选AI 评分68

Google DeepMind 发布 FACTS Benchmark Suite 评测大模型事实准确性

FACTS Benchmark Suite: Systematically evaluating the factuality of large language models

AI 导读

Google DeepMind 联合 Kaggle 发布 FACTS Benchmark Suite,包含参数知识、搜索、多模态及 Grounding v2 共 3,513 个公开测试用例。评测显示 Gemini 3 Pro 以 68.8% 的 FACTS Score 领先,所有模型准确率均低于 70%,多模态任务得分最低。

推荐理由

提供涵盖参数知识、搜索与多模态的4项事实性评测基准,揭示前沿模型在事实准确性上的具体短板。

来源:Google DeepMind · deepmind.google