一句话摘要:斯坦福 CRFM 的 HELM 指出,在它之前主流模型平均只被评估了 17.9% 的核心场景,有些知名模型之间连一个共同场景都没有;它提出用”场景 × 指标”矩阵做整体评估,一次性测 7 个指标而不只是准确率。

原始标题

Holistic Evaluation of Language Models (HELM)(Stanford CRFM,arXiv:2211.09110,TMLR 2023)

来源:https://arxiv.org/abs/2211.09110


它解决什么问题

2022 年之前,“模型好不好”基本等于”在几个榜单上准确率多少”。HELM 的作者认为这是选择性报告:只报自己强的那几个榜。他们要建立一套无法挑食的评估体系。

关键要点

1. 先给”场景 × 指标”建分类学

“First, we taxonomize the vast space of potential scenarios (i.e. use cases) and metrics (i.e. desiderata) that are of interest for LMs. Then we select a broad subset based on coverage and feasibility, noting what’s missing or underrepresented (e.g. question answering for neglected English dialects, metrics for trustworthiness).”

第一步是给「场景(用例)× 指标(期望)」这个巨大的空间建立分类学,然后按覆盖度和可行性挑一个尽量广的子集,并明确指出哪些被漏掉了(比如被忽视的英语方言问答、可信度类指标)。

2. 多指标:准确率只是 7 个之一

“we adopt a multi-metric approach: We measure 7 metrics (accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency) for each of 16 core scenarios when possible (87.5% of the time). This ensures metrics beyond accuracy don’t fall to the wayside, and that trade-offs are clearly exposed.”

我们采用多指标方案,对每个核心场景测 7 个指标,确保准确率之外的指标不被丢在一边,并且把取舍清楚地暴露出来——最后这半句是 HELM 的设计目标,不是副作用。

7 个指标含义
accuracy准确率
calibration置信度是否与实际正确率匹配
robustness输入扰动下是否稳定
fairness群体间表现是否均衡
bias偏见
toxicity有害输出
efficiency推理成本

3. 覆盖度数据——最惊人的一条

“Prior to HELM, models on average were evaluated on just 17.9% of the core HELM scenarios, with some prominent models not sharing a single scenario in common. We improve this to 96.0%.”

这句话是”跑分不能横向比较”的方法论根源:两个模型可能根本没在同一张卷子上考过。

4. HELM 官方页面上按能力分组的场景组

Visual perception、Reasoning、Knowledge、Bias、Fairness、Toxicity、Safety、Robustness、Multilinguality。

5. “直面取舍”是设计目标而非副产品

HELM 官方 plots 页专门提供 “Accuracy as a function of other metrics”、“Correlation between metrics” 等图表——即把”换准确率换来了什么”作为一等公民展示。

重要引用(英文原文)

“We conduct a large-scale evaluation of 30 prominent language models (spanning open, limited-access, and closed models) on all 42 scenarios, 21 of which were not previously used in mainstream LM evaluation.”

我们对 30 个主流语言模型(涵盖开源、受限访问和闭源三类)在全部 42 个场景上做了大规模评估,其中 21 个场景此前从未被用于主流评测。

边界与争议

  • HELM 极其昂贵(要测 7 个指标 × 42 个场景 × 30 个模型),所以更新慢,实际影响力不如简单榜单。
  • 它自己也无法消除权重之争:把 7 个指标合成一个总分时,权重怎么定仍然是个价值判断。
  • 2026 年的第三方机构(如 Artificial Analysis)继承了它的思路:把评测按 skill(能力) 与 knowledge(知识域) 两个维度打标签,其中 skill 细分为 Reasoning、Agentic、Tool Use、Coding、Instruction Following、Long Context、Writing、User Interaction、Faithfulness、Multilingual 共 10 类。

与本文其他页面的关系

  • 能力维度 —— 概念页
  • MMLU —— 只测其中一维的单一榜单
  • 评测污染 —— 榜单可信度的第二个问题