一句话摘要:斯坦福 CRFM 的 HELM 指出,在它之前主流模型平均只被评估了 17.9% 的核心场景,有些知名模型之间连一个共同场景都没有;它提出用”场景 × 指标”矩阵做整体评估,一次性测 7 个指标而不只是准确率。
原始标题
Holistic Evaluation of Language Models (HELM)(Stanford CRFM,arXiv:2211.09110,TMLR 2023)
来源:https://arxiv.org/abs/2211.09110
它解决什么问题
2022 年之前,“模型好不好”基本等于”在几个榜单上准确率多少”。HELM 的作者认为这是选择性报告:只报自己强的那几个榜。他们要建立一套无法挑食的评估体系。
关键要点
1. 先给”场景 × 指标”建分类学
“First, we taxonomize the vast space of potential scenarios (i.e. use cases) and metrics (i.e. desiderata) that are of interest for LMs. Then we select a broad subset based on coverage and feasibility, noting what’s missing or underrepresented (e.g. question answering for neglected English dialects, metrics for trustworthiness).”
第一步是给「场景(用例)× 指标(期望)」这个巨大的空间建立分类学,然后按覆盖度和可行性挑一个尽量广的子集,并明确指出哪些被漏掉了(比如被忽视的英语方言问答、可信度类指标)。
2. 多指标:准确率只是 7 个之一
“we adopt a multi-metric approach: We measure 7 metrics (accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency) for each of 16 core scenarios when possible (87.5% of the time). This ensures metrics beyond accuracy don’t fall to the wayside, and that trade-offs are clearly exposed.”
我们采用多指标方案,对每个核心场景测 7 个指标,确保准确率之外的指标不被丢在一边,并且把取舍清楚地暴露出来——最后这半句是 HELM 的设计目标,不是副作用。
| 7 个指标 | 含义 |
|---|---|
| accuracy | 准确率 |
| calibration | 置信度是否与实际正确率匹配 |
| robustness | 输入扰动下是否稳定 |
| fairness | 群体间表现是否均衡 |
| bias | 偏见 |
| toxicity | 有害输出 |
| efficiency | 推理成本 |
3. 覆盖度数据——最惊人的一条
“Prior to HELM, models on average were evaluated on just 17.9% of the core HELM scenarios, with some prominent models not sharing a single scenario in common. We improve this to 96.0%.”
这句话是”跑分不能横向比较”的方法论根源:两个模型可能根本没在同一张卷子上考过。
4. HELM 官方页面上按能力分组的场景组
Visual perception、Reasoning、Knowledge、Bias、Fairness、Toxicity、Safety、Robustness、Multilinguality。
5. “直面取舍”是设计目标而非副产品
HELM 官方 plots 页专门提供 “Accuracy as a function of other metrics”、“Correlation between metrics” 等图表——即把”换准确率换来了什么”作为一等公民展示。
重要引用(英文原文)
“We conduct a large-scale evaluation of 30 prominent language models (spanning open, limited-access, and closed models) on all 42 scenarios, 21 of which were not previously used in mainstream LM evaluation.”
我们对 30 个主流语言模型(涵盖开源、受限访问和闭源三类)在全部 42 个场景上做了大规模评估,其中 21 个场景此前从未被用于主流评测。
边界与争议
- HELM 极其昂贵(要测 7 个指标 × 42 个场景 × 30 个模型),所以更新慢,实际影响力不如简单榜单。
- 它自己也无法消除权重之争:把 7 个指标合成一个总分时,权重怎么定仍然是个价值判断。
- 2026 年的第三方机构(如 Artificial Analysis)继承了它的思路:把评测按 skill(能力) 与 knowledge(知识域) 两个维度打标签,其中 skill 细分为 Reasoning、Agentic、Tool Use、Coding、Instruction Following、Long Context、Writing、User Interaction、Faithfulness、Multilingual 共 10 类。