Lilian Weng(翁荔)
类型:AI 研究者 / 技术写作者 经历:OpenAI 前应用 AI 研究负责人、安全系统团队负责人;现为 Thinking Machines Lab 联合创始人 博客:Lil’Log(https://lilianweng.github.io/)
在本知识库中的角色
她是”框架(harness)能放大模型,但不能替代模型”这个平衡立场的主要来源。
1. Harness Engineering for Self-Improvement(2026-07-04)
对 harness 的定义:
“A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results.”
框架就是围在基座模型外面的那一层系统。它负责安排执行流程,并决定模型怎么思考和规划、怎么调用工具和行动、怎么感知和管理上下文、怎么保存产物、怎么评估结果。
对模型与框架的关系:
“the layer between the raw model and the real-world context seems to be as important as the model’s raw intelligence”
夹在原始模型和真实世界上下文之间的那一层,看起来和模型本身的智能一样重要。
但她同时给出了关键限定:
“STOP improved mean downstream performance across iterations with GPT-4 but degraded with weaker models like GPT-3.5 and Mixtral. Recursive structure alone is not enough. The base model must be capable enough to improve the mechanism. This implies that harness improvement enables better deployment of the model but intelligence is still the core.”
同一种自我改进(STOP)方法,配 GPT-4 会越迭代越好,配 GPT-3.5 和 Mixtral 反而越迭代越差。光有递归结构不够,基座模型必须「足够有能」,这套机制才起作用。
2. 她引述的”两个独立的轴”(Lin et al. 2026, arXiv 2605.30621)
“harness-updating refers to the capability of producing useful harness edits and harness-benefit denotes the capability of utilizing the updated harness… a range of model of different sizes and core intelligence, from Qwen3.5-9B to Claude Opus 4.6, were observed to show similar harness updating capability; the 9B harness proposer/evolver is able to write a skill procedurally isomorphic to Opus.”
harness-updating 指「写出有用的框架改动」的能力,harness-benefit 指「把改好的框架用起来」的能力。研究发现:从 90 亿参数的 Qwen3.5-9B 到顶级模型 Opus 4.6,前一项能力基本持平——9B 模型写出的技能在程序结构上和 Opus 写的是同构的。
“Main results: (A) harness updating capability is measured flat across a range of models from Qwen2-32B to Opus 4.6; (B) harness benefit capability is non-monotonic where middle tier models benefit the most.”
这条是”参数大 ≠ 工具调用强”的机制解释,也是本知识库中”能力维度”论点的关键支撑。
3. 她给出的 harness 演进序列
“The progression in the object being optimized in the harness system is roughly: instruction [prompts] → structured context → workflow → harness code → optimizer code.”
框架系统里「被优化的对象」大致经历了这样一条演进路线:提示词 → 结构化的上下文 → 工作流 → 框架代码 → 优化器代码。
为什么她的立场重要
在”框架比模型重要”这个 2026 年的流行叙事里,她是少数坚持双向限定的声音:
“Eventually it is possible that many harness improvements will be internalized into core model behavior, but the interface with external context and tools should remain.”