判别式模型(Discriminative Model)

一句话定义:判别式模型(Discriminative Model)只做”判断”(给定输入,输出它属于哪一类 / 哪个值),不”生成”新内容;生成式模型(Generative Model)才负责”写出”文本。TypeSafe 明确把自己放在判别这一边,官方措辞是”决策模型(decision model)”。

官方 AI primer 副标题:“Why TypeSafe trains decision models with calibrated probabilities instead of optimizing for generated text.”

中文:为什么 TypeSafe 训练的是”带校准概率的决策模型”,而不是去优化”生成的文本”。

⚠️ 术语提示:官方文档通篇使用的词是决策模型(decision model)与分类器(classifier);“判别器(discriminator)“只在讲 GAN 模式崩溃的一处类比里出现。本页的”判别式 / 生成式”是这条分工的通用叫法。


1. 两条后训练路径,两种输出合同

官方把预训练语言模型的后续演化列成三条支线(原文卡片刻字):

路径全称造出来的东西
RLHFReinforcement Learning from Human Feedback(人类反馈强化学习)把预训练模型变成聊天机器人——“It trains models to produce responses people prefer.”
RLVRReinforcement Learning with Verifiable Rewards(可验证奖励强化学习)造出擅长数学等任务的推理模型,“but slower and more expensive”
RLCDReinforcement Learning for Calibrated Decisions(面向校准决策的强化学习)TypeSafe 自己的路径——“trains TypeSafe to return decisions and calibrated probabilities instead of generated text”

所以判别 / 生成的分工,本质是输出合同不同:

“RLCD optimizes for a different output contract:

  • The model does not generate text.
  • It returns decisions and probabilities.
  • Higher probability should correspond to a greater chance that the answer is correct.”

RLCD 优化的是另一份输出合同:模型不生成文本;它返回决策与概率;更高的概率应对应”答案更可能正确”。

2. 判别 vs 生成:官方给的完整分工

发布博客的对照表逐项把两类模型分开(原样摘录):

维度生成式(现有 LLM)判别式(系统一 / Jev)
优化目标人类偏好:人类评分者更喜欢的写作与聊天回复校准决策:在系统一任务上给出认知诚实的概率
输入侧重非结构化文本,强调顺序消息非结构化文本,强调结构化程序状态
输出字符串;“can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values”类型安全的结构化值;“Possible outputs and structure are defined in advance. The model never makes type errors.”
采样顺序式(Sequential),一次一个 token并行式(Parallel),一次查询产出全部输出
与人/机的关系人类在环(human-in-the-loop):聊天机器人、copilot、编程智能体AI 驱动的软件工作流 / “smart if-statements”:分类、路由、打分、抽取、分支

判别模型的输入依然是自然语言(这一点和 LLM 相同),差别只在它被要求交出的东西是”值”而不是”文字”:

“Like an LLM, a System One model understands natural-language input. It returns typed decisions and probabilities rather than generated text.”

3. Jev 站在哪一边

站在判别 / 决策这一边。 官方几处表述:

  • 首页:“We took the opposite research direction — not chat.”(我们选了相反的研究方向——不是聊天。)RLHF 让人喜欢,但有模式丢弃(mode dropping)、过度自信、不可靠等固有缺陷,“These flaws mean that LLMs require humans-in-the-loop.”
  • 官方把目标场景定为机器对机器:“large-scale AI automation will be closer to 99% machine-to-machine interactions and 1% human interaction.” 他们给这件事起的名字是机器原生智能(Machine Native Intelligence)——“AI with software-like properties such as structure, reliability, observability, testability, speed, consistency, and low cost.”
  • 意图路由(Intent Routing)模式里说得最直白:“TypeSafe can sit in front of all of these as a fast, cheap classifier that determines which handler to invoke.”(TypeSafe 可以坐在所有处理器前面,充当一个又快又便宜的分类器,决定该调用哪个处理器。)
  • 结构化格式修复(autoformat)cookbook 里把问题与判据本身称为分类器的完整规格:“…are the entire specification of the classifier. There is no other logic.”

需要澄清一点:判别模型并不等于”另一个更小的 LLM”。官方 FAQ 专门列了”Is Jev just a smaller LLM?”这一问,答案落在训练目标与采样方式上——Jev 用 RLCD 训练、用并行采样器出结果,而不是把一个生成模型缩小。

4. 为什么判别可以更小更快

这是判别路线最实际的好处,链条在官方材料里是完整的:

  1. 不生成文本 → 没有逐个 token 的顺序步骤(见 自回归生成)。
  2. 一次查询产出全部输出 → 并行采样,硬件利用率高。官方称”a new model architecture, parallel sampler for maximum efficiency”(见 并行采样)。
  3. 输出免费 → 定价只按输入词元计费,42 / 十亿词元),“Output tokens: FREE (too cheap to meter)”。
  4. 学习任务更窄 → 判别只学”在给定选项上分配概率”,不需要学”写出任意文本”这种极宽的能力。

官方给出的整体速率比:“This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries.”(在系统一形态的查询上,同等前沿智能水平下快 40~200 倍。)

社区也提供了”判别模型可以很小”的旁证。开源复现 Luce 的做法是:“a task description, an LLM teacher that writes the data, then LoRA plus a decision head on Qwen3-4B-Base returning calibrated choice, score, and boolean probabilities on a 12 GB GPU.”(一个任务描述 + 一个写数据的 LLM 教师,在 Qwen3-4B-Base 上挂 LoRA 加一个决策头,就能在 12 GB 显卡上返回校准的 Choice / Score / Noul 概率。)另一个开源项目 Laya 则是”Open local Choice, Score, and Noul decision models with published checkpoints”。

5. 判别路线的边界

判别模型做不了生成,这不是暂时缺陷,而是路线选择的结果。官方在缺陷清单里单列一条:

“jev-1.13 is not trained to generate text. While you can force it to by chaining choices, this will not work well and will be very slow.”

Instead: when the answer space is bounded, turn extraction into a Choice over the options rather than asking for the value itself. If you really need to generate text… there are other models for that.

要文本生成,就用生成模型;判别模型负责把生成模型的外围收紧——过滤、路由、把关、打分(见 系统一模型 与 类型安全)。

相关来源

  • 官方文档《AI primer》:processed/jev-原始资料/官方-文档/introduction__machine-learning-primer.md(https://docs.typesafe.ai/introduction/machine-learning-primer)
  • 官方发布博客《Introducing System One Models & Jev》:processed/jev-原始资料/官方-网页/o04-blog-sysone.txt(https://typesafe.ai/blog/system-one)
  • 官网首页:《We took the opposite research direction — not chat》:processed/jev-原始资料/官方-网页/o01-home.txt(https://typesafe.ai)
  • 官方文档《Introduction》:processed/jev-原始资料/官方-文档/introduction.md
  • 官方文档《How to build with TypeSafe》《Intent routing》《Jev 1.13 jaggedness》:processed/jev-原始资料/官方-文档/concepts__how-to-build-with-system-one.md、patterns__intent-routing.md、model-jaggedness__jev-1.13.md
  • 社区精选(Luce / Laya 开源复现):processed/jev-原始资料/社区精选/awesome-typesafe.md