自回归生成(Autoregressive Generation)

一句话定义:自回归生成(Autoregressive Generation)是大模型产出文本的方式——一次只生成一个词元(Token),且每一个词元都以前一个词元为条件,因此必须顺序进行。它强大、通用,但天生慢、天生贵,这也是 Jev 要绕开它的原因。

官方发布博客原话:“Sequential. Generates one token at a time, each conditioned on the last.”

中文:顺序式。一次生成一个 token,每个 token 都以前一个为条件。


1. 机制:为什么必须一步一步来

官方对这条路径的描述只有一句,但把关键全说了——每个 token 以前一个为条件(each conditioned on the last)。要产出 N 个 token,就必须走 N 个有先后依赖的步骤:不知道第 k 个 token,就无法算第 k+1 个。这是自回归的”自”——用自己上一步的输出当作下一步的输入。

Jev 走的是另一条路,官方把二者直接并置:

“Jev outputs all probabilities in parallel instead of autoregressively generating by token.”

一边是”逐 token 顺序生成”,一边是”全部概率并行输出”(见 并行采样)。

2. 为什么慢、为什么贵

慢。 官方给出的端到端响应时间是:

“End-to-end response time is 3 to 329 seconds for frontier models. Fast enough for interfacing with humans, but a big bottleneck when integrated in code.”

(前沿模型的端到端响应是 3 到 329 秒。对人交互够快,但集成进代码里就是大瓶颈。)对照 Jev:

“End-to-end response time is 70ms-500ms for TypeSafe. This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries.”

大模型Jev
回答一次要多久3~329 秒70~500 毫秒(多数约 100 毫秒)
生成的粒度一个一个 token 往外写所有答案一次算完

贵。 大模型按”输入的字 + 输出的字”双向收费,而且输出更贵:

“Cost — Input tokens: from 10 / MTok. Output tokens: ~5x more expensive than input tokens.”

(输入词元 10 / MTok;输出词元比输入贵约 5 倍。)Jev 只有输入收费:42 / 十亿词元),输出免费。官方称其输入价比 Claude Fable 5.1 低 238 倍。

首页那组工作流对照更直观:Jev 0.013880 / 8.566s(按这两个数相除,约贵 170 倍、慢 75 倍;官方对同一批 workflow 给出的综合口径是 “193.6x Faster, 444.6x Cheaper.”)。

3. 生成字符串的隐性代价

官方对”字符串”的定位很清醒——不是不好,而是贵且有风险:

“Strings are extremely powerful and general, but costly. ‘Giving up’ strings actually gives us a lot of superpowers!”

放弃生成能力,换来的是快、便宜、类型安全。对系统集成者来说,真正的痛点在”下游摩擦”:

“Strings / generated text. Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails.”

(字符串灵活到可以是任何东西:聊天回复、代码、幻觉、拒答,甚至类型安全的结构化值。要被软件使用,回复必须先解析再校验;而且总存在”跑偏”的风险。)这就是为什么会有”让模型按 JSON 格式输出”的纠结——官方在 Introduction 里称之为一种错配:

“When you need a model to make a judgment that your code will consume, that creates a mismatch: you are coercing a text-generation system into outputting structured decisions, then parsing the results back into something your code can depend on.”

(你需要模型做一个”代码要消费”的判断时,就产生了错配:你在强迫一个文本生成系统输出结构化决策,再把结果解析回代码能依赖的东西。)

4. 多问几个问题,生成模型要付出什么

对自回归模型,“多问一个问题”往往意味着要么更长的一次生成、要么更多轮往返,两者都要串行地花时间与输出词元。而 Jev 因为并行采样,“Adding questions barely changes the response time… Asking a question you might not need is close to free.”官方 cookbook 实测:13 个问题一次问 vs 拆 13 次问,后者要 2.71 秒、0.000497(便宜 12.2 倍、快 10.0 倍)。

这也是官方”快慢分工”论点的由来:让生成模型只负责它不可替代的”写”,把大量判断交给判别式模型(见 判别式模型)。

5. 与 Jev 的关系:分工,而不是取代

Jev 明确不做生成,官方在缺陷清单里单列一条并给出替代方案:

“jev-1.13 is not trained to generate text. While you can force it to by chaining choices, this will not work well and will be very slow.”

Instead: when the answer space is bounded, turn extraction into a Choice over the options rather than asking for the value itself. If you really need to generate text… there are other models for that.

值得注意:强行让 Jev”逐 Choice 串起来当生成器用”——本质就是在判别模型上模拟一次自回归生成——结果依然是又慢又差。这反过来印证了自回归的成本来自机制本身,而不是某个模型实现得不好。

正确的用法是混搭。官方智能家居 demo 就是范例:

“When TypeSafe determines that the user query is a request for general information or conversation, the system calls an LLM to generate a freeform response… The initial TypeSafe response is so fast compared to the LLM response that it adds negligible latency to the overall system.”

(当 TypeSafe 判定用户是在问信息或闲聊时,系统才调用 LLM 去生成自由文本回复;TypeSafe 那一步相对 LLM 快到可以忽略,几乎不给整体延迟加负担。)分工原则:判别模型做过滤、路由、把关、打分;生成模型做写作。

相关来源

  • 官方发布博客《Introducing System One Models & Jev》(Sampling / Cost / Speed 对照表):processed/jev-原始资料/官方-网页/o04-blog-sysone.txt(https://typesafe.ai/blog/system-one)
  • 官网首页(193.6x / 444.6x 与 0.114s vs 8.566s):processed/jev-原始资料/官方-网页/o01-home.txt(https://typesafe.ai)
  • 官方文档《Introduction》(“mismatch” 段):processed/jev-原始资料/官方-文档/introduction.md(https://docs.typesafe.ai/introduction)
  • 官方文档《AI primer》(RLHF / RLVR / RLCD):processed/jev-原始资料/官方-文档/introduction__machine-learning-primer.md
  • 官方文档《Primitives (Questions)》:processed/jev-原始资料/官方-文档/primitives.md
  • 官方文档《Jev 1.13 jaggedness · Generation》:processed/jev-原始资料/官方-文档/model-jaggedness__jev-1.13.md
  • 官方文档《Smart home assistant demo》:processed/jev-原始资料/官方-文档/demos__smart-home.md