并行采样(Parallel Sampling)

一句话定义:并行采样(Parallel Sampling)指模型在一次查询里把全部输出一起算出来,而不是像大模型那样一个 token 一个 token 地顺序生成。它是 Jev 能快到几十毫秒、且”多问问题几乎不增加耗时”的机制根源。

官方发布博客原话:“Jev outputs all probabilities in parallel instead of autoregressively generating by token.”

中文:Jev 并行输出全部概率,而不是按 token 自回归生成。

官方给这两条路线的命名很直接:

大模型Jev
采样(Sampling)Sequential(顺序式)。“Generates one token at a time, each conditioned on the last.”Parallel(并行式)。“Generates all outputs in a single query. Incredibly efficient and hardware-aware.”

1. 它为什么会诞生

官方说这不是把现成模型调快,而是换了一整套栈:

“We built a new stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD).”

首页的表述一致:“We’re building with a new architecture, a new sampler, and a new training algorithm: RLCD.”——采样器(sampler)是被单独列出来的三大创新之一。

关键前提是 Jev 不生成文本:既然输出只是”在给定选项上的一份概率分布”,就没有”要先写第 1 个 token、才能写第 2 个”的强制顺序。所有问题的答案可以一次性算出。

“Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.”

(把 Jev 想成一次”前沿智能的函数调用”:非结构化状态进,类型化概率决策出。)

2. 最实用的一条:加问题几乎不增加耗时

因为同一请求里的问题是并行求值的,多问几个问题的边际代价极小:

“System One models evaluate every question in a request in parallel. Adding questions barely changes the response time and costs only the tokens for the extra questions, which are cheap. Asking a question you might not need is close to free.”

“Jev ingests the state once and evaluates every question against it in parallel.”

“Decomposition does not require more round trips. Questions over the same state run in parallel.”

还有一条容易被忽略的好处:问题之间互不污染。官方在对比生成模型时说:“Each question is evaluated independently, so adding more questions does not create context-rot.”——多问问题不会制造上下文腐坏(Context Rot),因为一个问题的结果不会变成另一个问题的隐藏上下文。

官方据此给出一个反直觉的实践模式——投机式扇出(Speculative Fan-out, 也叫 speculative fan-out):

“Because TypeSafe supports sending many questions in a single API call, we recommend putting all of the questions your system needs in a single request, and then using code to decide what is relevant after the fact.”

(把系统可能用到的全部问题一次全问出去,事后再用代码挑有用的。)例:工单分诊时,“bug 严重度”只有判为 bug 时才用得上,但仍一起问——“We include all upfront because additional questions usually have little effect on response time.”

3. 量化对照

官方 cookbook《Parallel questions》用一个实跑实验给出数字:把 GDPR 维基条目 + 13 个问题,一次批量问 vs 拆成 13 次分别问(结果取 5 次运行均值):

batching                 calls        cost  total time
one call, all 13             1   $0.000497       0.27s
13 calls, one each          13   $0.006090       2.71s

batching: 12.2x cheaper, 10.0x faster

即一次问完便宜 12.2 倍、快 10.0 倍,且答案完全一致(no change in answers)。同一份材料,13 个问题一次调用只花 $0.000497、0.27 秒。

⚠️ 数字口径提示:官方《Primitives》正文把同一实验写成 “11.5x cheaper and 9.6x faster”;而 cookbook 实跑输出为 12.2x / 10.0x。两个数字都出现在官方材料里,此处以 cookbook 的实跑表为准并保留原样。

官网首页另有一组端到端对照(同一次工作流):Jev 0.013880 / 8.566s,官方据此打出 “193.6x Faster, 444.6x Cheaper.”(并注明是基于系统一任务的 workflow 数据)。

4. 与自回归生成的对照

维度自回归生成(Autoregressive Generation)并行采样(Parallel Sampling)
生成粒度一次一个 token,每步以前一步为条件一次查询产出全部输出
端到端耗时前沿模型 3~329 秒TypeSafe 70ms~500ms(多数约 100ms)
加问题通常意味着更多轮次 / 更长输出”barely changes the response time”
输出计费输出词元约为输入的 5 倍输出免费
硬件关系逐 token 串行,天然受限”Incredibly efficient and hardware-aware.”

官方对同水平智能下的整体速率比的表述是 40x~200x(“This can range from 40x-200x faster for the same levels of frontier intelligence for System One shaped queries.”)。机制层面见 自回归生成。

5. 边界与代价

  • 并行只发生在”同一请求、同一状态”内。 如果后一个问题依赖前一个答案,就必须发第二次请求。官方原则:“If a later judgment depends on an earlier answer, make a second request in code… Two requests are the exception, not the rule.”(Skill suggestion、autoformat、hierarchical classification 三个 cookbook 是有正当理由发第二请求的例外。)
  • 约束仍在。 每请求 64k 词元;其中 state 加单个最长问题合计不超过 32k。
  • 并行不等于无限。 Choice 上限 255 个选项;文档提到在超高基数选择上官方用”先独立打分再显式选择”的两阶段做法,会偶尔变慢。
  • 速度受环境影响的坦白:官方承认 published evals 大多是在西海岸的笔记本上跑的,服务目前也在该地。

在 Harness工程 语境的用法是:让便宜快的判别模型做大量、高频、并行的”反射式”判断,把串行、慢、贵的生成模型只留给真正需要深思的环节(见 系统一模型)。

相关来源