llama.cpp 草稿提速 42 倍:本地编码模型的一次免费升级

·阅读约11分钟·Evergreen Tools Team

2026 年 9 月 26 日,工程师 Hayder Tirmazi 公布了一组针对 llama.cpp 的改动:把 prompt lookup 解码(prompt lookup decoding)中的「草稿」阶段提速最高 42 倍,同时把内存占用最多降低 2.6 倍。在 541 MB 语料下,草稿一个 token 的耗时从上游的约 165 微秒降到约 4 微秒。重要的是:模型没换、硬件没换、算法也没换。prompt lookup 是「投机解码」的一个特例,它用一个 n-gram 模型代替第二个神经网络来打草稿,因此是加速重复性工作(比如代码编辑)的一种廉价手段。

一、prompt lookup 解码到底在做什么

先讲清它是什么。使用 prompt lookup 解码时,推理引擎会用一条规则来「猜」接下来的 k 个 token:把当前上下文里已有的 n-gram 拿出来,看历史上这个 n-gram 后面最常跟着哪个 token,然后把那个 token 当作草稿。设当前 token 序列为 x_1…x_t,一个 n-gram 就是 n 个连续 token。引擎为每个 n 选出得分最高的候选 token,再依据两个阈值决定是否真的把它当草稿:该 n-gram 至少出现过 a_n 次,且该候选词在后继中占比至少 p_n。llama.cpp 会依次尝试 n=4、3、2、1,采用第一个通过阈值的候选。因为草稿来自「文中已经出现过的东西」,所以它在代码编辑这类高重复场景里命中率很高。

# Build a static n-gram cache and then benchmark prompt lookup decoding.
# Both tools ship in llama.cpp's own lookup example, which is why the
# numbers in the write-up are reproducible rather than vendor claims.

# 1) build the static cache from a corpus (here: WikiText-103)
./llama-lookup-create -m ./models/model.gguf -f ./corpus/wiki103.txt -o w103.cache

# 2) replay a text file as if it were model output and measure drafting
./llama-lookup-stats -m ./models/model.gguf \
  --lookup-cache-static w103.cache \
  --ctx-size 4096 \
  -f ./corpus/wiki103.test.txt

# The benchmark reports: drafted tokens matched, time to draft, and the
# static cache load time. Those three numbers are what the four changes move.
本地推理与命令行

在 Apple M4 Pro(14 核 / 48 GB)上测得

二、llama.cpp 的三类 n-gram 缓存

llama.cpp 维护三类 n-gram 缓存。上下文缓存(context cache)保存当前正在处理 token 的 1 至 4 元组,并随着生成新 token 而更新;动态缓存(dynamic cache)保存此前运行的 n-gram,例如更早的对话;静态缓存(static cache)保存来自静态语料库的 2 元组,用 llama-lookup-create 构建。打分时有一个容易被忽略的细节:当某个候选词同时被静态缓存认可时,它的权重会被乘上 100;也就是说静态缓存不只是「兜底」,它还会放大其它缓存里的一致候选。理解这一点,才能理解为什么缓存实现方式对性能影响如此之大。

# The acceptance-rate thresholds are hard-coded in llama.cpp and they are
# worth knowing before you try to "tune" anything. Drafting only proposes a
# token when the n-gram appeared often enough and the same follower is
# dominant enough. Lower these and you draft more but accept less.

# context cache (n-grams of size 1..4 from the tokens being generated)
CTX_A = (2, 2, 1, 1)          # minimum occurrences  a_n
CTX_P = (0.66, 0.5, 0.5, 0.5) # minimum dominance      p_n

# dynamic cache (n-grams from previous runs of the model)
DYN_A = (4, 3, 2, 2)
DYN_P = (0.75, 0.66, 0.66, 0.66)

# The engine tries n = 4, 3, 2, 1 and drafts the first candidate that passes.
# A token y* is drafted only when: f(X_n, y*) >= p_n * F(X_n) and F(X_n) >= a_n.

三、四个改动分别做了什么

四个改动依次展开。第一,是「别再拷贝 map」:llama.cpp 的 n-gram 缓存实现为嵌套的 std::unordered_map,而内层 map 在每一个草稿步骤上都被不必要地拷贝了多次,改成按引用读取后,草稿立即提速 4.5 至 25.6 倍。第二,把外层 map 从 std::unordered_map 换成扁平哈希表(作者选择了 ankerl::unordered_dense),因为标准库实现用链式冲突解决、对缓存不友好。第三,把每个 n-gram 的后继词存成有序数组,用定长二分查找取计数,代价可预测。第四,把静态缓存的外层 map 换成已验证的 constmap,因为静态缓存在加载后永不改变。

# Reproduce the scoring rule in Python before touching C++. Understanding
# the weight w(y) matters: the static cache is not just a fallback, it
# multiplies candidates from the other caches by 100 when it agrees.

def score(ctx, dyn, static, X_n, y, vocab):
    base = ctx(X_n, y) if ctx(X_n, y) else dyn(X_n, y)
    w = 100 * static(X_n, y) if static(X_n, y) > 0 else 1
    return base * w

def draft(ctx, dyn, static, X_n, vocab, a_n, p_n):
    F = sum(ctx(X_n, y) for y in vocab)          # occurrences of the n-gram
    best = max(vocab, key=lambda y: score(ctx, dyn, static, X_n, y, vocab))
    if F >= a_n and score(ctx, dyn, static, X_n, best, vocab) >= p_n * F:
        return best                              # accept the draft
    return None                                  # fall through to n-1
n-gram 缓存相关的代码

四个改动,接受率保持不变

四、真实数字:165 微秒到 4 微秒

数字来自作者公开的基准:在 0、25、50、100、200、541 MB 语料下,上游每次草稿的耗时分别为 8.54、45.61、59.73、83.46、113.46、165.48 微秒;四项改动全部应用后为 0.89、3.06、3.25、3.32、3.47、3.98 微秒。静态缓存加载时间从 3.76 秒降到 0.23 秒(541 MB 语料);峰值内存从 1.71 GB 降到 1.31 GB,降幅最多 1.30 倍。测试环境是 Apple M4 Pro、14 核、48 GB 内存,取三次运行的中位数。最关键的一条:改动后每个语料上的接受率与原始实现完全一致——也就是说,变的是开销,不是行为。

# The last change is the interesting one: the static cache becomes a
# read-only constmap. Because it never changes after loading, the followers
# of every n-gram can live in one contiguous array, and the map stores a
# single 64-bit value holding a position and a count.

# value = (position << 24) | count          for up to 16.7M followers
POSITION = value >> 24
COUNT    = value & 0xFFFFFF

# lookup: one map probe, then a short fixed-length binary search in the array
entry_ptr, n_pairs = constmap_lookup(static_map, ngram_tokens)
count = binary_search(entry_ptr, n_pairs, candidate_token)

# Result from the write-up, 541 MB corpus:
#   static cache load time   3.76 s  ->  0.23 s   (6.3x to 16.1x faster)
#   peak memory              1.71 GB ->  1.31 GB  (up to 1.30x less)
#   acceptance rate          identical on every corpus

五、它对哪类工作负载真正有用

对工作负载的意义要诚实地说。作者明确没有改动 prompt lookup 的算法,因此改动影响的是「每次草稿的延迟」「静态缓存加载时间」「静态缓存内存」这三项,接受率不变。这意味着收益完全取决于你的文本有多重复:在代码编辑、重构、模板化输出里,上下文里大量内容会重复出现,草稿命中率高、收益明显;在自由对话里,重复少,草稿更多是纯开销,这时应当把 prompt lookup 关掉,而不是期待它带来加速。换句话说,这是一次对本地编码模型非常友好的免费升级,但并不是一个万能旋钮。

# Practical checklist for adopting the upstream PRs, in the order that
# pays off. The first two drop straight into any build; the last two change
# the static cache format, so they only matter if you use a corpus cache.

1. Read inner n-gram maps by reference, never by value.  # PR: stop copying
   -> 4.5x to 25.6x faster drafting depending on corpus size
2. Swap the outer std::unordered_map for a flat hash map (unordered_dense).
   -> removes cache-unfriendly pointer chasing on every lookup
3. Store each n-gram's followers as a sorted vector, search with a
   fixed-length binary search.                          # deterministic cost
4. Replace the static cache's outer map with a verified constmap.
   -> load 6.3x to 16.1x faster, memory down to about the file size

# Sanity check before shipping: acceptance rate must be identical on every
# corpus, otherwise you changed behaviour, not just performance.
草稿延迟随语料规模变化

165 微秒 → 4 微秒,来自作者公开的基准

六、怎么试、上线前的自检

要试用它,路径很短。llama.cpp 仓库自带 prompt lookup 的示例,其中 llama-lookup-create 用于从语料构建静态缓存,llama-lookup-stats 用于基准测试:它把一个文本文件当作「模型输出」逐 token 回放,记录草稿匹配数、草稿耗时与静态缓存加载时间。作者的代码与结果都在其公开仓库中,相关改动分别以 PR 形式提交。需要注意仓库已经从 ggerganov/llama.cpp 迁到 ggml-org/llama.cpp。示例 5 给出了一份采纳清单:前两项可以直接享受收益,后两项只有在使用语料静态缓存时才有意义。上线前的自检只有一条硬标准:接受率必须在每个语料上与原始实现一致。

📌 常见问题 FAQ

提速 42 倍是什么意思?

指的是草稿阶段的延迟最多下降 42 倍,而不是整体生成速度提升 42 倍;作者的改动不改变接受率,因此端到端的收益取决于文本的重复程度。

prompt lookup 解码是什么?

投机解码的一个特例:用一个 n-gram 模型代替第二个神经网络作为草稿模型,从上下文中已出现过的 n-gram 里猜下一个 token。

四个改动分别是什么?

内层 map 改为按引用读取、外层 map 改用扁平哈希表、内层后继词改为有序数组加定长二分查找、静态缓存改用 constmap。

数字是多少?

541 MB 语料下草稿耗时从 165.48 微秒降至 3.98 微秒;静态缓存加载从 3.76 秒降至 0.23 秒;峰值内存从 1.71 GB 降至 1.31 GB;接受率保持不变。

在什么硬件上测的?

Apple M4 Pro,14 核,48 GB 内存,模型上下文长度按 4096 token 计,结果为三次运行的中位数。