llama.cpp 草稿提速 42 倍:本地编码模型的一次免费升级
💡 工具推荐:AI Token 计数器、AI 代码解释器、AI 单元测试生成
2026 年 9 月 26 日,工程师 Hayder Tirmazi 公布了一组针对 llama.cpp 的改动:把 prompt lookup 解码(prompt lookup decoding)中的「草稿」阶段提速最高 42 倍,同时把内存占用最多降低 2.6 倍。在 541 MB 语料下,草稿一个 token 的耗时从上游的约 165 微秒降到约 4 微秒。重要的是:模型没换、硬件没换、算法也没换。prompt lookup 是「投机解码」的一个特例,它用一个 n-gram 模型代替第二个神经网络来打草稿,因此是加速重复性工作(比如代码编辑)的一种廉价手段。
一、prompt lookup 解码到底在做什么
先讲清它是什么。使用 prompt lookup 解码时,推理引擎会用一条规则来「猜」接下来的 k 个 token:把当前上下文里已有的 n-gram 拿出来,看历史上这个 n-gram 后面最常跟着哪个 token,然后把那个 token 当作草稿。设当前 token 序列为 x_1…x_t,一个 n-gram 就是 n 个连续 token。引擎为每个 n 选出得分最高的候选 token,再依据两个阈值决定是否真的把它当草稿:该 n-gram 至少出现过 a_n 次,且该候选词在后继中占比至少 p_n。llama.cpp 会依次尝试 n=4、3、2、1,采用第一个通过阈值的候选。因为草稿来自「文中已经出现过的东西」,所以它在代码编辑这类高重复场景里命中率很高。
# Build a static n-gram cache and then benchmark prompt lookup decoding.
# Both tools ship in llama.cpp's own lookup example, which is why the
# numbers in the write-up are reproducible rather than vendor claims.
# 1) build the static cache from a corpus (here: WikiText-103)
./llama-lookup-create -m ./models/model.gguf -f ./corpus/wiki103.txt -o w103.cache
# 2) replay a text file as if it were model output and measure drafting
./llama-lookup-stats -m ./models/model.gguf \
--lookup-cache-static w103.cache \
--ctx-size 4096 \
-f ./corpus/wiki103.test.txt
# The benchmark reports: drafted tokens matched, time to draft, and the
# static cache load time. Those three numbers are what the four changes move.在 Apple M4 Pro(14 核 / 48 GB)上测得
二、llama.cpp 的三类 n-gram 缓存
llama.cpp 维护三类 n-gram 缓存。上下文缓存(context cache)保存当前正在处理 token 的 1 至 4 元组,并随着生成新 token 而更新;动态缓存(dynamic cache)保存此前运行的 n-gram,例如更早的对话;静态缓存(static cache)保存来自静态语料库的 2 元组,用 llama-lookup-create 构建。打分时有一个容易被忽略的细节:当某个候选词同时被静态缓存认可时,它的权重会被乘上 100;也就是说静态缓存不只是「兜底」,它还会放大其它缓存里的一致候选。理解这一点,才能理解为什么缓存实现方式对性能影响如此之大。
# The acceptance-rate thresholds are hard-coded in llama.cpp and they are
# worth knowing before you try to "tune" anything. Drafting only proposes a
# token when the n-gram appeared often enough and the same follower is
# dominant enough. Lower these and you draft more but accept less.
# context cache (n-grams of size 1..4 from the tokens being generated)
CTX_A = (2, 2, 1, 1) # minimum occurrences a_n
CTX_P = (0.66, 0.5, 0.5, 0.5) # minimum dominance p_n
# dynamic cache (n-grams from previous runs of the model)
DYN_A = (4, 3, 2, 2)
DYN_P = (0.75, 0.66, 0.66, 0.66)
# The engine tries n = 4, 3, 2, 1 and drafts the first candidate that passes.
# A token y* is drafted only when: f(X_n, y*) >= p_n * F(X_n) and F(X_n) >= a_n.三、四个改动分别做了什么
四个改动依次展开。第一,是「别再拷贝 map」:llama.cpp 的 n-gram 缓存实现为嵌套的 std::unordered_map,而内层 map 在每一个草稿步骤上都被不必要地拷贝了多次,改成按引用读取后,草稿立即提速 4.5 至 25.6 倍。第二,把外层 map 从 std::unordered_map 换成扁平哈希表(作者选择了 ankerl::unordered_dense),因为标准库实现用链式冲突解决、对缓存不友好。第三,把每个 n-gram 的后继词存成有序数组,用定长二分查找取计数,代价可预测。第四,把静态缓存的外层 map 换成已验证的 constmap,因为静态缓存在加载后永不改变。
# Reproduce the scoring rule in Python before touching C++. Understanding
# the weight w(y) matters: the static cache is not just a fallback, it
# multiplies candidates from the other caches by 100 when it agrees.
def score(ctx, dyn, static, X_n, y, vocab):
base = ctx(X_n, y) if ctx(X_n, y) else dyn(X_n, y)
w = 100 * static(X_n, y) if static(X_n, y) > 0 else 1
return base * w
def draft(ctx, dyn, static, X_n, vocab, a_n, p_n):
F = sum(ctx(X_n, y) for y in vocab) # occurrences of the n-gram
best = max(vocab, key=lambda y: score(ctx, dyn, static, X_n, y, vocab))
if F >= a_n and score(ctx, dyn, static, X_n, best, vocab) >= p_n * F:
return best # accept the draft
return None # fall through to n-1四个改动,接受率保持不变
四、真实数字:165 微秒到 4 微秒
数字来自作者公开的基准:在 0、25、50、100、200、541 MB 语料下,上游每次草稿的耗时分别为 8.54、45.61、59.73、83.46、113.46、165.48 微秒;四项改动全部应用后为 0.89、3.06、3.25、3.32、3.47、3.98 微秒。静态缓存加载时间从 3.76 秒降到 0.23 秒(541 MB 语料);峰值内存从 1.71 GB 降到 1.31 GB,降幅最多 1.30 倍。测试环境是 Apple M4 Pro、14 核、48 GB 内存,取三次运行的中位数。最关键的一条:改动后每个语料上的接受率与原始实现完全一致——也就是说,变的是开销,不是行为。
# The last change is the interesting one: the static cache becomes a
# read-only constmap. Because it never changes after loading, the followers
# of every n-gram can live in one contiguous array, and the map stores a
# single 64-bit value holding a position and a count.
# value = (position << 24) | count for up to 16.7M followers
POSITION = value >> 24
COUNT = value & 0xFFFFFF
# lookup: one map probe, then a short fixed-length binary search in the array
entry_ptr, n_pairs = constmap_lookup(static_map, ngram_tokens)
count = binary_search(entry_ptr, n_pairs, candidate_token)
# Result from the write-up, 541 MB corpus:
# static cache load time 3.76 s -> 0.23 s (6.3x to 16.1x faster)
# peak memory 1.71 GB -> 1.31 GB (up to 1.30x less)
# acceptance rate identical on every corpus五、它对哪类工作负载真正有用
对工作负载的意义要诚实地说。作者明确没有改动 prompt lookup 的算法,因此改动影响的是「每次草稿的延迟」「静态缓存加载时间」「静态缓存内存」这三项,接受率不变。这意味着收益完全取决于你的文本有多重复:在代码编辑、重构、模板化输出里,上下文里大量内容会重复出现,草稿命中率高、收益明显;在自由对话里,重复少,草稿更多是纯开销,这时应当把 prompt lookup 关掉,而不是期待它带来加速。换句话说,这是一次对本地编码模型非常友好的免费升级,但并不是一个万能旋钮。
# Practical checklist for adopting the upstream PRs, in the order that
# pays off. The first two drop straight into any build; the last two change
# the static cache format, so they only matter if you use a corpus cache.
1. Read inner n-gram maps by reference, never by value. # PR: stop copying
-> 4.5x to 25.6x faster drafting depending on corpus size
2. Swap the outer std::unordered_map for a flat hash map (unordered_dense).
-> removes cache-unfriendly pointer chasing on every lookup
3. Store each n-gram's followers as a sorted vector, search with a
fixed-length binary search. # deterministic cost
4. Replace the static cache's outer map with a verified constmap.
-> load 6.3x to 16.1x faster, memory down to about the file size
# Sanity check before shipping: acceptance rate must be identical on every
# corpus, otherwise you changed behaviour, not just performance.165 微秒 → 4 微秒,来自作者公开的基准
六、怎么试、上线前的自检
要试用它,路径很短。llama.cpp 仓库自带 prompt lookup 的示例,其中 llama-lookup-create 用于从语料构建静态缓存,llama-lookup-stats 用于基准测试:它把一个文本文件当作「模型输出」逐 token 回放,记录草稿匹配数、草稿耗时与静态缓存加载时间。作者的代码与结果都在其公开仓库中,相关改动分别以 PR 形式提交。需要注意仓库已经从 ggerganov/llama.cpp 迁到 ggml-org/llama.cpp。示例 5 给出了一份采纳清单:前两项可以直接享受收益,后两项只有在使用语料静态缓存时才有意义。上线前的自检只有一条硬标准:接受率必须在每个语料上与原始实现一致。
📌 常见问题 FAQ
提速 42 倍是什么意思?
指的是草稿阶段的延迟最多下降 42 倍,而不是整体生成速度提升 42 倍;作者的改动不改变接受率,因此端到端的收益取决于文本的重复程度。
prompt lookup 解码是什么?
投机解码的一个特例:用一个 n-gram 模型代替第二个神经网络作为草稿模型,从上下文中已出现过的 n-gram 里猜下一个 token。
四个改动分别是什么?
内层 map 改为按引用读取、外层 map 改用扁平哈希表、内层后继词改为有序数组加定长二分查找、静态缓存改用 constmap。
数字是多少?
541 MB 语料下草稿耗时从 165.48 微秒降至 3.98 微秒;静态缓存加载从 3.76 秒降至 0.23 秒;峰值内存从 1.71 GB 降至 1.31 GB;接受率保持不变。
在什么硬件上测的?
Apple M4 Pro,14 核,48 GB 内存,模型上下文长度按 4096 token 计,结果为三次运行的中位数。
🔧 推荐工具
📚 参考资料
- Hayder Tirmazi — 42x faster prompt lookup drafting in llama.cpp (Sep 26, 2026)
- ggml-org/llama.cpp — lookup example (llama-lookup-create, llama-lookup-stats)
- ggml-org/llama.cpp — PR #5479, the change that added the static n-gram cache (benchmark method)
- jadidbourbaki/llama.cpp — reference pull requests for the four changes
- Merity et al. — WikiText-103, the corpus used for the benchmark