42x Faster Drafting in llama.cpp: A Free Win for Local Coding Models

·11 min read·Evergreen Tools Team

On September 26, 2026, engineer Hayder Tirmazi published a set of changes that make drafting for prompt lookup decoding in llama.cpp up to 42 times faster while using up to 2.6 times less memory. On a 541 MB corpus, the time to draft a single token falls from roughly 165 microseconds to roughly 4. Nothing about the model, the hardware, or the algorithm changed. Prompt lookup decoding is a special case of speculative decoding that uses an n-gram model instead of a second neural network as the draft model, which makes it a cheap way to accelerate repetitive work such as code edits.

1. What Prompt Lookup Decoding Actually Does

Start with what it is. With prompt lookup decoding, the inference engine guesses the next k tokens using a simple rule: take the n-grams already present in the current context, find the token that most often followed that n-gram historically, and propose it as a draft. Let the current token sequence be x_1 through x_t; an n-gram is n consecutive tokens. The engine picks the highest-scoring candidate for each n, then applies two thresholds before accepting it as a draft: the n-gram must have appeared at least a_n times, and the candidate must account for at least a fraction p_n of those occurrences. llama.cpp tries n equals 4, then 3, 2, 1, and drafts the first candidate that passes. Because drafts come from text that already appeared, hit rates are high in repetitive material such as code.

# Build a static n-gram cache and then benchmark prompt lookup decoding.
# Both tools ship in llama.cpp's own lookup example, which is why the
# numbers in the write-up are reproducible rather than vendor claims.

# 1) build the static cache from a corpus (here: WikiText-103)
./llama-lookup-create -m ./models/model.gguf -f ./corpus/wiki103.txt -o w103.cache

# 2) replay a text file as if it were model output and measure drafting
./llama-lookup-stats -m ./models/model.gguf \
  --lookup-cache-static w103.cache \
  --ctx-size 4096 \
  -f ./corpus/wiki103.test.txt

# The benchmark reports: drafted tokens matched, time to draft, and the
# static cache load time. Those three numbers are what the four changes move.
Local inference on the command line

Measured on an Apple M4 Pro, 14 cores, 48 GB

2. llama.cpp's Three n-gram Caches

llama.cpp maintains three kinds of n-gram cache. The context cache stores n-grams of sizes 1 to 4 from the tokens currently being processed, and updates as the model generates. The dynamic cache stores n-grams from previous runs, such as earlier conversations. The static cache stores 2-grams from a fixed text corpus, built with llama-lookup-create. One easily missed detail in the scoring rule: when the static cache also agrees with a candidate, that candidate's weight is multiplied by 100. The static cache is therefore not merely a fallback, it reweights candidates from the other caches. That detail explains why the cache implementation has such an outsized effect on performance.

# The acceptance-rate thresholds are hard-coded in llama.cpp and they are
# worth knowing before you try to "tune" anything. Drafting only proposes a
# token when the n-gram appeared often enough and the same follower is
# dominant enough. Lower these and you draft more but accept less.

# context cache (n-grams of size 1..4 from the tokens being generated)
CTX_A = (2, 2, 1, 1)          # minimum occurrences  a_n
CTX_P = (0.66, 0.5, 0.5, 0.5) # minimum dominance      p_n

# dynamic cache (n-grams from previous runs of the model)
DYN_A = (4, 3, 2, 2)
DYN_P = (0.75, 0.66, 0.66, 0.66)

# The engine tries n = 4, 3, 2, 1 and drafts the first candidate that passes.
# A token y* is drafted only when: f(X_n, y*) >= p_n * F(X_n) and F(X_n) >= a_n.

3. What Each of the Four Changes Does

The four changes, in order. First, stop copying maps: llama.cpp stores its n-gram caches as nested std::unordered_maps, and the inner maps were being copied unnecessarily on every drafting step. Reading them by reference instead made drafting 4.5 to 25.6 times faster immediately. Second, swap the outer map from std::unordered_map to a flat hash map; the author chose ankerl::unordered_dense because the standard library resolves collisions with linked lists, which is cache-unfriendly. Third, store each n-gram's followers as a sorted vector and use a fixed-length binary search for counts, which makes the cost predictable. Fourth, replace the static cache's outer map with a verified constmap, which is safe precisely because the static cache never changes after loading.

# Reproduce the scoring rule in Python before touching C++. Understanding
# the weight w(y) matters: the static cache is not just a fallback, it
# multiplies candidates from the other caches by 100 when it agrees.

def score(ctx, dyn, static, X_n, y, vocab):
    base = ctx(X_n, y) if ctx(X_n, y) else dyn(X_n, y)
    w = 100 * static(X_n, y) if static(X_n, y) > 0 else 1
    return base * w

def draft(ctx, dyn, static, X_n, vocab, a_n, p_n):
    F = sum(ctx(X_n, y) for y in vocab)          # occurrences of the n-gram
    best = max(vocab, key=lambda y: score(ctx, dyn, static, X_n, y, vocab))
    if F >= a_n and score(ctx, dyn, static, X_n, best, vocab) >= p_n * F:
        return best                              # accept the draft
    return None                                  # fall through to n-1
Code around the n-gram caches

Four changes, with acceptance rates unchanged

4. The Real Numbers: 165 Microseconds to 4

The numbers come from the author's published benchmark. Upstream takes 8.54, 45.61, 59.73, 83.46, 113.46 and 165.48 microseconds per drafted token for corpora of 0, 25, 50, 100, 200 and 541 MB. With all four changes applied, the same figures are 0.89, 3.06, 3.25, 3.32, 3.47 and 3.98 microseconds. Static cache load time drops from 3.76 seconds to 0.23 seconds on the 541 MB corpus, and peak memory falls from 1.71 GB to 1.31 GB, up to 1.30 times lower. The test machine is an Apple M4 Pro with 14 cores and 48 GB of memory, and every figure is the median of three runs. The most important line is this: the acceptance rate is identical to the original implementation on every corpus, so what changed is overhead, not behaviour.

# The last change is the interesting one: the static cache becomes a
# read-only constmap. Because it never changes after loading, the followers
# of every n-gram can live in one contiguous array, and the map stores a
# single 64-bit value holding a position and a count.

# value = (position << 24) | count          for up to 16.7M followers
POSITION = value >> 24
COUNT    = value & 0xFFFFFF

# lookup: one map probe, then a short fixed-length binary search in the array
entry_ptr, n_pairs = constmap_lookup(static_map, ngram_tokens)
count = binary_search(entry_ptr, n_pairs, candidate_token)

# Result from the write-up, 541 MB corpus:
#   static cache load time   3.76 s  ->  0.23 s   (6.3x to 16.1x faster)
#   peak memory              1.71 GB ->  1.31 GB  (up to 1.30x less)
#   acceptance rate          identical on every corpus

5. Which Workloads Actually Benefit

Be honest about what this means for a workload. The author made no algorithmic changes to prompt lookup decoding itself, so the work affects three metrics only: latency per drafted token, static cache load time, and static cache memory. Acceptance rates are unchanged. The payoff therefore depends entirely on how repetitive your text is. In code edits, refactors and templated output, much of the context repeats, drafts are accepted often, and the gain is real. In free-form chat, repetition is low and drafting is mostly pure overhead, so you should turn prompt lookup off rather than expect a speedup. In short, this is a free upgrade that suits local coding models very well, but it is not a universal knob.

# Practical checklist for adopting the upstream PRs, in the order that
# pays off. The first two drop straight into any build; the last two change
# the static cache format, so they only matter if you use a corpus cache.

1. Read inner n-gram maps by reference, never by value.  # PR: stop copying
   -> 4.5x to 25.6x faster drafting depending on corpus size
2. Swap the outer std::unordered_map for a flat hash map (unordered_dense).
   -> removes cache-unfriendly pointer chasing on every lookup
3. Store each n-gram's followers as a sorted vector, search with a
   fixed-length binary search.                          # deterministic cost
4. Replace the static cache's outer map with a verified constmap.
   -> load 6.3x to 16.1x faster, memory down to about the file size

# Sanity check before shipping: acceptance rate must be identical on every
# corpus, otherwise you changed behaviour, not just performance.
Drafting latency by corpus size

165 µs to 4 µs, from the author's published benchmark

6. How to Try It, and the One Check Before Shipping

Trying it is short work. llama.cpp ships an example for prompt lookup decoding, including llama-lookup-create to build a static cache from a corpus and llama-lookup-stats to benchmark. The latter treats a text file as if it were model output, replays it, and records how many drafted tokens matched, how long drafting took, and how long the static cache took to load. The author's code and results are in a public repository, and the changes were submitted as pull requests. Note that the project has moved from ggerganov/llama.cpp to ggml-org/llama.cpp. Code sample 5 gives an adoption checklist: the first two items pay off immediately, while the last two only matter if you use a corpus static cache. Before shipping, there is exactly one hard check: the acceptance rate must match the original implementation on every corpus.

📌 Frequently Asked Questions

What does 42x faster mean?

It refers to drafting latency falling by up to 42 times, not end-to-end generation speed. Acceptance rates are unchanged, so the end-to-end gain depends on how repetitive the text is.

What is prompt lookup decoding?

A special case of speculative decoding that uses an n-gram model instead of a second neural network as a draft model, guessing the next token from n-grams already present in the context.

What are the four changes?

Reading inner n-gram maps by reference, replacing the outer map with a flat hash map, storing followers in sorted vectors with a fixed-length binary search, and replacing the static cache's outer map with a verified constmap.

What are the measured numbers?

On a 541 MB corpus, drafting falls from 165.48 to 3.98 microseconds, static cache load falls from 3.76 seconds to 0.23 seconds, peak memory falls from 1.71 GB to 1.31 GB, and acceptance rates are unchanged.

On what hardware was this measured?

An Apple M4 Pro with 14 cores and 48 GB of memory, at a model context size of 4096 tokens, reporting the median of three runs.