Field guide · Local inference
Qwen3.8-27B and its Swift-1.5 and Pi fine-tunes on a 24 GB RTX 3090: measured settings, speeds and energy
Copy-paste server configurations, the speed and power each one delivers, and the measurement behind every recommendation. Three files — the anchor (unsloth UD-IQ4_XS of the base model), Swift IQ4_XS and Swift IQ3_XXS — are compared in one sweep (2026-10-04); Swift IQ4_XS carries quality scores on the 175-prompt suite, while Swift IQ3_XXS is ranked on perplexity, speed and appetite only (§15). A fourth, bartowski’s IQ4_XS of bytkim’s Pi-agent fine-tune Qwen3.8-27B-pi, was measured against Swift IQ4_XS on 2026-10-06/07 (Appendix B). Settings for other cards — RTX 30/40/50, DGX Spark (GB10), Intel Arc Pro B50/B70, Intel Arc B390-class iGPU — are calculated from the same arithmetic and marked as calculated.
Evidence tier. Everything here was measured on one machine — a 24 GB RTX 3090, driver 596.36, Windows 11 Pro 26200, i5-13600KF, llama.cpp build 10502 (commit 0adcc3bb5) — across 2026-08-21 through 2026-08-28, 2026-10-03 through 2026-10-05 and 2026-10-06 through 2026-10-07. Nine files of this model were ranked on three independent instruments; every claim comparing one file with another is paired on identical items. The October comparison sweep spent 11.41 GPU hours, failed attempts and re-runs included (one stopped attempt of about 3 minutes left no end record), on three files with the anchor in the same sweep — appetite, speed probes, drafter regimes, deep fills, perplexity, the scored suite, agent probes, vision and loaded-idle power; the knee sweep before it ran 69 server loads with 9.0 hours of context filling. A Pi-agent fine-tune, bytkim’s Qwen3.8-27B-pi, was measured on 2026-10-06/07 against its Swift-1.5 twins in its own sweep, with no anchor arm (Appendix B). Measured by Chin Keong Ang on one machine he owns, and written from those measurements by Claude models (the footer names the writers of the October blocks). Every correction, with the measurement that forced it, is listed in §15. Instruments, run logs and what was never measured are also in §15.
Full provenance — dates, hours, rounds, confidence tiers
GPU time. 2026-08-23 alone spent about 12.9 hours across five completed rounds: a two-hour independent re-measurement of this page's own figures, run without sight of them; a 1.7-hour follow-up round; an 8.48-hour benchmark sweep of 525 generations over seven benchmark sets; a 43.5-minute power matrix of 19 configurations; and two zero-GPU energy analyses. A sixth round, a quantization ladder, ran from 14:13 on 08-23 to 12:35 on 08-24 for about six hours more, and a seventh pass used no GPU at all — the blind judge panel of §09.
The ladder. It ranked nine files of this model on 294,912 scored token positions each, from 4.40 down to 1.83 bits per weight — how many bits a file spends on each of the model's numbers, averaged over the whole file — with automated checks beside every file for the failures a score cannot see: an empty answer, output that repeats until it hits the cap, a broken chat template. Eight of those files were then put through a second instrument on 2026-08-24 — a frozen three-benchmark suite scored on the identical 75 items per file, so that every claim comparing one file with another is paired, meaning only the questions the two answered differently count. A tenth file joined both instruments afterwards, on 2026-08-25 and 08-26 under those same conditions: sdkyuan’s quantisation-aware-trained QAT-Q2_0, which §08 prints as a counterexample rather than as a rung of this vendor’s ladder.
The 08-25 rounds. Three further rounds measured what the ladder could not: the drafter on and off on both candidate files; all three candidate configurations deep-filled to 218,233 real tokens at the model's full native window; and a requirement sweep of three files × three windows × the drafter on and off, each window deep-filled to about 90% of itself, which is the source of §08's memory-and-speed table and of the finding that the fastest file changes with the window. Those results fill §08, and §15's run log dates each round. The 08-21 and 08-22 rounds that produced the perplexity, quantization and drafter tables were not hour-logged.
The 10-03 to 10-05 rounds. A knee sweep (2026-10-03 to 2026-10-04) deep-filled eight launcher picks across 69 server loads (9.0 hours of filling). The comparison sweep (2026-10-04 to 2026-10-05) spent 11.41 GPU hours, failed attempts and re-runs included (one stopped attempt of about 3 minutes left no end record), on three files with the anchor in the same sweep: speed probes at three depths, drafter regimes and perplexity on all three files; appetite at xhigh on all three and at low and medium on the Swift files; the scored suite on the anchor and Swift IQ4_XS; deep fills on Swift IQ4_XS; vision on both Swift files; one agent probe (Pi) on Swift IQ3_XXS; and loaded-idle power on the anchor. Swift IQ3_XXS was cut from the suite for time when the recipes were fixed (§15); the needle tests did not run; the reasoning-regime speed group is void (§06). The provenance chain in §17 records what each sweep measured and what the anchor returned in it.
The 10-06 to 10-07 rounds. The Pi fine-tune study (2026-10-06 23:07 to 2026-10-07 10:34) ran bartowski’s IQ4_XS and IQ3_XXS of bytkim’s Qwen3.8-27B-pi with Swift IQ4_XS and Swift IQ3_XXS beside them and no anchor arm: server checks and a Pi smoke test on the IQ4_XS file; the 175-prompt suite at xhigh and medium on it, paired with the anchor’s and Swift IQ4_XS’s 2026-10-04 cells; speed probes and drafter regimes against Swift IQ4_XS (1,780 s and 2,388 s); deep fills at three windows in the order Pi, Swift, Swift, Pi (08:58 to 10:08); and a sampled drafter follow-up (1,300 s) that was not pre-registered. The other steps were pre-registered on 2026-10-06 before any measurement; the plan’s fixed 36 GB job memory cap was replaced by a recorded per-step cap after the first start was refused. The server checks, the Pi smoke test and the start of the xhigh suite overlapped a model download. No speed cell was void. One start of the server checks was refused by the memory-commit guard before it measured anything, and the suite waited from 01:06 to 06:19 for commit room. ALPACA and MT-Bench were not judged for the Pi file (Appendix B).
Confidence tiers. Solid, meaning measured more than once or on more than one file: the VRAM budget model (17 server loads with the whole model and its cache on the card, all reproduced within 127 MiB), the drafter's speed and its energy, the depth curve, decode energy per token, and the perplexity ranking. Smoke-tier, meaning one run or a small sample decided it: every quality judgement (n=1 or n=2 per effort level; the 08-22 pages were graded blind by independent reviewers, the 08-23 single runs were not), the 20-question and 25-question benchmark cells, and the whole vision section. That same rule binds the quantization ladder: its accuracy column — the ladder being nine versions of this one model, largest file to smallest — is n=25 per benchmark set, which is exactly the sample size this page tells readers not to rank files with — so it is used to locate where the model breaks and never to order one file against another, and every “tie” or “worse” in §08 comes from a paired test on identical items instead.
Not measured at all, and named rather than estimated: any card that is not this 3090 — so every statement about how fast a 16 or 12 GB card runs this model is derived, and always will be, while how much memory a given file and window need is a property of the model and the flags rather than of the card and is measured here (§08) — and any separate draft model beyond DFlash2 (which was measured and lost); the built-in ngram drafters are measured (2026-08-26). §15’s register itemises what is measured and what is not.C32 Two slots with the drafter on (n10/p0.5) gain 22.0% of aggregate throughput and cost 35.4% of each user's own speed (measured 2026-08-25). Board power with concurrent drafter-on slots is unmeasured as a decode-only figure: the October knee loads at two slots logged board power only as a window mean that includes the pauses between probe rounds, and this page does not publish it.
One instrument is not a measurement of this machine at all: the two open-ended benchmark sets are scored by a blind panel of three Claude Opus 5 judges reading the kept transcripts (§09) — a judgement, labelled as one wherever it appears.
-c 32768, 86.3 at a 1,458-token fill after the card had settled (§11)C27xhigh reasoning stream at that same depth-c 32768, 8 probes of 700 tokens, text only): ×1.154 (Pi), ×1.153 (Swift) der. 5-set suite composite (ALPACA and MT-Bench not judged) 80.6 (xhigh) and 81.2 (medium) against 81.3 (anchor) and 82.7 (Swift IQ4_XS) meas: no difference the suite can detect (Appendix B).This page tells you how to run one model family — Qwen3.8-27B and two fine-tunes of it, Swift-1.5 and bytkim’s Qwen3.8-27B-pi — on one graphics card, a 24 GB RTX 3090, and what to expect when you do. Every setting it recommends was measured on that machine, and the measurement that decides each recommendation is printed beside it. It also carries settings for other cards, marked clearly as calculated rather than measured. It is not a review, and not a set of numbers you should expect to match exactly on different hardware. It does compare three files of one architecture side by side in one sweep (2026-10-04) — the anchor (unsloth UD-IQ4_XS of the base model), Swift IQ4_XS and Swift IQ3_XXS — and every difference it finds belongs to the file (fine-tune, quantizer and size together), not to Swift-1.5 as a model (§17). A fourth file, the Pi fine-tune’s IQ4_XS, was measured on 2026-10-06/07 against Swift IQ4_XS in its own sweep and against the 2026-10-04 suite cells (Appendix B). If all you want is a working configuration, go to §03 and copy the block for your card.
The limits, stated plainly: one machine, one model family, one operating system. Every measured row on this page is the same RTX 3090 on Windows 11 with driver 596.36. Nothing here has been checked on Linux, on AMD, on Apple silicon, or on a second copy of the same card. Where a number was calculated instead of measured, it says so in the cell.
Three labels, and nothing on this page is a fourth thing. measured means it came off this machine on a stated date, with its conditions attached. derived means it was calculated from measured numbers by arithmetic printed nearby. cited means it came from somebody else's document, with a link. A single-sample figure carries n=1 beside it. Where a published block or cell differs in any field from the command actually measured — a different port, a different -c, a different effort level — the difference is named in that cell together with the reason it does not change the answer.
Conditions travel with numbers. A speed on this page is meaningless without four things: the file, the drafter flags, the token regime (see §02), and the prompt depth. A watt figure is meaningless without its instrumentation tier. Both are printed with every table.
Speed. One configuration was probed four times, each probe fired immediately after its prefill, and read 18.27, 18.82, 19.21 and 26.60 t/s (2026-08-23). The highest is 45.6% above the lowest; half of that spread is ±22.8%, which this page rounds to a ±25% band for any single probe taken right after a long prefill. The "45% swing" in §11 and the "±25%" band are therefore the same measurement stated two ways, and there is only one speed noise floor on this page. Under the cooled protocol (§02, documented in §11) — throw away the probe that comes back with the prefill, wait 30–45 s, take the median of several short repeats — repeat probes on this machine agree to 0.4%. The drafter-off decode floor holds inside 3.6% across five contents and both token regimes, which is why 3.6% is the pass band on the reproduction check in §15.
Speed, on the sampler you will actually use: a 25% spread from
one request to the next. Every speed number on this page was measured
greedy — temperature 0, top_k 1 — because that
holds the generated text still so the machine is the only thing changing.
Qwen's own model card recommends temperature 1.0, top_p 0.95,
top_k 20, and with the drafter on that produces different text on every
request, which changes mean draft length, which is what decides speculative
throughput. Measured 2026-08-25, 25 alternating pairs of one prompt in one
server load on a host-load-gated quiet machine
(-c 32768, n-max 4 / p-min 0.75, 700 tokens):
greedy read 73.93 t/s with a
3.8% spread, the recommended sampler
72.17 t/s with a 25.5% spread
— 59.1 to 77.5 t/s.
The extra scatter attributable to sampling alone is 5.6%, and
the whole measurement was run three separate times: greedy
0.99 / 0.76 / 0.77% and the recommended sampler
5.76 / 5.61 / 5.68%, the third being the one whose data file
is published. Within that run, decode speed tracks mean draft length at
r = 0.923 — the mechanism, not just the spread. So do not judge your setup
on one request. A single generation can land anywhere in that band
with nothing wrong; the average over several is what is stable, and the
recommended sampler also runs about 2.5% slower on average
than greedy (95% CI 0.2–3.5%). This does not affect the
reproduction check in §15, which pins
temperature 0 and top_k 1 precisely so that the reader is measuring
the machine rather than the dice.
Energy. One configuration measured three times in the 2026-08-23 power matrix reproduces joules per decode token to 2.9% on the slow arms and 5.6% on energy-delay product, tightening to 0.03% on the fast speculating arms. Separately, board power drifted ±6% between arms measured hours apart at constant throughput, so a 6% energy gap between two arms run at different times is the instrument, not the setting.
What survives those bands: shapes, ratios and rankings. Individual levels do not. Do not believe a slow-arm energy gap under about 3%, and do not believe a speed difference between two single post-prefill probes under about 25%. Printed precision on this page follows the band, with one deliberate exception. Single probes are rounded; cooled and multi-probe speeds are given to one decimal. The energy tables print four significant figures (8.104, 3.743, 3.210 and the rest) because that is what the integrator reports and because the ratios between them are the finding — but the energy noise floor is 2.9%, so the third digit of any of those is not meaningful and only their ratios and ranking are. Two figures are printed to four and five places on purpose, and neither is a level claim: 81.71 t/s and perplexity 6.5956 are reproduction matches between two independent runs, quoted at that precision because matching to the decimal is the point.
Comparisons between files in one sweep. In the speed probes at depth (§06), a speed ratio between two files is published only when its two passes agree within 3% of their mean — three to four times the greedy run-to-run scatter printed above (0.76 to 0.99% across its three runs, as a coefficient of variation; the 3.8% is one run’s full range). A group whose passes disagree by more is printed void. The one-prompt drafter comparisons (§06, §08) were not held to that band — their two passes differ by up to 4.1% — so read only differences far larger than that there. Perplexity and quality differences between files are read against their own printed errors and intervals (§08, §09).
The anchor (unsloth UD-IQ4_XS of the base model, 14.25 GB / 13.27 GiB), Swift IQ4_XS (15.48 GB / 14.41 GiB) and Swift IQ3_XXS (12.32 GB / 11.47 GiB) are measured side by side in one sweep (2026-10-04). All three files share 866 tensors with identical names and shapes meas and a shared tokenizer (gate passed 2026-10-04), so their perplexities sit on one scale: a quant rank between the two Swift files, but only drift from the corpus between a Swift file and the anchor, whose weights differ (§08). Ratios are to the anchor, because it is the file the base guide measured: a Swift figure is set beside the anchor only inside one sweep, and is carried across sweeps only as a ratio to it. The quantizer and file size differ between the anchor (unsloth) and the Swift files (bartowski), so every difference mixes fine-tune, quantizer and size — it belongs to the file, not to Swift-1.5 as a model. Run Swift IQ4_XS for single-turn coding and reasoning tasks at xhigh — it generated 0.446 of the anchor’s total tokens on the 175-prompt suite (greedy, xhigh, drafter off, single-turn, 2026-10-04; 0.52 over the 171 items both finished) der at no detectable quality cost n=25; the token saving is greedy and was not shown at the cards’ temperature-1.0 sampler n=2 (§17, §09, Appendix A).
Provenance. The anchor file hashes to 40fac4050e94… (full value in §17). All three files ran on llama.cpp build 10502 (commit 0adcc3bb5).
bartowski’s IQ4_XS of bytkim’s Qwen3.8-27B-pi (15.48 GB / 14.41 GiB), a fine-tune of this model for the Pi coding agent, was measured on 2026-10-06/07 against Swift IQ4_XS, whose tensor layout and chat template it shares. On the 175-prompt suite (greedy, drafter off) its 5-set composite is 80.6 at xhigh and 81.2 at medium (ALPACA and MT-Bench not judged), no difference the suite can detect against the anchor or Swift IQ4_XS n=25; with the drafter off it decodes at Swift IQ4_XS’s speed (1.001) der; in Pi it passed a three-probe smoke test n=1. Nothing measured shows an agent-task benefit from the fine-tune. Its launch command, the n-max 4 / p-min 0 drafter that command uses (re-measured 2026-10-08, §06), and what was not measured are in Appendix B.
The first several screens of this page are a buying guide. The instrument is below them. Two complete agentic runs on one RTX 3090 were instrumented end to end — board power and clocks by NVML at 0.5 s, the clock-event mask, per-request prompt depth and KV reuse, the server’s own prefill-against-decode split, and host CPU, syscalls and context switches on the same clock. Nineteen figures, each carrying its conditions on the figure itself.
Go to the figures: §16 — Power and performance has all nineteen inline, ordered by the question they answer.
Or read the standalone reports: UD-IQ4_XS agentic run · UD-Q2_K_XL agentic run — self-contained, every figure embedded, nothing fetched.
Run Swift IQ4_XS for single-turn coding and reasoning tasks at xhigh. It generated 0.446 of the total tokens of the anchor (unsloth UD-IQ4_XS of the base model) on this page’s 175-prompt suite (greedy, xhigh, drafter off, single-turn, 2026-10-04) der — a total that leans on the anchor’s capped items: 0.52 over the 171 items both finished, median per-item ratio 0.87, fewer tokens on 139 of 175 items — with no quality difference larger than the detectable gap at 25 questions per set n=25, a smoke-test scale (§10): the 5-set composite is 82.7 for Swift IQ4_XS against 81.3 for the anchor meas, difference +1.49 der, paired 95% bootstrap interval −0.2 to +3.9 (10,000 resamples, seed 42) including zero. With the drafter off it decodes at the same speed at 1.5k depth (ratio 1.005 der), so derived time over the suite is 0.44 of the anchor’s der (about 0.86 for the typical item) and measured energy 0.47 der. For the four agents this page tested only on the base model (aider, OpenCode, Qwen Code, DeepSeek Harness), the high effort level and the whole August evidence base, stay on the anchor’s default card. For two users or the longest single-slot Swift window, take Swift IQ3_XXS, at the cost of a perplexity gap against Swift IQ4_XS on their shared tokenizer (§08); it is ranked on perplexity, speed and appetite only — its quality suite arm was cut (§15).
With the drafter on (n4/p0.75, the setting the cards shipped until 2026-10-08; the Swift cards now run n3/p0 with one slot and n4/p0 with two, §06), both Swift files’ drafted thinking was accepted less on one prompt n=4 — reasoning-regime decode 59.96 t/s for Swift IQ4_XS and 50.61 for Swift IQ3_XXS against the anchor’s 76.23 meas — and the reasoning-regime measurement across depths is void (its passes disagreed beyond the 3% band in the sweep and again in a re-run on 2026-10-05), so with the drafter on part of the time saving may be given back while the model reasons. A third drafter setting measured on 2026-10-07, n-max 3 / p-min 0, decoded Swift IQ4_XS’s reasoning tokens ×1.153 faster than n4/p0.75 at the cards’ sampler (4 prompts × 2 seeds, -c 32768, text only) and ×1.087 greedy on the same one prompt der; the anchor did not run in that sweep, so how much of the gap to it this recovers is not measured (§08). The token saving is greedy: at the cards’ temperature-1.0 sampler, two appetite runs per file on one task did not show Swift IQ4_XS generating fewer tokens than the anchor n=2 (ranges, §15). Every difference in this section belongs to the file (fine-tune, quantizer and size), not to Swift-1.5 as a model (sameness, limits).
For the Pi coding agent there is one more file to consider, bytkim’s Qwen3.8-27B-pi, and nothing measured makes it the better choice. bartowski’s IQ4_XS of it shares Swift IQ4_XS’s tensor layout and chat template; on the same 175 prompts it shows no difference the suite can detect against the anchor or Swift IQ4_XS (5-set composite 80.6 at xhigh, 81.2 at medium; ALPACA and MT-Bench not judged) n=25, with the drafter off it decodes at Swift IQ4_XS’s speed, and in Pi it passed a three-probe smoke test n=1. No agent-task benefit was measured, and the author’s agent results are cited, not reproduced. Its launch command and its limits are in Appendix B.
Choosing a file and card by the job
Find your job in the table; each row names the file, the card to copy and the measurement behind the pick. The token, score and speed figures are from one sweep, 2026-10-04, and the two window figures from the knee sweep of 2026-10-03/04 and Appendix A’s arithmetic; speed bands are greedy-only (temperature 0) and are not measured at the temperature-1.0 sampler the cards ship (§15).
| Job | File | Card | The number that decides it | Scope |
|---|---|---|---|---|
| Single-turn coding or reasoning task at xhigh (not an agent loop) | Swift IQ4_XS | text · vision | 0.446 of the anchor’s total tokens (median item 0.87) der | 175-prompt suite, greedy, drafter off, single-turn; among the Swift files, Pi ran end to end on Swift IQ3_XXS only (§14) |
| Vision screenshot loop | Swift IQ4_XS or Swift IQ3_XXS | IQ4_XS vision · IQ3_XXS vision | 7 of 7 with the image, 0 of 7 without, both files meas n=1 per question | seven questions per file at the full image budget; an instrument check, not a quality ranking of the files against each other or the anchor (§12) |
| Long-context reading | Swift IQ3_XXS | text 1 slot | 216,064-token window der | fully resident at a 1,796 MiB desktop; quality at depth not verified (§15) |
| Two users at once | Swift IQ3_XXS | text 2 slots | 123,904 tokens per slot der | desktop up to 976 MiB |
| Pi coding agent, if you prefer bytkim’s Pi fine-tune | Qwen3.8-27B-pi IQ4_XS | Pi vision | no difference the suite can detect against the anchor or Swift IQ4_XS (5-set composite 80.6 at xhigh) n=25 | its own sweep, 2026-10-06/07, no anchor arm; a three-probe Pi smoke test only, no agent-task benefit measured (Appendix B) |
| A 12 or 16 GB card | base model: 16 GB · 12 GB cards; Swift files: Appendix A | Swift files: not measured — see §15 | Swift files: derived memory only | |
Three files, one architecture
Three GGUF files of the same architecture run every comparison in this section; the anchor is named throughout as “the anchor” (files hashed and headers read 2026-10-04).
| File | Quantizer and repository | Bytes meas | GiB der | bpw der | MTP draft head meas | Projector used (bytes) meas |
|---|---|---|---|---|---|---|
| Anchor (UD-IQ4_XS) | unsloth/Qwen3.8-27B-GGUF | 14,252,845,984 | 13.27 | 4.1735 | 334.75 MiB (F32, Q6_K, Q8_0) | 931,145,856 (BF16, lmstudio-community) |
| Swift IQ4_XS | bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF | 15,475,951,936 | 14.41 | 4.5316 | 227.91 MiB (F32, Q4_0) | 927,607,552 (F16) |
| Swift IQ3_XXS | bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF | 12,320,168,256 | 11.47 | 3.6076 | 227.91 MiB (F32, Q4_0) | 927,607,552 (F16) |
Every tensor name and shape is identical across all three files — 866 tensors each meas, with one presence-only array in the Swift files, so the architecture and the per-token memory of the window are measured once and apply to all three; file size and the drafter head differ per file (sameness). The three files share a tokenizer (gate passed 2026-10-04), so raw perplexity across files is a legal drift measure (§08).
What was held constant
The three files ran on one machine, one llama.cpp build (10502, commit 0adcc3bb5), one driver (596.36) and the same flags per instrument, in one sweep (2026-10-04) with the anchor interleaved between Swift arms wherever it ran (the knee sweep before it, 2026-10-03/04, ran its picks in a fixed order and carries no cross-model claim) — so a difference in any table below belongs to the file, not to the rig. Two quantities were not held constant: the quantizer and file size differ between files (limits), and the chat templates differ: they render the suite’s single-turn prompts identically, and the Swift template rejects high (sameness).
| Held quantity | Value | Artefact (paths are in the measured-inference repository, under results/swift-1.5-qwen3.8-27b/ except scripts/; cell.json is data/rule21/cells/*/cell.json; window-knee.py is scripts/bench/window-knee.py; agent-probe.py is work/agent-probe.py) | Where a card differs |
|---|---|---|---|
| Machine | NVIDIA GeForce RTX 3090, 24,576 MiB board VRAM, 31.8 GB host RAM, Windows 11 (build 26200) | machine.json | — |
| Backend and build | CUDA; build 10502, commit 0adcc3bb5, Clang 20.1.8; binary 9,216 B, mtime 2026-08-19. machine.json carries a stale tag, d7bd3bf, from an older install record; the build of record is the one llama-server.exe --version reported on 2026-10-04 | data/stage0/build.json | — |
| Driver | 596.36 | machine.json | — |
| Desktop reserve | 264 MiB at 09:38 on 2026-10-04 (n=8); fit tests used 258–855 MiB from 72 loads and one direct reading (mean 358.2, sd 121.0 MiB; fit limit 23,600.0 MiB of dedicated memory at depth) | machine.json, data/deepfill/g5-gate.json | — |
| Flags common to every arm | q8_0 KV cache (key and value), every instrument; --load-mode mmap explicit in the arm files, the server default in window-knee.py and bench.py; vision steps load --mmproj with --image-min-tokens 1024 --image-max-tokens 10580, other instruments omit it | arms/*.json, window-knee.py, cell.json | — |
| Sampler per instrument | greedy temperature 0 for speed probes, drafter regimes, knee and deep fills; greedy temperature 0.0, top_k 1, seed 42 for the suite; temperature 1.0, top_p 0.95, top_k 20, min_p 0.0 for appetite | arms/*.json defaults, cell.json | the speed band is greedy-only (temperature 0); speed at the cards’ temperature-1.0 sampler is not measured as a band (§15) |
| Caps and windows per instrument | suite: 32,768-token cap on a 65,536-token server window, max_prompt_tokens 8,192; appetite: -c 131072, uncapped (n_predict −1); speed probes: n_predict 300 and 400; drafter regimes: n_predict 700 | cell.json, arms/*.json | — |
| Effort per instrument | xhigh for the suite and appetite; low and medium only in the low/medium appetite runs; thinking off and thinking on as separate arms in the speed probes; knee and deep fills: --reasoning off | cell.json, arms/*.json, window-knee.py | — |
| Suite | hash 1cdf54f8eb9d3f8f, seed 42, 25 samples per benchmark | cell.json | — |
| Ports | 1238 (arm sweeps), 1236 (suite), 1241 (knee sweep and deep fills), 1234 (agent probe) | arms/*.json, cell.json, agent-probe.py | no measured quantity depends on the port |
| Job memory cap | 30–37 GB per queued GPU step on 2026-10-04 (the knee sweep’s loads: 38 GB, 32 for six, one at the 23.9 GB default); formula min(38, preflight), floor 30 GB (34 for the needle step); MEASURED_INFERENCE_MEM_CAP_GB | data/queue-conditions.jsonl | the published cards do not set a cap |
| Power instrument | NVML at 500 ms; power logger from 09:37, throttle loggers from 09:38 and 17:35 (in-band GPU board telemetry) | data/power/loggers.json | — |
| Host state | quiet for speed probes, drafter regimes and deep fills (CPU p95 13.19–15.32%), except the 12:40–13:00 download overlap during the speed-probe retry; not quiet for the suite (CPU p95 45.23–58.92%, available RAM down to 249 MB), whose wall times are host-conditioned | data/telemetry/swift-campaign-host.csv; the download overlap: §15 | — |
| Arm order | alternate in every arm file: pass 1 in order, the anchor first in each group, and pass 2 in reverse, so no arm owns the warm end; the low and medium appetite runs are Swift only; the suite reverses its two arms on every second benchmark | arms/*.json, cell.json | — |
Four departures from the plan apply to this sweep; the first two bear on every Swift table below. The quantizer and file size differ between the anchor (unsloth UD-IQ4_XS) and the Swift files (bartowski IQ4_XS and IQ3_XXS), so every speed or quality difference mixes fine-tune, quantizer and size. The scored suite ran with the drafter off, while the published cards ship it on: the drafter’s effect on speed is measured in the drafter-regime arms, its effect on scores is not. Every GPU step ran under a job memory cap (MEASURED_INFERENCE_MEM_CAP_GB); a step that needs more is stopped, not slowed, so the cap changes no measured quantity. Appetite ran before the speed probes, drafter regimes and deep fills, so its token counts could size the suite’s plan; the order changes no measured quantity either. The limits that follow are listed below.
Where the three files are the same
The model’s architecture, tokenizer arrays, KV cache geometry, suite prompt rendering and sampling defaults are identical across all three files (header diff, 2026-10-04), and the per-token memory slope measured on Swift IQ4_XS matches Appendix A’s within the 2% check. Where the base guide measured one of these on the anchor, that measurement applies to both Swift files without re-running it.
| The same in all three files (so the base guide’s measurement applies) | The evidence, and the section it lets you use |
|---|---|
No architecture keys differ between the anchor and either Swift file. One array, qwen35.attention.recurrent_layers (65 entries), is present in both Swift files and absent from the anchor — a presence-only addition, not a structural change. All 866 tensor names and shapes are identical across the three files. |
Header diff, 2026-10-04. Tensor architecture: §04. |
| The three tokenizer arrays (merges, token_type, tokens) are byte-identical across all three files, so raw perplexity across these files is legal — the corpus gate confirmed it on 1,290,590 bytes of corpus. | Header diff, 2026-10-04; corpus gate, 2026-10-04. Perplexity: §08. |
| KV bytes per token: 65,536 B at f16, 34,816 B at q8_0, 18,432 B at q4_0 — identical in all three files der. Native context length: 262,144 tokens. The formula: 16 full-attention layers × 4 KV heads × (256 + 256) = 32,768 values per token. | Derived from header shape. §05. |
| Per-token memory slope: 46.0 KiB/token der n=2 on Swift IQ4_XS — 0.28% from Appendix A’s 45.87 KiB/token, within the 2% check. | Two deep fills on Swift IQ4_XS: 139,264 and 159,744 tokens per slot (at load: dedicated 21,268 and 22,148 MiB, shared 398 and 438 MiB; 920 MiB over 20,480 tokens), 2026-10-04. Appendix A. |
| All 175 suite prompts render byte-identically through both templates. The three effort levels (low, medium, xhigh — where xhigh is the default) render identically in both templates. | Header diff template render, 2026-10-04. §09. |
| Sampling defaults in the header: temperature 1.0, top_p 0.95, top_k 20 — identical across all three files. meas | Header (general.sampling), all three files. |
What does not transfer. The MTP draft head is quantized to Q4_0 in both Swift files (227.91 MiB) against Q6_K and Q8_0 in the anchor (334.75 MiB), so drafter VRAM and drafter acceptance do not carry over — this page measured both for Swift (§08, §06). The token_embd and output tensors differ in type and byte count: the anchor carries token_embd Q3_K (546,304,000 B) and output Q5_K (874,086,400 B); Swift IQ4_XS carries IQ4_XS (675,430,400 B) and Q6_K (1,042,944,000 B); Swift IQ3_XXS carries Q4_K (715,161,600 B) and Q5_K (874,086,400 B) — yet with the drafter off Swift IQ4_XS, 8.58% larger than the anchor, decodes at the anchor’s speed (ratio 1.005), so decode speed is measured per file rather than predicted from size. Two tokenizer scalars differ at the header: padding_token_id is 248055 in the anchor and 248044 in both Swift files, and add_bos_token (false) is present only in the Swift files (header diff, 2026-10-04). The chat template diverges outside the suite’s single-turn prompts, and the high effort level, which the Swift template rejects, does not transfer (§14).
The provenance chain: each sweep is read against the anchor that ran in it
Absolute decode speeds do not survive across sweeps (§06), so a Swift figure on this page is set beside the anchor only inside one sweep and one build, and is carried across sweeps only as a ratio to the anchor. A new model that wants to join the comparison runs the anchor beside it, in its own sweep, and the anchor’s reproduction of its known figures is what ties that sweep back to the August data in §06–§09. The chain below records what each sweep measured, what the anchor returned in it, and what the sweep therefore licenses. The August sweeps’ own absolutes live in §15; they are not restated here as October numbers. The Pi fine-tune’s study (2026-10-06/07) is tied in differently: its speed and window steps ran Swift IQ4_XS (and, for one window, Swift IQ3_XXS) beside it, not the anchor, so its speed figures are ratios to Swift inside that sweep, and its suite cells are paired with the anchor’s and Swift IQ4_XS’s cells of row R7, reused rather than re-run (row R9).
| Row | Sweep | Dates | The anchor’s own absolute | Arms | What it licenses |
|---|---|---|---|---|---|
| R0 | August sweeps | 2026-08-21 to 08-28 | Cited by section from §15; not restated as October numbers | See §06–§10 | Same-sweep comparison |
| R1 | Knee sweep | 2026-10-03 21:58 to 2026-10-04 09:25 | The anchor’s four launcher configurations decode at their verified windows: 37.62 t/s at -c 180224 (vision on, n4/p0.75), 52.2 t/s at -c 188416 (n10/p0.5), 16.02 t/s at -c 262144 (drafter off), 43.81 t/s at -c 155648 (vision on, n10/p0.5) meas — all answer tokens, --reasoning off, temp 0, 400 predicted, --parallel 1. Swift IQ3_XXS’s four configurations (text and vision, one and two slots): 2, 7, 2 and 8 loads. 69 loads total, 32,431.9 s of fill. No build line in the server logs; the binary is the same 9,216-byte llama-server.exe as August, build 10502 (commit 0adcc3bb5) | Anchor + Swift IQ3_XXS | The anchor’s windows; the Swift IQ3_XXS loads are single-arm, no cross-model claim |
| R2 | Appetite (xhigh, low, medium) | 2026-10-04 10:06 to 12:10 | Generated tokens n=2: 22,718–80,213 (anchor, xhigh) — temperature 1.0, -c 131072, uncapped | Anchor + both Swift files at xhigh; Swift files only at low and medium | Same-sweep comparison at xhigh, n=2, reported as ranges |
| R3 | Speed probes | 2026-10-04 12:12 to 13:40 | Anchor against August, greedy (temperature 0): answer regime at 1.5k 79.14 t/s (August 86.3, ratio 0.917, the lower of this machine’s two levels), at 28k 72.97 (August 80.2, ratio 0.91), at 91k 60.07 (August 64.76, ratio 0.928) meas. The reasoning-regime group is void: its first pass overlapped downloads on the test host, the anchor’s included, and a re-run on 2026-10-05 was void as well | Anchor + Swift files | Same-sweep comparison (answer-depth group); the reasoning-regime group is void and licenses no ratio |
| R4 | Drafter regimes | 2026-10-04 13:40 to 14:24 | Anchor decode at -c 32768, greedy, 700 tokens: n10/p0.5 thinking off 100.81 t/s, n10/p0.5 thinking on 83.53, n4/p0.75 thinking off 87.42, n4/p0.75 thinking on 76.23, no drafter 42.88 meas | Anchor + Swift files | Same-sweep comparison |
| R5 | Swift IQ4_XS deep fills | 2026-10-04 14:24 to 14:37 | No anchor arm in this step | Swift IQ4_XS only | Single-arm characterisation; no cross-model claim |
| R6 | Perplexity | 2026-10-04 14:53 to 15:10 | 6.5956 ± 0.04453 meas — reproduced to the printed digit (exact reproduction, 0.0% drift) | Anchor + Swift files | Same-sweep comparison |
| R7 | 175-prompt suite | 2026-10-04 15:10 to 21:15 | Composite (32k view) 81.3, composite (16k view) 77.3 meas; 171 of 171 comparable answers byte-identical to August (response text and token count). Swift IQ3_XXS did not run a suite arm (§15) | Anchor + Swift IQ4_XS | Same-sweep comparison |
| R8 | Agent probe, vision, needles | 2026-10-04 21:15 to 21:47 | No anchor arm for these steps. The needle tests did not run: the commit-headroom wait expired (§15) | Swift IQ3_XXS + Swift IQ4_XS | Single-arm characterisation; no cross-model claim |
| R9 | Pi fine-tune study | 2026-10-06 23:07 to 2026-10-07 10:34 | No anchor arm. Its suite cells are paired with row R7’s anchor and Swift IQ4_XS cells, reused read-only; its speed and window steps ran the Swift twins beside the Pi files | Pi IQ4_XS and IQ3_XXS + Swift IQ4_XS and IQ3_XXS | Ratios to Swift within the sweep; suite comparisons paired with R7 (Appendix B) |
Rows R2–R8 are one sweep, 2026-10-04, whose reasoning-regime probes were re-run, void again, on 2026-10-05; row R1 is the knee sweep before it, 2026-10-03/04. Row R9 is the Pi fine-tune’s own sweep, 2026-10-06/07, with no anchor arm. The speed absolutes of rows R2–R8 meet the August sweeps of row R0 only as the anchor’s ratio in row R3 (§06).
The anchor file is 14,252,845,984 bytes with sha256 40fac4050e940397dbf13087afd50f4734a11805bf9d65ef8ddd7483470e6199, hashed 2026-10-04 meas. Its perplexity reproduces at 6.5956 to the printed digit (0.0% drift, exact reproduction). Its transcripts reproduce: 171 of 171 comparable (non-truncated) anchor items are byte-identical in response text and generated token count to the August transcripts. August recorded no sha256, so the identity of the anchor across the two sweeps rests on byte size, perplexity and 171 transcript matches, not on a hash recorded in August.
The single-arm characterisations — the Swift IQ4_XS deep fills, vision on both Swift files, the needle tests (which did not run; §15), and the Swift IQ3_XXS knee loads — license no cross-model claim. Appetite is a same-sweep comparison at n=2, reported as ranges: the anchor generated 22,718–80,213 tokens, Swift IQ4_XS 55,621–60,991, Swift IQ3_XXS 52,929–67,455 (xhigh, temperature 1.0, -c 131072, uncapped, 2026-10-04).
What this comparison cannot say
Every difference measured in this sweep mixes a fine-tune, a quantizer and a file size. The list below names the claims that do not follow from the data, with the reason each one fails and, where one exists, the measurement in the register (§15) that would close it.
- “Swift-1.5 thinks less” or “Swift-1.5 is faster” as a property of the fine-tune. The fine-tune, the quantizer and the bit-width all differ between the files: the anchor (unsloth UD-IQ4_XS) packs token_embd as Q3_K (546,304,000 B) against IQ4_XS (675,430,400 B) in Swift IQ4_XS, and output as Q5_K (874,086,400 B) against Q6_K (1,042,944,000 B); Swift IQ4_XS is 1,223,105,952 B larger (+8.58%). A same-quantizer pair, or the Swift BF16 weights, would separate them.
- The anchor’s drafter VRAM or acceptance rate applied to the Swift files. The MTP draft heads differ in precision and size: the anchor’s head is Q8_0 + Q6_K + F32 at 334.75 MiB; the Swift heads are Q4_0 + F32 at 227.91 MiB. Drafter VRAM and acceptance are measured per file (§06, §08).
- An August absolute set beside an October absolute. This machine shows two throughput levels across sessions (§03); the anchor sat on the lower one in October. Only ratios to the anchor within one sweep are comparable.
- The divergence (KL) of these files against the Swift-1.5 BF16 weights. The BF16 file is 54.66 GB; this machine has 31.8 GB of RAM. ukisai’s card KLD is for ukisai’s files, not bartowski’s. A KLD run against a BF16 or Q8_0 reference would close it (§15).
- Perplexity as a quality ranking between the base model and the fine-tune. The tokenizer gate passed, so raw perplexity is comparable as drift from wikitext, but across fine-tunes with different training data it is not a quality rank.
- “The same quality” beyond single-turn tasks of this suite. The 175-prompt suite has no read-then-edit task and no multi-turn agent loop. A read-then-edit task would narrow it (§15).
- The suite’s quality scores with the drafter on, as the cards ship. The scored arms ran drafter off; on the one drafter prompt with reasoning off, greedy text matched across drafter settings for Swift IQ4_XS (evidence, not proof) and differed for the anchor and Swift IQ3_XXS (cannot say either way). Suite runs with the drafter on would close it (§15).
- The judged pair as an independent witness. The judge seats are Claude, as is this page’s author. A second-vendor or human judge would close it (§15).
- Judged scores compared across judge sessions. The same 25 MT-Bench answers scored 79.7 and 87.0 in two sessions; on ALPACA the same answers moved +0.59 and +1.33 against the two earlier passes. Only paired comparisons inside one session are read (§15).
- A thinking-length difference at the cards’ own sampler. Appetite is n=2 per file at temperature 1.0 on one task; the token saving is greedy-only. More appetite samples would narrow it (§15).
- Swift IQ3_XXS results applied to another quantizer’s IQ3_XXS. bartowski’s quantizer and recipe differ from other quantizers’ IQ3_XXS files.
- Swift speed at the cards’ temperature-1.0 sampler as a band. Every Swift speed band on this page is greedy; the temperature-1.0 decode is a cross-reference from two appetite runs per file, not a band. A sampler band would close it (§15).
- The reasoning-regime speed of any file in this sweep. That probe group is void in both runs: in the sweep its first pass overlapped downloads on the test host, and a re-run on 2026-10-05 failed the same 3% agreement between passes in five of its six Swift cells. More passes per cell would let the spread be measured rather than tested against the band (§15).
- The energy per token of the drafter-on cards beyond one prompt, and at the pairs they carry now. The drafter-regime energy (n10/p0.5 and n4/p0.75, 2026-10-04) rests on one prompt of 700 generated tokens (14.5–33.1 s per pass), and the board sat at its software power cap in 97–100% of active samples, so joules per token restates throughput. A drafter-regime sweep over more prompts, with the power logger running, would close it.
- The drafter settings as optimal for the Swift head — narrowed 2026-10-08. The n/p sweep ran a grid at the cards’ sampler and
xhigh, shallow and after a 57,540-token prefix, at-c 81920: on Swift IQ4_XS no drafter, n2–n6 and n10 at p-min 0, n4 at p-min 0.25, 0.5 and 0.75, and n10/p0.5; on Swift IQ3_XXS n3, n4 and n5 at p-min 0 and n4/p0.75 with one slot, and n3/p0, n4/p0 and n4/p0.75 with two slots busy. Best score: n3/p0 with one slot (shallow reasoning tokens 1.189 and 1.097 of n4/p0.75; after the prefix 1.052 and 1.007, intervals including 1) and n4/p0 with two (1.337 per slot, second pass, one load per setting) der (§06). Still open: the optimum at each card’s own deep-filled window with an image, n2 on Swift IQ3_XXS, other prompts and agent use, quality with the drafter on, and n-gram or draft-model drafting (§15). - Swift behaviour in aider, OpenCode, Qwen Code or DeepSeek Harness. None has been tested; the Swift template differs from the anchor’s (unsloth) template and rejects
highas a reasoning effort. It is byte-identical to the template in lmstudio-community’s Q4_K_M of the base model (8,952 characters, header read 2026-10-06), so the difference is the unsloth file’s, not the base model’s.C40 A probe per agent would close it (§15). - An agent benefit from the Pi fine-tune. bytkim’s Qwen3.8-27B-pi ran only a three-probe smoke test in Pi (n=1 each); its suite scores show no difference the suite can detect, and the author’s agent results are cited, not reproduced. An agent-task benchmark with repeats would close it (Appendix B, §15).
- ukisai’s Swift files against bartowski’s on quality. The two repos carry the same recipe in 21 of their 22 shared tiers, with a different imatrix (Appendix A). A KLD pass per tier for each repo’s file against the same reference would close it (§15).
- The quality of the 12 and 16 GB Swift files in Appendix A. No Swift file below IQ3_XXS ran any quality instrument. A suite run per file would close it (§15).
- Swift IQ3_XXS’s quality on the 175-prompt suite, and quality at depth for any Swift card. Swift IQ3_XXS was cut from the suite for time when the recipes were fixed, and ranks on perplexity, speed and appetite only; the needles did not run, so the deep-window cards are speed-verified but quality-unverified at depth. The register carries both gaps (§15).
These are the technical words this page uses without stopping to explain them again. Each is defined once, here, with one of this campaign's own numbers attached wherever a number makes it concrete. Every later section links back to this box instead of defining anything again.
- token
- A piece of a word. The model reads and writes tokens, not letters. English text runs roughly 3–4 characters per token, so a 700-token answer is about half a page. Speeds on this page are tokens per second, written t/s.
- context window
- The total number of tokens the server will hold at once — your prompt, the model's thinking, and its answer, all in one pool. You set it with the
-cflag. This model can go up to 262,144; what actually limits you is memory. The recipes here ship 122,880 as the everyday value. Depth on this page means how much of that window is already filled when a measurement is taken: "64.8 t/s at 91k of depth" means the prompt already held 91,000 tokens. - slot
- One conversation the server will work on at a time, set with
--parallel. One slot means one request is answered at a time and it gets the whole window and the whole card; two slots serve two people at once, splitting both. Measured here (UD-IQ4_XS, drafter at n10/p0.5, 2026-08-25), two slots deliver 22.0% more tokens in total and 35.4% fewer to each person. - VRAM
- The memory on the graphics card itself. The reference card has 24,576 MiB of it. Your card's own memory is reported as dedicated; memory the card borrows from the system is reported as shared. When the model plus its context window need more than the card has, the extra spills into ordinary system memory and everything slows down (see spill below).
- resident
- Living in the card's own memory rather than in system memory. A file's resident size is what it occupies once loaded, and it is not the size printed on its download page: UD-IQ4_XS is 14.25 GB there and 13.3 GiB resident. A window is fully resident when every byte of its cache is on the card too. Past that point the server still starts and still answers a short prompt — which is exactly what makes it dangerous (§05).
- the three memory counters
- Three different numbers describe the same card, they do not agree, and this page always says which one it means. The server's own report at load counts only the server process — budget with this one; it reproduced to 127 MiB across 17 loads. Board VRAM, from
nvidia-smi, is the whole card, so it also contains your desktop and your browsers. Dedicated GPU memory in Task Manager is what the card holds in its own memory, and it cannot see a spill; the shared figure beside it is the one that can. This page writes board VRAM and board power in full, and never lets the word "board" stand alone as a number. (Board here is the graphics card itself — the circuit board with the chip, the memory and the fans on it.) - headless
- No graphical session is using this card, so the model gets all of its memory. Unplugging a monitor does not achieve this on Windows: the Desktop Window Manager keeps running and keeps its share, measured here at 1,179–1,669 MiB in direct no-server readings. What does achieve it: close everything that draws, drive the display from a second graphics card, or connect from another machine with nobody logged in locally. This page uses the word for that condition and for nothing else — a browser or a coding agent running with no window on screen is a different thing entirely, and is written out in plain words wherever it appears.
- desktop VRAM share
- How much of the card's memory your own screen is using: the window manager, the browser, anything else drawing. Measured here at 1,179–1,669 MiB in direct no-server readings on one machine. It is a quantity that depends on what is on screen, not a switch you can throw, and the three ways to shrink it are the three listed under headless above.
- slack
- The VRAM left on the card once the server has taken its share. It is what your desktop, your browser and load-to-load variation have to fit into. This page judges slack against one threshold, 1,796 MiB der, and it is important to be exact about what that number is. Two measurements go into it: the desktop's own share of board VRAM, which measured 1,179–1,669 MiB in direct no-server readings depending on what was on screen, and 127 MiB of load-to-load variation in the server's own report. The 1,796 itself is derived — the worst case of the first plus the second — and it is this page's own threshold for calling a configuration usable with a screen attached. The desktop does not “need” 1,796 MiB. It needed anywhere from 1,179 to 1,669 on this machine, and 1,796 is the ceiling this page plans against so that a bad day still fits. Subtract your own desktop instead if you know it: the number that matters is what your screen holds, not this one's worst case.C13
- bandwidth
- How fast a card can read its own memory, in gigabytes per second. It is the single number that decides generation speed here, because the card must read the whole model once per token. The reference card is 936 GB/s; a laptop's integrated graphics is nearer 150.
- projector
- The extra file (
mmproj) that lets the model see images. It is optional at start-up. Measured here it occupies 1,138 MiB of VRAM and costs nothing in speed. - quantization
- Storing the model's numbers with fewer bits so the file is smaller and faster to read. A 4-bit file of this model is about 13–15 GiB instead of 54. It costs a little quality: measured here, the gap between the two 4-bit files this page recommends is 0.9% of perplexity. Bits per weight is the unit: how many bits the file spends on each of the model's numbers, averaged over the whole file. It is measured from the file rather than read off its name — a file called
Q2_K_XLmeasures 2.912 (§08). Two families appear on this page and the filename tells you which is which: a name containing IQ (UD-IQ4_XS, UD-IQ3_XXS, UD-IQ2_S) is an IQ-quant; a name containing K without IQ (Q4_K_M, UD-Q3_K_XL, UD-Q2_K_XL, Q6_K) is a K-quant. The two decode at different efficiencies, which is the one place the distinction changes a number you work out for yourself (§04). - KV cache
- What the model remembers about the tokens already in the window, kept in VRAM so it does not have to re-read them. It grows with the window. Measured on this model, a server allocates 39,936 bytes per window token with no drafter and 45,056 with one.
- prefill and decode
- Two different phases with two different speeds. Prefill is reading your prompt: measured here at about 1,000 t/s on short prompts, slowing to 816 t/s on a 92,679-token one, which took 1.9 minutes. Decode is writing the answer, one token at a time: 40–90 t/s here. Every speed on this page is decode unless it says prefill.
- prose, and novel code
- Prose is ordinary written English — sentences and paragraphs, the kind of text in documentation, an article or an email. Code is programming text. This page keeps saying which of the two a speed was measured on, and that is not padding: the drafter guesses far better on code than on prose, because code is predictable. After
function foo(almost certainly comes); indentation repeats; brackets close in known ways. The next word of an English sentence could be any of hundreds. So the same drafter setting that wins by 12 % on code loses by 10 % on prose (§06), and a speed figure without its content type is not a number you can plan with. “Novel code” means a program the model has not effectively memorised — textbook algorithms like a red-black tree inflate the drafter’s hit rate, so this page’s drafter measurements deliberately ask for something less familiar. - wikitext
- A standard, freely available body of English text (Wikipedia articles) used as the fixed reading material for perplexity scoring, so that every file on this page is judged on identical words. It is prose, which is why the perplexity ranking and the speed rankings can disagree: they are measured on different kinds of text.
- acceptance, and mean draft length
- Two numbers the server reports about the drafter, and neither alone predicts speed. The drafter proposes several tokens at once; the full model then checks them and keeps the ones it agrees with. Acceptance is the share it kept. Mean draft length is how many tokens it proposed per attempt. High acceptance is not the same as fast — if the drafter only dares propose two tokens at a time, it can be right almost always and still save you little. And long drafts are not the same as fast either — if the drafter proposes ten tokens and discards half, every discarded token was still computed. Measured on this page: a setting at 70.5% acceptance and 2.803 mean draft length ran 8.4% faster than one at 51.6% and 4.035, because the first wasted less work per accepted token.C24 The quantity that ranks configurations is accepted tokens per unit of drafting work, not accepted length alone.
- answer tokens vs reasoning tokens
- With thinking switched on, this model writes two things: a private working-out pass, then the reply you actually read. Reasoning tokens are the working-out; answer tokens are the deliverable. They run at different speeds — answers about 1.7× faster than reasoning on code — so a “tokens per second” figure means two different things depending on which it counted. Every speed band on this page says which. If you are waiting at a keyboard for a reply, the answer-token rate is your experience; the reasoning rate is what decides how long you wait first.
- drafter (speculative decoding)
- A small helper that guesses several tokens ahead so the full model only has to check them in one pass. Every guess is checked before it counts, so it is meant to change only the speed; on one prompt on this build (greedy, thinking off), the text with the drafter on was not always identical to the text without it (§08).C38 This model has one built in, called an MTP head, switched on with
--spec-type draft-mtp. Measured here it roughly doubles decode speed and cuts energy per token by 2.52× (n10/p0.5 against no drafter, 2026-08-23). Two numbers describe how well it is doing: draft acceptance is the share of guesses that survive checking, and mean draft length is how many tokens it dares to guess per pass. Acceptance tells you whether the guessing was right; neither it nor mean draft length alone ranks settings that differ in n-max (§08, §08). Two flags tune it. n-max (--spec-draft-n-max) is how many tokens the helper may guess before the full model checks them — guessed one after another; the recipes the 2026-10-08 sweep covered ship 3 with one slot and 4 with two slots and for the Pi fine-tune, the others 4 (4 or 10 before then). p-min (--spec-draft-p-min) is a confidence floor: the helper stops at the first guess it is less sure of than that, its sureness being its top pick’s probability renormalised over its ten best candidates, which the request’s sampler never moves — 0 on every drafter-on recipe card, where it never ends a draft (measured on the swept recipes; derived on the others since 2026-10-09 der: p-min costs no memory, and on the four files swept p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot; 0.75 or 0.5 in older measurements; §06). The page writes the pair as n-max 4 / p-min 0.75, shortened to n4/p0.75. - token regime
- Which kind of token is being counted. Reasoning tokens are the model thinking; answer tokens are what you keep. This model thinks by default, so most speed numbers published anywhere are reasoning tokens. On code they run about 1.7× slower than answer tokens on the same machine with the same flags. A speed without its regime cannot be checked.
- effort
- A dial on this model —
low,mediumorxhigh— that sets how long it thinks before answering. Measured across a 175-prompt benchmark suite, it moved wall-clock from 1.0 to 2.7 hours and the composite score by 1.6 points, which is a tie at that sample size. - perplexity
- A quality measure for a model file, not a search engine. Run the model over real text and score how well it predicted each next token. A perplexity of 6.5 means it was on average as uncertain as choosing between about 6.5 equally likely words. Lower is better, and it can only be compared between files of the same model family scored on the same text with the same tokenizer.
- spill
- When the card runs out of VRAM, the driver quietly moves part of the model or its cache into system memory instead of refusing. Nothing errors. The signature is shared GPU memory rising during a run while dedicated sits at its ceiling (§11). Measured here: on a 91k-token document, the same file decoded 30.6 t/s in a window that fits on the card and 8.0 t/s in a
-c 262144window that spilled 3,502 MiB at load.C29 - percentage point vs percent
- A point is an absolute unit of a score: 94% to 95% is one point. "Percent better" is relative and means something else. Points and questions are chained: in a scored run of n=20 each question is worth 5 points, at n=25 each is worth 4 points, and at n=200 each is worth 0.5. That is why a 25-question benchmark cell cannot separate two healthy files (§10).
- board power, in-band, J/token, tokens/kWh, EDP
- Board power here means what the graphics card itself draws. It is read in-band — from the card's own sensor (NVML), not from a meter at the wall. So it is not wall power: the power supply, the processor and the rest of the machine are excluded and were never measured. J/token is joules of board energy per decode token — measured here from 3.21 (n10/p0.5, 2026-08-23; energy at the pairs the cards carry now is not measured) to 9.29 (worst quant, no drafter). tokens/kWh is the same fact upside down: one kilowatt-hour buys 387,000 to 1,121,000 tokens on this card. EDP is energy-delay product, energy multiplied by how long it took, in joule-seconds — it is the number that punishes a slow setting twice.
- arm
- One configuration inside a comparison run. Two arms differ in exactly one setting, so whatever separates them can be blamed on that setting. A 19-arm power matrix is nineteen such configurations measured one after another.
- rung, and the ladder
- The ladder is this page's series of one model squeezed to different sizes, largest file to smallest; each file in it is a rung. §08 measures eight of them.
- GGUF
- The single-file format llama.cpp loads. One
.gguffile holds the whole model, including the draft head on the files this page recommends. - prefix cache
- The server keeps the processed form of a prompt it has already read, so a later request whose text starts identically skips re-reading that part. Measured here, reading a 92,679-token prompt takes 1.9 minutes and is 90.7% of that request's energy (§05) — a reused prefix pays none of it.
- cooled protocol
- How this page takes a speed measurement that repeats. Fire the prompt, throw away the decode timing that comes back with it, wait 30–45 s, then send several short repeat requests and report the median of those. The first probe after a long prompt reads up to 25% low, because the card is still raising its clock (§11); under this protocol repeat probes agree to 0.4%.
- blind
- Used on this page for judging only: a judge who cannot see which setting wrote the answer being rated (§09). A second kind of blinding also happened here, and it is deliberately called something else — the 2026-08-23 round that re-measured this page's own figures without sight of them is called the independent re-measurement throughout, so the two can never be confused.
- functional detector
- An automated check for the failures a score cannot see: an answer that comes back empty, output that repeats until it hits the cap, a broken chat template. In §08 they are the empty-answer and truncation columns printed beside perplexity and accuracy.
- empty answer · at cap · silent
- Three words this page's failure columns use, and they name different failures. An empty answer is a reply with no characters in it: you asked, and nothing came back. At cap means that reply stopped because it ran out of its token budget and was cut off — the server records that as a truncation, so it shows up in a log. Silent means the opposite: the model ended the reply by itself, in the ordinary way, and handed back nothing. A silent empty is the dangerous one, because no truncation counter can see it and nothing in your logs reports a problem at all. Measured here, empty answers are exactly zero down to 2.912 bits per weight and then rise; five of the five at 1.994 bits are silent (§08). All of those counts are greedy-decoding measurements. Under the sampler this page's recipes actually ship, 300 generations per file produced no blank answers at all on the three files tested (§08).
- paired test, discordant, and
p - Paired means the same questions were put to both files and compared one by one, so only the questions they answer differently count; those are the discordant ones.
pis how often a difference this large would turn up by chance if the two were really equal —p=1.00 means constantly,p=0.0001 means almost never. - greedy decoding, pass@1, ROUGE-L F1
- Three scoring words that appear in this page's benchmark tables. Greedy: always take the most likely next token, so with the drafter off the same prompt gives the same answer every time; with the drafter on, this build did not always repeat itself (§08).C38 pass@1: the share of problems whose first generated program actually runs and passes its tests. ROUGE-L F1: how much of a reference summary's wording a generated summary reproduces, scored out of 100.
- Ampere, Blackwell, Battlemage
- Names for generations of graphics chips, used here because the vendors and the other documentation you will read use them. Ampere is NVIDIA's RTX 30 series — this page's own 3090 is an Ampere card. Blackwell is the RTX 50 series, GB10 and B200. Battlemage is Intel's Arc B series, including the B580 and the Arc Pro B50 and B70.
xhigh run whose thinking wants 61,000–76,000 tokens fails by arithmetic inside a 49,152-token window, and returns nothing (2026-08-22 and 2026-08-23, RTX 3090).One recipe per class of card. Find yours below, copy the block. The reasoning behind every flag is in §04 to §09.
| Card class | File | Jump to recipe |
|---|---|---|
| NVIDIA 24–32 GB (3090 / 4090 / 5090) | UD-IQ4_XS | default · speed · big window · Q4_K_M |
| NVIDIA 24 GB (3090) | Swift IQ4_XS | text · vision |
| NVIDIA 24 GB (3090) | Swift IQ3_XXS | text 1 slot · text 2 slots · vision 1 slot · vision 2 slots |
| NVIDIA 16 GB (5080 / 4080 / 4070 Ti S / 5060 Ti) | UD-Q2_K_XL | recipe |
| 12 GB (3060 / 5070 / Arc B580) | Q4_K_M + CPU offload | recipe |
| Intel Arc Pro B70 · 32 GB | Q6_K | recipe |
| Intel Arc Pro B50 · 16 GB | UD-Q2_K_XL | recipe |
| DGX Spark · GB10, 128 GB | Q4_K_M | recipe |
| Intel Core Ultra iGPU | Q4_K_M | recipe |
24 GB readers: two files compete. Take the default (UD-IQ4_XS) unless you want the highest measured quality at a 122k window, in which case take the Q4_K_M and accept that it leaves only about 0.3 GiB for your desktop. The two are statistically tied on quality and the default is smaller and faster in every comparison — details at the Q4_K_M recipe.
NVIDIA 24–32 GB · 3090 / 4090 / 5090 — THE DEFAULT · UD-IQ4_XS (vision, speed, big window)
:: TAKE THIS ONE for everyday use: quality statistically TIED with Q4_K_M (6.596 vs :: 6.535, +0.9%, smaller than the +/-0.063 combined error bar), measured FASTER :: everywhere (decode 42.97 vs 39.99 with no drafter and 93.9 vs 81.7 at n10/p0.5 at -c 32768, :: both from one matched sweep; prefill 1365 vs 1303), 2.1 GiB smaller — and :: this block ships VISION on. The Q4_K_M block below is kept so that an option you :: have already heard of is answered rather than absent; take it only if you want the :: highest measured quality at a 122k window and can free the VRAM your desktop is :: using - it leaves only about 0.3 GiB. Caveat: IQ-format decode (files whose name contains IQ) was measured on :: CUDA only - verify on Vulkan or Metal before trusting these speeds there. :: file: unsloth/Qwen3.8-27B-GGUF (UD-IQ4_XS, 14.25 GB = 13.3 GiB) + BF16 mmproj llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 :: min 1024 is a FLOOR only — it lifts small images and never :: shrinks a big one (measured: 720p 922 -> 1,077 tokens, 1440p :: unchanged at 3,602). max 10580 = 4K-class detail. Per-image :: token costs: section 12. The projector costs 1,138 MiB of :: VRAM and 0% of decode (0.04-0.09% at a 91k fill, section 06) -c 122880 :: 20,658 MiB of server VRAM by the measured budget model :: (section 05), leaving 2.7-3.7 GiB for a desktop that itself :: measured 1,179-1,669 MiB idle, no server. The largest window :: whose every byte still lives on the card in THIS configuration :: is 163,840; 122,880 is the daily value, chosen to leave a :: desktop room -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: -ngl 99 is not a guess: llama.cpp counts the output layer as :: layer 65, so -ngl 64 leaves the output layer on the CPU and :: costs 29.5% of decode with NO memory signature (section 11) --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 :: n3/p0 since :: 2026-10-08: the best score for this file at this card's :: sampler and xhigh (n/p sweep, -c 81920, text only, two :: sessions): shallow reasoning tokens 1.111x n4/p0.75 :: (paired 95% 1.075-1.156), answers 1.120x; after a :: 57,540-token prefix 1.061x, interval including 1. :: n10/p0.5 read 0.961x and 1.049x, and 1.079x after the :: prefix, intervals overlapping (section 06). :: About 150 MiB LESS than n4, so this window only gains :: margin. n10/p0.5 was the August grid's peak on GREEDY, :: thinking-off code: 93.9 vs 83.5 t/s at -c 32768 (section 06). :: On PROSE (the August greedy probe), drop speculation: :: 1.16x for ~1.8 GiB, measured at n4/p0.75 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: all three levels fit this window ONLY IF you leave :: room for them. The window is one POT: prompt + thinking + :: answer share it. xhigh wants 61,500-75,800 tokens of :: THINKING, so your prompt and history must stay under about :: 47,000 of these 122,880 - 38%% of the window. Past that, :: xhigh truncates and returns NOTHING (section 09). medium :: when you are waiting at the keyboard, and whenever the :: context is deep :: FLAGS DELIBERATELY OMITTED HERE: :: --parallel 2 — serving one person, one slot is faster: a second slot costs 35.4% :: of per-slot throughput (85.79 -> 55.41 t/s) while buying 22.0% of aggregate :: (82.98 -> 101.25 t/s), measured 2026-08-25 on the matched drafter-on pair. :: Raise it only to serve two users at once (section 03's axis table). :: -ctk q4_0 — measured at +0.693% perplexity against fp16 (q8_0 costs +0.309%). :: A knowing trade for a bigger window, not a default (section 08). :: nvidia-smi -pl — MEASURED 2026-08-25, re-measured 2026-08-28 in a saturating :: regime (SwPowerCap residency 99-100%). Capping the board is a real :: efficiency win: 300 W costs 4.1% of throughput for 13.1% less board power :: and 9.4% less energy per token; 250 W costs 14.6% for 27.1% less power :: and 14.7% less energy. The non-saturating sweep (2026-08-25) read :: 5.0%/6.5% and 12.8%/13.7% — understated the deal at 300 W because its :: stock arm sat 32 W under its own cap. Left out of the recipe because it is a persistent :: hardware setting and needs an elevated shell, not because it does not :: work (section 11). :: text-only variants — drop the --mmproj/--image-* line and the window grows: :: 2026-10-08: text-only, keep the n3/p0 drafter above: it holds about :: 7 x 150 MiB less than n10 at the same window (derived, not filled). :: The ceilings below were measured with the drafters they name. :: -c 180224 + n-max 10 / p-min 0.5: the measured text-only ceiling WITH that :: drafter - 23,729 MiB of board VRAM, 847 MiB of slack left for a desktop. :: (196,608 leaves 415 MiB and 212,992 leaves 280, with decode already :: sagging 56.6 -> 51.9 -> 49.2 on a :: SHORT probe. The 847 MiB at 180,224 does not clear the 1,796 MiB reserve.) :: -c 262144 THE FULL NATIVE CONTEXT - but ONLY with --spec-type none AND with no :: graphical session using this card ("headless", section 02: closing everything :: that draws, a second card driving the screen, or a remote login with nobody :: logged in locally - NOT just unplugging the monitor). :: Drafter off it derives to 23,216 MiB (22.7 GiB) at load, leaving 1,360 MiB of :: board VRAM - 309 MiB SHORT of the 1,669 MiB desktop worst case, so it does not :: fit with a screen attached. WITH the n-max 4 / p-min 0.75 :: drafter on it does not fit at all: 2,364 MiB measured living in system RAM :: (text-only). With the projector loaded as well, 3,502 MiB spilled (same :: drafter) and a 91k-token document then decodes at 8.0 t/s (section 05). :: The rule: a ceiling belongs to the whole configuration — file + drafter flags + :: projector + the desktop you run. Quote all four or the ceiling is not portable.
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 122880 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"Why this pick never carried n-max 10 — measured 2026-08-25
2026-10-08: this pick now runs n3/p0, the n/p sweep’s best score on this file at the card’s sampler — shallow reasoning tokens 1.111 and answer tokens 1.120 of n4/p0.75, against n10/p0.5’s 0.961 and 1.049 (after a 57,540-token prefix n10/p0.5 read 1.079 against 1.061, intervals overlapping) — and it needs about 150 MiB less than n4 and about 1 GB less than n10 (§06). The 2026-08-25 record below is kept as measured: the wide drafter never went on this pick because it does not clear the 1,796 MiB desktop reserve with a real image in flight at this window. With a 1440p screenshot on 105,397 tokens, board VRAM peaks at 23,450 MiB and keeps only 1,126 — 670 MiB short, and the same configuration varied 397 MiB between two loads on the same day, so the shortfall is not a rounding error on a stable reading. The failure mode is a spill into system memory that raises no error and costs half your speed — the load prints one start-up memory warning, easy to miss (§11).C31 At the card’s sampler neither detour is needed for speed any more, because n3/p0 beat n10/p0.5 on this file on shallow reasoning and answer tokens. The August routes, for the record: text-only, pick 2 carried the wider drafter; vision with the wide drafter clears at -c 98304, by 522 (the former speed pick below).C1C2
The investigation behind this verdict
The wide drafter is not free at this window. At this same 122,880 window, deep-filled to 112,735 real tokens with the projector loaded, it delivers 62.47 t/s against 53.33 — +17.1 % — and leaves 1,776 MiB, which is 20 MiB SHORT of the 1,796 MiB reserve.C14
Shrinking the window is not the loss it looks like. At 98,304 the wide drafter measured 87.04 against 87.49 t/s — the same speed, inside both arms’ own probe spread — so those 24,576 tokens buy no throughput. What they buy is the reserve. With a real image in flight the same drafter keeps 1,126 MiB at 122,880 and 2,318 MiB at 98,304, and only the second clears. On text-only work the wide drafter won outright in this August greedy-code comparison, which is why pick 2 carried it until 2026-10-08, when it moved to n3/p0 at the cards’ sampler (§06) — pick 3 does not, and cannot: it ships --spec-type none because the full 262,144 window has no room for a drafter at all.
Raising --spec-draft-n-max on its own is not a throughput lever. A sweep at the recipe this page shipped on 2026-08-26, varying only --spec-draft-n-max and nothing else, found 74.13 t/s at n-max 4, 73.52 at 6 and 75.59 at 8, against a run-to-run band of 0.77% (§06). Those three do not rise with the cap, and the largest difference among them is +2.0% rather than the roughly 12% the matched sweep reported. Until they are reconciled, treat the wider drafter as something to measure on your own content rather than as faster by construction.
Four routes were measured on the same image-at-depth test meas:
| route to n10/p0.5 with vision | peak | slack | verdict |
|---|---|---|---|
as shipped, -c 122880 | 23,450 | 1,126 | no — 670 short |
plus -ctkd q8_0 -ctvd q8_0 (quantise the draft cache, which -ctk never reaches) | 23,321 | 1,255 | no — 541 short |
-c 98304 keep vision, keep the wide drafter, spend 24,576 tokens of window | 22,258 | 2,318 | YES — clears by 522 |
| headless (no desktop, so no reserve to clear) | 23,450 | 1,126 | YES |
So the answer is -c 98304. Verified with a 1440p screenshot on 77,194 tokens, board VRAM sampled every 0.5 s — the model still read the 15 px table cell correctly at that depth. Quantising the draft cache saved 129 MiB where the per-token arithmetic predicted roughly 300, which is a reminder that the drafter’s memory is not all cache.
Before you take it, check you want it. The wide drafter is the code peak, not the speed peak in general. On prose the conservative setting wins — 48.4 against 43.8 t/s — and speculation is worth only 1.16× there at all, so writing work is better served by --spec-type none and the memory back. And the advantage shrinks with depth: +33.2 % on an empty window, +17.1 % at 112k of real fill. A long agent session gets about half the figure that sells the setting.
Two things worth keeping from that run. The model answered a fine-detail question about the screenshot correctly at 105,397 tokens of depth — reading a 15 px table cell — so vision is not degraded by context pressure, only memory is. And the wide drafter’s advantage measured +33.2 % on an empty window against +17.1 % at real depth: a shallow probe overstates speculation by about double.
The file: unsloth/Qwen3.8-27B-GGUF, UD-IQ4_XS — 14.25 GB on the download page, 13.3 GiB resident (download pages count in decimal GB, memory budgets in binary GiB; the ~7% gap is real and matters whenever memory is tight). Plus the BF16 mmproj projector from either repository if you want vision. The window: -c 122880. The flags: all layers on the GPU, one slot, an 8-bit KV cache, and the n-max 3 / p-min 0 drafter, the best score for this file at this card’s sampler (§06). The effort levels it offers: all three — low, medium and xhigh. Nothing is excluded here: xhigh's measured thinking appetite is 61,500–75,800 tokens and a 122,880-token window holds its upper tail with room for a prompt and an answer. Expected speed, measured before 2026-10-08 with the drafters named (n3/p0 was measured only as ratios, §06): 76–79 t/s of answer tokens on short-context code in most sessions (20 of 23 sessions fell in this band on 2026-08-28, UD-IQ4_XS, n10/p0.5, -c 32768, -ctk/-ctv q8_0, -ngl 99, --parallel 1, -fa on, greedy, 700 tokens, fresh server per session, one warmup discarded, RTX 3090 at 350 W, build 10502). Three of those 23 sessions reached 88–89 t/s, about 13% higher, with nothing in between and nothing predicting which level a session will land on. The published 86.91 t/s in §08 sits inside the high band; the figure a reader will usually see is in the low band. Seven hypotheses for the gap have been tested and ruled out; the cause is unmeasured (§15, C27). Then 64.8 at 91k of depth (n4/p0.75), and 37–39 t/s on the xhigh reasoning stream at 91k (measured at n4/p0.75 and n10/p0.5) — which is where a long agent run actually spends its wall clock. VRAM at the top of this window: 20,658 MiB of server process by the measured budget model, leaving 2.7–3.7 GiB of board VRAM depending on what your desktop is holding. Room left for your desktop: 2.7–3.7 GiB — this is the row chosen so that a browser and an agent's web interface do not push it over.
Twenty-three sessions of one identical configuration land on two distinct speed levels about 13% apart, and nothing predicts which level a given session will reach. All 23 used UD-IQ4_XS, n-max 10 / p-min 0.5, -c 32768, -ctk/-ctv q8_0, -ngl 99, --parallel 1, -fa on, reasoning off, greedy (temperature 0, top-k 1), 700 predicted tokens, fresh server per session, one warmup discarded, one RTX 3090 at its stock 350 W limit, llama.cpp build 10502.
- 20 sessions: 75.71 to 78.65 t/s
- 3 sessions: 88.76, 88.18, 88.21 t/s
- Nothing in between. About 13% apart.
Everything tested has been ruled out, and the list matters because it is what makes “unexplained” a measurement rather than a shrug:
| Ruled out | How |
|---|---|
| the workload | greedy output is bit-identical, probe for probe |
| the drafter | acceptance and mean draft length bit-identical across sessions |
| the llama.cpp build | every compute library dated 2026-08-19, unchanged |
| the core clock | a high session ran 49 MHz slower than a low one |
| board temperature | 75 to 83 °C in every session alike |
| the memory clock | pinned at 9501 MHz in every probe ever recorded |
| prior load | tested and refuted — see the paired test below |
The prior-load test. The one session in fifteen that had reached the published figure followed a sustained heavy run; every other followed idle. So four alternating pairs were run — COLD (150 s idle first) against HOT (180 s sustained burn on a different prompt first), five probes each, arms alternated so drift falls on both equally:
- COLD: 77.33, 88.21, 78.33, 78.05 — mean 80.48
- HOT: 88.18, 77.41, 78.27, 76.67 — mean 80.13
- difference −0.43%, ranges overlapping
The high sessions fall on one arm each. Prior load is not the cause. The first pair alone read 77.33 against 88.18, a 14% gap with no overlap, which looks exactly like a confirmed hypothesis — and the second pair reversed it.
The rule this gives a reader: plan for 76–79 t/s, and treat the 86–89 t/s level as something that happens unpredictably in about one session in eight. Any single speculative throughput figure on this page is drawn from a band at least this wide.C27
NVIDIA 24–32 GB · UD-IQ4_XS at 98,304 with vision · the former speed pick (n3/p0 since 2026-10-08)
It no longer earns a separate pick. At the cards’ temperature-1.0 sampler with thinking on (which this command leaves to the client), the n/p sweep measured this card’s old drafter, n10/p0.5, at 0.961 of n4/p0.75 on this file’s shallow reasoning tokens and 1.049 on answer tokens, and n3/p0 at 1.111 and 1.120; after a 57,540-token prefix n10/p0.5 read higher, 1.079 against 1.061, intervals overlapping (§06). Since 2026-10-08 the command below carries n3/p0, which makes it the default recipe at a smaller window: take the default unless you want the desktop slack that 24,576 fewer tokens buys. Everything else — file, projector, image budget, KV width — is the default’s. If your client decodes greedily (temperature 0), the sweep’s greedy probe ranked n4/p0 and n5/p0 above n3/p0 on this file (§06). The +32.5% that named this card was greedy novel JavaScript with thinking off at short context, and the 2026-08-25 record follows as measured.
What it buys, measured 2026-08-25. Both arms carry the projector; both are 700-token probes on novel JavaScript at short context; three settled probes each after a discarded warmup, so the spread below is this rig’s own scatter and not an error bar on the difference:
| arm, projector loaded | decode | spread | against the default |
|---|---|---|---|
the default until 2026-10-08 — n4/p0.75 at -c 122880 | 65.67 t/s | 5.2% | — |
n10/p0.5 at the same -c 122880 | 87.49 t/s | 3.5% | +33.2% — but see the reserve below |
n10/p0.5 at -c 98304 — this recipe until 2026-10-08 | 87.04 t/s | 2.1% | +32.5% |
Read those two ways round, because only one of them is about the window. The wide drafter is worth about a third at a fixed window. Shrinking the window from 122,880 to 98,304 then costs 0.5%, which is inside both arms’ own spread. The window is what pays for the drafter; it is not itself a speed setting — and the common guess that a smaller -c decodes faster is not supported by anything on this page. What costs throughput is depth of fill: 86.3 t/s at 1.5k, 80.2 at 28k, 64.8 at 91k, all at one -c and at n4/p0.75.
Why 98,304 and not the default 122,880. The wide drafter costs 898 MiB, and with the projector loaded that is more than 122,880 can spare once an image is in flight. Sampled every 0.5 s across the whole request, against the 1,796 MiB reserve: at -c 122880 with a 1440p shot on 105,397 tokens it peaks at 23,450 MiB and keeps 1,126 — 670 short; at -c 98304 with a shot on 77,194 tokens it peaks at 22,258 and keeps 2,318 — clears by 522. So the window is where the drafter and the projector stop competing, not a round number.
Not measured at this window: depth. The +17.1% the wide drafter shows on a 112,735-token fill was taken at -c 122880, and this window cannot hold that fill. Acceptance rises with depth on this head, so the wide drafter is expected to keep leading — but that is an expectation and this page does not publish it as a number here. And read §06 before treating the wide drafter as settled: three sweeps have measured n-max on its own at p-min 0.75, all at -c 32768, and they disagree — 80.99 → 80.32 flat, 83.50 → 90.84 climbing, 74.13 → 75.59 flat. The pairing measured here is n10/p0.5, which the August greedy, thinking-off grids put at the top, though the 2026-08-28 retest put n4/p0 above it (C24); the 2026-10-08 sweep at the cards’ sampler with thinking on put it below n3/p0 on shallow reasoning tokens on all three files that ran both, and above n3/p0 after a 57,540-token prefix on UD-IQ4_XS and the Pi file, intervals overlapping (§06).
If you do not need images, the August launcher’s pick 2 ran the same n10/p0.5 drafter text-only at -c 180224 with no trade, because the 1,138 MiB the projector would have taken is more than that drafter costs. Since 2026-10-08 that row runs n3/p0 too, which scored higher at the cards’ sampler and needs about 1 GB less than n10 (§06).
:: 2026-10-08: n3/p0 now, the n/p sweep's best score for this file at the cards' :: temperature-1.0 sampler with thinking on (section 06), so this is the default :: recipe at a smaller window. The record that named it, measured 2026-08-25: :: THE SPEED PICK: the default recipe with the WIDE drafter, bought with 24,576 :: tokens of window. 87.04 t/s against the default's 65.67 (+32.5%), both with the :: projector loaded, 700-token novel-JS probes at short context, three settled :: probes after a discarded warmup. Peaks 22,258 MiB with a 1440p shot on 77,194 :: tokens and keeps 2,318 - clearing this page's 1,796 MiB desktop reserve by 522. :: The SAME drafter at -c 122880 keeps only 1,126 and is 670 SHORT, which is why :: the window moved. The August serve-qwen.bat shipped this as pick [0]. :: file: unsloth/Qwen3.8-27B-GGUF (UD-IQ4_XS, 14.25 GB = 13.3 GiB) + BF16 mmproj llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 98304 :: 24,576 fewer than the default. It paid for the n10/p0.5 :: drafter until 2026-10-08; it is not itself a speed setting. -ngl 99 :: all 65 layers - the output layer counts as one (section 05) --parallel 1 :: one slot: the whole window goes to one conversation -fa on -ctk q8_0 -ctv q8_0 :: q8_0 KV REQUIRES flash attention; -fa off refuses to load --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 :: n3/p0 since 2026-10-08: shallow reasoning 1.111x and :: answers 1.120x n4/p0.75 on this file at the cards' :: sampler, against n10/p0.5's 0.961x and 1.049x (section 06). :: With a greedy client (temperature 0) the sweep's greedy :: probe ranked n4/p0 and n5/p0 above n3/p0. :: About 7 x 150 MiB lighter than the n10/p0.5 this :: window was measured with (derived), so it only gains :: margin. Until 2026-10-08: n10/p0.5, 898 MiB over :: n4/p0.75, which is why the window moved. --jinja :: required for the chat template and for reasoning_content --host 127.0.0.1 --port 1234
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 98304 -ngl 99 --parallel 1 -fa on -ctk q8_0 -ctv q8_0 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234
NVIDIA 24–32 GB — THE BIG-WINDOW PICK · UD-Q2_K_XL (vision and the drafter and 196,608 tokens)
The file: unsloth UD-Q2_K_XL — 9.83 GB, 9.154 GiB resident. The window: -c 196608. The flags: the universal set, the n-max 4 / p-min 0 drafter der, and the full image budget. This file was not in the 2026-10-08 n/p sweep, and its window was measured with n4/p0.75; p-min costs no memory, so the window still fits. On the four files that sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot and read 0.90–1.03 of it after a 57,540-token prefix; on UD-IQ4_XS deep answers it was slightly faster (n4/p0 read 0.973 of it, interval 0.969–0.996; §06). n-max 3 scored best with one slot on the three files run at xhigh and is about 150 MiB lighter, so n3/p0 here is a candidate to measure, not a measured setting. VRAM at depth: 22,014 MiB meas — peak, with a 1440p screenshot in flight on 163,124 tokens, sampled every 0.5 s. Room left for your desktop: 2,562 MiB, clearing this page’s 1,796 MiB reserve by 766. Effort: medium is what this pick ships, and the reason is the window rather than the clock. The window is one pot — prompt, thinking and answer share it. An ordinary xhigh answer only wants about 2,217 tokens of thinking, which fits here easily; but this pick exists to fill 196,608, and at the 163,124 tokens it was measured at only 33,484 remain. A complete-build task — the one thing measured wanting 61,500–75,800, four August samples of one task (§09) — would truncate there and return nothing. So: xhigh for questions of a full window, medium when you are asking for a whole program. Effort changes how much the model writes, not how long it takes to read your prompt.C3
Why this pick exists, added 2026-08-25. A reader wanted the largest window they could get while keeping both vision and the drafter, and this page had no answer for them — the 4-bit file cannot do it, and the full 262,144 fails in every drafter state once an image is in flight. The 2-bit file can, and three separate measurements say it costs less than its name suggests: it ties the 4-bit reference on accuracy (one of 75 paired items different, p=1.00), it reads images identically (7/7 against 7/7 on a generated detail target), and its window genuinely retrieves — 5 of 5 needles found at every depth out to 241,655 tokens, with a clean control (§08). What you actually trade is decode speed at short windows, where the 4-bit file drafts better.
196,608 is the ceiling, measured. -c 229376 peaks at 23,529 MiB and keeps only 1,047 — it is under the reserve at load, before an image is even sent. So this window is not a round number chosen for looks; it is the largest that survived the test. One honest note on that failing run: its answer to the image question came back as prose rather than the bare number, which the scoring here cannot grade either way. It is reported as a memory failure only, because that is what was measured.
The cost that matters more than VRAM: prefill. Filling this window took 278 seconds meas, against 210 s at -c 163840 and roughly 8 minutes at the full 262,144. Prefill is compute-bound and does not get cheaper. If your agent appends to its context you pay this once and the prefix cache carries you; if it rebuilds context every turn you pay it every turn, and -c 163840 saves you a minute for 32,768 fewer tokens. Choose on that, not on the bigger number.
:: THE BIG-WINDOW PICK: vision + drafter + 196,608 tokens, which no 4-bit :: configuration on this page can hold. MEASURED 2026-08-25 with a 1440p image :: in flight at 163,124 tokens: peak 22,014 MiB, 2,562 MiB left for a desktop. :: 196,608 IS THE CEILING, and it is measured rather than assumed: :: -c 196608 peak 22,014 MiB 2,562 MiB slack CLEARS by 766 :: -c 229376 peak 23,529 MiB 1,047 MiB slack FAILS by 749 :: -c 262144 peak 22,877 MiB 1,699 MiB slack FAILS by 97 (drafter OFF!) :: 229,376 is already under the reserve AT LOAD (1,516) before an image lands. :: file: unsloth/Qwen3.8-27B-GGUF (UD-Q2_K_XL, 9.83 GB = 9.154 GiB) llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 :: full image budget - the 1024 cap makes the model :: confidently MISREAD fine detail (section 05) -c 196608 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: 22,014 MiB peak MEASURED at depth with an image, not :: at load. Drop to -c 163840 for 3,852 MiB of slack and :: a minute less prefill --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: derived: :: not in the 2026-10-08 n/p sweep. This window was :: measured at n4/p0.75, and p-min costs no memory, so :: it still fits. On the four files swept, p-min 0.75 :: ran 0.87-0.97x p-min 0 at n-max 4 on shallow reasoning :: and 0.90-1.03x after a 57.5k prefix (one slot; :: section 06). n3/p0, ~150 MiB lighter, is a :: candidate to measure here, not a measured :: setting. n10/p0.5 costs ~900 MiB more and is not :: measured at this window - do not assume it fits --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: the window is one POT :: an ordinary xhigh answer averages 2,217 tokens (175 benchmark :: prompts), which the 33,484 left at the 163,124-token depth above :: holds 15 times over. medium is shipped because a COMPLETE BUILD :: is the one thing measured wanting 61,500-75,800 (four August samples of :: one task, section 09) and that would truncate here. Ask questions :: of a full window at xhigh; ask for a whole program at medium. :: This has nothing to do with the 278 s prefill: effort changes how :: much the model WRITES, not how long it takes to READ your prompt.
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 196608 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}"NVIDIA 24–32 GB — THE ONE YOU HAVE HEARD OF · Q4_K_M (no measurable quality advantage, 2.1 GiB larger, ~0.3 GiB left for your desktop)
Q4_K_M is the best-known 4-bit GGUF and that is the main reason it is still here. It is kept so that an option you have already heard of is answered rather than absent. But it is not a close second, and it does not have a measurable quality advantage.C19
The perplexity gap is 0.061 (6.535 against 6.596). This page's own yardstick for comparing one file with another is a combined error of ±0.063 (§08) — and that passage names this very pair when it says so. The gap is inside the bar. It is not resolved. The only other evidence, a 200-question GSM8K comparison that scored Q4_K_M 94.0% against UD-IQ4_XS 93.0%, cannot be used: its grader compared the whole answer line instead of the number — a bug worth 8 points, and one that only bites the arm that writes units — and it has never been re-scored (§08). So the perplexity gap is unresolved, and the GSM8K result is unusable.
What it costs to keep believing it: 2.1 GiB, which is either the vision projector or about 50,000 more tokens of window; roughly 0.3 GiB left for your desktop, against the 1,796 MiB this page reserves for one; no room for the wider drafter; and about 8% more energy per token (8.853 against 8.198 J per decode token, measured with the drafter off, 2026-08-23). Take the default above unless you have a specific reason not to, and if you already downloaded this file, you have lost nothing worth measuring.
The file: lmstudio-community/Qwen3.8-27B-GGUF, Q4_K_M — 16.5 GB on the download page, 15.4 GiB resident. The window: -c 122880. The flags: the same universal set as the default recipe, with the n-max 4 / p-min 0 drafter der. Q4_K_M was not in the 2026-10-08 n/p sweep and this row was measured at n4/p0.75, but p-min costs no memory, so the row still fits; on the four files the sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot and read 0.90–1.03 of it after a 57,540-token prefix, and on UD-IQ4_XS deep answers it was slightly faster (n4/p0 read 0.973 of it; §06). n10/p0.5 does not fit here, and at the cards’ sampler it was slower than n3/p0 on UD-IQ4_XS on shallow reasoning and answer tokens anyway. The effort levels it offers: all three, for the same reason as the default recipe. An ordinary xhigh answer averages 2,217 tokens of thinking across 175 benchmark prompts; the 61,500–75,800 figure quoted elsewhere on this page is a ceiling from four August samples of one task — a complete build — and a 122,880-token window holds even that, leaving about 47,000 for your prompt. Expected speed, measured with these flags but at n4/p0.75 (n4/p0 is not measured on this file): 69.8 t/s of answer tokens on short-context code and 57.9 t/s on the reasoning stream; the 81.7 t/s figure this file is famous for on this page belongs to the wider drafter, which this row cannot afford — it is a cross-reference, not this recipe's speed. VRAM at the top of this window: about 23.7 GiB of board VRAM with a desktop up, measured, leaving roughly 0.3 GiB. Room left for your desktop: about 0.3 GiB — less than a browser typically holds. §11 measured this very configuration spilling about 1.9 GB with two browsers open and falling to 20–35 t/s.
:: TAKE THIS ONE only if you believe an unresolvable 0.9% perplexity edge is :: worth 2.1 GiB AND you can free the VRAM your desktop uses - this row leaves :: about 0.3 GiB, less than a browser typically holds. The edge is 0.061 of :: perplexity against this page's own +/-0.063 comparison error: INSIDE THE BAR. :: The GSM8K half of the old argument is unaudited. THE DEFAULT ABOVE is better :: on context, speed, vision and desktop headroom, and no worse on any :: measurement this page can resolve. :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M, 16.5 GB = 15.4 GiB) :: before first run: set "Prefer No Sysmem Fallback" (NVIDIA Control Panel) and :: never write -ngl 64 — both explained in section 11 llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 122880 :: 24 GB: the largest window whose every byte still lives on the :: card is ~131k, so this leaves slack (section 05's two :: ceilings). 5090 32 GB: -c 262144 :: fits — budget ~11 GiB of KV at the MEASURED slope, not the :: 8.5 GiB the q8 arithmetic suggests -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: derived: :: this row's ~23.7 GiB was measured at n4/p0.75; p-min costs :: no memory, so it still fits. The 2026-08-23 matched sweep :: ranked n10/p0.5 first on BOTH files (81.7 vs 69.8 here at :: n4/p0.75), but that was greedy code with thinking off. Q4_K_M :: was not in the 2026-10-08 n/p sweep; on the four files it ran, :: p-min 0.75 ran 0.87-0.97x p-min 0 at n-max 4 on shallow :: reasoning, one slot, and 0.90-1.03x after a 57.5k prefix (section :: 06). n3/p0 is a candidate to measure; it would free ~150 MiB. :: BUDGET the drafter: 1,008 MiB fixed + 11.4% of your window :: (section 05). --spec-type none frees all of it; on prose it is :: worth only ~1.16x, so that is the easiest 1.8 GiB to reclaim :: vision: add --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 :: --image-max-tokens 10580 when you need it. The projector costs 1,138 MiB :: once loaded - not the 0.867 GiB its file weighs - and with this file at -c 122880 :: that is the measured ~23.7 GiB reference configuration: it fits, but leaves only :: ~0.3 GiB of the card for your desktop. With a desktop running, drop to -c 65536, :: or -c 49152 if you keep browsers open. :: Budget from the MEASURED slope: each 32k of window costs ~1.375 GiB :: with the drafter on and ~1.22 GiB with it off; the 1.06 GiB q8 arithmetic is a :: FLOOR. Vision + a desktop + a BIG window = the default recipe above. :: On the 1 GiB-larger UD-Q4_K_XL, projector + full context needs more than 24 GB - :: use -c 98304 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: quality-first default; :: medium when the wait matters (section 09's price list) :: FLAGS DELIBERATELY OMITTED: the same three as the default recipe — --parallel 2 (+22.0% :: aggregate but -35.4% per slot with the drafter on, measured 2026-08-25 — slower :: for one user), -ctk q4_0 (+0.693% perplexity, a window trade), and nvidia-smi -pl :: (measured: 300 W costs 4.1% of throughput for 9.4% less energy, 250 W costs 14.6% :: for 14.7%; left out because it is a persistent hardware setting needing an elevated :: shell, not because it does not work).
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 122880 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"Swift-1.5 recipes — no detectable quality cost at IQ4_XS, wider windows at IQ3_XXS
Take the default card unless you have a reason to choose a Swift file. Swift IQ4_XS scored indistinguishably from the anchor on this page’s 175-prompt suite at smoke-test scale (25 questions per set, greedy, drafter off) and generated fewer tokens there (§09). Swift IQ3_XXS holds the longer window of the two Swift files and, with vision at two slots, more tokens per slot (§05, §08), at the cost of a perplexity gap against Swift IQ4_XS on their shared tokenizer (§08). Five of the six Swift cards carry the window that is fully resident at this page’s 1,796 MiB desktop threshold, measured where a deep fill shows it and derived from the per-card arithmetic otherwise; the per-card arithmetic has no two-slot text configuration, so the Swift IQ3_XXS two-slot text card carries the October launcher’s window, fully resident with a desktop of up to 976 MiB rather than 1,796. Each card’s speed band was measured on the lower of this machine’s two speed levels (the two-levels note); its derived higher-level line divides each figure by the anchor’s 0.917 (at 1.5k) or 0.928 (at 91k) of its August speed. The card for bytkim’s Qwen3.8-27B-pi, a fine-tune whose bartowski IQ4_XS shares Swift IQ4_XS’s tensor layout and template, is in Appendix B.
NVIDIA 24 GB · 3090 — Swift IQ4_XS (text, 1 slot, 159,744-token window)
:: Swift IQ4_XS, text, 1 slot, 159,744 tokens per slot. :: One sweep 2026-10-04; not comparable with the August :: tables in section 06. :: Choose this for text-only on the Swift fine-tune; :: for vision on the same file choose :: the Swift IQ4_XS vision card, for the base model choose :: the default card; for a longer window choose :: the Swift IQ3_XXS text card. :: file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF :: (IQ4_XS, 15,475,951,936 bytes) llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ4_XS.gguf --alias qwen/qwen3.8-27b -c 159744 :: measured: deep fill to 156,261 tokens :: left 1,880 MiB, fully resident at this :: page's 1,796 MiB threshold. This config :: is not in the October launcher. :: Server footprint at depth (dedicated + :: shared): 22,212 + 484 MiB (fill to :: 156,261 tokens, 44.57 t/s at depth) -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 :: the :: n3/p0 drafter since 2026-10-08: the n/p :: sweep's best score for this file at this :: sampler and xhigh (-c 81920, text only, two :: sessions): shallow reasoning 1.189x n4/p0.75 :: (paired 95% 1.158-1.224), 1.052-1.057x after :: a 57,540-token prefix, answers 1.062-1.064x; :: n2/p0 faster shallow but 0.932x after the :: prefix (section 06). ~150 MiB LESS than n4, :: so the window, filled at n4/p0.75, only :: gains margin. :: Earlier, one prompt (-c 32768, greedy, :: reasoning off): n4/p0.75 acceptance 0.917, :: 3.348 accepted tokens per verify pass; :: n10/p0.5 acceptance 0.641, 5.667 per pass. :: Board VRAM at -c 32768: n4 costs :: 1,092-1,124 MiB; n10 adds 898 MiB (one pass :: read 851 on a board counter that includes the :: desktop, section 05). :: Greedy thinking-off text is byte-identical :: across none/n4/n10 for this file on that :: one prompt: evidence, not proof, that :: drafter-off scores transfer (section 08) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: low, medium, xhigh offered; high is not a :: level of this template (section 14). :: Turns per window: section 09 :: speed (greedy-only; not the card's temperature-1.0 sampler), :: 2 passes, -c 131072, one slot, n4/p0.75, 2026-10-04: :: floor (drafter off, reasoning off, 1.5k): 42.88 t/s (first :: pass inside the downloads; passes agree within 3%) :: answer regime, 1.5k: 87.2 t/s (higher level, derived: 95.1) :: answer regime, 91k: 61.14 t/s (higher level, derived: 65.9) :: reasoning regime, 1.5k and 91k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: ceiling (n10/p0.5, reasoning off, -c 32768, one prompt): :: 98.1 t/s :: temperature-1.0 decode (n=2, n4/p0.75): 61.11-65.9 t/s :: (section 09) :: FLAGS DELIBERATELY OMITTED: :: --parallel 2 — unmeasured for this file (section 03) :: -ctk q4_0 — unmeasured for this file (the q8_0 cache :: measured 6.7708 perplexity against f16's 6.7742, :: section 08) :: nvidia-smi -pl — not measured for this file; a persistent hardware :: setting that needs an elevated shell. :: invalidation trigger: build 10502 / 0adcc3bb5, :: driver 596.36, a new upload of the file :: fields that differ from the measured command: port (no :: effect on inference); sampler and effort :: (the speed probes ran greedy; the temperature-1.0 line :: above ran the card's sampler at xhigh); the deep fill :: ran --reasoning off at temperature 0 (KV memory per :: token is the same for prompt and generated tokens); :: every measurement loaded the file memory-mapped and the :: window fills turned the host-RAM prompt cache off :: (--cache-ram 0); the card loads with --load-mode none :: (system RAM and load time differ, not throughput: 43.0 :: against 43.4 t/s on the anchor, section 15) and keeps the prompt cache :: on (its RAM use at depth is unmeasured). :: Quality at depth: not verified (the needle test did :: not run)
Copy-paste version — the same command with the comments removed
llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ4_XS.gguf --alias qwen/qwen3.8-27b -c 159744 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"The file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF, IQ4_XS — 15.48 GB on the download page, 14.41 GiB (15,475,951,936 bytes). The window: -c 159744 meas — deep-filled to 156,261 tokens at an 846 MiB desktop and fully resident; the fill left 1,880 MiB for a desktop der, which clears this page’s 1,796 MiB threshold. This configuration is not in the October launcher. The flags: all layers on the GPU, one slot, an 8-bit KV cache, the n-max 3 / p-min 0 drafter, and --reasoning-preserve. The effort levels it offers: low, medium and xhigh; high is not a level of this template (§14); turns per window in §09. Speed (greedy-only; not the card’s temperature-1.0 sampler) n=2, n4/p0.75, one sweep 2026-10-04: 87.2 t/s answer tokens at 1.5k of depth (higher speed level, derived: 95.1), 61.14 t/s at 91k (higher level: 65.9); floor 42.88 t/s with no drafter (its first pass also ran inside the downloads; the passes agree within 3%), ceiling 98.1 t/s at n10/p0.5 on one prompt at -c 32768; on 2026-10-07, n-max 3 / p-min 0 decoded reasoning tokens ×1.153 faster than this card’s n4/p0.75 under its own sampler (4 prompts × 2 seeds, -c 32768, text only), with answer tokens level on one prompt and 152 MiB less VRAM (§08); since 2026-10-08 the card runs n3/p0, the n/p sweep’s best score for this file (shallow reasoning tokens 1.189 of n4/p0.75, §06); its window was deep-filled at n4/p0.75, and n-max 3 needs about 150 MiB less; the reasoning-regime probes are void (their passes disagreed beyond 3%, in the sweep and in a re-run on 2026-10-05). Not comparable with the August tables in §06. Temperature-1.0 decode n=2, n4/p0.75: 61.11–65.9 t/s (§09). VRAM at the top of this window: server footprint at depth 22,212 MiB dedicated + 484 MiB shared meas (deep fill to 156,261 tokens, 2026-10-04). Room left for your desktop: 1,880 MiB — clears this page’s 1,796 MiB threshold.
NVIDIA 24 GB · 3090 — Swift IQ4_XS (text + vision, 1 slot, 128,000-token window, derived)
UNVERIFIED — DERIVED CONFIG :: Swift IQ4_XS, text + vision, 1 slot, 128,000 tokens :: per slot. :: One sweep 2026-10-04; not comparable with the August :: tables in section 06. :: Choose this for vision on the Swift fine-tune; for :: text-only on the same file (wider window) choose :: the Swift IQ4_XS text card, for the base model choose :: the default card; for a longer window choose :: the Swift IQ3_XXS vision card. :: file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF :: (IQ4_XS, 15,475,951,936 bytes) + f16 mmproj :: (927,607,552 bytes) llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 128000 :: DERIVED from per-card arithmetic to :: clear this page's 1,796 MiB desktop :: threshold. The October launcher ships :: -c 139264: fully resident with an :: 841 MiB desktop (deep fill to 126,838 :: tokens, 2026-10-04); it left 1,330 MiB, :: below the 1,796 MiB threshold. :: Server footprint at depth from the :: -c 139264 fill with vision (dedicated :: + shared): 22,802 + 444 MiB; at the :: card's -c 128000 the footprint derives :: lower -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 :: the :: n3/p0 drafter since 2026-10-08: the n/p :: sweep's best score for this file at this :: sampler and xhigh (-c 81920, text only, two :: sessions): shallow reasoning 1.189x n4/p0.75 :: (paired 95% 1.158-1.224), 1.052-1.057x after :: a 57,540-token prefix, answers 1.062-1.064x; :: n2/p0 faster shallow but 0.932x after the :: prefix (section 06). ~150 MiB LESS than n4, :: so the window, filled at n4/p0.75, only :: gains margin. :: The sweep ran with no projector and no :: image; how image tokens move the best :: n-max was not measured. :: Earlier, one prompt (-c 32768, greedy, :: reasoning off): n4/p0.75 acceptance 0.917, :: 3.348 accepted tokens per verify pass; :: n10/p0.5 acceptance 0.641, 5.667 per pass. :: Board VRAM at -c 32768: n4 costs :: 1,092-1,124 MiB; n10 adds 898 MiB (one pass :: read 851 on a board counter that includes the :: desktop, section 05). :: Greedy thinking-off text is byte-identical :: across none/n4/n10 for this file on that :: one prompt: evidence, not proof, that :: drafter-off scores transfer (section 08) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: low, medium, xhigh offered; high is not a :: level of this template (section 14). :: Turns per window: section 09 :: speed (greedy-only; not the card's temperature-1.0 sampler), :: 2 passes, -c 131072, one slot, no projector, n4/p0.75, :: 2026-10-04: :: floor (drafter off, reasoning off, 1.5k): 42.88 t/s (first :: pass inside the downloads; passes agree within 3%) :: answer regime, 1.5k: 87.2 t/s (higher level, derived: 95.1) :: answer regime, 91k: 61.14 t/s (higher level, derived: 65.9) :: reasoning regime, 1.5k and 91k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: ceiling (n10/p0.5, reasoning off, -c 32768, one prompt): :: 98.1 t/s :: temperature-1.0 decode (n=2, n4/p0.75): 61.11-65.9 t/s :: (section 09) :: deep fill decode at -c 139264 with vision: 41.03 t/s :: on 126,838 tokens (n=1, 2026-10-04); the text-only :: fill at the same window: 47.42 t/s (n=1) :: re-filled 2026-10-07 (Pi fine-tune study; another :: sweep, not comparable with 41.03): 45.14 / 45.55 :: t/s (n=2), 1,217 MiB desktop, the same footprint :: of 22,802 + 444 MiB :: n-max 3 / p-min 0 (2026-10-07, -c 32768, no projector): :: reasoning tokens 1.153x n4/p0.75 at this card's :: sampler, 1.087x greedy; answer tokens level on one :: prompt; 152 MiB less VRAM. Since 2026-10-08 this :: card runs n3/p0 (n/p sweep: shallow reasoning :: 1.189x n4/p0.75 at xhigh, -c 81920, with no :: projector and no image) :: (section 08) :: FLAGS DELIBERATELY OMITTED: :: --parallel 2 — unmeasured for this file (section 03) :: -ctk q4_0 — unmeasured for this file (the q8_0 cache :: measured 6.7708 perplexity against f16's 6.7742, :: section 08) :: nvidia-smi -pl — not measured for this file; a persistent hardware :: setting that needs an elevated shell. :: invalidation trigger: build 10502 / 0adcc3bb5, :: driver 596.36, a new upload of the file :: fields that differ from the measured command: port (no :: effect on inference); sampler and effort :: (the speed probes ran greedy; the temperature-1.0 line :: above ran the card's sampler at xhigh); the deep fill :: ran --reasoning off at temperature 0 (KV memory per :: token is the same for prompt and generated tokens); :: every measurement loaded the file memory-mapped and the :: window fills turned the host-RAM prompt cache off :: (--cache-ram 0); the card loads with --load-mode none :: (system RAM and load time differ, not throughput: 43.0 :: against 43.4 t/s on the anchor, section 15) and keeps the prompt cache :: on (its RAM use at depth is unmeasured). :: Quality at depth: not verified (the needle test did :: not run)
Copy-paste version — the same command with the comments removed
llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 128000 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"The file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF, IQ4_XS — 15.48 GB on the download page, 14.41 GiB (15,475,951,936 bytes), plus the f16 mmproj projector (927,607,552 bytes). The window: -c 128000 der — derived from the per-card arithmetic to clear this page’s 1,796 MiB desktop threshold; the October launcher ships -c 139264, which is fully resident with an 841 MiB desktop (deep fill to 126,838 tokens, 2026-10-04) but leaves only 1,330 MiB for the desktop, below the 1,796 MiB threshold. The flags: all layers on the GPU, one slot, an 8-bit KV cache, the n-max 3 / p-min 0 drafter, the full image budget (1,024–10,580 tokens), and --reasoning-preserve. The effort levels it offers: low, medium and xhigh; high is not a level of this template (§14); turns per window in §09. Speed (greedy-only; not the card’s temperature-1.0 sampler) n=2, one slot, no projector, n4/p0.75, one sweep 2026-10-04: the band of the text card, because the probes ran without the projector — 87.2 t/s answer tokens at 1.5k (higher speed level, derived: 95.1), 61.14 t/s at 91k (higher level: 65.9); floor 42.88 (its first pass also ran inside the downloads; the passes agree within 3%), ceiling 98.1 at n10/p0.5; n-max 3 / p-min 0 (2026-10-07) decoded reasoning tokens ×1.153 faster than n4/p0.75 under the card’s sampler, with answer tokens level on one prompt and 152 MiB less VRAM (§08); since 2026-10-08 the card runs n3/p0 (§06), which needs about 150 MiB less than the n4/p0.75 its window was filled at; the reasoning-regime probes are void (their passes disagreed beyond 3%, in the sweep and in a re-run on 2026-10-05). Deep fill decode at -c 139264 with vision, n4/p0.75: 41.03 t/s on 126,838 tokens n=1, against 47.42 t/s for the text-only fill at the same window n=1 — not a paired test of the projector: the vision fill also held a 1440p image and drew lower drafter acceptance (0.83 against 0.88, the mean of two probes each), and the paired probe in §06 found the projector itself decode-free (0.04–0.09% at a 91k fill). A second fill on 2026-10-07, in the Pi fine-tune study (window-knee.py, in the order Pi, Swift, Swift, Pi, the same 126,838 tokens, projector loaded), held at 45.14 and 45.55 t/s meas n=2, with a 1,217 MiB desktop before load, a board VRAM peak of 24,019 MiB and 444 MiB in shared memory at depth, which leaves the same 22,802 MiB dedicated footprint as the 2026-10-04 fill der; its absolute t/s are not comparable with the 2026-10-04 figure (Appendix B). Not comparable with the August tables in §06. Temperature-1.0 decode n=2, n4/p0.75, no projector: 61.11–65.9 t/s (§09). VRAM at the top of this window: the fill at -c 139264 with vision measured 22,802 MiB dedicated + 444 MiB shared server footprint at depth meas (fill to 126,838 tokens, 2026-10-04); at the card’s -c 128000 the footprint derives lower. Room left for your desktop: at the launcher’s -c 139264 the fill left 1,330 MiB (below this page’s 1,796 MiB threshold); the derived -c 128000 targets 1,796 MiB by construction.
NVIDIA 24 GB · 3090 — Swift IQ3_XXS (text, 1 slot, 216,064-token window, derived)
UNVERIFIED — DERIVED CONFIG :: Swift IQ3_XXS, text, 1 slot, 216,064 tokens (derived from the per-card :: arithmetic at this page's 1,796 MiB desktop threshold). :: Choose this for one user and the longest window of the Swift cards; :: choose the Swift IQ3_XXS two-slot text card for two users, :: the Swift IQ3_XXS vision card for images, :: the Swift IQ4_XS text card for the lower-perplexity file, :: the default card for the base model. :: file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF :: (IQ3_XXS, 12,320,168,256 bytes) llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b -c 216064 :: DERIVED, not measured: the per-card :: arithmetic's fully resident window at :: 1,796 MiB desktop (Appendix A). The :: October launcher carries 262,144, :: decode-verified at 253,952 (38.46 t/s) :: and 262,144 (36.23 t/s) with a 415 MiB :: desktop, and fully resident with a :: desktop of up to 976 MiB (35 to spare, :: dedicated memory only); on the server :: footprint the 262,144 fill left 327 MiB -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: VRAM at the top of this window: not measured at 216,064. The knee :: rows above it measured a server footprint (dedicated + shared) of :: 23,213 + 668 MiB at 253,952 and 23,565 + 684 at 262,144, both :: fully resident with a desktop of up to 976 MiB; 216,064 is the :: per-card arithmetic's window at a 1,796 MiB desktop (derived, :: Appendix A). :: Speed band (greedy-only; not the card's temperature-1.0 sampler): :: one slot, n4/p0.75, -c 131072, two passes, one sweep 2026-10-04; not :: comparable with the August tables in section 06. :: Floor (drafter off, reasoning off, 1.5k): 44.12 t/s (43.74, 44.49); :: first pass inside the downloads, passes within 3% :: Short context, answer regime, 1.5k: 79.31 t/s (79.17, 79.45); :: higher speed level, derived: 86.5 t/s :: Short context, reasoning regime, 1.5k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: At depth, answer regime, 91k: 56.34 t/s (56.35, 56.33); :: higher speed level, derived: 60.7 t/s :: At depth, reasoning regime, 91k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: Ceiling (n10/p0.5, reasoning off, -c 32768, one prompt): :: 99.53 t/s :: Temperature-1.0 decode (the card's sampler, n4/p0.75, -c 131072, :: xhigh): 63.86-63.98 t/s :: Effort levels offered: low, medium, xhigh (high is not a level of :: this template, section 14). Turns per window: :: section 09. :: Drafter: on the one prompt (JavaScript, -c 32768, greedy, reasoning off): :: n4/p0.75 acceptance 0.927, 3.142 accepted tokens per verify pass :: n10/p0.5 acceptance 0.686, 6.071 accepted tokens per verify pass :: Drafter board VRAM at -c 32768: n4 costs 1,092-1,098 MiB, n10 adds 898 MiB :: 2026-10-08 n/p sweep (this file, one slot, -c 81920, text only, xhigh): :: n3/p0 shallow reasoning 1.097x n4/p0.75 (paired 95% 1.065-1.130), answers :: 1.054x; n4/p0 1.034x, n5/p0 0.940x. n3 keeps one fewer recurrent row :: than n4 (149.625 MiB predicted; not resolved on this file, section 06). :: FLAGS DELIBERATELY OMITTED: :: --parallel 2 — the two-slot text card. :: -ctk q4_0 — not measured for this file. :: nvidia-smi -pl — not measured for this file; a persistent hardware :: setting that needs an elevated shell. :: Invalidation trigger: build 10502 / 0adcc3bb5, driver 596.36, :: a new upload of the bartowski Swift-1.5 IQ3_XXS file. :: Fields that differ from the measured command: :: Port (no effect on inference). Sampler and effort :: (the speed probes ran greedy; the temperature-1.0 line ran the :: card's sampler at xhigh). The knee rows ran --reasoning off at :: temperature 0 (KV memory per token is the same for prompt and :: generated tokens); every measurement loaded the file :: memory-mapped and the window fills turned the host-RAM prompt :: cache off (--cache-ram 0); the card loads with --load-mode none :: (system RAM and load time differ, not throughput: 43.0 against :: 43.4 t/s on the anchor, section 15) and keeps the prompt cache on (its RAM :: use at depth is unmeasured). Quality at depth: not verified (the :: needle test did not run).
Copy-paste version — the same command with the comments removed
llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b -c 216064 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"The file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF, IQ3_XXS — 12.32 GB on the download page, 11.47 GiB (12,320,168,256 bytes). The window: -c 216064 der — derived from the per-card arithmetic at this page’s 1,796 MiB desktop threshold (Appendix A); not measured at this window. The October launcher carries -c 262144, decode-verified at 253,952 (38.46 t/s) and 262,144 (36.23 t/s) with a 415 MiB desktop meas, and fully resident with a desktop of up to 976 MiB, 35 MiB to spare, by the fully resident test, which counts dedicated memory only; on the server footprint (dedicated + shared) the 262,144 fill left 327 MiB der (§05). The flags: all layers on the GPU, one slot, an 8-bit KV cache, the n-max 3 / p-min 0 drafter (§06), and --reasoning-preserve. The effort levels it offers: low, medium and xhigh; high is not a level of this template (§14); turns per window in §09. Speed (greedy-only; not the card’s temperature-1.0 sampler) n=2, one slot, n4/p0.75, one sweep 2026-10-04: 79.31 t/s answer tokens at 1.5k of depth (two pass means 79.17–79.45; higher speed level, derived: 86.5), 56.34 t/s at 91k (two pass means 56.33–56.35; higher level: 60.7); floor 44.12 t/s with no drafter (43.74–44.49; its first pass also ran inside the downloads), ceiling 99.53 t/s at n10/p0.5 on one prompt at -c 32768; the reasoning-regime probes are void (their passes disagreed beyond 3%, in the sweep and in a re-run on 2026-10-05). Not comparable with the August tables in §06. Temperature-1.0 decode n=2, n4/p0.75: 63.86–63.98 t/s (§09). VRAM at the top of this window: not measured at -c 216064; the knee rows measured a server footprint of 23,213 MiB dedicated + 668 MiB shared at 253,952 tokens and 23,565 + 684 at 262,144 meas (2026-10-04). Room left for your desktop: 1,796 MiB — the per-card arithmetic’s target at this page’s desktop threshold der (Appendix A).
NVIDIA 24 GB · 3090 — Swift IQ3_XXS (text, 2 slots, 123,904 tokens per slot, desktop up to 976 MiB)
:: Swift IQ3_XXS, text, 2 slots, 123,904 tokens per slot: the October :: launcher's window. The per-card arithmetic has no two-slot text :: configuration, so this page derives no window at its 1,796 MiB :: desktop threshold for this card; this one is fully resident with a :: desktop of up to 976 MiB, not at 1,796 MiB. :: Choose this to serve two users on text; choose :: the Swift IQ3_XXS text card for one user and a longer :: window, the Swift IQ3_XXS two-slot vision card for images, :: the default card for the base model. :: file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF :: (IQ3_XXS, 12,320,168,256 bytes) llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b -c 247808 :: 123,904 per slot x 2, derived: the :: launcher's knee rule, between loaded :: steps at 116,736 and 124,928 per slot :: that both passed the fully resident :: test with a desktop up to 976 MiB :: (section 05). Decode held to 133,120 per :: slot (29.2 t/s per slot) and collapsed :: at 135,168 (14.04 t/s per slot). The :: knee loads ran at a 356-479 MiB desktop. -ngl 99 --parallel 2 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: VRAM at the top of this window: the knee row at 124,928 per slot, :: just above the card's window, measured a server footprint :: (dedicated + shared) of 23,546 + 424 MiB, deep-filled with both :: slots decoding (2026-10-04). :: Speed band (greedy-only; not the card's temperature-1.0 sampler): :: one slot, n4/p0.75, -c 131072, two passes, one sweep 2026-10-04; not :: comparable with the August tables in section 06. :: Floor (drafter off, reasoning off, 1.5k): 44.12 t/s (43.74, 44.49); :: first pass inside the downloads, passes within 3% :: Short context, answer regime, 1.5k: 79.31 t/s (79.17, 79.45); :: higher speed level, derived: 86.5 t/s :: Short context, reasoning regime, 1.5k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: At depth, answer regime, 91k: 56.34 t/s (56.35, 56.33); :: higher speed level, derived: 60.7 t/s :: At depth, reasoning regime, 91k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: Ceiling (n10/p0.5, reasoning off, -c 32768, one prompt): :: 99.53 t/s :: Two slots decoding at once (knee row, 124,928 per slot, answer :: regime, n4/p0.75): 28.97 t/s per slot :: Temperature-1.0 decode (the card's sampler, n4/p0.75, -c 131072, :: xhigh, one slot): 63.86-63.98 t/s :: Effort levels offered: low, medium, xhigh (high is not a level of :: this template, section 14). Turns per window: :: section 09. :: Drafter: on the one prompt (JavaScript, -c 32768, greedy, reasoning off): :: n4/p0.75 acceptance 0.927, 3.142 accepted tokens per verify pass :: n10/p0.5 acceptance 0.686, 6.071 accepted tokens per verify pass :: Drafter board VRAM at -c 32768: n4 costs 1,092-1,098 MiB, n10 adds 898 MiB :: 2026-10-08 n/p sweep at the October launcher's two-slot vision shape :: (123,904 x 2, projector loaded, both slots busy, xhigh; second pass, :: one load per setting, no interval): per slot, n4/p0 decoded :: reasoning tokens 1.337x n4/p0.75 and answers 1.484x (n3/p0: :: 1.185x and 1.233x). Depth with two slots: not resolved. Same :: n-max, same memory (section 06). The text-only two-slot shape was not run. :: FLAGS DELIBERATELY OMITTED: :: -ctk q4_0 — not measured for this file. :: nvidia-smi -pl — not measured for this file; a persistent hardware :: setting that needs an elevated shell. :: Invalidation trigger: build 10502 / 0adcc3bb5, driver 596.36, :: a new upload of the bartowski Swift-1.5 IQ3_XXS file. :: Fields that differ from the measured command: :: Port (no effect on inference). Sampler and effort :: (the speed probes ran greedy; the temperature-1.0 line ran the :: card's sampler at xhigh). The knee rows ran --reasoning off at :: temperature 0 (KV memory per token is the same for prompt and :: generated tokens); every measurement loaded the file :: memory-mapped and the window fills turned the host-RAM prompt :: cache off (--cache-ram 0); the card loads with --load-mode none :: (system RAM and load time differ, not throughput: 43.0 against :: 43.4 t/s on the anchor, section 15) and keeps the prompt cache on (its RAM :: use at depth is unmeasured). Quality at depth: not verified (the :: needle test did not run).
Copy-paste version — the same command with the comments removed
llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b -c 247808 -ngl 99 --parallel 2 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"The file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF, IQ3_XXS — 12.32 GB on the download page, 11.47 GiB (12,320,168,256 bytes). The window: -c 247808, 123,904 per slot × 2 der — the October launcher’s window, set by its knee rule between two loaded steps, 116,736 and 124,928 per slot, that both passed this page’s fully resident test with a desktop of up to 976 MiB meas (§05); decode held to 133,120 per slot (29.2 t/s per slot) and collapsed at 135,168 (14.04 t/s per slot), both at n4/p0.75. This page’s per-card arithmetic has no two-slot text configuration, so the page derives no window at its 1,796 MiB desktop threshold for this card. The flags: all layers on the GPU, two slots, an 8-bit KV cache, the n-max 4 / p-min 0 drafter (§06), and --reasoning-preserve. The effort levels it offers: low, medium and xhigh; high is not a level of this template (§14); turns per window in §09. Speed (greedy-only; not the card’s temperature-1.0 sampler) n=2, one slot, n4/p0.75, one sweep 2026-10-04: 79.31 t/s answer tokens at 1.5k of depth (two pass means 79.17–79.45; higher speed level, derived: 86.5), 56.34 t/s at 91k (two pass means 56.33–56.35; higher level: 60.7); floor 44.12 t/s with no drafter (43.74–44.49; its first pass also ran inside the downloads), ceiling 99.53 t/s at n10/p0.5 on one prompt at -c 32768; the reasoning-regime probes are void (their passes disagreed beyond 3%, in the sweep and in a re-run on 2026-10-05). With both slots decoding at 124,928 per slot, the knee row measured 28.97 t/s per slot at n4/p0.75 meas. Not comparable with the August tables in §06. Temperature-1.0 decode n=2, one slot, n4/p0.75: 63.86–63.98 t/s (§09). VRAM at the top of this window: the knee row at 124,928 per slot measured a server footprint of 23,546 MiB dedicated + 424 MiB shared meas (2026-10-04). Room left for your desktop: up to 976 MiB by the fully resident test, which counts dedicated memory only; on the server footprint (dedicated + shared) the knee row at 124,928 per slot left 606 MiB der — both below this page’s 1,796 MiB threshold.
NVIDIA 24 GB · 3090 — Swift IQ3_XXS (text + vision, 1 slot, 195,584-token window, derived)
UNVERIFIED — DERIVED CONFIG :: Swift IQ3_XXS, text + vision, 1 slot, 195,584 per slot (derived from the :: per-card arithmetic at this page's 1,796 MiB desktop threshold). :: Choose this for vision on Swift IQ3_XXS with one slot; choose :: the Swift IQ3_XXS two-slot vision card for two users at a :: shorter window, the Swift IQ4_XS vision card for the :: lower-perplexity file, the Swift IQ3_XXS text card for :: text only, the default card for the base model. :: file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF :: (IQ3_XXS, 12,320,168,256 bytes) + F16 mmproj :: (mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf, 927,607,552 bytes) llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 195584 :: DERIVED, not measured: the per-card :: arithmetic's fully resident window at :: 1,796 MiB desktop (Appendix A). The :: October launcher carries 262,144, :: decode-verified at 253,952 (31.14 t/s) :: and 262,144 (30.84 t/s) with a 366-415 :: MiB desktop; not fully resident: it :: spills at load from its first step -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: VRAM at the top of this window: not measured at 195,584. The knee :: rows above it measured a server footprint (dedicated + shared) of :: 23,837 + 1,516 MiB at 253,952 and 23,847 + 1,874 at 262,144, both :: spilling; 195,584 is the per-card arithmetic's fully resident :: window at a 1,796 MiB desktop (derived, Appendix A). :: Speed band (greedy-only; not the card's temperature-1.0 sampler): :: one slot, no projector, n4/p0.75, -c 131072, two passes, one sweep 2026-10-04; :: not comparable with the August tables in section 06. :: Floor (drafter off, reasoning off, 1.5k): 44.12 t/s (43.74, 44.49); :: first pass inside the downloads, passes within 3% :: Short context, answer regime, 1.5k: 79.31 t/s (79.17, 79.45); :: higher speed level, derived: 86.5 t/s :: Short context, reasoning regime, 1.5k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: At depth, answer regime, 91k: 56.34 t/s (56.35, 56.33); :: higher speed level, derived: 60.7 t/s :: At depth, reasoning regime, 91k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: Ceiling (n10/p0.5, reasoning off, -c 32768, one prompt): :: 99.53 t/s :: Temperature-1.0 decode (the card's sampler, n4/p0.75, -c 131072, :: xhigh): 63.86-63.98 t/s :: Effort levels offered: low, medium, xhigh (high is not a level of :: this template, section 14). Turns per window: :: section 09. :: Drafter: on the one prompt (JavaScript, -c 32768, greedy, reasoning off): :: n4/p0.75 acceptance 0.927, 3.142 accepted tokens per verify pass :: n10/p0.5 acceptance 0.686, 6.071 accepted tokens per verify pass :: Drafter board VRAM at -c 32768: n4 costs 1,092-1,098 MiB, n10 adds 898 MiB :: 2026-10-08 n/p sweep (this file, one slot, -c 81920, text only, xhigh): :: n3/p0 shallow reasoning 1.097x n4/p0.75 (paired 95% 1.065-1.130), answers :: 1.054x; n4/p0 1.034x, n5/p0 0.940x. n3 keeps one fewer recurrent row :: than n4 (149.625 MiB predicted; not resolved on this file, section 06). :: No projector or image in the sweep; how images move the best n-max: not measured. :: FLAGS DELIBERATELY OMITTED: :: --parallel 2 — the two-slot vision card. :: -ctk q4_0 — not measured for this file. :: nvidia-smi -pl — not measured for this file; a persistent hardware :: setting that needs an elevated shell. :: Invalidation trigger: build 10502 / 0adcc3bb5, driver 596.36, :: a new upload of the bartowski Swift-1.5 IQ3_XXS file. :: Fields that differ from the measured command: :: Port (no effect on inference). Sampler and effort :: (the speed probes ran greedy; the temperature-1.0 line ran the :: card's sampler at xhigh). The knee rows ran --reasoning off at :: temperature 0 (KV memory per token is the same for prompt and :: generated tokens); every measurement loaded the file :: memory-mapped and the window fills turned the host-RAM prompt :: cache off (--cache-ram 0); the card loads with --load-mode none :: (system RAM and load time differ, not throughput: 43.0 against :: 43.4 t/s on the anchor, section 15) and keeps the prompt cache on (its RAM :: use at depth is unmeasured). Quality at depth: not verified (the :: needle test did not run).
Copy-paste version — the same command with the comments removed
llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 195584 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"The file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF, IQ3_XXS — 12.32 GB on the download page, 11.47 GiB (12,320,168,256 bytes), plus the F16 mmproj (927,607,552 bytes). The window: -c 195584 der — derived from the per-card arithmetic at this page’s 1,796 MiB desktop threshold (Appendix A); not measured at this window. The October launcher carries -c 262144, decode-verified at 253,952 (31.14 t/s) and 262,144 (30.84 t/s) with a 366–415 MiB desktop; not fully resident: it spills at load from its first step. The flags: all layers on the GPU, one slot, an 8-bit KV cache, the vision projector, the n-max 3 / p-min 0 drafter, and --reasoning-preserve. The effort levels it offers: low, medium and xhigh; high is not a level of this template (§14); turns per window in §09. Speed (greedy-only; not the card’s temperature-1.0 sampler) n=2, one slot, no projector, n4/p0.75, one sweep 2026-10-04: 79.31 t/s answer tokens at 1.5k of depth (two pass means 79.17–79.45; higher speed level, derived: 86.5), 56.34 t/s at 91k (two pass means 56.33–56.35; higher level: 60.7); floor 44.12 t/s with no drafter (43.74–44.49; its first pass also ran inside the downloads), ceiling 99.53 t/s at n10/p0.5 on one prompt at -c 32768; the reasoning-regime probes are void (their passes disagreed beyond 3%, in the sweep and in a re-run on 2026-10-05). Not comparable with the August tables in §06. Temperature-1.0 decode n=2, n4/p0.75, no projector: 63.86–63.98 t/s (§09). VRAM at the top of this window: not measured at -c 195584; the knee rows measured a server footprint of 23,837 MiB dedicated + 1,516 MiB shared at 253,952 tokens and 23,847 + 1,874 at 262,144, both spilling meas (2026-10-04). Room left for your desktop: 1,796 MiB — the per-card arithmetic’s target at this page’s desktop threshold der (Appendix A).
NVIDIA 24 GB · 3090 — Swift IQ3_XXS (text + vision, 2 slots, 93,184 tokens per slot, derived)
UNVERIFIED — DERIVED CONFIG :: Swift IQ3_XXS, text + vision, 2 slots, 93,184 per slot (derived from :: the per-card arithmetic at this page's 1,796 MiB desktop threshold). :: Choose this to serve two concurrent users with vision on Swift :: IQ3_XXS; choose the Swift IQ3_XXS one-slot vision card for a :: longer window when serving one user, :: the Swift IQ3_XXS two-slot text card for text only, :: the default card for the base model. :: file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF :: (IQ3_XXS, 12,320,168,256 bytes) + F16 mmproj :: (mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf, 927,607,552 bytes) llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 186368 :: 93,184 per slot x 2. DERIVED, not :: measured: the per-card arithmetic's :: fully resident window at 1,796 MiB :: desktop (Appendix A). The October :: launcher carries 123,904 per slot :: (247,808 total), decode-verified at :: 122,880/slot (24.9 t/s per slot) and :: 133,120 (23.12 t/s per slot) with a :: 338-539 MiB desktop; not fully resident :: (Appendix A's knee table puts the spill :: onset at 122,880 per slot) -ngl 99 --parallel 2 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: VRAM at the top of this window: not measured at 93,184 per slot. :: The knee rows above it measured a server footprint (dedicated + :: shared) of 23,691 + 870 MiB at 114,688/slot and 23,737 + 1,530 at :: 122,880/slot; 93,184 is the per-card arithmetic's fully resident :: window at a 1,796 MiB desktop (derived, Appendix A). :: Speed band (greedy-only; not the card's temperature-1.0 sampler): :: one slot, no projector, n4/p0.75, -c 131072, two passes, one sweep 2026-10-04; :: not comparable with the August tables in section 06. :: Floor (drafter off, reasoning off, 1.5k): 44.12 t/s (43.74, 44.49); :: first pass inside the downloads, passes within 3% :: Short context, answer regime, 1.5k: 79.31 t/s (79.17, 79.45); :: higher speed level, derived: 86.5 t/s :: Short context, reasoning regime, 1.5k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: At depth, answer regime, 91k: 56.34 t/s (56.35, 56.33); :: higher speed level, derived: 60.7 t/s :: At depth, reasoning regime, 91k: void (its group failed :: 3% pass agreement in the sweep and in a 2026-10-05 re-run) :: Ceiling (n10/p0.5, reasoning off, -c 32768, one prompt): :: 99.53 t/s :: Two slots decoding at once (knee row, 114,688 per slot, answer :: regime, n4/p0.75): 25.2 t/s per slot :: Temperature-1.0 decode (the card's sampler, n4/p0.75, -c 131072, :: xhigh, one slot): 63.86-63.98 t/s :: Effort levels offered: low, medium, xhigh (high is not a level of :: this template, section 14). Turns per window: :: section 09. :: Drafter: on the one prompt (JavaScript, -c 32768, greedy, reasoning off): :: n4/p0.75 acceptance 0.927, 3.142 accepted tokens per verify pass :: n10/p0.5 acceptance 0.686, 6.071 accepted tokens per verify pass :: Drafter board VRAM at -c 32768: n4 costs 1,092-1,098 MiB, n10 adds 898 MiB :: 2026-10-08 n/p sweep at the October launcher's two-slot vision shape :: (123,904 x 2, projector loaded, both slots busy, xhigh; second pass, :: one load per setting, no interval): per slot, n4/p0 decoded :: reasoning tokens 1.337x n4/p0.75 and answers 1.484x (n3/p0: :: 1.185x and 1.233x). Depth with two slots: not resolved. Same :: n-max, same memory (section 06). This card's own 93,184 x 2 window is derived. :: FLAGS DELIBERATELY OMITTED: :: -ctk q4_0 — not measured for this file. :: nvidia-smi -pl — not measured for this file; a persistent hardware :: setting that needs an elevated shell. :: Invalidation trigger: build 10502 / 0adcc3bb5, driver 596.36, :: a new upload of the bartowski Swift-1.5 IQ3_XXS file. :: Fields that differ from the measured command: :: Port (no effect on inference). Sampler and effort :: (the speed probes ran greedy; the temperature-1.0 line ran the :: card's sampler at xhigh). The knee rows ran --reasoning off at :: temperature 0 (KV memory per token is the same for prompt and :: generated tokens); every measurement loaded the file :: memory-mapped and the window fills turned the host-RAM prompt :: cache off (--cache-ram 0); the card loads with --load-mode none :: (system RAM and load time differ, not throughput: 43.0 against :: 43.4 t/s on the anchor, section 15) and keeps the prompt cache on (its RAM :: use at depth is unmeasured). Quality at depth: not verified (the :: needle test did not run).
Copy-paste version — the same command with the comments removed
llama-server.exe -m ukisai_Swift-1.5-Qwen3.8-27b-IQ3_XXS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 186368 -ngl 99 --parallel 2 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"The file: bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF, IQ3_XXS — 12.32 GB on the download page, 11.47 GiB (12,320,168,256 bytes), plus the F16 mmproj (927,607,552 bytes). The window: -c 186368 (93,184 per slot × 2) der — derived from the per-card arithmetic at this page’s 1,796 MiB desktop threshold (Appendix A); not measured at this window. The October launcher carries 123,904 per slot (247,808 total), decode-verified at 122,880 per slot (24.9 t/s per slot) and 133,120 (23.12 t/s per slot), both at n4/p0.75, with a 338–539 MiB desktop; not fully resident (Appendix A’s knee table puts the spill onset at 122,880 per slot). A fill at the launcher’s own window on 2026-10-07 (123,904 × 2, projector loaded, 112,652 and 112,647 tokens in the two slots, both decoding) held at 24.48 and 24.81 t/s per slot at n4/p0.75 meas n=2, with 572–634 MiB desktops and 1,620–1,722 MiB in shared memory at depth; that is another sweep, so its absolute t/s are not comparable with the 2026-10-04 figures (Appendix B). The flags: all layers on the GPU, two slots, an 8-bit KV cache, the vision projector, the n-max 4 / p-min 0 drafter, and --reasoning-preserve. The effort levels it offers: low, medium and xhigh; high is not a level of this template (§14); turns per window in §09. Speed (greedy-only; not the card’s temperature-1.0 sampler) n=2, one slot, no projector, n4/p0.75, one sweep 2026-10-04: 79.31 t/s answer tokens at 1.5k of depth (two pass means 79.17–79.45; higher speed level, derived: 86.5), 56.34 t/s at 91k (two pass means 56.33–56.35; higher level: 60.7); floor 44.12 t/s with no drafter (43.74–44.49; its first pass also ran inside the downloads), ceiling 99.53 t/s at n10/p0.5 on one prompt at -c 32768; the reasoning-regime probes are void (their passes disagreed beyond 3%, in the sweep and in a re-run on 2026-10-05). With both slots decoding at 114,688 per slot, the knee row measured 25.2 t/s per slot at n4/p0.75 meas. Not comparable with the August tables in §06. Temperature-1.0 decode n=2, one slot, n4/p0.75, no projector: 63.86–63.98 t/s (§09). VRAM at the top of this window: not measured at 93,184 per slot; the knee rows measured a server footprint of 23,691 MiB dedicated + 870 MiB shared at 114,688 per slot and 23,737 + 1,530 at 122,880 per slot meas (2026-10-04). Room left for your desktop: 1,796 MiB — the per-card arithmetic’s target at this page’s desktop threshold der (Appendix A).
Every number above is tied to llama.cpp build 10502, commit 0adcc3bb5, and NVIDIA driver 596.36. Re-measure the speed band and the VRAM ceiling when any of these changes: a new llama.cpp build (the MTP implementation moved twice in two days of the August 2026 measurements), a driver update, a new quantization of this model, a change to the model’s own draft head, or a new upload of the bartowski Swift-1.5 files. The cheapest re-check is the one command in §15; it takes about two minutes and tells you, on the anchor file, whether the build, the driver or the card moved; it does not check a Swift file.
The full menu for a 24 GB card, losers included
The reference machine’s August launcher offered eight configurations (rows 1–8), and this is all of them — including the ones that lost, each with the number that beat them, so that an option you have already heard of is answered rather than absent. Rows 9–14 are the Swift cards (§03): their speeds come from one sweep on 2026-10-04 and are not comparable with rows 1–8, and five of their six windows are derived. Hover or tap a row to see the flags that produce it. Since 2026-10-08 rows 1, 2 and 9–14 carry the n/p sweep’s drafters (n3/p0 with one slot, n4/p0 with two); rows 4–6 run files the sweep did not measure and carry n4/p0 since 2026-10-09 der (p-min costs no memory, so their windows still fit; on the four files the sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot and read 0.90–1.03 of it after a 57,540-token prefix); rows 3, 7 and 8 keep the August flags (drafter off, an external drafter, and the August comparison run that row 8 reproduces); every Expect figure keeps the drafter it was measured at (§06).
| # | Configuration | Window | Expect (regime stated) | When to pick it — or why not |
|---|---|---|---|---|
| 1 | UD-IQ4_XS + hi-res vision (max-tokens 10580 · n-max 3 / p-min 0) | 122,880 | 76–79 t/s answer tokens, short contextC27 · 64.8 at 91k · 37–39 on xhigh reasoning at 91k · 2.7–3.7 GiB slack | the daily default — screenshot loops to 4K detail (§12); a full xhigh cycle plus about ten 1440p shots fit the window. Runs n3/p0 since 2026-10-08, the best score for this file at the cards’ sampler (§06); the speeds in this row were measured at n10/p0.5 and n4/p0.75, and why it never carried n10/p0.5 is in the box below |
| 2 | UD-IQ4_XS text-only (n-max 3 / p-min 0; n10/p0.5 was the August greedy-code peak) | 180,224 | 73.94 t/s answer tokens at -c 180224 (three settled probes, 1.91% spread) — −3.1% against 76.32 t/s at -c 32768 in the same probeC23 (a different probe on novel JavaScript read 95.4; matched sweep peak 93.9 t/s at -c 32768) · 582 MiB of slack at -c 180224; 847 MiB measured at the same window days earlier — neither clears the 1,796 MiB reserve | the largest text-only window measured resident with n10/p0.5, this row’s August drafter, and the row where the longer guesses paid for themselves on greedy code (+12% on code on this file, +17% on Q4_K_M). At 196,608 the slack measures only 415 MiB. 2026-10-08: this row’s Expect figures, its window and the gains above were measured at n10/p0.5, on greedy code with thinking off; at the cards’ temperature-1.0 sampler with thinking on, n3/p0 beat it on this file on shallow reasoning and answer tokens (1.111 and 1.120 against 0.961 and 1.049 of n4/p0.75) and needs about 1 GB less, so the row now carries n3/p0; after a 57,540-token prefix n10/p0.5 read higher, intervals overlapping (§06) |
| 3 | UD-IQ4_XS text-only + --spec-type none | 262,144 | ~43 t/s (no drafter) at load · 15.96 t/s and 755 MiB left when actually filled to 218,233 tokens · drafter left on (n-max 4) = 2,364 MiB spilled at load (text-only) · 3,502 MiB spilled at load and 8.0 t/s at a 91k fill with the projector loadedC29 | full native context on this file — drafter off, card free of any graphical session (§02), and filled to 218,233 real tokens it keeps only 755 MiB, which is 1,041 MiB short of the 1,796 MiB this page reserves for a desktop. If you actually need this window, use UD-Q2_K_XL instead — same -c 262144, drafter on at n-max 4 / p-min 0 der (21.33 t/s at the same depth with 1,717 MiB of slack, measured at p-min 0.75; p-min costs no memory), and it ties this file on accuracy (§08). The price is a seven-minute prefill |
| 4 | Q4_K_M + vision | 122,880 | 69.8 t/s answer tokens · 57.9 reasoning · ~0.3 GiB slack (n4/p0.75) | maximum measured quality — leaves about 0.3 GiB for your desktop, less than a browser typically holds (browsers spill it to 20–35 t/s, §11); with a desktop up, drop -c or go text-only |
| 5 | UD-Q4_K_XL text-only | 122,880 | ~55–65 t/s reasoning | loser: perplexity 6.682 against Q4_K_M's 6.535 (+2.3%) while being 1 GiB bigger. On 24 GB the extra gigabyte gains you nothing |
| 6 | UD-Q4_K_M + vision | 122,880 | ~55–65 t/s reasoning | loser: measured worse than plain Q4_K_M at the same size — perplexity +1.8%. Skip |
| 7 | DFlash2 external drafter | — | ~46 t/s on real code (best at n-max 2) | loser: beaten by the built-in MTP head's 57.9 on the same content, and it costs a 1.14 GB download plus a source build. Historical interest (§06) |
| 8 | NVFP4-HIGH | — | ~48–54 t/s reasoning | loser on this card's chip generation (Ampere — the RTX 30 series, which the reference 3090 belongs to): software dequantization fallback, perplexity +4.4%, and 13.4% more energy per token than UD-IQ4_XS. It runs natively only on Blackwell — RTX 50, GB10, B200 (§08) |
| 9 | Swift IQ4_XS text (1 slot · n3/p0) | 159,744 | 87.2 t/s answer tokens at 1.5k, 61.14 at 91k n=2 (greedy-only; not the card’s temperature-1.0 sampler; n4/p0.75, 2026-10-04) · reasoning regime: void | text-only on the Swift IQ4_XS file (card) |
| 10 | Swift IQ4_XS + vision (1 slot · n3/p0) | 128,000 der | 87.2 t/s answer tokens at 1.5k, 61.14 at 91k n=2 (greedy-only; not the card’s temperature-1.0 sampler; n4/p0.75, no projector, 2026-10-04) · reasoning regime: void | vision on the Swift IQ4_XS file (card) |
| 11 | Swift IQ3_XXS text (1 slot · n3/p0) | 216,064 der | 79.31 t/s answer tokens at 1.5k (two pass means 79.17–79.45), 56.34 at 91k (two pass means 56.33–56.35) n=2 (greedy-only; not the card’s temperature-1.0 sampler; n4/p0.75, 2026-10-04) · reasoning regime: void | one user, longest single-slot Swift window (card) |
| 12 | Swift IQ3_XXS text, 2 slots (n4/p0) | 123,904 × 2 der (desktop ≤ 976 MiB) | 79.31 t/s answer tokens at 1.5k, 56.34 at 91k, one slot n=2 (greedy-only; not the card’s temperature-1.0 sampler; n4/p0.75, 2026-10-04) · reasoning regime: void | two concurrent text users (card) |
| 13 | Swift IQ3_XXS + vision (1 slot · n3/p0) | 195,584 der | 79.31 t/s answer tokens at 1.5k (two pass means 79.17–79.45), 56.34 at 91k (two pass means 56.33–56.35) n=2 (greedy-only; not the card’s temperature-1.0 sampler; n4/p0.75, no projector, 2026-10-04) · reasoning regime: void | vision on Swift IQ3_XXS, one slot (card) |
| 14 | Swift IQ3_XXS + vision, 2 slots (n4/p0) | 93,184 × 2 der | 79.31 t/s answer tokens at 1.5k, 56.34 at 91k, one slot n=2 (greedy-only; not the card’s temperature-1.0 sampler; n4/p0.75, no projector, 2026-10-04) · reasoning regime: void | two concurrent users with vision on Swift IQ3_XXS (card) |
Every "Expect" band carries its token regime, because answer tokens run about 1.7× the speed of reasoning tokens on code and a band without its regime is not a number you can plan with. Speed also falls with depth, not only with spill: at n4/p0.75 (the flags shipped before 2026-10-08), answer tokens measure 86.3 → 80.2 → 64.8 t/s at 1.5k / 28k / 91k of fill under the cooled protocol, while the reasoning stream on the same file and flags runs 51.2 → 47.1 → 36.6. An agent session holding 10–30k of context therefore lives near 80 t/s of deliverable; the same session left on the xhigh default over a 91k document spends most of its wall clock at 37–39 t/s, at either measured drafter setting (n4/p0.75, n10/p0.5; n3/p0 was not measured at 91k). Slack figures are the board VRAM measured left over with little on screen; the desktop's own share swung 1,179–1,669 MiB in direct no-server readings on one machine (§05), so a row keeping under about 1 GiB has room for a near-idle desktop and nothing beyond it.
The August launcher behind rows 1–8, its picks 1 and 2 moved to n3/p0 on 2026-10-08 and picks 4–6 to n4/p0 on 2026-10-09 — the reference machine's own serve-qwen.bat, with paths genericized (tap to expand)
2026-10-08 and 2026-10-09: this is the August file with two changes, marked by dated rem lines inside it: picks 1 and 2 (UD-IQ4_XS) now take --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0. Its August split — n10/p0.5 on the text-only picks, n4/p0.75 wherever the projector is loaded — is superseded for its own sampler (temperature 1.0, thinking on), where the n/p sweep found n3/p0 the best score on that file with one slot (§06); n10/p0.5 was the August grid’s peak on greedy, thinking-off code. Picks 4–6 run files the sweep did not measure and, since 2026-10-09, take n4/p0 (SPEC_N4) der: p-min costs no memory, so their windows still fit, and on the four files the sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot and read 0.90–1.03 of it after a 57,540-token prefix; the t/s the file quotes keep the drafter they were measured at.
@echo off
rem ============================================================================
rem serve-qwen.bat - llama.cpp server for Qwen3.8-27B on RTX 3090 24GB
rem ============================================================================
rem Measured 2026-08-22, corrected 2026-08-23. Every pick carries the
rem measurement that justifies it, because this file travels to other
rem machines without the guide.
rem - -ngl 99 always (llama.cpp counts the output layer as layer 65; -ngl 64
rem leaves the output layer on CPU: 25.7 vs 39.7 t/s on Q4_K_M,
rem 29.8 vs 42.3 on UD-IQ4_XS, with NO memory signature)
rem - MTP flags are PER PICK, and the split is NOT per file. A matched sweep
rem (2026-08-23, both quants, novel code, thinking OFF) peaks at n-max 10 /
rem p-min 0.5 on BOTH files: IQ4_XS 93.9 t/s at -c 32768 (2.18x over its
rem 43.0 no-drafter floor; at -c 180224 the same drafter measures -3.1%%),
rem Q4_K_M 81.7 (2.04x over 40.0). The shipped n-max 4 / p-min 0.75
rem reads 83.5 and 69.8 on the same probe - 11%% / 15%% off the peak, and
rem 898 MiB cheaper. So, in August: n10/p0.5 on the text-only picks where
rem that VRAM is spare, n4/p0.75 wherever the projector is loaded or the
rem slack is thin.
rem 2026-10-08: SUPERSEDED for this launcher's own sampler (temp 1.0,
rem thinking on). The n/p sweep found n3/p0 best on UD-IQ4_XS with one
rem slot: shallow reasoning 1.11x and answers 1.12x n4/p0.75, against
rem n10/p0.5's 0.96x and 1.05x; n-max 3 is ~150 MiB lighter than 4 and
rem ~1 GB lighter than 10. So picks 1 and 2 now take SPEC_N3. The figures
rem above are greedy code with thinking off, where longer drafts still
rem lead. Picks 4-6 run files the sweep did not measure (guide, section 06).
rem Acceptance is a property of the MTP head, not of the quant - the two
rem files land within 1.6 points at six of seven configs (3.7 at the seventh).
rem - the drafter is also the cheapest ENERGY setting here: 3.21 J per decode
rem token at n10/p0.5 against 8.10 with --spec-type none (2.52x), and
rem the board draws the SAME power while decoding either way (344.6 / 341.7 /
rem 341.0 W): the entire win is throughput, not wattage.
rem - TOKEN REGIME decides the number. Answer tokens (thinking off) run ~1.7x
rem reasoning tokens, and mean DRAFT LENGTH is what predicts speed, not
rem acceptance: at 91k depth the reasoning stream drafts 2.99 tokens per
rem pass (36.6 t/s) and the answer stream 4.31 (62.0) - at the same ~0.90
rem acceptance. Highest-acceptance config (n2/p0.75, 96.5%%) is the slowest.
rem - Depth, answer tokens, n4/p0.75, cooled probes: 86.3 t/s at 1.5k, 80.2 at
rem 28k, 64.8 at 91k. On the xhigh REASONING stream at 91k expect 37-39 t/s
rem (n10/p0.5 buys only 5.6%% there - not worth 898 MiB).
rem - UD-IQ4_XS (13.3 GiB): quality tied with Q4_K_M (perplexity 6.596 vs 6.535;
rem the GSM8K half of that comparison is unaudited - see the guide),
rem FASTER everywhere measured, and 8%% cheaper per token in energy.
rem Measured ceilings for a window whose every byte still lives on the card:
rem -c 180224 text-only WITH n10/p0.5; the full native -c 262144 ONLY with
rem --spec-type none (1,360 MiB of board VRAM left at load; filled to
rem 218,233 tokens it keeps 755 MiB, 1,041 short of the 1,796 reserve -
rem needs no graphical session on the card at all); -c 163840 with vision
rem at n4/p0.75.
rem 122880 + vision is the daily default, chosen to leave a desktop room.
rem Filling a window is the only test that counts - a short
rem probe at 262144 with the drafter on looks fine and delivers 8 t/s on a
rem 91k document.
rem - Q4_K_M (15.4 GiB): the quality leader, by one metric and narrowly; +mmproj
rem at -c 122880 (~0.3 GiB of slack - room for a near-idle desktop, no more).
rem - the BF16 mmproj costs 1,138 MiB of VRAM and 0%% of decode: measured at
rem 90,862 tokens of fill, projector on vs off, 0.04-0.09%% apart.
rem - KV cache: q8_0 costs +0.309%% perplexity, q4_0 +0.693%% (measured
rem 2026-08-23, super-linear in bits). q8_0 is what these recipes ship.
rem - reasoning_effort is THE wall-clock knob: ~4x medium on one long
rem authoring task, 1.8x across a 175-prompt benchmark suite where the
rem three levels scored within 1.6 points of each other. llama-server
rem ignores per-request effort, so it must be set here.
rem ----------------------------------------------------------------------------
rem USAGE: serve-qwen.bat [low^|medium^|xhigh] [context] [1-8]
rem arg1 = reasoning effort, default xhigh (quality-first)
rem arg2 = context override (otherwise each choice's safe default)
rem arg3 = model choice 1-8, skips the menu (for scripts)
rem No args: menu below, auto-picks [1] after 8 seconds, so an
rem unattended restart never blocks on a human.
rem ----------------------------------------------------------------------------
set EFFORT=%1
if "%EFFORT%"=="" set EFFORT=xhigh
if /i "%EFFORT%"=="low" goto effort_ok
if /i "%EFFORT%"=="medium" goto effort_ok
if /i "%EFFORT%"=="xhigh" goto effort_ok
echo Unknown reasoning effort "%EFFORT%". Usage: serve-qwen.bat [low^|medium^|xhigh] [context] [1-8]
pause
exit /b 1
:effort_ok
set CTXARG=%2
set PICK=%3
set MODELS=C:\path\to\your\models
set MMPROJ=%MODELS%\mmproj-Qwen3.8-27B-BF16.gguf
rem the drafter is PER PICK (it is a load-time flag; llama-server ignores the
rem per-request speculative fields). 2026-10-08: SPEC_N3 is the n/p sweep's
rem best for UD-IQ4_XS with one slot and serves picks 1 and 2 (header note).
set SPEC_N3=--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0
rem August: n10/p0.5 is the greedy-code peak on both files but costs 898 MiB,
rem so it rode the text-only pick. Superseded 2026-10-08, kept for comparison,
rem used by no pick. SPEC_SAFE (n4/p0.75) likewise since 2026-10-09: picks 4-6 take SPEC_N4.
set SPEC_FAST=--spec-type draft-mtp --spec-draft-n-max 10 --spec-draft-p-min 0.5
set SPEC_N4=--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0
set SPEC_SAFE=--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75
if not "%PICK%"=="" goto pick_%PICK%
echo.
echo Qwen3.8-27B - pick a model (effort: %EFFORT%):
echo [1] UD-IQ4_XS + HI-RES vision -c 122880 n3/p0 ~3 GiB slack (DEFAULT)
echo ~83-86 t/s answer tokens on code,
echo ~65 at a 91k fill, 37-39 on xhigh
echo reasoning at 91k (n4/p0.75 / n10/p0.5);
echo screenshots to ~4K detail (1080p=2.0k /
echo 1440p=3.6k / 4K=8.2k ctx tokens each);
echo full xhigh cycle + ~10 shots fit
echo [2] UD-IQ4_XS text-only -c 180224 n3/p0 at n10/p0.5 (August): 93.9 t/s
echo at -c 32768, -3.1%% at -c 180224 (73.94,
echo 582 MiB slack), the measured text-only
echo ceiling WITH n10/p0.5: little on screen
echo [3] UD-IQ4_XS text-only -c 262144 DRAFTER OFF - required, not optional
echo full native - needs the card free of any
echo graphical session;
echo no drafter = ~43 t/s short-context floor
echo [4] Q4_K_M + vision -c 122880 n4/p0 highest measured quality,
echo reference config
echo ~0.3 GiB slack - a near-idle desktop only
echo [5] UD-Q4_K_XL text-only -c 122880 n4/p0 quality LOSES to Q4_K_M
echo (perplexity +2.3%%) and is 1 GB bigger:
echo gains nothing
echo [6] UD-Q4_K_M + vision -c 122880 n4/p0 measured worse than plain Q4_K_M
echo (perplexity +1.8%%) - skip
echo [7] DFlash2 build ~46 t/s on real code (loses to built-in MTP)
echo [8] NVFP4 HIGH ~48-54 t/s dequant fallback, lower quality,
echo +13%% J/token vs IQ4_XS
echo.
echo t/s above are ANSWER tokens (thinking off) at short context unless said
echo otherwise. Reasoning tokens run ~1.7x slower on the same server, and
echo decode falls with DEPTH: 86 shallow / 80 at 28k / 65 at 91k (answer),
echo 51 / 47 / 37 (reasoning). Acceptance actually RISES with depth
echo (0.80 -^> 0.92): the cost is KV reads, not the drafter failing.
echo.
choice /C 12345678 /T 8 /D 1 /M "Choice"
goto pick_%ERRORLEVEL%
:pick_1
set MODEL=%MODELS%\Qwen3.8-27B-UD-IQ4_XS.gguf
set CTX=122880
set MM=--mmproj "%MMPROJ%" --image-min-tokens 1024 --image-max-tokens 10580
rem August: n4/p0.75 on purpose - this pick already spends 1,138 MiB on the
rem projector, and the n10 flags would take another 898 MiB out of the slack
rem that absorbs the desktop (measured 1,179-1,669 MiB). They also buy least
rem here - 5.6%% on the xhigh reasoning stream this default decodes at depth.
rem 2026-10-08: n3/p0 (n/p sweep, this file at xhigh): shallow reasoning 1.11x
rem and answers 1.12x n4/p0.75; ~150 MiB lighter than n4, so the slack grows.
set SPEC=%SPEC_N3%
goto launch
:pick_2
set MODEL=%MODELS%\Qwen3.8-27B-UD-IQ4_XS.gguf
set CTX=180224
set MM=
rem August: text-only, the 898 MiB was spare here, so it took the greedy-code
rem peak n10/p0.5 (+12%%). 2026-10-08: n3/p0 instead - at this launcher's
rem sampler it beat n10/p0.5 on shallow reasoning (1.11x against 0.96x of
rem n4/p0.75) and on answers (1.12x against 1.05x), and needs ~1 GB less.
set SPEC=%SPEC_N3%
goto launch
:pick_3
set MODEL=%MODELS%\Qwen3.8-27B-UD-IQ4_XS.gguf
set CTX=262144
set MM=
rem drafter OFF is required, not a preference: with the n-max 4 / p-min 0.75
rem drafter on this window needs ~24.9 GiB and 2,364 MiB was measured living
rem in system RAM at that setting.
set SPEC=--spec-type none
echo.
echo WARNING: -c 262144 fits ONLY with the drafter off. Filled to 218,233
echo tokens it keeps 755 MiB, 1,041 short of the 1,796 reserve. At load:
echo 23,216 MiB derived, 1,360 MiB of board VRAM left - 309 MiB SHORT once a
echo worst-case desktop takes its measured 1,669 MiB. WITH the n-max 4 /
echo p-min 0.75 drafter on it does not fit at all: 2,364 MiB spilled text-only,
echo 3,502 MiB with the projector loaded and 8.0 t/s on a 91k-token fill.
echo A browser or agent
echo web UI will spill it either way. Run this with NO graphical session using
echo the card: close everything that draws, drive the screen from a second card,
echo or log in from another machine. Unplugging the monitor does not do it.
echo.
goto launch
:pick_4
set MODEL=%MODELS%\Qwen3.8-27B-Q4_K_M.gguf
set CTX=122880
set MM=--mmproj "%MMPROJ%" --image-min-tokens 1024
rem ~0.3 GiB slack with the projector loaded: no room for the n10 flags here.
set SPEC=%SPEC_N4%
goto launch
:pick_5
set MODEL=%MODELS%\Qwen3.8-27B-UD-Q4_K_XL.gguf
set CTX=122880
set MM=
rem kept in the menu as an answered loser: perplexity 6.682 vs Q4_K_M's 6.535.
set SPEC=%SPEC_N4%
goto launch
:pick_6
set MODEL=%MODELS%\Qwen3.8-27B-UD-Q4_K_M.gguf
set CTX=122880
set MM=--mmproj "%MMPROJ%" --image-min-tokens 1024
rem answered loser: perplexity 6.654, +1.8%% over plain Q4_K_M at the same size.
set SPEC=%SPEC_N4%
goto launch
:pick_7
rem DFlash2 needs its own build (llama.cpp PR 27342) and a separate 1.14 GB
rem drafter; best on real code measured 46.5 t/s at n-max 2, below MTP's 57.9.
call serve-qwen-dflash2.bat %EFFORT%
exit /b
:pick_8
rem NVFP4 on Ampere runs through a software fallback: slower, +4.4%% perplexity,
rem +13%% J/token. Kept for Blackwell users and for the record.
call serve-qwen-nvfp4.bat HIGH %EFFORT%
exit /b
:launch
if not "%CTXARG%"=="" set CTX=%CTXARG%
if not exist "%MODEL%" (
echo Model not found: %MODEL%
pause
exit /b 1
)
echo Serving %MODEL%
echo effort=%EFFORT% ctx=%CTX%
echo spec=%SPEC%
llama-server.exe ^
-m "%MODEL%" ^
%MM% ^
--alias qwen/qwen3.8-27b ^
-c %CTX% ^
-ngl 99 ^
--parallel 1 ^
--load-mode none ^
--api-key dummy ^
-ctk q8_0 -ctv q8_0 ^
%SPEC% ^
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 ^
--reasoning-preserve ^
--chat-template-kwargs "{\"reasoning_effort\":\"%EFFORT%\"}" ^
--jinja ^
--host 127.0.0.1 ^
--port 1234
Three structural differences from the file that actually runs on the reference machine, named so that nobody assumes this is a byte-for-byte copy. The model paths are genericized to %MODELS%; picks 7 and 8 hand off to sibling scripts whose paths are local; llama-server.exe is unqualified here where the real file gives an absolute path, so the real file's matching "binary not found" guard is dropped with it. The eight picks, their windows, the timed default and every other flag are identical. The launcher on disk was synced to this listing on 2026-08-23 (comments and echo text only; behavior verified identical across all eight picks). The later differences are the drafters of picks 1 and 2, moved to n3/p0 in this listing on 2026-10-08, and of picks 4–6, moved to n4/p0 on 2026-10-09 through a new SPEC_N4 der, with SPEC_SAFE (n4/p0.75) kept for comparison and used by no pick, as in the October launcher (its dated rem lines, §06); the reference machine now runs the October launcher, whose picks each carry their measured drafter.
NVIDIA 16 GB · 5080 / 4080 / 4070 Ti S / 5060 Ti
The file: unsloth UD-Q2_K_XL — 9.83 GB, 9.154 GiB resident. The window: -c 65536. The flags: the universal set plus the n4/p0 drafter der: its 13,982 MiB was measured at n4/p0.75, and p-min costs no memory, so the window still fits; on the four files the 2026-10-08 n/p sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot and read 0.90–1.03 of it after a 57,540-token prefix, and n3/p0, about 150 MiB lighter, scored best with one slot at xhigh; neither is measured on this file or card (§06). VRAM at the top of this window: 13,982 MiB meas. Room left for your desktop: 606 MiB against the 14,588 MiB this page budgets for a 16 GB card — enough for a light desktop, not for a browser full of tabs; drop to -c 49152 if you need more. The effort levels it offers: low and medium. xhigh is not offered: its measured thinking appetite is 61,500–75,800 tokens against a 65,536 window, so it would fit on its shortest runs and truncate on its longest — which is worse than a clean no, because you would not know which you got. Expected speed: 25–50 t/s drafter-off, derived from the bandwidth formula in §04 der; the drafter adds to that. No speed on a 16 GB card is measured on this page and none can be, from a machine that owns one 24 GB card. The memory figure is measured and transfers to any card.
Why not UD-Q3_K_XL: it needs 16,906 MiB at -c 65536 — more than the whole card — and ~16,218 MiB at -c 49152, still 1,142 MiB over what a 16 GB card has to give (§08). UD-Q2_K_XL at the same -c 65536 gives twice the window inside less memory and ties the 4-bit reference on accuracy (§08).C15
:: best daily experience on 16 GB, MEASURED 2026-08-25: the 2-bit file, not the :: 3-bit one. UD-Q2_K_XL at -c 65536 and n4/p0.75 needs 13,982 MiB; UD-Q3_K_XL :: at the same window needs 16,906 and does not fit. See section 08's requirement table. :: Q4-with-CPU-offload preserves more fidelity but crawls at ~8-12 t/s (section 07). :: file: unsloth/Qwen3.8-27B-GGUF (UD-Q2_K_XL, 9.83 GB = 9.154 GiB) llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: 13,982 MiB MEASURED on the reference 3090 at this exact :: window at n4/p0.75 (p-min costs no memory) - not arithmetic. Against :: the 14,588 MiB this page budgets for a 16 GB card that :: leaves 606 MiB for your desktop. Drop to -c 49152 if :: you run a browser, or --spec-type none to buy back ~1.3 GiB. :: A >131k window cannot fit resident on 16 GB - that needs :: the Q4 offload path (section 07) --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: p-min 0 derived; start at 4 and :: sweep on your own card (section 06): the optimum moved with :: content and hardware in every run measured here :: On the 3090 (2026-10-08) p-min 0 read higher than 0.75 :: at n-max 4 on shallow reasoning on every file swept, and :: n3/p0 scored best at one slot at xhigh: include both :: (unmeasured on this card and file) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: xhigh NOT OFFERED: :: measured appetite 61.5-75.8k tokens against a 65,536 window, :: so it would fit on short runs and truncate on long ones. You :: would not know which you got - see section 09's ceiling table :: FLAGS DELIBERATELY OMITTED: --parallel 2 (settled 2026-08-25: +22.0% aggregate, :: -35.4% per slot with the drafter on - slower for one user), :: -ctk q4_0 (+0.693% perplexity; UD-Q2_K_XL at -c 65536 fits with 606 MiB :: spare, so the trade buys no window this recipe needs).
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}"12 GB class · RTX 3060 / RTX 5070 / Arc B580
The honest framing first: this class runs a 27B model poorly in every configuration measured or derived here, and if the exact model is not a requirement, a 14B-class Qwen at 4-bit fits resident and is several times faster. The file, if it must be this model: Q4_K_M with most layers on the processor. The window: -c 112640. The effort levels it offers: all three fit the window — but the recipe ships medium, and here the binding constraint is wall clock, not the window. How much wall clock depends entirely on what you ask for. At 6–8 t/s an ordinary xhigh answer — 2,217 tokens of thinking, averaged over 175 benchmark prompts — takes about 5 minutes. A complete build, the one task measured wanting 61,500–75,800, takes about three hours against one for medium. The three-hour figure is the cost of a complete build and overstates ordinary work by roughly thirty times.C4 Ask questions at xhigh freely here; save the overnight run for whole deliverables. Expected speed: 6–8 t/s, derived, and set by your system memory bandwidth rather than by the graphics card. Room left for your desktop: nearly all of the card — almost none of the model is on it. For the Swift-1.5 files' 12 GB windows — which files fit, and how large a window each holds — see Appendix A.
:: honest framing: this class runs the 27B poorly in every configuration. If the :: exact model is not required, a 14B-class Qwen at Q4 fits resident and flies. :: If it must be this model: Q4 with most layers on CPU (~6-8 t/s; speed is set by :: your system RAM. The ~6-8 t/s assumes TWO sticks of DDR5 in dual channel :: (~90 GB/s); one stick = single channel = half the bandwidth = half the t/s). :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M — on a bandwidth-bound :: offload path the smaller Q4 file means fewer GB re-read per token, and the :: measured perplexity in section 08 favours it too) llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 112640 -ngl 28 :: ~28 of the 65 offloadable layers (64 + the output :: layer — section 11); raise until <500 MB VRAM is free --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: CPU layers need ~10 GB of real RAM :: no MTP flags here on purpose: speculation is UNMEASURED on the CPU-offload path. :: The measurement that would justify adding them is one paired probe with and :: without --spec-type draft-mtp at this -ngl, on your machine. It costs ten minutes. --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: unlike the 16 GB card :: this is NOT a window limit — the 112k window (mostly in system RAM) holds xhigh's :: 61.5-75.8k appetite fine. It is a PATIENCE limit: at ~6-8 t/s an xhigh answer runs :: ~3 HOURS against ~1 h for medium (section 09). Switch to xhigh for unattended runs. :: Arc B580: use the Vulkan build (same SYCL caveat as the Arc Pro cards - they are :: all Battlemage, Intel's Arc B series)
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 112640 -ngl 28 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}"Intel Arc Pro B70 · 32 GB
The file: Q6_K — 22.4 GB, 20.9 GiB resident, near-8-bit quality, which the extra VRAM affords. The window: -c 229376 with the drafter, or the full 262144 with --spec-type none. The effort levels it offers: all three; the window holds xhigh's appetite several times over. Expected speed: 18–26 t/s, derived from bandwidth and unmeasured on this card's chip generation (Battlemage — Intel's Arc B series). VRAM at the top of the window: about 30.5 GiB derived with the drafter off. Room left for your desktop: not measured — every figure for this card is derived, so treat the last gigabyte as unverified.
:: VULKAN llama.cpp build - currently beats SYCL on Battlemage, Intel's Arc B series, :: which this card belongs to (SYCL has known perf :: bugs there; IPEX-LLM is archived — do not use it). Re-benchmark after driver updates. :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q6_K, 22.4 GB = 20.9 GiB) llama-server.exe -m Qwen3.8-27B-Q6_K.gguf --alias qwen/qwen3.8-27b -c 229376 -ngl 99 :: NOT the full 262144 with a drafter loaded: at the :: MEASURED 45,056 B/token the native window costs ~11 GiB of :: KV, not the 8.5 the q8 arithmetic suggests, and Q6_K + MTP + :: that needs more than 32 GB. For the FULL 262144 here, add --spec-type :: none and drop the MTP line below (~30.5 GiB). Both figures :: are DERIVED from the 3090's constants — unmeasured on Battlemage --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: verify KV-quant support in your build --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: p-min 0 derived; start at 4, sweep (section 06) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: all three fit, with room to spare at THIS window (229,376). :: An ordinary xhigh answer averages 2,217 tokens of thinking :: (175 benchmark prompts). Only a COMPLETE BUILD has been :: measured wanting 61,500-75,800 - four August samples of one task, :: section 09 - and even that leaves ~153,000 here for prompt :: and history. Earlier editions of this comment quoted 122,880 :: and ~47,000: copied from the 24 GB recipe, never re-done :: against this window. SUPERSEDED 2026-08-25 :: FLAGS DELIBERATELY OMITTED: --parallel 2 (+22.0% aggregate, -35.4% per slot with :: the drafter on, measured 2026-08-25 — see section 03's axis table).
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-Q6_K.gguf --alias qwen/qwen3.8-27b -c 229376 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"Intel Arc Pro B50 · 16 GB
The file: UD-Q2_K_XL — 9.83 GB, 9.154 GiB resident. A 4-bit file does not fit, and at 224 GB/s this card is very slow as soon as any layer has to sit on the processor, so every byte must be resident. The window: -c 65536. VRAM: 13,982 MiB meas on the reference 3090 at this window at n4/p0.75 (p-min costs no memory, so n4/p0 needs the same der) — a memory requirement is a property of the model and the flags, so it transfers to this card even though no speed here does. Room left for your desktop: 606 MiB of the 16 GB. The effort levels it offers: low and medium. xhigh is not offered, on the same window constraint as the NVIDIA 16 GB recipe: 61,500–75,800 tokens of appetite against a 65,536-token window, so it fits on short runs and truncates on long ones. Expected speed: 8–12 t/s der from this card’s 224 GB/s and the bandwidth formula in §04 — no Intel card was ever measured for this page.
UD-Q3_K_XL does not fit on this card with the drafter on (§08).C15
:: Q4 does not fit, and at 224 GB/s any layer left on the CPU is very slow - so run :: the 2-bit file with every byte on the card, Vulkan build. MEASURED 2026-08-25: :: UD-Q2_K_XL at -c 65536 needs 13,982 MiB; UD-Q3_K_XL at that window needs 16,906 :: and does not fit at all. Section 08 has the full requirement table. :: file: unsloth/Qwen3.8-27B-GGUF (UD-Q2_K_XL, 9.83 GB = 9.154 GiB) llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: 13,982 MiB measured, leaving 606 MiB of a 16 GB card for :: your desktop. --spec-type none buys back ~1.3 GiB if you :: would rather have headroom than the drafter. A >131k window :: cannot fit resident on this card --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: p-min 0 derived; start at 4, sweep (section 06) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: xhigh NOT OFFERED: 61.5-75.8k :: appetite against a 65k window - a WINDOW limit (section 09)
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}"DGX Spark · GB10, 128 GB unified memory
The file: Q4_K_M — 16.5 GB, 15.4 GiB resident. With 128 GB of unified memory, capacity is not the constraint; bandwidth is, at 273 GB/s, so the smallest strong file wins. The window: the full -c 262144. The effort levels it offers: all three — but note the wall clock: the reference 100k-token run derives to about 2–2.8 hours here, so xhigh is an unattended-run setting. Expected speed: 10–15 t/s, derived from bandwidth. Room left for your desktop: not applicable — this machine has 128 GB of memory shared between the chip and the system.
:: llama-server (ARM64 Linux build — no .exe here), same stack as every other card. :: Bandwidth (273 GB/s) is the ceiling, so the smallest high-quality file wins: :: Q4_K_M (16.5 GB) over UD-Q4_K_XL (17.6 GB) is ~6% more t/s on a bandwidth-bound :: box — and the measured perplexity favours it too (section 08). :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M, 16.5 GB = 15.4 GiB) llama-server -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 262144 -ngl 99 :: 128 GB unified: full native context is trivial (~11 GiB of :: KV at the MEASURED slope with the drafter on, not the 8.5 :: the arithmetic suggests) --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: p-min 0 derived; start at 4, sweep :: (section 06) — UNMEASURED on GB10; the 3090 gained +45% on :: real code at n4/p0.75; there, p-min 0 read higher than :: 0.75 at n-max 4 on shallow reasoning (2026-10-08) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: window holds it; wall clock :: is the cost — ~2-2.8 h for the 100k reference run (section 07) :: faster vendor alternative: Blackwell-native NVFP4 on vLLM (~1.5x BF16 speed, FP8 KV, :: 92-97% accuracy) — vllm serve unsloth/Qwen3.8-27B-NVFP4 --max-model-len 131072. :: llama.cpp cannot load that release (FP8 lm_head); switching stacks buys the FP4 path
Copy-paste version — the same command with the comments removed
llama-server -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 262144 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"Intel Core Ultra iGPU · Arc B390-class
The file: Q4_K_M — 16.5 GB, 15.4 GiB resident; every gigabyte re-read per token hurts at 153.6 GB/s. The window: -c 65536 as an interactive default. This is not a memory limit — an integrated graphics chip borrows system memory — but a practical one. The effort levels it offers: low and medium at this window; xhigh is not offered here for two independent reasons, and both need naming because they have different fixes: the 65,536-token window cannot hold a 61,500–75,800-token thought (raise -c to fix that), and at 5–7 t/s an xhigh answer takes three to four hours (only patience fixes that). Expected speed: 5–7 t/s, derived, and it scales with your memory bandwidth. Room left for your desktop: ample — an integrated graphics chip borrows system memory — but leave the processor and operating system some working memory.
:: llama-server.exe, VULKAN build (SYCL has known perf bugs on current Intel GPUs; :: IPEX-LLM is archived). RAM is the spec that matters: the B390 tier (Core Ultra X9 :: 388H) ships dual-channel LPDDR5X-9600 = 153.6 GB/s, up to 96 GB shared. :: Intel's formal Arc-branding floor for B390/B370 is 7,467 MT/s — slower RAM relabels :: the iGPU as generic "Intel Graphics" — and other Core Ultra configs may ship slower :: RAM or ONE module (single channel = half the bandwidth = half the t/s). The ~5-7 t/s :: scales with YOUR bandwidth via section 05's formula; the model fits easily. :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M, 16.5 GB = 15.4 GiB) llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 :: interactive default, NOT a memory limit: an iGPU borrows :: system RAM, ~57% by default. To raise it: Intel Graphics :: Software >=25.26.1602.2 with driver >=32.0.101.6974 -> the :: GPU's General tab -> "Shared GPU Memory Override" slider :: (~87%; up to ~93% on B390/B370 with current Arc Pro :: drivers, RAM-dependent) -> reboot. Needs >=10 GB RAM, select :: Core Ultra systems. With 96 GB RAM even -c 262144 fits --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: verify KV-quant support in the Vulkan build --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: p-min 0 derived; start at 4, sweep :: (section 06); UNMEASURED on iGPUs — drop the spec flags if :: a paired probe shows no gain on your machine; on the 3090, :: p-min 0 read higher than 0.75 at n-max 4 on shallow reasoning (2026-10-08) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: xhigh NOT OFFERED, for TWO :: separate reasons with two separate fixes: 61.5-75.8k of appetite against this 65k :: WINDOW (fix: raise -c, you have the RAM), and 3-4 HOURS per answer at 5-7 t/s :: (fix: nothing but patience). Raise -c and effort together for overnight runs only. :: possibly-faster vendor alternative: Intel's OpenVINO Model Server on the iGPU with :: OpenVINO/Qwen3.8-27B-int4-ov (OVMS LLM quickstart; OpenAI-compatible /v3 endpoint) — :: Intel tunes that stack for its own silicon; benchmark both if the t/s matters
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}"Standardized industry metrics
What each recipe costs in electricity, in the units the industry compares on. Read the two plain-language conclusions first: turning the drafter on is the single cheapest energy decision on this page — it cut the energy per token by 2.52× at n10/p0.5 and 2.17× at n4/p0.75 (energy at the n3/p0 and n4/p0 the cards carry now is not measured) — and a slower setting is a more expensive setting in exactly the same proportion, because while the model is generating, this card draws essentially the same power whatever it is doing, so the joules per token are just those watts divided by the tokens per second. The whole energy win is throughput. Nothing here saves power; it saves time at the same power, which is the same thing arriving by a less interesting route.
Instrumentation tier for every figure below: in-band GPU board power — read from the card's own sensor rather than from a meter at the wall (NVML, nvidia-smi --query-gpu=power.draw). Inside the number: the graphics chip, its memory, the board's voltage regulators and fans. Excluded and never measured on this machine: power-supply conversion loss, the processor, system memory, drives, chassis fans, the display, and any datacentre overhead. Do not call any of these figures system power or wall power, and do not divide an electricity bill by them.
The idle pair, dated. Settled, with little on screen, first 60 seconds of each idle window discarded (2026-08-23 power matrix): 29.9 W with no server running, 34.1 W with the model resident and answering nothing. A resident model cost +4.2 W in that pair — small, and in the physically sensible direction. The 34.1 W figure is what the idle-subtracted columns below remove.C5
Sustained board power is a range, not a constant. Across 2.5 hours of drafter-off benchmark work the board drifted 305.5 → 341.1 W (+11.7%) at constant throughput, constant 77–79 °C and constant memory clock, tracking its own SM clock from 1,453 to 1,606 MHz. Decode at one request at a time is limited by memory bandwidth, so the extra clock bought no tokens and cost about 6% more energy per token. Plan with the band — 306–341 W drafter-off, 338–345 W on the long drafter-on authoring runs, 302–339 W across the short power-matrix arms — or with your own run's mean.C6
| Recipe | Tier | Mean W whole window / during decode | J / decode token, gross | J / decode token, net of 34.1 W idle | J / prompt token (prefill) | tokens / kWh | EDP (J·s) | Wh per 700-token answer, gross | Wh, net |
|---|---|---|---|---|---|---|---|---|---|
| [1] UD-IQ4_XS + vision · n4/p0.75 energy arm B2; the recipe runs n3/p0 since 2026-10-08, its energy not measured | in-band NVML | 308.4 / 341.7 meas | 3.743 meas | 3.369 der | 0.144 meas | 961,846 der | 20,091 meas | 0.786 meas | 0.699 der |
| [2] UD-IQ4_XS text-only · n10/p0.5 energy arm B3; the recipe runs n3/p0 since 2026-10-08, its energy not measured | in-band NVML | 302.4 / 341.0 meas | 3.210 meas | 2.889 der | 0.146 meas | 1,121,380 der | 14,811 meas | 0.683 meas | 0.606 der |
| [3] UD-IQ4_XS text-only · no drafter energy arm B1 | in-band NVML | 325.2 / 344.6 meas | 8.104 meas | 7.302 der | 0.099 meas | 444,218 der | 93,380 meas | 1.616 meas | 1.447 der |
| [4] Q4_K_M + vision · n4/p0.75 energy arm C2; the recipe runs n4/p0 since 2026-10-09, its energy not measured whole row measured drafter-OFF — not this recipe's flags | in-band NVML | 326.9 / 345.8 meas | 8.853 meas | 7.980 der | 0.102 meas | 406,644 der | 111,069 meas | 1.763 meas | 1.579 der |
| Every recipe on a card that is not this 3090 | Not measured. No power instrumentation exists on any other machine in this campaign, and board power does not scale with bandwidth the way decode speed does. Do not derive it | ||||||||
| E_comm (interconnect energy) | N/A — single GPU, no interconnect. Every recipe here serves from one card, so no energy is spent moving activations between accelerators. On a multi-GPU machine this row would have to be measured or split out | ||||||||
Conditions, and the ways each row differs from the recipe it prices. All four energy arms ran on 2026-08-23 at -c 32768 with a 1,458-token prompt and 700 generated tokens, thinking off, no projector, three requests per arm, 500 ms power sampling, 100% coverage. Row [1] and [2]: the recipes serve at 122,880 and 180,224 tokens of window with the projector loaded. Window allocation does not change joules per decode token; prompt depth does, and the depth series is its own set of arms with its own shallow reference point — 3.919 J/token at a 1.5k fill, 4.453 at 28k, 5.598 at 91k. Read that series against its own 3.919, not against this row's 3.743: the two shallow points sit 4.7% apart in energy and 6.4% apart in throughput, both outside the 2.9% floor, because they are different arms on different prompts. The projector is decode-free (0.04–0.09% at a 91k fill), so it changes nothing here. Rows [1] and [2] both decoded faster than the matched sweep did at the same flags — 91.3 against 83.5 for row [1] (9.3% apart) and 106.2 against 93.9 for row [2] (13% apart). This gap is real, still open, and larger than any single session can see. Three twenty-probe runs of the identical configuration (UD-IQ4_XS, n10/p0.5, -c 32768, -ctk/-ctv q8_0, greedy, 700 tokens, one server per run, one warmup discarded, RTX 3090 at 350 W, build 10502, all 2026-08-28) gave session means of 88.76, 77.35 and 77.15 t/s. Runs 1 and 2 do not overlap at all (ranges 86.22–94.00 and 73.44–84.21); their means differ by 14.8%. The spread within any one session — standard deviation (typical spread from the mean) 2.70–3.22 — is much smaller than the gap between sessions. The band is set between sessions, not within one; twenty probes in one sitting cannot measure it. Within a session, throughput correlates (tracks together) with mean draft length — r = +0.683, +0.597, +0.790, +0.861, +0.894, +0.240 across six sessions, every one positive. Between sessions, draft length cannot be the mechanism: six further sessions (six probes each, greedy, same configuration, one server per session torn down between, all 2026-08-28) produced bit-identical acceptance and mean-draft-length sequences probe for probe (acceptance 0.6128, 0.6032, 0.6185, 0.6036, 0.6398, 0.6459; mean draft length 7.792, 7.683, 7.603, 7.754, 8.080, 8.117) while session means spanned 75.71–78.65 t/s. Across all nine sessions measured on 2026-08-28 the means span 75.71 to 88.76 t/s, a 16.6% range, while the drafter does precisely the same work in each. What is ruled out for the between-session gap: the workload (bit-identical), the drafter (bit-identical), the llama.cpp build (every compute library dated 2026-08-19, unchanged), the core clock (run 2 held a mean SM clock 49 MHz higher than run 1 while delivering 12.9% less throughput), and heat (75–83 °C in every session). The between-session cause is unmeasured.C26 Any single speculative-decoding throughput figure on this page is a point drawn from a band at least this wide; plan accordingly.C25 Use these rows for energy ratios, which survive because every arm in the matrix shares the same spread. Row [4]: this arm ran with the drafter off while the recipe ships n4/p0 since 2026-10-09 (n4/p0.75 before). On UD-IQ4_XS n4/p0.75 cut joules per token by 2.17× (8.104 → 3.743), so recipe [4]'s own figure at that drafter is expected near 4.1 and is unmeasured; no file was measured for energy at n4/p0 — the printed 8.853 is this file's no-drafter cost and is the right number for comparing files, not for pricing this recipe. "Wh per 700-token answer" means exactly that: every arm generated 700 tokens and stopped, so these are not complete-answer figures. Complete-answer energy, for one real authoring task, is in §09. Net columns subtract the dated 34.1 W loaded idle over the covered seconds and are derived; the campaign's own runner subtracted 31.0 W instead, a uniform shift of about 1% that changes no ranking.
Decode energy for the anchor and both Swift files at n4/p0.75, each in the answer regime (thinking off) and the reasoning regime (thinking on), measured 2026-10-04. Loaded idle for this sweep: 12.97 W (anchor, n4 drafter, -c 32768, first 60 s dropped, 240 s window); no-server idle: 13.17 W. One sweep, 2026-10-04; not comparable with the August rows above.
| Recipe | Tier | Mean W whole window | J / decode token, gross | J / decode token, net of 12.97 W idle | J / prompt token (prefill) | tokens / kWh | EDP (J·s) | Wh per 700-token answer, gross | Wh, net |
|---|---|---|---|---|---|---|---|---|---|
| Anchor (UD-IQ4_XS) · n4/p0.75, thinking off | in-band NVML | 300.7 meas | 3.512 meas n=2 | 3.364 der | not measured §15 | 1,025,431 der | 19,673 meas | 0.692 meas | 0.663 der |
| Anchor (UD-IQ4_XS) · n4/p0.75, thinking on | in-band NVML | 292.9 meas | 3.925 meas n=2 | 3.755 der | not measured §15 | 917,369 der | 25,164 meas | 0.773 meas | 0.739 der |
| Swift IQ4_XS · n4/p0.75, thinking off | in-band NVML | 289.7 meas | 3.136 meas n=2 | 2.999 der | not measured §15 | 1,147,875 der | 16,269 meas | 0.619 meas | 0.591 der |
| Swift IQ4_XS · n4/p0.75, thinking on | in-band NVML | 304.0 meas | 5.217 meas n=2 | 4.998 der | not measured §15 | 690,037 der | 44,160 meas | 1.024 meas | 0.980 der |
| Swift IQ3_XXS · n4/p0.75, thinking off | in-band NVML | 293.3 meas | 3.518 meas n=2 | 3.366 der | not measured §15 | 1,023,735 der | 20,314 meas | 0.694 meas | 0.663 der |
| Swift IQ3_XXS · n4/p0.75, thinking on | in-band NVML | 308.6 meas | 6.176 meas n=2 | 5.920 der | not measured §15 | 583,019 der | 59,708 meas | 1.210 meas | 1.159 der |
| E_comm (interconnect energy) | N/A — single GPU, no interconnect. Every recipe here serves from one card, so no energy is spent moving activations between accelerators. On a multi-GPU machine this row would have to be measured or split out | ||||||||
Conditions, and the ways each row differs from the recipe it prices. All six arms ran on 2026-10-04 at -c 32768 with one prompt (a JavaScript red-black tree) sent three times per load; the first send is discarded, and pass 2 repeated pass 1’s text. Each request generated 700 tokens; each pass’s two kept requests took 15.3–28.4 s. The thinking-on arms set no effort level; the template’s default renders the same prompt as xhigh, the cards’ setting. Power sampling at 500 ms, coverage 1.0, means of 2 passes n=2. The board sat at its software power cap in 97–100% of active samples in every window, so J/token restates throughput. J per prompt token is not measured: the long prefills sit in the discarded rows (§15). In-band GPU board power (NVML); PSU, wall and PUE excluded. Net columns subtract the 12.97 W loaded idle over the covered seconds.
Both Swift files cost more energy per decode token than the anchor in the reasoning regime (n4/p0.75, one prompt, n=2), in step with fewer accepted draft tokens per verify pass while thinking — 1.479 and 0.974 against the anchor’s 2.684 (acceptance 0.837 and 0.772 against 0.901).
Energy per setting: the J/token table
Every knob this page recommends on, priced in energy, so that a setting chosen for speed can be checked for electricity too. The pattern underneath all of it is simple and was verified with no residual left over: joules per decode token equals the watts drawn while decoding divided by the decode tokens per second. That qualifier matters — the mean-watt column in the August table above spans each request's whole window, including a short low-power prefill, so it runs 6–11% below the decode figure and dividing it gives the wrong answer. Anything that raises throughput at constant power lowers energy per token in the same proportion. That is why the drafter is an energy feature and why depth is an energy problem.
| Axis | Arms compared (in-band NVML; the August rows 2026-08-23 unless a row says otherwise) | J / decode token | What it says |
|---|---|---|---|
| Drafter | B1 --spec-type none · B2 n4/p0.75 · B3 n10/p0.5, same server, same prompt | 8.104 → 3.743 → 3.210 meas | 2.52× less energy per token at n10/p0.5, 2.50× the throughput, 6.3× better EDP. Power while decoding is unchanged — 344.6 → 341.7 → 341.0 W — so the energy saving is the throughput gain, exactly and with nothing left over. The whole-window means fall further (325.2 → 308.4 → 302.4 W), but that is the low-power prefill segment taking a larger share of a shorter run, not the card using less power |
| Quant | C1 UD-IQ4_XS · C2 Q4_K_M · C3 NVFP4-HIGH, drafter off | 8.198 · 8.853 · 9.293 meas | Q4_K_M costs +8.0% and NVFP4-HIGH +13.4% against UD-IQ4_XS — both outside the 2.9% noise floor, so both are real. Choosing a file is choosing an energy bill. (The campaign's summary line prints +7.8% and +13.1%. Those are not reproducible from either the gross or the idle-subtracted columns of the arm table, both of which give +8.0% and +13.4%; the ratios above are re-derived from the gross cells printed here) |
| KV cache precision | D1 -ctk/-ctv f16 · D2 q8_0 | 8.284 · 8.341 meas | 0.7% apart — a clean null, well inside the 2.9% floor. Half the KV memory for no measurable energy, which independently supports the q8_0 default that every recipe ships |
| Token regime | E1 thinking on · E2 thinking off, one server, one prompt | 6.066 · 3.744 meas | 1.62×, against the 1.69× throughput ratio the mean-draft-length measurement produced on a different instrument entirely (§06). Two independent measurements of the same mechanism |
| Depth | F1 1.5k · F2 28k · F3 91k of prompt fill, same 700-token answer | 3.919 · 4.453 · 5.598 meas | Decode energy rises 43% across the span — but the real story is prefill: at a 91k fill 90.7% of the arm's joules are prefill, and the same 700-token answer costs 0.83 Wh at 1.5k against 11.71 Wh at 91k, a factor of 14. Caveat: F3 decoded 10.4% slower than the cooled reference ladder at the same depth, so its level is suspect while its shape is not |
| Effort level | One authoring task, drafter on, temp 1.0, n=1 per level — low · medium · xhigh truncated · xhigh completed | 4.26 · 5.18 · 6.60 · 6.13 meas | Rising energy per token with effort — but this is a throughput effect, not an effort effect. On 25-question arms with the drafter off, where all three levels decode at the same speed, energy per token is flat to 0.3% (7.921 at xhigh against 7.947 at medium). Effort changes joules per answer, not joules per token (§09) |
--parallel 1 vs 2 | G1 · G2, drafter off, UD-IQ4_XS, measured 2026-08-23 — and the matched drafter-on pair that replaced them, 2026-08-25 | 8.59 → 5.19 superseded | Ship --parallel 1. Measured 2026-08-25, UD-IQ4_XS, drafter on at n10/p0.5, three reps each with the first post-prefill probe discarded: 82.98 → 101.25 t/s aggregate (+22.0%) and 85.79 → 55.41 t/s per slot (−35.4%), acceptance flat at 0.618 → 0.620. One user gets their tokens about 55% faster with a single slot (85.79 against 55.41 t/s per slot). Two slots are worth taking only when you are waiting on parallel sub-agents rather than reading the output. The drafter is already saving most of the repeated reading of the model that batching would otherwise save, leaving batching far less to win. The drafter-off arm in this matrix measured +60.3% aggregate, −39.6% J, −62% EDP; an earlier drafter-on measurement gave ~+11%; both are superseded by the matched pair. The energy column (8.59 → 5.19) was measured with the drafter off; the matched pair measured throughput, not board power, so the 5.19 figure is retired rather than corrected and no drafter-on concurrency energy figure exists (§06) |
Power cap (nvidia-smi -pl) | 350 · 300 · 250 W, one server load each | 4.033 · 3.771 · 3.479 superseded | The one setting on this page that lowers the board’s draw at all. At saturating load (SwPowerCap residency 99–100%, measured 2026-08-28): −4.1% t/s, −13.1% W, −9.4% J at 300 W and −14.6%, −27.1%, −14.7% at 250 W — capping to 300 W is a better deal than the earlier non-saturating arm showed.C22 That earlier arm (measured 2026-08-25C7) never reached its own cap: capping to 300 W cost 5.0 % of throughput and saved 11.2 % of power; 250 W cost 12.8 % and saved 24.8 %; energy per token improved 6.5 % and 13.7 %. Every other lever here saves energy only by finishing sooner at the same wattage; this one lowers the wattage. The mechanism is visible in the clock: throughput falls less than the SM clock does (−5.0 against −5.8 %, then −12.8 against −22.1) because decode is partly memory-bandwidth-bound and memory clock is untouched. And the stock arm never reaches its own cap — 305.4 W mean against a 350 W limit — which is why the first 50 W costs so little: it removes headroom the workload was not using. Decode only; prefill is compute-bound and would plausibly lose more (§11) |
One sweep, 2026-10-04: the Swift files beside the anchor; not comparable with the August rows above. The drafter and quant rows are one prompt at -c 32768; the depth row is the speed probes at -c 131072. | |||
| Drafter (Swift IQ4_XS, thinking off) | --spec-type none · n4/p0.75 · n10/p0.5, 2026-10-04, one prompt | 7.592 → 3.136 → 3.005 meas n=2 | The drafter cuts energy per token in the answer regime, as on the anchor. The board sat at its software power cap in 97–100% of active samples, so J/token restates throughput |
| Drafter (Swift IQ4_XS, thinking on) | n4/p0.75 · n10/p0.5, 2026-10-04, one prompt | 5.217 → 6.448 meas n=2 | n10/p0.5 costs more J/token than n4/p0.75 while thinking — the reverse of the answer-regime pattern. Acceptance: 0.837 at n4, 0.404 at n10; accepted draft tokens per verify pass 1.479 and 1.657. The anchor in the same sweep gains at n10/p0.5 while thinking (3.925 → 3.604). Power cap active in 97–100% of active samples, so J/token restates throughput |
| Drafter (Swift IQ3_XXS, thinking off) | --spec-type none · n4/p0.75 · n10/p0.5, 2026-10-04, one prompt | 7.265 → 3.518 → 2.929 meas n=2 | The drafter cuts energy per token in the answer regime, as on the anchor. Power cap active in 97–100% of active samples, so J/token restates throughput |
| Drafter (Swift IQ3_XXS, thinking on) | n4/p0.75 · n10/p0.5, 2026-10-04, one prompt | 6.176 → 7.022 meas n=2 | n10/p0.5 costs more J/token than n4/p0.75 while thinking, as on Swift IQ4_XS. Acceptance: 0.772 at n4, 0.424 at n10; accepted draft tokens per verify pass 0.974 and 1.389. Power cap active in 97–100% of active samples, so J/token restates throughput |
| Quant (drafter off) | Anchor · Swift IQ4_XS · Swift IQ3_XXS, --spec-type none, thinking off, 2026-10-04, one prompt | 7.702 · 7.592 · 7.265 meas n=2 | Swift IQ3_XXS, the smallest file, costs the least energy per token with the drafter off; Swift IQ4_XS, the largest, reads 7.592 against the anchor’s 7.702, so file size alone does not order the three. Power cap active in 97–100% of active samples, so J/token restates throughput |
| Depth (answer regime, n4/p0.75) | 1.5k · 28k · 91k of prompt fill, a 400-token answer to the same coding task, -c 131072, 2026-10-04, two passes | anchor 3.825 · 4.136 · 5.001; Swift IQ4_XS 3.391 · 3.605 · 4.955; Swift IQ3_XXS 3.729 · 4.036 · 5.562 meas n=2 | Decode energy rises with depth on all three files. These answer-regime probes’ first pass ran before the host’s download overlap and the second began in its last two minutes; the board sat at its power cap, so J/token restates throughput. The reasoning-regime energy from the same probes is not printed: that group is void |
KV cache precision, --parallel, effort, power cap | Swift: not measured, 2026-10-04 | — | Not measured on the Swift files (§15). The August rows above cover these axes on the anchor. Prefill energy at depth on the Swift files is not measured either (§15) |
Five things about how this model is built explain almost every recommendation in §03. It stores far less memory per token of context than its size suggests, so long contexts are cheap. Its context limit is 262,144 tokens, and what you will actually run out of is card memory, not model capability. It thinks before answering by default, and that thinking comes out of the same pool as your prompt. Vision is optional at start-up and costs memory rather than speed. And it wants a specific sampling setup that some clients quietly override, which causes the model to repeat itself. Everything below is those five facts with their numbers.
One more fact governs speed rather than settings, and it is worth having in plain words before the arithmetic: with speculation off, a bigger file is a slower file. To write each token, the card has to read the whole model out of its memory once. A file that is 2 GiB larger takes proportionally longer to read, every single token, forever.
The drafter breaks that rule: with speculation on, a smaller file is not necessarily a faster file. meas Turn speculation on and the ordering can invert: UD-IQ4_XS goes 42.34 → 86.91 t/s (2.05×) while the smaller UD-Q2_K_XL goes 45.66 → 77.01 (1.69×) — so the file that is faster drafter-off is slower drafter-on, by 12.9%. The draft head is quantized along with the model, so a more damaged file drafts worse and gets less back from speculation. The window moves the ordering too (§08). Since every recipe here ships with the drafter on, “smallest that holds its quality” is the right rule for memory and the wrong one for speed.
- Hybrid attention. Most of its 64 layers use Gated DeltaNet, a linear-attention design whose memory use stays constant no matter how long the context gets; only 16 of the 64 are standard attention layers (and those use grouped-query attention with 4 KV heads). The result: the KV cache costs just 64 KiB per token at fp16 — 4.0 GiB at 122k context with an 8-bit cache, 7.5 GiB unquantized — where a dense 27B storing KV in every layer would need exactly 4× as much. Those are the arithmetic, and arithmetic is a floor: what a real server allocates measured 15–29% steeper (§05). Huge contexts are cheap in memory here, and speed falls off gently as they fill rather than collapsing.
- Native 262,144 context. No rope or YaRN tricks are needed below that; the limit you will actually hit is VRAM. The model card lists extension to 1M tokens via YaRN, which is out of scope here: the KV cache alone would run about 34 GiB even at
q8_0, by the same arithmetic. - Thinking is always on, and defaults to
xhigh. Reasoning tokens share the context window with your prompt and your output, so a 100k-token thinking run needs a window bigger than 100k (§09). This is also why every speed number on this page states its token regime. - Vision is optional at serve time — and it costs context, not speed. The BF16
mmprojprojector occupies 1,138 MiB of VRAM once loaded — not the 0.867 GiB its file weighs. Two independent measurements (2026-08-23, same machine, UD-IQ4_XS) of the resident cost agreed to the megabyte, and a follow-up pair at 91k of depth reproduced the same 1,138 twice more. At the measured window slope that is about 26,500 tokens of 8-bit window with the coding drafter loaded, 29,900 without it (§05). What it does not cost is speed: at 90,862 tokens of fill, decode with and without the projector differed by 0.04–0.09% (§06) — the only question the projector asks is about window. Serving text-only converts vision you are not using directly into context, but not all the way to the native maximum: with the drafter on, UD-IQ4_XS's measured text-only ceiling is 180,224, and 262,144 fits only with the drafter off, and then only just (§03 prints both variants). - Official thinking-mode sampling: temperature 1.0 · top-p 0.95 · top-k 20 · min-p 0. Clients that silently send temperature 0 — aider does — cause repetition loops. Override them; §14 shows how, per agent.
config.json; what a server really allocates is measured in §05 and runs 15–29% higher.The arithmetic that predicts decode speed
Generation speed comes down to one ratio: how fast your card can read its memory, divided by how much it must read per token. Every token requires reading the entire model — about 16.5 GB for a 4-bit K-quant, 14.25 GB for UD-IQ4_XS. So take your card's bandwidth in GB/s, divide by your file's GB, and multiply by an efficiency constant, and that is your speed ceiling in tokens per second. That constant is format-specific: on CUDA it measures about 0.70 for K-quants and about 0.65 for IQ-quants (not a single ~0.7 for both). You can tell which family a file belongs to from its name: a name containing IQ — UD-IQ4_XS, UD-IQ3_XXS, UD-IQ2_S — is an IQ-quant, so use 0.65; a name containing K without IQ — Q4_K_M, UD-Q3_K_XL, UD-Q2_K_XL, Q6_K — is a K-quant, so use 0.70. The independent re-measurement measured 0.649 on UD-IQ4_XS — 936 GB/s ÷ 14.25 GB = 65.7 t/s theoretical against 42.6 measured — and Q4_K_M lands at 0.700: 936 ÷ 16.5 = 56.7 against 39.7 measured. Those are one file of each family: on this card with the drafter off, bartowski’s Swift IQ4_XS and IQ3_XXS read 0.71 and 0.58 of the same arithmetic der (42.9 and 44.1 t/s, 2026-10-04, beside the anchor’s 42.7 in the same sweep, which is 0.65 of its own arithmetic; §06), so the constant belongs to the file as much as to its family — §08’s 2.9-bit K-quant, UD-Q2_K_XL, reads 0.48 against the K-quant 0.70 (45.66 t/s at 9.83 GB; C39). This one calculation predicts every row of the matrix in §07, and it doubles as a fault detector: land far below your format's constant at short context and something is wrong (§11).
A 24 GB card has two ceilings, not one, and the difference between them is the most expensive misunderstanding on this page. The first ceiling is how large a window you can set and still have every byte of it live on the card — fully resident, in the words of §02: about 131,000 tokens for the reference configuration. The second is how large a window you can set before the server refuses to start at all: about 213,000. Between those two numbers the server starts happily, answers a short prompt at full speed, and then collapses to a fifth of that speed the first time somebody actually fills the window. Set your window under the first ceiling, not the second, and never trust a ceiling that was checked with a short prompt. The rest of this section is the memory arithmetic that tells you where your own first ceiling is.
The second law of this section: speed against context is a cliff, not a curve. It is flat while the weights and the cache fit in VRAM, and it collapses once anything spills across the PCIe bus — with the magnitude set by how much spilled: about 10× when the weights file itself lands in system memory; 1.1–2× from a 1.9 GB spill of weights answering a short prompt (§11); and 2.5–3.8× once the spilled pages are cache that a deep prompt actually reads (the table below). What decides the size is not how much spilled but how often the spilled pages are touched. Short probes on the reference 3090, showing where an unfilled window stops being fast — the table after the chart shows where a filled one collapses:C33
q8_0 KV, drafter on, projector loaded); the solid segments only join them and do not plot the ten readings between 122,880 and 212,992 (45.1–54.9 t/s). The solid run-in to the left of 122,880 is drawn without a probe because the 2026-08-22 sweep started at -c 122880; the dashed tail to 256k is an older run with no drafter where the 4-bit weights file itself spilled, so it is a different configuration and is drawn dashed for that reason. The 256k point is not a verdict on 256k — a smaller resident file gets closer to full native context on the same card. The dashed vertical at 131k marks where dedicated VRAM fills: everything to the right of it is fast only while the deep context stays untouched, which is exactly what a short probe never tests.Allocated is not the same as used: a big context window with a small prompt in it costs VRAM, not speed. And as long as you stay below the cliff, shrinking the window gains you nothing — there is no hidden optimal point to search for.
That corollary now has a measured face. Probed 2026-08-22 with Q4_K_M, a reconstruction (the record does not name the file; q8_0 KV, drafter on, projector loaded — the 1 GiB-larger UD-Q4_K_XL shifts every ceiling about 30k tokens lower, while dropping --mmproj frees 1,138 MiB and raises them about 26,500) by stepping -c upward with short temperature-0 wall-clock probes (prefill included, first request after each load, completion tokens divided by the stopwatch time of the whole request): every reading from -c 122880 through -c 212992 lands between 45.1 and 55.4 t/s, all above the sweep's 41.6 t/s pass line; 217,088 falls just under that line (41.2 t/s); at 221,184 the short probe reads 19.5 t/s.C33 But nvidia-smi tells the other half of the story: dedicated VRAM hits its ~24 GB ceiling already at -c 131072 (measured 24,052 MiB). Past that point the KV cache is overcommitted, and the run stays fast only while a short prompt leaves the deep pages untouched. So: about 131k is the practical fully-resident ceiling — fast even when a long run really fills it, though it sits right at the limit, which is why the recipe ships 122,880 — while up to about 213k will load and answer a short prompt. Past that ceiling, a window that is actually filled does not degrade gradually — it collapses, 30.6 t/s on -c 131072 → 8.0 t/s on -c 262144 at 90,886 tokens (UD-IQ4_XS, 2026-08-23), a 3.82× penalty — as the table below measures.
An overcommitted window does not slow gradually as real tokens fill the spilled pages — it collapses, and only once those pages are actually read, which is exactly what a short probe never does. Measured 2026-08-23 (same machine, UD-IQ4_XS, projector on, drafter at n-max 4 / p-min 0.75, identical prompts, only -c differing) with both configurations filled to depth rather than probed shallowly:
| Prompt actually filled | -c 131072 · fully resident | -c 262144 · projector loaded, ~3.5 GiB in system RAMa different configuration from the text-only 262,144 row in the budget table, which spills 2,364 MiB | Penalty |
|---|---|---|---|
| 1,531 tok | 53.5 t/s | 42.9 t/s | 1.25× |
| 29,837 tok | 47.3 t/s | 18.6 t/s | 2.54× |
| 90,886 tok | 30.6 t/s | 8.0 t/s | 3.82× |
Prefill collapses with it — 819 → 183 t/s at 90.9k, 1,085 → 408 at 29.8k — while draft acceptance is unchanged (0.907 against 0.849 at 90.9k), which rules out any drafting explanation: this is purely the KV cache being read across the PCIe bus. The practical reading: a reader who trusted the shallow ceiling gets 26% of the promised speed the first time they open a 90k-token document — 8 t/s where the fully resident -c 131072 window does 30.6 at the same 90,886-token depth, greedyC29 — and the one line the server log carries about it is the start-up no-fit warning (§11), easy to miss.C31 Never label a window resident or safe on the strength of a short probe; fill it once, deeply, before you believe it.
One caveat on the table's own numbers: every cell in it is a single decode probe fired immediately after its prefill, and that is the one moment this card is not at its boost clock — the ±25% band declared in §01 and measured in §11. Re-measured under the cooled protocol, the resident 90.9k row reads 36.6 t/s rather than 30.6, about 20% higher. The spilled column was taken the same way and carries the same uncertainty, so read the penalty ratios as approximate. Nothing about the verdict moves: a 2.5–3.8× collapse is an order of magnitude outside a ±25% band.
-cThe same re-measurement round measured the same -c 163840 configuration at 22,418 MiB dedicated with the n-max 4 drafter and 23,189 MiB with n-max 10 — a 771 MiB swing from a flag, on top of a 474 MiB swing in board VRAM from nothing but what was on screen. (That 771 and the 898 MiB in the budget table are the same quantity read two ways: 898 is the constant fitted across all 17 resident loads, 771 is what this one pair of loads happened to show. They differ by 127 MiB, which is exactly the budget model's worst residual — so the disagreement is the model's stated error, not a second finding. Budget with 898.) A window is resident or not as a property of the whole configuration: file + drafter flags + projector + desktop. Quote all four whenever you quote a ceiling — reports of "same command, different ceiling" are almost always one of those four moving.
The VRAM budget table
The chart above says where a short probe stops being fast; the collapse table says where a filled window fails; this budget table says why.C33 C8 Budget from llama-server's own dedicated-VRAM counter, not from arithmetic: a fixed base plus a per-window-token slope, fitted across 17 fully-resident server loads (2026-08-23, same machine, UD-IQ4_XS), reproducing all 17 to within 127 MiB. (Those 17 are the fully-resident subset of 26 loads in total; the other nine were overcommitted configurations, which the model does not describe. Later measurement rounds loaded further configurations that have never been checked against it.)
| Line item | Measured | What it means |
|---|---|---|
| Per window token, drafter off | 39,936 B meas | +15% over the 34,816 arithmetic — alignment and padding that the config.json maths cannot see |
| Per window token, drafter on | 45,056 B meas | +29% over arithmetic. The 5,120 B/token difference is the draft path: speculation costs 11.4% of your window |
| Weights + buffers (UD-IQ4_XS, text-only, no drafter) | 13,232 MiB meas | the fixed base. Q4_K_M runs about 2.1 GiB above it, UD-Q4_K_XL about 3.1 |
Vision projector (BF16 mmproj) | 1,138 MiB meas | ≈26,500 tokens of window with the drafter on, 29,900 without. The file is 0.867 GiB — quoting the file size under-books it by about 0.25 GiB |
--image-max-tokens 1024 — shrinking a screenshot | saves ~2,600 tokens of window meas | Measured 2026-08-25, and it is not the free win it looks like. On a 1440p dashboard it cuts the image from 3,602 tokens to 1,010 — 3.57× (3,635 to 1,043 with the question).C32 The model still reads the layout, the headline and both 4-digit values. It loses the smallest type and does not know it: a latency of 207 came back as 287, and a serial QUB-85731-3D3B came back as QLN-15731-3053. It does not refuse and it does not hedge — a confident misread, which is worse than a blank, because nothing downstream can see it. A control run with the image withheld scored 0/7, so the readings above are perception and not guesswork. Use it for layout questions, never for reading numbers off a screenshot. n=1 per question, 7 questions, greedy |
--spec-type draft-mtp — turning the drafter on | 1,008 MiB fixed meas | 1,008 MiB fixed, plus 5,120 B per window token (the per-token line above) — 11.4% of your window. The head's own weights are 334.7 MiB; the rest is draft-path buffers. --spec-type none loads none of it — the server logs unused tensor blk.64.* |
--spec-draft-n-max 10 instead of 4 | +898 MiB meas | a flag moves your ceiling by about 0.9 GiB — roughly 20,900 tokens of window — without -c changing at all |
| The desktop's own share of board VRAM | 1,179–1,669 MiB meas | measured in direct no-server readings on one machine, and it depends on what is on screen. Not yours to spend (§02, §11) |
Conditions: RTX 3090, 24,576 MiB, UD-IQ4_XS, q8_0 KV, -ngl 99, --parallel 1, llama.cpp build 10502. Worked example derived — §03's the default recipe with vision at -c 122880, if you put the n-max 10 coding drafter on it: 13,232 + 1,138 + 1,008 + 898 + (122,880 × 45,056 ÷ 1,048,576 = 5,280) = 21,556 MiB of server-process VRAM, leaving 3,020 MiB (2.95 GiB) of board VRAM for the desktop and slack. Drop the 898 and the same row keeps 3,918 MiB (3.83 GiB) — which is one reason §03’s vision recipe never carried n-max 10: the wider drafter setting fits here, it just spends the slack that absorbs a desktop measured anywhere from 1,179 to 1,669 MiB. Since 2026-10-08 the recipe ships n-max 3 / p-min 0, lighter than the n-max 4 the 1,008 above was measured at (§06).
Two rules follow, and both change how you read every ceiling in this guide. First, budget on the server's own report, not on the board VRAM total — the three counters and what each can see are in §02. The nvidia-smi board VRAM figure includes your window manager and your browsers; the campaign watched one identical configuration read 474 MiB apart on the board VRAM — about 11,000 tokens of window — with nothing changing but what was on screen, while llama-server's own number moved by at most 127 MiB. That spread is the whole explanation for "same command, different ceiling". Second, the drafter is a line item, not a freebie. At -c 163840 with vision, its fixed and per-token costs together book roughly 1.8 GiB — more than the vision projector. On a pure prose workload, where speculation was measured worth only about 1.16× (§06), that is the easiest 1.8 GiB on the card to reclaim: --spec-type none.
With those constants in hand, what fits on a 24 GiB card by file and window — the KV column stated as the q8_0 arithmetic floor, so add 15% (drafter off) or 29% (drafter on) for what a server really allocates:
| Context | KV type | KV size (arithmetic floor) | Largest unsloth file that fits |
|---|---|---|---|
| 262,144 | fp16 | 16.0 GiB der | nothing — even the smallest file (UD-IQ1_S, 5.8 GiB) misses once buffers are counted. That file is now measured, and it is not a fallback: at 1.83 bits per weight it is +35% perplexity against this page's default and it is the worst of four rungs that fail a functional detector (§02) — 2, 3, 5 and 28 empty replies out of 75 at 2.481, 2.153, 1.994 and 1.835 bits per weight (§08) — and the only one that also stops terminating, at a median 932 tokens against 424. Four rungs fail a functional detector, not one.C16 Its code does not run either, and neither does that of the two rungs above it (§08) |
| 262,144 | q8_0 | 8.5 GiB der | UD-IQ4_XS (13.3 GiB) — text-only, no projector, and only with --spec-type none: at the measured drafter-off slope the window costs 9,984 MiB (9.75 GiB), not 8.5, for 23,216 MiB total — leaving 1,360 MiB of board VRAM, which is 309 MiB short once a desktop takes its measured worst case of 1,669 MiB.C36 Marginal, and only with no graphical session using the card (§02). With the drafter at n-max 4 / p-min 0.75 on it does not fit at all (the n-max 10 coding drafter costs 898 MiB more): 25,504 MiB (24.9 GiB) required against 24,576 available, and 2,364 MiB measured living in system RAM at load with that drafter (2026-08-23).C29 C9 Applying the same measured slope, UD-Q3_K_XL (12.2 GiB) lands near 22.2 GiB — about 1 GiB more comfortable, though unmeasured — and UD-IQ3_XXS (10.15 GiB) fits with real room to spare |
| 262,144 | q4_0 | 4.5 GiB der | UD-Q4_K_M would fit — and a q4_0 cache costs +0.693% perplexity against fp16 (§08). Small, but nothing here argues for spending it on a 24 GB card: q8_0 already reaches every ceiling this page publishes. On a 16 GB card it is the one place the trade earns its keep |
| 131,072 | q8_0 | 4.25 GiB der | Q4_K_M (15.4 GiB) — 19.7 GiB by this table's own two columns, and that is the floor: apply the measured constants and the same file at -c 122880 with the projector and the n-max 4 drafter books about 22.3 GiB of server VRAM, which is why §03's the Q4_K_M recipe measures ~23.7 GiB of board VRAM with a desktop up. This is the reference recipe's row; budget it from the constants, not from this column. The alternative UD-Q4_K_XL (17.6 GB = 16.39 GiB) also fits when VRAM is spare |
| 65,536 | q8_0 | 2.125 GiB der | UD-Q5_K_S (~17.4 GiB) — the quality play when you do not need long thinking runs |
Two footnotes the table earns. First, the q4_0 KV flag. A q4_0 K/V cache costs +0.693% perplexity against fp16 — 6.6413 against 6.5956 (2026-08-23, same corpus and flags as the q8_0 run) — where q8_0 costs +0.309%. The degradation is super-linear in bits: the second halving of the cache costs slightly more than the first (+0.383% on top of q8_0), which is the shape the mechanism predicts and the reason the cut stops paying. 0.69% is small — smaller than the perplexity gap between the two 4-bit files this page recommends (0.93%). The mechanism concern — that at long context every new token attends over hundreds of thousands of perturbed entries, and retrieval gets unreliable at exactly the 200k-plus lengths a q4_0 cache would enable — has been tested on UD-Q2_K_XL, where q4_0 scored identically to f16 and q8_0 at depths up to 119,435 tokens (§15); the same test on the 4-bit files at the 200k-plus lengths the mechanism targets has not been run. q8_0 remains what every recipe ships because it already buys the context ceilings here; q4_0 is a window-extension trade to take knowingly (§08 carries the row and its caveat). Second, prefill: these budgets make a 200k-plus prompt fit, and the table below measures what filling one costs, across a 60× span of prompt lengths on a fully-resident -c 131072 server:
| Prompt filled | Prefill | Wall to first token | Prefill energy at this depth |
|---|---|---|---|
| 1,533 tok | 1,082 t/s meas | 1.4 s | 0.164 J/prompt-token meas |
| 10,658 tok | 1,198 t/s meas | 8.9 s | — |
| 28,423 tok | 1,097 t/s meas | 25.9 s | 0.306 J/prompt-token meas |
| 56,840 tok | 951 t/s meas | 59.8 s | — |
| 92,679 tok | 816 t/s meas | 1.9 min | 0.421 J/prompt-token meas |
Conditions: UD-IQ4_XS, -c 131072 fully resident, q8_0 KV, drafter at n-max 4 / p-min 0.75, greedy decoding, no projector, prompts prefixed with a unique identifier so the prefix cache (§02) cannot be reused (2026-08-23). The energy column comes from a separate 2026-08-23 arm at 1.5k / 28k / 91k of fill, not from these five prompts, and its depths are close but not identical. Use about 1,000 t/s as the planning figure: prefill decays only 25% over a 60× longer prompt, because the quadratic term lives on just 16 of the 64 layers (§04) — the hybrid dilutes it. A genuinely full 100k prompt is about 2 minutes; extrapolating the same slope, 200k lands nearer 5 minutes. Overcommit the window and this number collapses with everything else: 183 t/s at 91k in the correction above.
With a short prompt and a long answer, reading the prompt costs almost nothing: measured at 0.12–0.18 J per prompt token against 4.3–6.6 J per generated token, which is 0.06–0.38% of an answer's energy. That flips completely at depth. At a 91,000-token fill, 90.7% of the whole request's joules are prefill, and the same 700-token answer costs 0.83 Wh at 1.5k of depth against 11.71 Wh at 91k — fourteen times more for identical output. The energy per prompt token itself climbs 0.164 → 0.306 → 0.421 J across those three depths, which is attention's quadratic term showing up as electricity. Practical consequence for anyone running agents: reusing a cached prefix — starting a follow-up request with exactly the same opening text, so the server does not have to read it again (§02) — is not a speed optimization, it is an energy one, and it is the largest single saving available at depth. (All figures 2026-08-23, in-band board power; the 91k arm decoded 10.4% slower than the cooled reference ladder at the same depth, so treat its level as soft and its shape as solid.)
The Swift files use the same memory per token of context; only their weights differ
Swift IQ4_XS and Swift IQ3_XXS are bartowski’s quantizations of ukisai’s Swift-1.5, a fine-tune of this model with the same architecture (Appendix A). They share the same memory per token of context as the anchor (unsloth UD-IQ4_XS of the base model). For Swift IQ4_XS it is measured across its two text-only deep fills at 139,264 and 159,744 tokens (one sweep, 2026-10-04), where dedicated memory grows 44.0 KiB per token meas — the anchor’s 45,056 B in the budget table above — and the server footprint (dedicated + shared) grows 46.0 KiB per token meas, 0.3% der above the 45.87 KiB one-slot footprint slope of the per-card arithmetic and inside the ±2% tolerance set before the fills ran. For Swift IQ3_XXS the 45.87 KiB slope is itself an August-to-October transplant from the anchor onto knee pick 5 (Appendix A). Only the weights term changes. By file arithmetic (file size less the token embedding, which llama.cpp keeps in host RAM, and the header), Swift IQ4_XS (15,475,951,936 B meas) puts 1,043.3 MiB more weight on the card than the anchor der — 23,291 tokens of window at 45.87 KiB per token der — and Swift IQ3_XXS (12,320,168,256 B meas) puts 2,004.2 MiB less der. Swift IQ4_XS held decode, fully resident, at -c 159744, the largest window filled n=1 (one deep fill, text-only, n4/p0.75, one slot, 2026-10-04; the fill left 1,880 MiB for a desktop between the card’s 24,576 MiB and the server footprint at depth, dedicated + shared der, above this page’s 1,796 MiB threshold (§02); no larger window was tried). Swift IQ3_XXS held decode, fully resident, at -c 262144 in text at one slot n=1 (knee pick 5, 2026-10-04, 257,415 tokens filled at a 415 MiB desktop); that fill left 327 MiB for a desktop on the same basis der, so at this page’s 1,796 MiB threshold its text window is 216,064 der (Appendix A).
The deep fills on Swift IQ4_XS, one sweep, 2026-10-04; not comparable with the August rows above in this section:
| Configuration | Per-slot window | Fill (tokens) meas | Server footprint at load (MiB) meas | Server footprint at depth (MiB) meas | Board VRAM peak (MiB) meas desktop inside | Decode (t/s) meas | Desktop before (MiB) meas | Fully resident test |
|---|---|---|---|---|---|---|---|---|
| Swift IQ4_XS · text · n4/p0.75 · 1 slot | 139,264 | 127,453 | 21,666 | 21,776 | 22,172 | 47.42 | 855 | fit only lowest step n=1 |
| Swift IQ4_XS · vision · n4/p0.75 · 1 slot | 139,264 | 126,838 | 22,802 | 23,246 | 23,647 | 41.03 | 841 | pass n=1 |
| Swift IQ4_XS · text · n4/p0.75 · 1 slot | 159,744 | 156,261 | 22,586 | 22,696 | 23,057 | 44.57 | 846 | pass n=1 |
Conditions: 2026-10-04, NVIDIA GeForce RTX 3090, q8_0 KV, -ngl 99, --parallel 1, thinking off, greedy decoding, decode the mean of two 400-token probes after the fill; the vision row loads bartowski’s F16 projector and has a 1440p image in flight. Each row is one load n=1. Server footprint is dedicated + shared; board VRAM peak is memory.used with the desktop inside. Fully resident test: fit (dedicated memory at depth + the 855 MiB desktop maximum + the 121.0 MiB spread ≤ 24,576 MiB), spill and knee, as defined under the next table; the vision row is judged against its text twin at the same window. The 139,264 text row is tested for fit only, because no lower text step within 32,768 tokens was loaded.
Ceilings by kind, per Swift configuration, from two October runs that are not comparable with each other or with the August rows above in this section: the Swift IQ4_XS rows are the deep fills of 2026-10-04 (desktops 841–855 MiB), the Swift IQ3_XXS rows the knee sweep of 2026-10-03/04 (desktops 338–539 MiB), so decode is not compared between them. Each row names its scope — file, drafter, projector, slots — and the desktop the fit test assumed (73 desktop readings, 72 loads and one direct reading: max 855 MiB, sd 121.0 MiB, range 258–539 MiB from the knee sweep plus 841–855 from the deep fills plus a 264 MiB direct reading, 2026-10-04). The knee chart is in Appendix A.
| Configuration | Fully resident fit + spill + knee, loaded steps only | Decode-verified last held − 0.7 GiB, or the native 262,144 where decode held to it; the October launcher’s | Last window whose decode held (t/s) meas | Collapse point knee meas | Scope |
|---|---|---|---|---|---|
| Swift IQ4_XS · text · n4/p0.75 · 1 slot | 159,744 meas | not swept not in the October launcher | 159,744 (44.57) n=1 | none to 159,744 no larger window loaded | no projector; desktop 855 + 121.0 MiB; RTX 3090 |
| Swift IQ4_XS · vision · n4/p0.75 · 1 slot | 139,264 meas | not swept the October launcher ships 139,264 from the per-card table, not from a knee | 139,264 (41.03) n=1 | none to 139,264 no larger window loaded | projector loaded; desktop 855 + 121.0 MiB; RTX 3090 |
| Swift IQ3_XXS · text · n4/p0.75 · 1 slot | 262,144 meas | 262,144 der | 262,144 (36.2) n=1 | none to 262,144 | no projector; desktop 855 + 121.0 MiB; RTX 3090 |
| Swift IQ3_XXS · text · n4/p0.75 · 2 slots | 124,928 meas next step loaded: 133,120, fails | 123,904 der | 133,120 (29.2) n=1 | 135,168 (14.0) n=1 | no projector; desktop 855 + 121.0 MiB; RTX 3090 |
| Swift IQ3_XXS · vision · n4/p0.75 · 1 slot | none: the lowest step loaded, 253,952, fails fit and spill; lower windows untested | 262,144 der | 262,144 (30.8) n=1 | none to 262,144 | projector loaded; desktop 855 + 121.0 MiB; RTX 3090 |
| Swift IQ3_XXS · vision · n4/p0.75 · 2 slots | 16,384 meas fit only next step loaded: 114,688, fails | 123,904 der | 133,120 (23.1) n=1 | 135,168 (11.6) n=1 | projector loaded; desktop 855 + 121.0 MiB; RTX 3090 |
Fully resident: the highest loaded window at which every loaded step up to it passes three tests (dedicated memory at depth + the 855 MiB desktop maximum of the deep fills + the 121.0 MiB spread of 73 desktop readings (72 loads and one direct reading) ≤ 24,576 MiB; shared memory at load up by no more than 100 MiB over the next-lower step; decode within 15% of the next-lower step’s); a step with no loaded step of its configuration within 32,768 tokens below it is tested for fit only, unless it is a vision step with a text-only twin at the same window, which it is then judged against for spill and decode; steps between loaded steps were not tested. That allows 976 MiB der for the desktop, not this page’s 1,796 MiB threshold (§02). Judged against 1,796 MiB on the basis of the per-card tables in Appendix A — what each fill left between the card’s 24,576 MiB and the server footprint at depth, dedicated + shared — the Swift IQ4_XS text fill left 1,880 MiB, so its window stays at 159,744, and the vision fill left 1,330 MiB, so its window is 128,000 der; on dedicated memory alone, the counter of the fit test above, the vision fill misses 1,796 MiB by 22. For Swift IQ3_XXS the per-card tables give 216,064 (text) and 195,584 (vision) at one slot and 93,184 per slot for vision at two der. The anchor’s counterpart of knee pick 7 in the same sweep is knee pick 1 (UD-IQ4_XS, vision, n4/p0.75, one slot: fully resident 180,224 meas, decode-verified 219,136 der; Appendix A); the per-card tables have no two-slot text configuration, so the two-slot text row has no 1,796 MiB figure. Decode-verified: the last step whose decode held with the window filled (to about 91%, 98% for knee picks 5 and 7), minus 0.7 GiB at the configuration’s per-token slope, floored to 1,024 — or the native 262,144 where decode held to it, as for knee picks 5 and 7 — at the knee loads’ 258–539 MiB desktops; it is the October launcher’s window, it is fully resident only where the table says so, and it can sit above the fully resident window (Appendix A). Collapse point: the knee, where decode falls 15% below the next-lower step. Where a ceiling kind has no backing measurement, the cell says “not swept”: no knee sweep was run on Swift IQ4_XS. Speed bands are greedy-only (temperature 0); speed at the temperature-1.0 sampler the cards ship is not measured as a band (§15).
The two Swift IQ4_XS windows were filled again on 2026-10-07 in the Pi fine-tune study (window-knee.py, in the order Pi, Swift, Swift, Pi, n=2 per file; Appendix B). Vision at 139,264 held at 45.14 and 45.55 t/s on the same 126,838 tokens (desktop 1,217 MiB, board peak 24,019 MiB, 444 MiB shared: the 2026-10-04 footprint of 22,802 MiB dedicated again der), and text at 159,744 held at 48.36 and 48.51 t/s at a 146,423-token fill (board peak 23,429 MiB, 484 MiB shared) meas. Their absolute t/s are not comparable with the 2026-10-04 rows, and they are still not knee sweeps.
The drafter bill per file (board VRAM at -c 32768, one slot, desktop inside, 2026-10-04): the anchor’s n4 step costs 1,138–1,198 MiB over drafter-off meas n=2, Swift IQ4_XS costs 1,092–1,124 MiB meas n=2, and Swift IQ3_XXS costs 1,092–1,098 MiB meas n=2. The n10 step over n4 costs 898–900 MiB in five of the six loads across the three files meas n=2 (the sixth, Swift IQ4_XS’s second pass, reads 851 on a board counter that includes the desktop), matching the 897.75 MiB expected from six recurrent-state rows at 149.625 MiB each der. The projector bill on Swift IQ4_XS is 1,136 MiB of server footprint at load and 1,470 MiB at depth meas n=1 (one load each; the at-depth figure has a 1440p image in flight; bartowski’s F16 file) — the anchor’s published projector cost, 1,138 MiB meas of llama-server’s dedicated memory for the BF16 projector on UD-IQ4_XS (§05, August 2026), is a separate measurement on another file, another projector file and another day.
The per-token arithmetic for the Swift files (server footprint = weights + a constant per configuration + slope × window, Appendix A): Swift IQ4_XS loads 14,104.4 MiB of GPU-resident weights with the drafter on der and Swift IQ3_XXS loads 11,056.9 MiB der, against the anchor’s 13,061.1 MiB der; the per-token slope of server footprint (dedicated + shared) is shared at 45.87 KiB per token at one slot der (confirmed at 46.0 KiB meas on the Swift IQ4_XS fills, 0.3% above; for Swift IQ3_XXS it is the transplant value itself). See Appendix A for every other card and file.
Six flags belong on every card, and each one is here because a measurement put it here. Put every layer on the graphics card with -ngl 99: writing -ngl 64 looks complete and is not, and it costs about 30% of your speed with no warning anywhere. Serve one request at a time with --parallel 1. Store the memory of the conversation in 8-bit with -ctk q8_0 -ctv q8_0: it halves that memory and costs 0.3% of perplexity. Turn the built-in drafter on with --spec-type draft-mtp: it roughly doubles your speed, changes nothing about what the model writes, and cuts the electricity per token by two and a half times (n-max 10 / p-min 0.5 against no drafter, 2026-08-23; 2.17× at n4/p0.75; not measured at the pairs the cards carry now). Choose --load-mode none if system memory is tight and --load-mode mmap if you restart the server often. And whatever number you end up quoting, say which kind of token you counted, because on this model that alone is worth a factor of 1.7.
- All layers on the GPU, or accept the cliff —
-ngl 99. llama.cpp counts the output layer as one more layer than the model has transformer layers, so on this 64-layer model-ngl 64reads as complete and leaves that output layer — a 5120 × ~151k-vocabulary matrix multiplication — on the processor, running for every generated token. Measured cost: 25.7 against 39.7 t/s on Q4_K_M, and independently 29.84 against 42.31 on UD-IQ4_XS — −29.5% — with no memory signature at all (§11). Partial offload is set by system memory bandwidth, not by the graphics card, so a 5080 and a 4060 Ti are predicted to offload at nearly the same 8–12 t/s der — that band is §04's formula run at system-memory bandwidth, not a measurement: neither of those cards has ever been in this machine, and the only offload figure measured here is the 25.7 t/s above, on this 3090. And this is a dense model — every part of it runs on every token — so the trick of parking a model's rarely-used parts on the processor has nothing to target here. - One slot —
--parallel 1: serving one person, one slot is the faster setting. Some interfaces default to 2, which halves the window each request gets — the KV cache stays the size-csets and is split between the slots, not allocated twice; a second slot at-c 32768added 134 MiB of board VRAM meas n=1, not a second 1,248 MiB der cache (2026-08-23, UD-IQ4_XS,q8_0KV, drafter off,kv_unifiedoff as logged).C28 Measured 2026-08-25 on UD-IQ4_XS with the drafter on at n-max 10 / p-min 0.5 — then the faster shipped setting — three reps each, first post-prefill probe discarded:--parallel 1gave 82.98 t/s aggregate and 85.79 per slot;--parallel 2gave 101.25 t/s aggregate and 55.41 per slot — +22.0% aggregate, −35.4% per slot. Draft acceptance is unchanged across the pair (0.618 → 0.620), so the second slot does not disturb the drafter at all; the whole cost is that each user waits about 55% longer for their own tokens — 12.6 s against 8.2 s for a 700-token answer. Once the drafter is already saving most of the repeated reading of the model, batching has far less left to save (§03). - 8-bit KV cache —
-ctk q8_0 -ctv q8_0, verified twice on two files. Measured wikitext perplexity with the 8-bit cache: 6.5498 against 6.5348 at fp16 on Q4_K_M — +0.23%, inside the error bars — and the re-measurement replicated the effect independently on UD-IQ4_XS at 6.6160 against 6.5956, +0.31%, likewise about one standard error. Half the KV memory for statistically nothing, on both files. It also costs nothing measurable in energy: 8.341 against 8.284 J per decode token, 0.7% apart and well inside the 2.9% floor. The one price that is not nothing is speed: fp16 KV measured 43.19 againstq8_0's 42.31 t/s at-c 32768, so 8-bit costs about 2% of decode — when VRAM is genuinely spare, fp16 is marginally the faster cache. The step below is measured too:q4_0K/V at 6.6413, +0.693% over fp16 (§08). - The drafter on —
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0with one slot on the files the 2026-10-08 sweep ran; n-max 4 / p-min 0 with two slots and for the Pi fine-tune atmedium. Those pairs scored best per file at the cards’ own sampler (UD-IQ4_XS, Swift IQ4_XS and Swift IQ3_XXS atxhigh; two slots on Swift IQ3_XXS only), and n3/p0 is also this build’s own default. Files, cards and slot counts the sweep did not run carry n4/p0 der, at the n-max 4 their windows were fitted at: p-min costs no memory, and on the four files the sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot, reasoning and answers at 0.67–0.75 per slot with two (Swift IQ3_XXS), and 0.90–1.03 after a 57,540-token prefix, where on UD-IQ4_XS deep answers it was slightly faster (n4/p0 read 0.973 of it, interval 0.969–0.996). The 262,144-token UD-IQ4_XS window runs with the drafter off; what n-max and p-min mean, and how the pairs were chosen, is in the n and p subsection. Every drafted token is checked by the full model before it counts; on one prompt (build 10502, greedy, thinking off, 2026-10-04) greedy text was not always identical token for token with the drafter on (§08).C38 Measured here it doubles decode on code and cuts energy per token by 2.52× at n10/p0.5 and 2.17× at n4/p0.75 (2026-08-23; energy at n3/p0 and n4/p0 is not measured). It is not free in memory — 1,008 MiB fixed plus 11.4% of your window, measured at n-max 4 (n-max 3 is about 150 MiB less; §05) — and on the August greedy prose probe it was worth only about 1.16×, the one workload measured where--spec-type noneand the memory back was the better answer; the 2026-10-08 sweep did not time prose on its own. The rest of this section is how far it goes and why. - Load mode —
nonefor tight RAM,mmapfor restart-heavy work. Without--load-mode none(which replaces the deprecated--no-mmap) the whole GGUF stays cached in system memory even after the weights are in VRAM — a measured 13,480 MiB for the 4-bit file meas. The cost is a slower model load: measured 10–30 s to first health check across dozens of restarts in the campaign proper, and it is storage-dependent. The independent re-measurement ran about 30 restarts on--load-mode mmapinstead and measured warm reloads at 6–10 s with decode unaffected (repeat probes agreed to 0.7%) — because at-ngl 99every weight lands in VRAM either way, so load mode cannot touch decode speed. So: if you are sweeping flags or restarting dozens of times a day,mmapis the better choice. Every recipe in §03 nonetheless shipsnone, and that is a deliberate default rather than a contradiction:noneis the safe assumption about a machine this page cannot see, because the failure it prevents (system memory exhausted by a cached model file) is worse than the one it causes (a load 20 seconds slower). Change it if you know your machine has the memory to spare. Measured 2026-08-26:noneholds a 1,275 MiB working set againstmmap’s 13,480, a saving of 12,205 MiB meas, confirmed by two measures agreeing to 3 MiB (§15). - Say which tokens you counted. This model thinks by default, so a probe labelled "code generation" can spend every timed token reasoning about code and return an empty answer field. On code, answer tokens run about 1.7× faster than reasoning tokens on the same server with the same flags. Every speed row below declares its regime.
Reasoning tokens (thinking on, the model's default) and answer tokens (enable_thinking:false, what a reader keeps) are not interchangeable — on code the answer tokens are the faster ones, by a near-constant 1.69–1.81× at every depth. This model thinks by default, so a probe labelled “code generation” can spend all 700 timed tokens reasoning about code and return an empty content field; the tokens are real and the server timed them correctly, but they are not the deliverable. Every speed row below declares its regime.
Flash Attention: you are already using it, and the recipes depend on it
None of the recipes on this page passes -fa, and none of them needs to. This build documents -fa, --flash-attn [on|off|auto] with a default of auto — so a recipe that omits the flag is not running with Flash Attention off, it is letting the server decide. On this machine it decides on meas, and that is not an inference: with -fa unset and -ctk q8_0 -ctv q8_0 the server loads fine, while the same configuration with -fa off refuses to start. Acceptance also reads identically across auto and on — 0.611 shallow and 0.571 deep, to three decimals — and deep decode differs by 0.1 %.
What Flash Attention actually is. It is a different way of computing attention, not a quality setting. Ordinary attention builds the whole attention matrix in memory and then reads it back, so its memory traffic grows with the square of how much context it is attending over. Flash Attention fuses those steps so the large intermediate is never written out at all. The arithmetic is the same; the memory traffic is far smaller, and the saving grows with depth.
Every recipe here ships -ctk q8_0 -ctv q8_0, which halves the cache and costs 0.23 % of perplexity. That setting does not work without Flash Attention. Asked for both, the server exits during model load:
llama_init_from_model: quantized V cache requires flash_attn to be enabled
srv load_model: failed to create_context with model ...
srv llama_server: exiting due to model loading error
This is the good kind of failure — loud, immediate, and it names its own cause. You cannot silently end up with a degraded cache. But it does mean the recipes have a dependency they never state: the KV setting they ship only works because Flash Attention is available and auto turns it on.
Where that dependency could actually fire. Not on this 3090. The risk is a backend where auto resolves differently — and this page carries recipes for Intel Arc Pro B50 and B70 under Vulkan that also ship q8_0 KV, measured on none of them. If Flash Attention is unavailable there, those recipes will stop at the line above rather than run slowly. If that happens to you, drop to -ctk f16 -ctv f16 and re-check the window against the budget table, because an f16 cache is roughly twice the size per token.
-ctk q8_0 does not quantise the drafter’s cache
Chasing the flag above turned up something no recipe on this page states. -ctk and -ctv apply to the target model only. The draft model has its own pair, -ctkd and -ctvd, and they default to f16. No recipe here passes them — so on every pick that runs a drafter, the draft cache is unquantised even though the main one is not.
This is visible in the server’s own log meas. With -fa unset, the two contexts resolve Flash Attention by different routes: the draft context reports flash_attn = auto and then runs a probe that turns it on, while the target context reports enabling flash_attn since it is required for quantized V cache — forced, because its cache is quantised. Both end enabled, which is why auto and on measure the same; but they get there differently, and only the target context is forced.
Whether -ctkd q8_0 -ctvd q8_0 is worth setting is UNMEASURED here and is not a recommendation. The arithmetic says it would roughly halve the drafter’s per-token window cost, which the budget table books at 5,120 B per token — on the order of 300 MiB at a 131k window. Whether it costs acceptance, and therefore speed, is exactly the kind of thing this page does not guess at. It is written down here so the next person measuring has somewhere to start.
So should you set it? No — leave it alone. Adding -fa on buys nothing over the default on this card (0.1 % at depth, inside noise), and pinning a flag you have not measured on your own backend removes the one mechanism that would have chosen correctly for you. The reason to know about it is diagnostic, not performance: if a recipe here refuses to load and mentions flash_attn, that is this dependency, and the fix is the KV width, not the flag.
One measurement from this run that is not trustworthy, stated so nobody uses it: the prefill figures. Probes reused a cached prefix, so the server reported a full prompt_n against a near-zero prompt_ms and the resulting “prefill t/s” is meaningless. Only the decode, acceptance, VRAM and load-success results from this arm are reported anywhere on this page.
Every speed this page measured, in one table
Read the band, not a number. The floor is set by how fast the card reads its own memory, and everything above that floor is the drafter, so the right row for you is the one whose content and token regime match your work.
| Workload | Regime | Acceptance | Decode | What it means |
|---|---|---|---|---|
| Ceiling — copying supplied text verbatim (IQ4_XS · n-max 10 / p-min 0.5 · thinking off) | answer | 0.99 | 148.7 t/s | the real ceiling. C10 |
| Novel JavaScript class (IQ4_XS · n-max 10 / p-min 0.5) | answer | 0.67 | 95.4 t/s | answer tokens on code beat reasoning tokens on the identical task (88.8) — written code is more predictable than free-form reasoning about it |
| Novel Python module + tests (same flags) | answer | 0.59 | 79.1 t/s | against 63.1 for the reasoning stream on the same task |
| Short maths, greedy (GSM8K — §08's scored table's build and workload, not this table's) | mixed | 0.90 | 63–70 t/s | short, predictable reasoning |
| Realistic code, temperature 0 (Q4_K_M · n-max 4 / p-min 0.75) | reasoning | 0.81 | 57.9 t/s | the tuned-flags benchmark number |
| Production coding, temperature 1.0 | mixed | — | 55–61 t/s | what a normal session feels like |
Long xhigh run, sustained across a 61–76k-token thought | reasoning | 0.84–0.88 | 48.8–55.8 t/s | KV deepens, thinking drafts worse |
| English prose explainer (IQ4_XS · thinking off) — the drafter's worst content | answer | 0.44 at n10/p0.5 0.80 at n4/p0.75 | 43.8 t/s at n10/p0.5 48.4 t/s at n4/p0.75 | speculation is nearly worthless here. The best configuration measured on this content in August is n4/p0.75 at 48.35 t/s, which is 1.16× a 41.55 floor; the wide n10/p0.5 tree manages only 43.80, or 1.05×. The drafter's ~1.8 GiB (§05) is hard to justify on a pure writing workload; --spec-type none buys that window back. Correction: earlier editions, and the campaign log's own summary line, attached the 1.16× to the 43.80 row. 43.80 ÷ 41.55 = 1.05; the 1.16× belongs to 48.35 |
| Floor — speculation off | either | — | 39.7–43.0 t/s | pure bandwidth (Q4_K_M 39.7–40.0; UD-IQ4_XS 41.46–42.97 across five contents and both token regimes — 41.46 copying, 42.17 novel JavaScript, 41.84 Python, 41.55 prose, 42.97 the rate-limiter prompt, plus 42.47–42.81 on the reasoning stream). Content moves it by ≤3.6%: without a drafter, what you are writing barely matters at all |
All rows measured 2026-08-21 to 2026-08-23 on the reference RTX 3090 at short context, from the server's own timings. The floor row is also the diagnostic threshold, and it is published as a band rather than a point for that reason: it was established across five contents and both token regimes and holds inside 3.6%.
How speed falls as the prompt grows
Speed drops as the conversation gets longer, gently and predictably, and the drop is about 25–30% between an empty window and a 91,000-token one. It is not a cliff — that only happens when the window overflows the card (§05). What matters far more than the depth is which kind of token you are counting: at the same 91k depth, the answers a reader keeps arrive at 64.8 t/s while the thinking that precedes them runs at 36.6. Plan an agent session on the first number and a long xhigh run on the second.
All three series below are UD-IQ4_XS at -c 131072, fully resident, with prompts prefixed by a unique identifier so nothing is served from the prefix cache. The first column is the one to plan against: it was measured 2026-08-23 under the cooled protocol (§02, documented in §11) — the probe that comes back with the prefill discarded, three probes per depth off the cached prefix — while the other two columns are single probes taken right after their prefill and therefore carry the ±25% band from §01.
| Prompt depth | Answer tokens · cooled ladder no projector · n-max 4 / p-min 0.75 · thinking off | Answer tokens · coding flags + vision vision on · n-max 10 / p-min 0.5 · single probes, ±25% | Reasoning tokens no projector · n-max 4 / p-min 0.75 · thinking on | Naive tokens ÷ wall-clock the WRONG number, shown so you recognise it |
|---|---|---|---|---|
| ~1.5k | 86.3 t/s · acc 0.89 | 94.5 t/s · acc 0.60 | 51.2 t/s · acc 0.80 | — |
| ~10.7k | — | — | 50.3 t/s · acc 0.89 | — |
| ~28–30k | 80.2 t/s · acc 0.93 | 81.0 t/s · acc 0.61 | 47.1 t/s · acc 0.92 | 9.2 t/s |
| ~57k | — | — | 41.1 t/s · acc 0.91 | — |
| ~91k | 64.8 t/s · acc 0.92 | 70.1 t/s · acc 0.65 | 35.8 t/s · acc 0.86 (36.6 cooled) | 2.4 t/s |
The last column is not a measurement of the model: it is the same reasoning-series runs divided the wrong way, tokens ÷ total wall-clock, which buries a 26-second or 105-second prefill in the denominator. It is printed here so that a reader who computes 9.2 t/s on a card that decodes at 47.1 recognises their own arithmetic instead of blaming their hardware. Quote the server's own timings, and quote prefill separately (§05). Why the acceptance columns differ so much: the cooled ladder re-used the campaign's red-black-tree prompt, a textbook algorithm the draft head predicts extremely well, while the middle column generated novel code. Acceptance ranks the content, not the protocol — and, as the next part measures, it does not rank speed at all.
Three things to take from it. First, plan an agent session on 65–70 t/s at 91k of depth — those are answer tokens, what an agent actually receives, and they run about 2× faster than the reasoning stream at the same depth. Second, acceptance rises with depth, 0.80 → 0.92 in the reasoning series and 0.60 → 0.65 in the answer series — two independent confirmations. Decode still falls 25–30% across the same span (86.3 → 64.8 on the cooled ladder), which settles the mechanism: the cost is per-token KV reads growing with depth, not the drafter failing. Depth and acceptance are independent axes. Third, the decline is gentle and it is not a cliff — as long as the window is genuinely resident. Overcommit it and the same depths give 18.6 and 8.0 t/s instead (§05).
The projector costs 1,138 MiB of VRAM and 0% of decode. Paired probe, 2026-08-23: byte-identical prompts filling 90,862 tokens, loads ordered A-B-B-A, cooled probes, the only difference being --mmproj: with the drafter off, with-projector 26.965 t/s against no-projector 26.975 (n=10 each) — a 0.04% difference against a 0.4% within-arm spread. Repeated with the drafter on: 62.71 against 62.655 t/s (n=8 each), 0.09% apart, with draft acceptance matching probe for probe. The projector costs exactly what §05 books for it and nothing else: 1,138 MiB of VRAM, 0% of decode — the VRAM delta reads 19,376 − 18,238 in one pair and 21,034 − 19,896 in the other, the same 1,138 MiB in every comparison.
The mechanism under the two regimes: mean draft length
Line up the cooled ladder against the reasoning column — same file, same flags, different protocols — and the reasoning side sits a near-constant 1.7× below at every depth: 86.3/51.2 = 1.69, 80.2/47.1 = 1.70, 64.8/35.8 = 1.81 — so call it 1.69–1.81×. A ratio that stable across a 60× span of depth is not clock noise, so it was isolated properly (2026-08-23): one server, one 90,894-token prompt, the same n-max 4 / p-min 0.75 flags, cooled probes, the only change being the chat template:
| Token stream at 91k | Decode | Draft acceptance | Mean draft length | Energy per decode token |
|---|---|---|---|---|
reasoning (enable_thinking on — the default) | 36.62 t/s | 0.8951 | 2.99 | 6.066 J (separate arm) |
answer (enable_thinking:false) | 62.02 t/s | 0.9073 | 4.31 | 3.744 J (separate arm) |
either, --spec-type none | 27.0 t/s | — | — | 8.104 J (separate arm) |
A 1.69× speed difference produced by nothing but the token stream — and acceptance says nothing about it (0.895 against 0.907, essentially tied). The mean-draft-length column is where the difference actually lives. On uncertain reasoning tokens the p-min 0.75 confidence gate keeps cutting the draft tree short — a mean of 2.99 tokens per pass against 4.31 on the answer stream — so acceptance stays high precisely because the gate is discarding the part that would have missed. That makes acceptance blind to the loss. Acceptance tells you whether a draft was right; mean draft length at a fixed n-max tells you how fast you will go. It is the same fact that makes the highest-acceptance configuration the slowest one in the sweeps below. But mean draft length alone does not rank configurations when n-max differsC24 — a drafter at n-max 4 / p-min 0.00 reaches 82.73 t/s at acceptance 70.5% and mean draft length 2.803, beating n-max 10 / p-min 0.5 at 76.32 t/s with 4.035, because the longer draft speculates up to ten tokens and discards about half, and every discarded token was still computed. The quantity that ranks these is accepted tokens per unit of drafting work, not accepted length alone (2026-08-28, UD-IQ4_XS, -c 32768, q8_0 KV, three settled probes each after a discarded warmup).
The energy column comes from a different pair of arms on a different day's prompt — 2026-08-23's power matrix, thinking on against thinking off at short context — and is printed beside these rows because it measures the same mechanism on a completely independent instrument: 1.62× in joules per token against 1.69× in throughput. Two instruments, one effect. Note also the floor row: 27.01 t/s thinking-on against 26.95 thinking-off — identical, which confirms once more that content dependence on this page is almost entirely a speculation effect: without a drafter, what you are writing moves decode by at most 3.6%.
The flags cannot buy the difference back. At the same depth with thinking on, --spec-type none gives 27.01, the then-shipped n4/p0.75 gives 36.62 (1.36×) and the August greedy-code peak, n10/p0.5, gives 38.67 (1.43×) — a 5.6% gain where the same flag change is worth 12% on code. So the practical number for the configuration this page shipped by default until 2026-10-08 (n4/p0.75; the n3/p0 it ships now was not measured at 91k, §06): anyone running xhigh over a 91,000-token document should plan on 37–39 t/s, not the 62–70 t/s the code probes advertise. The deliverable arrives at code speed; the thinking that precedes it does not.
What the built-in draft head is
Qwen3.8-27B ships with a small helper network inside the same file — a multi-token prediction (MTP) draft head — that guesses several tokens ahead so the full model only has to check them, which is much faster than generating each one. The standard K-quant GGUF files (§08's table) already include it, and the head's precision is set by whoever quantized the file; in UD-IQ4_XS it is held at 8-bit and 6-bit with F32 (Q8_0, Q6_K and F32, 334.75 MiB meas, read from the file header on 2026-10-04), and bartowski's IQ4_XS and IQ3_XXS quantizations of Swift-1.5, a fine-tune of this model with the same architecture, hold it at 4-bit with F32 (Q4_0 and F32, 227.91 MiB meas).C35 llama.cpp added support in July 2026, and every guess is verified by the full model before it is accepted, so the drafter is built to change only the speed — though on one prompt on this build (greedy, thinking off) its text was not always identical to the drafter-off text (§08).C38 It works on any GPU llama.cpp supports.
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 :: the best score with one slot on the three files swept at xhigh at the cards' sampler :: (2026-10-08); n-max 4 / p-min 0 with two slots and for the Pi fine-tune at medium; :: unswept files take n4/p0, derived, not measured. The best n-max moves with the file, the :: slot count and the REGIME: greedy code favours longer drafts (section 06, n and p). :: And it is not free: 1,008 MiB + 11.4% of your window at n-max 4, about 150 MiB per :: step of n-max, +898 MiB at n-max 10; p-min costs nothing (sections 05 and 06).
How much speed it adds depends heavily on hardware and workload — published reports range from +33% to +145%, and one RTX 5090 tester saw 148 average / 203 peak t/s on the NVFP4 tiers. On the reference 3090, generation went from 40.3 (the 08-21 build's cold-start baseline; the canonical 08-22 baseline is 39.7–39.9) to 58–60 t/s — +45% at the default-ish settings — and on the original benchmark prompt, tuning pushed it to 81.7 t/s, +105%, a figure a 2026-08-22 re-sweep appeared to overturn and a 2026-08-23 matched sweep reproduced to the decimal once the token regime was pinned. The warning all these numbers carry: benchmark on your own machine. That 5090 tester found n-max 4 fastest with 6–8 slower; this 3090 agreed on real code yet kept climbing to n-max 10 on its near-ideal prompt. The optimum moves with hardware and content, so there is no universal best value.
Content swings decode by 3.59× on this card at fixed flags, produced by nothing but what is being written. On answer tokens the spread is wider still: a 41.5–42.2 t/s floor without the drafter against 148.7 t/s copying at n10/p0.5. The earlier demonstration on the reasoning stream (2026-08-22): 51.5 t/s at acceptance 0.47 (novel code, mean accepted run 4.0 of 10 drafted) against 119.8 t/s at acceptance 0.93 (copying, mean run 9.9 of 10) — a 2.3× difference, narrower because it timed the model thinking about copying rather than copying, which is the caveat the band table at the top of this section attaches to the same figure. Which is why a speculative benchmark number quoted without its acceptance rate cannot be checked. With one refinement the sweep section below carries: acceptance ranks your content, but it does not rank your flags.
DFlash2, and why it lost
There is a second way to draft, and on this machine it was slower than the one already inside the file. DFlash2 predicts a whole seven-token block in a single pass, then picks the most plausible sequence out of the candidates it considered for each position. Like the built-in head, every draft is verified before it is accepted, so the output is identical to running the model alone. The costs are real: a separate 1.14 GB drafter to download (incoai/Qwen3.8-27B-DFlash2-GGUF, about 1.2 GB of VRAM) and — until llama.cpp PR 27342 is merged — building llama.cpp yourself from that branch. On its best benchmark prompt it reached 69–77 t/s; on realistic code its best was 46.5 t/s at n-max 2, below the built-in head's 57.9. That is the whole verdict.
| Configuration (hover or tap a row for its exact flags) | Decode | vs baseline | Draft acceptance |
|---|---|---|---|
| Q4_K_M, no speculation (08-21 cold-start build — the canonical 08-22 baseline is 39.7–39.9) | 40.3 t/s | — | — |
| Q4_K_M + built-in MTP (n-max 4, p-min 0.75 — the everyday pick in August; expect 69.8 t/s on answer tokens and 57.9 on the reasoning stream. The 58–60 / 89% in this row is the 08-21 build) | 58–60 t/s | +45% | 89% |
| Q4_K_M + MTP at n-max 10, p-min 0.5 (the peak, on answer tokens) | 81.7 t/s† | +105% | ≈80% (longer drafts) |
| Q4_K_M + DFlash2 (n-max 7; lost on this machine — best on realistic code was 46.5 t/s at n-max 2, below the built-in head’s 57.9, and costs a separate 1.14 GB drafter plus a source build from PR 27342. The 69–71 here is best-case benchmark-prompt content) | 69–71 t/s | +75% | 73% (7-token blocks) |
| NVFP4-HIGH + built-in MTP (2026-08-21 build — re-measured ordering in the row below) | 46.8 t/s | +16% | 90% |
| NVFP4-VERY-LOW + built-in MTP (2026-08-21 build — on 2026-08-22 re-measure at these flags VERY-LOW runs 54.6 against HIGH’s 48.2 t/s, both accepting 0.83, reversing this table’s order; both sit below Q4_K_M’s 57.9 — §08) | 43.0 t/s | +7% | 86% |
The exact llama-server command for each row (tap to expand)
:: shared flags in every run below: llama-server -c 32768 -ngl 99 --parallel 1 -ctk q8_0 -ctv q8_0 --jinja --port 1235 :: row 1 — baseline: just the model, no speculative flags -m Qwen3.8-27B-Q4_K_M.gguf :: row 2 — built-in MTP, n-max 4 / p-min 0.75 — the everyday pick in August (the head ships inside the main GGUF): -m Qwen3.8-27B-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: row 3 — MTP at n-max 10 / p-min 0.5 — the peak on answer tokens († = measured warm-cache): -m Qwen3.8-27B-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 10 --spec-draft-p-min 0.5 :: row 4 — DFlash2 (separate 1.14 GB drafter via -md; needs the PR 27342 build): -m Qwen3.8-27B-Q4_K_M.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 7 :: rows 5-6 — NVFP4 files with their embedded MTP head (same MTP flags): -m Qwen3.8-27B-NVFP4-MTP-HIGH.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -m Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: what each speculative flag means: :: --spec-type which drafting method: draft-mtp (built-in head) or draft-dflash (external drafter) :: -md / --model-draft path to the external drafter GGUF (DFlash2 only) :: --spec-draft-n-max how many tokens to draft per step, one after another - the main tuning knob, sweep 2-12 on your machine :: --spec-draft-p-min drafter confidence floor 0-1: the draft ends at the first guess whose top-1 probability (renormalised over the head's top 10) is below it; 0 never ends a draft (section 06)
Copy-paste version — the same command with the comments removed
llama-server -c 32768 -ngl 99 --parallel 1 -ctk q8_0 -ctv q8_0 --jinja --port 1235 -m Qwen3.8-27B-Q4_K_M.gguf -m Qwen3.8-27B-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -m Qwen3.8-27B-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 10 --spec-draft-p-min 0.5 -m Qwen3.8-27B-Q4_K_M.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 7 -m Qwen3.8-27B-NVFP4-MTP-HIGH.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -m Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75
That table says two things. First, speculation is real free speed: at the then-shipped everyday flags (n4/p0.75), decode measured 57.9 t/s on realistic code, and the temperature-1.0 xhigh sweep sustained 48.8 t/s across a 73,000-token thinking run — putting the 100k reference run at roughly 34 minutes, down from 45. Second, the NVFP4 rows show that the quantization and the drafter interact: speculative decoding makes the GPU check several drafted tokens at once, and that batch-checking is exactly the work NVFP4's software fallback does slowest. So on cards older than Blackwell, pair speculation with normal K-quants, not NVFP4.
n and p in plain words, and the best pair for each launcher pick (measured 2026-10-07/08)
On the files this sweep ran: n-max 3 / p-min 0 with one slot on UD-IQ4_XS, Swift IQ4_XS and Swift IQ3_XXS at xhigh; n-max 4 / p-min 0 with two slots busy (Swift IQ3_XXS, the one file run that way) and on the Pi fine-tune at medium. Each had the highest mean of shallow and after-prefix reasoning-token ratios for its file at the cards’ own temperature-1.0 sampler, among the settings timed at both depths (two slots: the best per-slot speed of three settings). On shallow reasoning tokens against n4/p0.75 in the same session, n3/p0 read 1.111 on UD-IQ4_XS, 1.189 on Swift IQ4_XS and 1.097 on Swift IQ3_XXS; after a 57,540-token prefix 1.061, 1.052 and 1.007, intervals that include 1. With two slots busy n4/p0 read 1.337 per slot (second pass only, one load per setting). On the Pi file n4/p0 read 1.038 of n3/p0, and 1.073 after the prefix der. The lead is clearest on shallow reasoning tokens: on answer tokens n4/p0 read the same as n3/p0 or higher (UD-IQ4_XS, Swift IQ4_XS), and after the prefix several longer settings read as high as n3/p0 or higher with overlapping intervals, n10/p0.5 among them on UD-IQ4_XS (on the Pi file it matched n3/p0 there, 1.021, but not the pick n4/p0, 1.073). Greedy code is where longer drafts lead most clearly (below). Files, cards and slot counts the sweep did not run take n4/p0 der, keeping the n-max 4 their windows were fitted at: p-min costs no memory, and at n-max 4 p-min 0.75 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot on all four files here, reasoning and answers at 0.67–0.75 per slot with two on Swift IQ3_XXS, and 0.90–1.03 after the prefix, where on UD-IQ4_XS deep answers it was slightly faster (n4/p0 read 0.973 of it). Efforts the sweep did not run keep each file’s pick, unmeasured there, except the Pi file at xhigh, where a later follow-up favoured n-max 3 on thinking tokens (Appendix B). No pick needs more memory than the drafter its window was filled at, and p-min costs nothing. Measured 2026-10-07 17:05 to 2026-10-08 10:04, 138 server loads meas.C42C43C44
To copy: --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 with one slot on those three files; --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 with two slots and for the Pi fine-tune at medium. The cards in §03 and Appendix B carry them already, and the per-pick table links each one.
n-max is how many tokens the helper may guess before the full model checks them. The built-in head guesses one token at a time, one small pass per guess, and from the second guess on it builds on its own last guess, so each later guess is right less often. It always guesses its own top pick, whatever n-max is: from the same starting point an n3 draft is the first three tokens of the n10 draft. By the source, with p-min 0 every step guesses exactly n-max tokens except near the end of a reply; measured on the Pi file’s reasoning tokens, 2.980 guesses per check at n3/p0 and 9.869 at n10/p0, close to n-max meas.
p-min is a confidence floor that can end a draft early. The confidence is the head’s probability for its top pick, renormalised over its ten best candidates at temperature 1, so your request’s sampler never touches it. The draft ends at the first guess below the floor (that guess is dropped, its pass already spent; the first guess is checked too). The test is strictly “below”, so p-min 0 never ends a draft. p-min moves where a draft is cut, never which tokens are guessed.
Both are load-time flags, and only n-max costs memory. llama-server silently ignores a request’s speculative.n_max and p_min (disabled since PR 22838): asking for n_max 1 on an n10/p0.5 load still drafted 8.65 tokens per check against 8.78 meas n=1, so every pair needs its own server start. Each +1 of n-max keeps one more recurrent-state row per slot, 149.625 MiB here. This build’s defaults are n-max 3 and p-min 0.00. The semantics come from llama.cpp’s source at commit 0adcc3bb5 (§15); the chips mark what was measured.
A wrong guess costs time, never quality. The full model checks every guess in one pass, keeps them up to the first it disagrees with and adds one token of its own. Greedy, a guess survives only if it is the full model’s own top token, so the text is identical in principle, though not bit for bit on this build.C38 At temperature 1.0 a guess survives with exactly the probability the full model gives it, so the output distribution is preserved in principle. A wrong guess costs its own pass, the passes of every later guess in that step and a row each in the check. From the source one expects sampling to favour shorter drafts than greedy decoding, because even a correct top-pick guess is rejected whenever the full model samples another token; this sweep did not isolate that, because its greedy probe also changed the prompt and turned thinking off. The larger difference it did measure was between reasoning text and answer text at the same sampler: at n10/p0 Swift IQ4_XS kept its third guess 0.258 of the time on reasoning tokens and 0.539 on answers, and on answers n4/p0 read higher than n3/p0 there (1.099 against 1.062). On the sweep’s one greedy probe every setting, drafter off included, gave the same text per file; quality was not scored at the new settings.
How often each guess survives, the curve under every result below: the share of draft steps whose first, second, third… guess was kept, from llama-server’s /metrics counter spec_decode_num_accepted_tokens_per_pos_total, at n10/p0, shallow, both passes pooled meas:
| File and tokens | 1st | 2nd | 3rd | 4th | 5th | 6th | 7th | 8th | 9th | 10th |
|---|---|---|---|---|---|---|---|---|---|---|
Pi IQ4_XS, reasoning (medium) | 0.800 | 0.629 | 0.475 | 0.354 | 0.273 | 0.204 | 0.153 | 0.114 | 0.087 | 0.063 |
| Pi IQ4_XS, answers | 0.821 | 0.684 | 0.550 | 0.429 | 0.347 | 0.254 | 0.193 | 0.160 | 0.126 | 0.093 |
| Pi IQ4_XS, greedy code (one prompt) | 0.953 | 0.837 | 0.698 | 0.558 | 0.512 | 0.442 | 0.302 | 0.302 | 0.186 | 0.140 |
Swift IQ4_XS, reasoning (xhigh) | 0.639 | 0.404 | 0.258 | 0.169 | 0.117 | 0.082 | 0.064 | 0.049 | 0.042 | 0.033 |
| Swift IQ4_XS, answers | 0.825 | 0.689 | 0.539 | 0.432 | 0.343 | 0.250 | 0.221 | 0.167 | 0.132 | 0.093 |
UD-IQ4_XS, reasoning (xhigh) | 0.654 | 0.435 | 0.296 | 0.206 | 0.157 | 0.120 | 0.093 | 0.076 | 0.062 | 0.050 |
| UD-IQ4_XS, answers | 0.829 | 0.655 | 0.499 | 0.404 | 0.324 | 0.242 | 0.201 | 0.165 | 0.126 | 0.104 |
The Pi file ran at medium, the others at xhigh, so effort and file differ together; a later follow-up ran the Pi file at xhigh, where n3/p0 read ahead of n4/p0 on reasoning tokens (1.080, Appendix B). The greedy row is one 256-token prompt. In the Pi file’s n10/p0 reasoning probes about 15% of the text was answer text.
Reading it. On reasoning tokens at xhigh, Swift IQ4_XS and UD-IQ4_XS keep their third guess 0.258 and 0.296 of the time against the Pi file’s 0.475; on answers the files are close. At n4/p0 the fourth guess survived 0.414 of steps on the Pi file’s reasoning and 0.198, 0.218 and 0.211 on Swift IQ4_XS, UD-IQ4_XS and Swift IQ3_XXS meas, consistent with n-max 4 being best on the Pi file and 3 on the one-slot xhigh files — a correlation; the cost of one guess was not measured. Survival barely depends on n-max (the Pi file’s first three positions: 0.806, 0.637, 0.500 at n3/p0). And p-min 0.75 cuts reasoning drafts short: at n-max 4 on UD-IQ4_XS it drafted 1.312 tokens per check and kept 1.048 (acceptance 0.799), against 3.975 and 1.605 at p-min 0 (acceptance 0.404) — higher acceptance, fewer tokens per check (above). After the prefix every position survived more often on reasoning tokens (UD-IQ4_XS fourth guess at n4/p0: 0.344 against 0.218) and less often on answer tokens (0.374 against 0.446), with the content changed too.
“With p-min 0, shouldn’t n-max 4 or 10 be faster than 3?” Under greedy decoding it often was: in the 2026-08-23 greedy re-measurement on novel JavaScript, n10/p0 decoded 87.3 t/s against n3/p0’s 82.0 (below). At the cards’ temperature-1.0 sampler with thinking on it was not, because each guess costs a head pass and a row in the check whether it survives or not, and late guesses rarely survive. On the Pi file’s reasoning tokens n10/p0 kept more tokens per check than n3/p0 (3.133 against 1.932) but paid for about ten guesses every step (9.869 against 2.980), its fifth to tenth guesses surviving only 0.273 down to 0.063 of the time, and it decoded 0.877 as fast meas. p-min 0.75 fails the other way, cutting drafts that would often have survived (the UD-IQ4_XS figures just above). For these files at this sampler the best cap sat between the two, at 3 or 4.
What the flags cost in memory, server VRAM in MiB meas:
| Drafter | Pi IQ4_XS | UD-IQ4_XS | Swift IQ4_XS |
|---|---|---|---|
| drafter off | 17,526 | 16,411 | 17,523 |
| n1/p0 | 18,486 one of two loads disturbed | — | — |
| n2/p0 | 18,536 | 17,488 | 18,536 |
| n3/p0 | 18,718 | 17,642 | 18,719 |
| n4/p0 | 18,872 | 17,792 | 18,872 |
| n5/p0 | 19,020 | 17,939 | 19,050 |
| n6/p0 | 19,170 | 18,084 | 19,167 |
| n8/p0 | 19,472 | — | — |
| n10/p0 | 19,772 | 18,690 | 19,820 |
| n4 at p0.25 / p0.5 / p0.75 | 18,872 / 18,924 / 18,871 | 17,791 / 17,786 / 17,792 | 18,872 / 18,870 / 18,871 |
| n10/p0.5 | 19,774 | 18,696 | 19,773 |
Board peak minus the desktop before the load, median of single-slot loads at -c 81920, --cache-ram 0, no projector; not the launcher windows’ footprints. From n2 on each step adds about 150 MiB (n3 to n10 on the Pi file: 1,054 against 1,047 predicted); p-min moved nothing beyond tens of MiB of noise. Turning the drafter on at all cost about 1 GB at this window (drafter off to n2/p0: 1,010 to 1,077 MiB on the three files), part of it the draft cache that grows with -c (§05); the Pi n1 median includes a load taken while the reference machine was in use. On Swift IQ3_XXS the n3 step was not resolved: its two n3/p0 loads read 15,731 and 15,831 MiB against 15,823 at n4 on every n4 load, so the n3 step is not visible there (n4 to n5: +148, as predicted).
The results, file by file. Decode-speed ratios inside one session against the block’s reference (above 1 = faster than the reference), paired 95% interval in grey der. Reasoning: thinking on at the pick’s effort, 6 prompts × 768 tokens × 2 passes; answers: thinking off, 4 × 384 × 2; after the prefix: 3 reasoning and 2 answer questions about 57,540 tokens of source code. The repeat block re-timed leading settings in fresh server starts against its own reference loads; the same setting and seed gave identical text, so it repeats the timing, not the sampling. Highlighted: what the launcher now runs.
bytkim’s Qwen3.8-27B-pi IQ4_XS at medium (pick [P]), against n3/p0. Only n4/p0 is above 1 on both reasoning columns in the main block, and it repeats. n4/p0.25 ties it (1.058 against 1.060 on the reasoning mean, repeat block) at the same n-max, so the pick is n4/p0, about 1.15× the old card drafter (1.038 ÷ 0.900, no interval) der.
| Drafter | Reasoning der | Reasoning after the prefix | Answers | Answers after the prefix |
|---|---|---|---|---|
| Main block | ||||
| drafter off | 0.549 0.510–0.591 | 0.462 0.397–0.539 | 0.543 0.495–0.592 | 0.549 0.470–0.591 |
| n1/p0 | 0.840 0.793–0.892 | — | 0.811 0.762–0.860 | — |
| n2/p0 | 0.985 0.944–1.030 | 0.913 0.868–0.952 | 0.972 0.934–1.005 | 0.944 0.888–0.978 |
| n3/p0 (the command until 2026-10-08) | 1 (the reference) | |||
| n4/p0 (the command now) | 1.038 1.000–1.092 | 1.073 1.041–1.116 | 1.030 0.985–1.075 | 0.978 0.946–1.147 |
| n5/p0 | 0.970 0.925–1.035 | — | 0.986 0.913–1.064 | — |
| n6/p0 | 0.913 0.861–0.988 | 0.978 0.889–1.123 | 0.911 0.831–0.996 | 0.876 0.840–1.079 |
| n8/p0 | 0.931 0.865–1.015 | — | 0.997 0.921–1.084 | — |
| n10/p0 | 0.877 0.807–0.964 | — | 0.947 0.870–1.033 | — |
| n4/p0.5 | 0.966 0.922–1.026 | — | 1.026 1.014–1.041 | — |
| n4/p0.75 (the old card drafter) | 0.900 0.875–0.930 | 0.965 0.904–1.057 | 0.939 0.912–0.970 | 0.893 0.867–1.004 |
| n10/p0.25 | 0.826 0.766–0.900 | — | 0.938 0.876–1.008 | — |
| n10/p0.5 (the August greedy-code peak) | 0.884 0.835–0.942 | 1.021 0.867–1.342 | 0.934 0.876–1.020 | 0.899 0.848–1.157 |
| n10/p0.75 | 0.874 0.830–0.931 | — | 0.922 0.875–0.979 | — |
| Repeat block | ||||
| n4/p0 | 1.038 1.001–1.091 | 1.082 1.060–1.117 | 1.029 0.984–1.074 | 0.975 0.942–1.139 |
| n4/p0.25 | 1.033 0.995–1.085 | 1.083 1.055–1.120 | 1.038 1.008–1.067 | 1.027 1.000–1.131 |
| n5/p0 | 0.971 0.927–1.035 | 1.058 0.999–1.146 | 0.989 0.918–1.067 | 0.930 0.898–1.091 |
Swift-1.5 IQ4_XS at xhigh (pick [9]), against n4/p0.75. n2/p0 leads shallow but drops to 0.932 after the prefix; n3/p0 has the best reasoning mean in both blocks (1.121), with n4/p0.25 in the 3% tie band (1.105). On answers alone n4/p0 reads higher, intervals overlapping. Drafter off read 0.782 on reasoning. One reference spread passed the pre-registered 3% line (main block, pass 2, after the prefix: 3.48%).
| Drafter | Reasoning der | Reasoning after the prefix | Answers | Answers after the prefix |
|---|---|---|---|---|
| Main block | ||||
| n2/p0 | 1.198 1.171–1.222 | — | 1.018 0.955–1.075 | — |
| n3/p0 (the cards now) | 1.189 1.158–1.224 | 1.052 0.999–1.099 | 1.062 0.992–1.124 | 1.068 1.057–1.070 |
| n4/p0 | 1.110 1.073–1.149 | 1.063 0.984–1.166 | 1.099 1.069–1.136 | 1.081 1.061–1.123 |
| n6/p0 | 0.913 0.877–0.953 | — | 0.968 0.917–1.033 | — |
| n10/p0 | 0.848 0.792–0.915 | — | 0.999 0.927–1.104 | — |
| n10/p0.5 (the August greedy-code peak) | 0.935 0.900–0.975 | 1.044 0.968–1.199 | 0.985 0.933–1.059 | 0.925 0.882–1.072 |
| n4/p0.75 (the cards until 2026-10-08) | 1 (the reference) | |||
| Repeat block | ||||
| n2/p0 | 1.195 1.169–1.217 | 0.932 0.833–1.035 | 1.011 0.949–1.068 | 0.991 0.906–1.021 |
| n3/p0 | 1.186 1.155–1.221 | 1.057 1.001–1.118 | 1.064 0.993–1.127 | 1.065 1.057–1.066 |
| n4/p0 | 1.111 1.074–1.148 | 1.027 0.970–1.112 | 1.100 1.071–1.136 | 1.076 1.054–1.119 |
| n5/p0 | 1.004 0.969–1.040 | 1.063 0.976–1.184 | 1.052 1.009–1.103 | 1.001 0.976–1.058 |
| n4/p0.25 | 1.096 1.038–1.151 | 1.114 1.046–1.208 | 1.092 1.062–1.132 | 1.091 1.068–1.114 |
| n4/p0.5 | 1.071 1.041–1.106 | — | 1.073 1.042–1.113 | — |
Qwen3.8-27B UD-IQ4_XS, the anchor, at xhigh (picks [1] [2] [4]), against n4/p0.75. n3/p0 has the best reasoning mean in both blocks (1.086, 1.087). n10/p0.5 is slower on shallow reasoning and answers and reads higher after the prefix (1.079 against 1.061), intervals overlapping. On deep answers n4/p0.75 was faster than n4/p0 in both blocks (n4/p0 0.973, two prompts). Drafter off read 0.788 on reasoning.
| Drafter | Reasoning der | Reasoning after the prefix | Answers | Answers after the prefix |
|---|---|---|---|---|
| Main block | ||||
| n2/p0 | 1.133 1.098–1.170 | — | 1.081 1.008–1.148 | — |
| n3/p0 (picks [1] [2] [4] now) | 1.111 1.075–1.156 | 1.061 0.954–1.161 | 1.120 1.063–1.168 | 1.071 0.975–1.135 |
| n4/p0 | 1.049 0.995–1.121 | 1.000 0.914–1.066 | 1.125 1.089–1.171 | 0.973 0.969–0.996 |
| n6/p0 | 0.863 0.803–0.952 | — | 0.989 0.924–1.073 | — |
| n10/p0 | 0.912 0.846–1.009 | — | 1.043 0.963–1.141 | — |
| n10/p0.5 (picks [2] [4] until 2026-10-08) | 0.961 0.911–1.019 | 1.079 1.004–1.267 | 1.049 0.961–1.162 | 0.998 0.987–1.039 |
| n4/p0.75 (pick [1] until 2026-10-08) | 1 (the reference) | |||
| Repeat block | ||||
| n2/p0 | 1.134 1.099–1.171 | 0.979 0.850–1.110 | 1.081 1.009–1.148 | 0.954 0.847–1.034 |
| n3/p0 | 1.110 1.073–1.155 | 1.064 0.956–1.164 | 1.121 1.063–1.172 | 1.071 0.981–1.133 |
| n4/p0 | 1.048 0.994–1.120 | 0.997 0.908–1.070 | 1.123 1.087–1.170 | 0.973 0.968–0.997 |
| n5/p0 | 0.956 0.897–1.040 | 0.947 0.864–1.058 | 1.068 1.016–1.138 | 0.958 0.927–1.047 |
| n4/p0.25 | 1.010 0.971–1.060 | 1.075 1.013–1.148 | 1.109 1.084–1.144 | 1.010 0.976–1.033 |
| n4/p0.5 | 1.047 1.013–1.091 | — | 1.098 1.060–1.151 | — |
Swift-1.5 IQ3_XXS, one slot, xhigh (picks [5] [7]), against n4/p0.75. n3/p0 leads (reasoning mean 1.052; n4/p0 1.015, n5/p0 0.961); its after-prefix interval includes 1, and n2 was not run.
| Drafter | Reasoning der | Reasoning after the prefix | Answers | Answers after the prefix |
|---|---|---|---|---|
| n3/p0 (picks [5] [7] now) | 1.097 1.065–1.130 | 1.007 0.935–1.069 | 1.054 1.008–1.094 | 1.077 1.006–1.095 |
| n4/p0 | 1.034 1.011–1.053 | 0.996 0.953–1.046 | 1.049 0.995–1.103 | 1.059 1.043–1.081 |
| n5/p0 | 0.940 0.917–0.962 | 0.982 0.949–1.046 | 0.997 0.951–1.052 | 1.005 0.988–1.035 |
| n4/p0.75 (until 2026-10-08) | 1 (the reference) | |||
Two slots, both busy, at pick [8]’s shape (123,904 per slot, projector loaded, no image): per-slot decode, second pass only, one load per setting and pass, no repeated reference, no interval der; the grey figures are ratios to n4/p0.75:
| Drafter | Reasoning tokens, per slot meas | Answer tokens, per slot | Reasoning after the prefix, per slot |
|---|---|---|---|
| n4/p0 (picks [6] [8] now) | 44.97 t/s ×1.337 | 64.38 t/s ×1.484 | 12.24 t/s not resolved |
| n3/p0 | 39.86 t/s ×1.185 | 53.51 t/s ×1.233 | 11.96 t/s not resolved |
| n4/p0.75 (until 2026-10-08) | 33.64 t/s | 43.39 t/s | 12.29 t/s not resolved |
Pass 1 is not quoted: its desktop changed between loads (1,094, 602 and 125 MiB); pass 2 ran at 159 MiB throughout. Both passes, in reversed order, put n4/p0 first (over n3/p0 by 1.135 in pass 1 and 1.128 in pass 2), but with one load per setting that gap is about the size of this rig’s level change between server starts (about 13%, §03); the gap to n4/p0.75 is well beyond it. The pre-registered 3% pass rule voids both challengers, whose ratios move with the pass-1 reference. The deep column is not a clean reading: the two slots’ deep probes decoded at very different speeds in every load and neither reused the cached prefix. Pick [6] (text only) was not run and inherits this result. Why two slots prefer n-max 4 was not measured.
On the sweep’s greedy probe the peak sat at n4–n5, never at n3. The probe (a Python rate limiter, 256 tokens, temperature 0, thinking off; one prompt, no interval der) read, on UD-IQ4_XS, n5/p0 1.118, n4/p0 1.112–1.115, n10/p0 1.091, n10/p0.5 1.069 and n3/p0 1.062–1.064 of n4/p0.75; on Swift IQ4_XS, n5/p0 1.212, n4/p0 1.165–1.169, n10/p0 1.169, n3/p0 1.127–1.130, n10/p0.5 1.081. So on those two files the greedy best is longer than the sampled-reasoning best, n3. On the Pi file greedy also peaked at n4/p0, its sampled pick: 1.058–1.059 of n3/p0, with n8/p0 1.047, n5/p0 1.029–1.030, n10/p0 1.010 and n10/p0.5 0.974. The probe changes the sampler, thinking and the prompt together, so it does not say which of them moves the peak. The page’s earlier greedy sweeps point the same way with thinking off — the 2026-08-23 matched sweep put n10/p0.5 first but tested no p-min 0 at n-max 4 or 5, and the 2026-08-28 retest put n4/p0 above n10/p0.5C24 — and were mixed with thinking on: the 2026-08-23 re-measurement on reasoning tokens favoured n10, the 2026-08-22 re-sweep n4/p0.75 (below). It ranks, it does not recommend: for greedy code, n4/p0 or n5/p0 is the candidate to measure.
The decision for each pick of the October launcher der. Pick numbers are the October launcher’s, not the menu’s rows; each row links the page’s card for that configuration, some at a different window:
| Pick | File, slots, window, effort | Drafter now (before) | What decided it (against the old drafter) | Server VRAM against before, -c 81920, no projector meas |
|---|---|---|---|---|
| [1] | UD-IQ4_XS + vision, 1 slot, 219,136, xhigh (card at 122,880) | n3/p0 (n4/p0.75) | shallow reasoning 1.111 / 1.110, after the prefix 1.061 / 1.064 (intervals include 1), answers 1.120 / 1.121 (two blocks) | −150 MiB |
| [2] [4] | UD-IQ4_XS, text / + vision, 1 slot, 196,608, xhigh (menu row 2 at 180,224) | n3/p0 (n10/p0.5) | n3/p0 1.111 and 1.120 of n4/p0.75 on shallow reasoning and answers, n10/p0.5 0.961 and 1.049; after the prefix n10/p0.5 read higher, 1.079 against 1.061, intervals overlapping, and nothing past 57,540 tokens was measured, so at these windows the depth trend is open | −1,054 MiB |
| [3] | UD-IQ4_XS, text, 1 slot, 262,144 (menu row 3) | off (unchanged) | the full window does not fit with the drafter on; not swept | — |
| [5] [7] | Swift IQ3_XXS, text / + vision, 1 slot, 262,144, xhigh (cards [5] at 216,064, [7] at 195,584) | n3/p0 (n4/p0.75) | shallow reasoning 1.097, after the prefix 1.007 (interval includes 1), answers 1.054; n4/p0 1.034, n5/p0 0.940; n2 not run | not resolved −150 predicted |
| [6] [8] | Swift IQ3_XXS, text / + vision, 2 slots, 123,904 × 2, xhigh (cards [6] at 123,904 × 2, [8] at 93,184 × 2) | n4/p0 (n4/p0.75) | [8]’s shape, both slots busy: 1.337 per slot on reasoning, 1.484 on answers (n3/p0 1.185 and 1.233); second pass, one load per setting, no interval; [6] not run | same n-max |
| [9] | Swift IQ4_XS + vision, 1 slot, 139,264, xhigh (card at 128,000) | n3/p0 (n4/p0.75) | shallow reasoning 1.189 / 1.186, after the prefix 1.052 / 1.057 (the first interval includes 1), answers 1.062 / 1.064; n4/p0.25 tied, smaller n-max chosen | −152 MiB |
| [P] | Pi IQ4_XS + vision, 1 slot, 139,264, medium (card) | n4/p0 (n3/p0) | shallow reasoning 1.038 / 1.038, after the prefix 1.073 / 1.082 (ranges over 3 prompts); n4/p0.25 tied at the same n-max; n5/p0 0.970–0.971 | +154 MiB, the n-max of its deep fill |
No window was re-filled at the new drafters; the argument is memory only (the same or a smaller n-max everywhere but [P], whose window was filled at n-max 4; p-min free), and no block had an image in flight. The page’s cards follow the same rule: the UD-IQ4_XS, Swift IQ4_XS and one-slot Swift IQ3_XXS cards, menu row 2 and the August listing’s picks 1 and 2 now carry n3/p0, the two-slot Swift IQ3_XXS cards and the Pi card n4/p0; the Q4_K_M, UD-Q2_K_XL, 16 GB and non-3090 cards, the August listing’s picks 4–6 and Appendix A’s other drafter-on cells were not swept and, since 2026-10-09, carry n4/p0 der: the n-max 4 their memory was measured with, and p-min 0, which costs no memory and at n-max 4 read higher than 0.75 on shallow reasoning on all four files swept (0.75 read 0.87–0.97 of it with one slot, 0.90–1.03 after the prefix).
How it was measured. One RTX 3090 (driver 596.36), build 10502. Each block ran its pick’s launcher arguments plus --cache-ram 0 --metrics at -c 81920, one slot, no projector (two-slot block: 123,904 × 2 with the projector), --load-mode none, q8_0 KV, the launcher’s sampler (temperature 1.0, top_p 0.95, top_k 20, min_p 0), medium for the Pi file and xhigh for the rest. Per load: a 20 s settle, a discarded warm-up, the greedy probe, 6 reasoning prompts (an ISO-8601 parser with tests, a time-zone interval merge with tests, a two-train word problem, an S3-sync design in prose, an off-by-one fix as a diff, a red-black tree in JavaScript), 4 of them with thinking off, and, on the depth subset, a 57,540-token prefix of four repository source files, prefilled once, and 3 + 2 questions about it. Rate = Σ generated tokens ÷ Σ decode time. Two passes per block with different seeds (100+i, 200+i), pass 2 in reverse order. In the one-slot blocks, reference loads ran at the start, middle and end of each pass and each load was divided by its own session’s reference (a pass splits into sessions at gaps over 30 min; shallow reasoning reference spread within a session 0.10–0.66%, one Pi session holding a single reference load); the two-slot block had one load per setting per pass and no repeated reference. Paired bootstrap over prompts within a session (B 2,000, seed 42), on 6, 3, 4 and 2 prompts for the four columns: read both after-prefix intervals as ranges. On the Pi file 8–17% of the reasoning probes’ text was answer text, depending on the setting. Never set an absolute t/s from this sweep beside another sweep’s: this rig has two speed levels about 13% apart between server starts (§03, §06). Instruments: §15.
The rule, written at 2026-10-07 17:05 before any result: take each file’s highest mean of the shallow and deep reasoning ratios, if its answer ratio is no more than 3% below the incumbent’s; settings within 3% of the best tie and the smaller n-max wins; a setting is void when its two passes differ by more than 3%. The 3% pass rule was mis-specified. It assumed identical probes, but the passes use different seeds by design, so it mixes the sampled text’s path with the speed level: it voided settings while the reference timeline was flat (on the Pi file’s reasoning: n2, n5, n8, n10, n10/p0.25, n10/p0.5, n4/p0.75), while re-timing the same texts in a second session reproduced the shallow reasoning ratios within 0.4% (after the prefix, up to 3.4%: Swift IQ4_XS n4/p0, 1.063 against 1.027). It is reported, not used; paired intervals replaced it (01:24). Taken literally it picks n4/p0.75 for UD-IQ4_XS in the main block (it voids every rival that scored higher; n10/p0.5 passed the filter and scored 0.961), n4/p0.25 in the repeat block, and n2/p0 for Swift IQ4_XS on shallow reasoning alone. Like-for-like scoring was added at 01:57, after Swift’s first pass showed n2/p0 leading with no deep reading: only settings with both readings compete, and n2/p0 with depth joined the repeat blocks before they ran; it lost after the prefix (0.932, 0.979). For UD-IQ4_XS the launcher departs from the pre-registered tie-break. The repeat block’s 3% band (threshold 1.054) holds n2/p0 at 1.057, so smallest-n-wins picks it, though it lost to n3/p0 after the prefix (0.979 against 1.064), on answers (1.081 against 1.121) and on deep answers (0.954 against 1.071), winning only shallow reasoning (1.134 against 1.110). At 08:15, before the Swift IQ3_XXS blocks finished, the decision was set to each file’s highest score, n3/p0 at 1.087 here; for the Pi file and Swift IQ4_XS that is also the tie-break’s pick. Survival was read directly from the per-position /metrics counter, found at the 17:12 pilot, rather than by differencing kept tokens across the p0 series as pre-registered. Thirteen loads were excluded because the reference machine was in use (from 18:07 its idle time fell from 36,991 s to 3 s and desktop VRAM rose from 681 to 1,204 MiB; and the first Swift load), kept on disk and re-run from 00:24 behind a gate that waits for 180 s without input. One kept load, Pi n1/p0 in pass 2, also saw input; n1 is in no pick. The sweep paused 18:43–00:10 for another GPU task.
What this does not show. Anything past the 57,540-token prefix, so not the launcher’s 139,264–262,144 windows; an image in flight; other efforts (the Pi file at medium, and at xhigh only n3/p0 and n4/p0 in a later follow-up, Appendix B; 2026-10-07’s ×1.154 was also at xhigh; the rest only at xhigh); n2 on Swift IQ3_XXS, n2 or n5 with two slots, or p-min between 0 and 0.25 (n4/p0.25 tied n4/p0 on the Pi file and sat in Swift IQ4_XS’s tie band: p-min 0 was never beaten outside a tie band, which does not mean every p-min above 0 hurts); quality, agent wall time or energy at the new settings; other files, slot counts, cards or builds (§15). Reading older text: where a table or sentence measured before 2026-10-08 says “shipped default”, “the cards’ drafter” or “the recipe this page ships”, it means n4/p0.75 (n10/p0.5 on the August text-only rows), and those measurements keep their conditions.
Tuning the drafter: what to set n-max and p-min to, and why
2026-10-08: the answer to “what should I set” moved, and it is in the n and p subsection above: on the files that sweep ran, n-max 3 / p-min 0 with one slot, and n-max 4 / p-min 0 with two slots and for the Pi fine-tune at medium, measured at the cards’ own temperature-1.0 sampler, thinking on and off.C42C43 What follows is the August evidence, measured greedy, mostly on code and mostly with thinking off, and its findings hold for that regime: there n-max 10 / p-min 0.5 led — about +12% over the then-shipped n4/p0.75, 898 MiB extra (~20,900 tokens of window), and the cheapest setting in energy at 3.210 J per decode token against 3.743 at n4/p0.75 and 8.104 with no drafter — and on reasoning tokens at depth the wider drafter recovered only 5.6% (38.7 against 36.6 t/s at 91k). On the August greedy prose probe, --spec-type none was the better answer. One qualification: a cap-only sweep at the recipe §03 shipped then (2026-08-26) measured 74.13 t/s at n-max 4, 73.52 at n-max 6 (−0.8%) and 75.59 at n-max 8 (+2.0%) against a run-to-run band of 0.77% — nothing like the ~12% the matched sweep reports, so treat the wider drafter as a setting to measure on your own content rather than one to assume. The advice is in the “So what should you set?” box at the end of this part; the space between here and there is the evidence, because the measurements went wrong twice on the way and both mistakes are ones any careful person could repeat.
Two knobs control drafting behaviour, and both mattered more than the defaults suggest.
The comparison table above used the settings each method's authors recommend. Then the campaign swept --spec-draft-n-max across both methods, same prompt, temperature 0, warm cache — a different run state from the cold-start table, so absolute numbers shift a few t/s in either direction: this sweep's baseline reads 39.8 against the table's 40.3, and comparisons belong inside one chart, not across them. On a warm cache and a near-ideal benchmark prompt (baseline 39.8 against the cold-start table’s 40.3 — compare within one chart, not across them), the built-in head kept gaining past the recommended n-max 4, peaking at 81.7 t/s at n-max 10 / p-min 0.5, and DFlash2 peaked at n-max 4–5 (77.6 / 77.5) rather than its own recommended 7 (74.0):
Three findings, all measured 2026-08-21:
- On this prompt — and only on prompts that draft like it — the tuned built-in head won outright: n-max 10 with p-min 0.5 gave 81.7 t/s, +105% over baseline, beating DFlash2's best (77.6 at n-max 4–5). No drafter download and no source build — but not free of VRAM. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) isolated the built-in head with an on/off pair and measured 1,008 MiB fixed, plus 5,120 B per token of window — 11.4% of whatever
-cyou set — plus a further 898 MiB for n-max 10 over n-max 4. At-c 163840that is roughly 1.8 GiB, more than the vision projector costs. The head's own weights are only 334.7 MiB; the rest is draft-path allocation.--spec-type nonefrees every byte of it, and the server then logsmodel has unused tensor blk.64.* -- ignoringto prove the head was never loaded. --spec-draft-p-minis the hidden second knob. At n-max 10, lowering it from 0.75 to 0.5 was worth +8 t/s (73.5 → 81.7); at n-max 4 the difference was small. A high confidence floor keeps pausing exactly the long drafts that a big n-max exists for.- DFlash2's recommended n-max 7 was not its own best setting here — 4 and 5 both beat it (77.6 and 77.5 against 74.0). Every recommendation in this space, including this page's, is one machine's measurement.
The 2026-08-22 re-sweep ran 10 configurations on a temperature-0 code-generation probe with thinking left on, which nobody recorded at the time. At those conditions n-max 4 / p-min 0.75 measured 57.9 t/s (acceptance 0.81) and n-max 10 / p-min 0.5 measured 48.3 — a ranking the 2026-08-23 matched sweep reversed once thinking was switched off, because reasoning tokens draft short (case study below). Both sets of numbers stand; the difference is the token regime. DFlash2 on the same real code measured its best at 46.5 t/s at n-max 2 — below the tuned built-in head, the reverse of its benchmark-prompt 69–77, and hard to justify against the extra download and source build it costs.
The highest-acceptance configuration is not the fastest one. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) swept eight configurations and found the two metrics pulling in opposite directions — the highest-acceptance configuration is not the fastest one:
| Config | Acceptance | Drafted per pass | Accepted per target pass | Decode |
|---|---|---|---|---|
| n3 / p0 | 87.1% | 2.97 | 2.59 | 82.0 t/s |
| n4 / p0.75 | 89.9% | 2.94 | 2.65 | 81.0 t/s |
| n6 / p0.5 | 69.0% | 5.22 | 3.61 | 80.2 t/s |
| n10 / p0 | 46.8% | 9.83 | 4.60 | 87.3 t/s |
| n10 / p0.5 | 59.2% | 7.69 | 4.56 | 87.4 t/s |
| n10 / p0.75 | 72.3% | 4.59 | 3.32 | 80.3 t/s |
| n16 / p0.5 | 47.1% | 10.36 | 4.88 | 80.3 t/s |
Conditions: UD-IQ4_XS, -c 32768, -ngl 99, q8_0 KV, --parallel 1, temperature 0 / top-k 1, an 83-token novel-JavaScript prompt, 700 predicted tokens, a fresh server per configuration, reasoning tokens. accepted per target pass = draft_n_accepted ÷ (predicted_n − draft_n_accepted), both fields from the server's own timings.
Rank configurations by accepted tokens per target pass, not by acceptance rate. n-max 4 / p-min 0.75 accepts a near-perfect 89.9% and still loses, because a shallow draft caps how much a single verification pass can deliver. But the metric is a ridge, not a straight line: n16 accepts the most per pass (4.88) and is 8% slower than n10, because drafting deeper costs real time. The same reversal shows up at the extreme: on the verbatim-copy row, n-max 4 accepted a flawless 100.0% (424 of 424) and ran 97.0 t/s while n-max 10 accepted 99.0% and ran 148.7 — 3.96 accepted per pass against 10.54.
n-max 10 / p-min 0.5 is fastest on both files — 93.86 t/s on UD-IQ4_XS and 81.71 on Q4_K_M — and n-max 2 / p-min 0.75 is the slowest speculating configuration on both, with the whole ranking matching between them. Seven configurations × two files × two probes each, same server flags, same 149-token prompt asking for a novel sliding-window rate limiter (not a textbook algorithm — those inflate acceptance), thinking off, so all 700 timed tokens are the deliverable:
| Config | UD-IQ4_XS | acceptance | Q4_K_M | acceptance |
|---|---|---|---|---|
--spec-type none | 42.97 t/s | — | 39.99 t/s | — |
| n-max 2 / p-min 0.75 | 73.41 | 96.5% | 65.76 | 96.7% |
| n-max 3 / p-min 0.75 | 79.57 | 93.3% | 66.74 | 94.2% |
| n-max 4 / p-min 0.75 (the August shipped default) | 83.50 | 89.7% | 69.82 | 88.6% |
| n-max 6 / p-min 0.5 | 87.78 | 75.8% | 66.38 | 72.1% |
| n-max 10 / p-min 0.5 | 93.86 · 2.18× | 61.4% | 81.71 · 2.04× | 60.0% |
| n-max 10 / p-min 0.75 | 90.84 | 76.8% | 78.15 | 75.4% |
Conditions: -c 32768 -ngl 99 --parallel 1 -ctk q8_0 -ctv q8_0 --jinja, enable_thinking:false, temperature 0 / top-k 1, 700 predicted tokens, a fresh server per configuration (the speculative flags are load-time) plus a discarded warm-up, two probes per server, agreeing to about 1%. Answer lengths varied by about 1% between configurations despite temperature 0 — the usual CUDA batch-shape reduction-order non-determinism, with every stream in a file sharing its opening 120 characters. "Identical output regardless of speculative flags" is true in principle and not quite literally true on this build.
Same winner, same loser, same shape on both files — n10/p0.5 is fastest on each, n2/p0.75 is the slowest speculating configuration on each, and the ranking between them matches. There is no per-file optimum. Acceptance is a property of the draft head, not of the quantization: at six of the seven configurations the two files land within 1.6 points of each other (96.5/96.7, 93.3/94.2, 89.7/88.6, 61.4/60.0, 76.8/75.4), and the seventh — n6/p0.5 at 75.8/72.1 — is 3.7 points apart. This page reports that exception rather than inheriting the tidier claim. The quantization does not change how well the head predicts — it changes how fast the target model verifies. UD-IQ4_XS is faster at every configuration, and gains more from a wider drafter (+7.5% with no drafter, +19.6% at n4, +14.9% at n10). And the then-shipped default (n4/p0.75) leaves 11% on the table on UD-IQ4_XS and 15% on Q4_K_M for this workload. The one genuine difference is of degree, not kind: UD-IQ4_XS climbs steadily to n6 while Q4_K_M flattens and dips there (69.8 → 66.4), because its verification step is expensive enough that a wide, low-acceptance tree stops paying sooner.
Q4_K_M at n-max 10 / p-min 0.5 measures 81.71 t/s on a deliberately novel code prompt with thinking off (2026-08-23 matched sweep). The same flags on the same file measure 48.3 t/s with thinking on, because reasoning tokens draft short — so the same configuration’s speed depends entirely on which token stream is being timed. The two decimal places are a reproduction match between two independent runs, not a claim that the level is stable to 0.01 t/s.
The mechanism: the 2026-08-22 re-sweep left thinking on, so it timed the model reasoning about code while its label said code generation, and reasoning tokens draft short (the mean-draft-length mechanism above). Switch the regime off and the figure reproduces to two decimal places. Three lessons, and the third is the expensive one. A speed number without its token regime cannot be checked — §10’s checklist line was added for exactly this. A disproof is a measurement and needs its conditions stated as carefully as the claim it overturns; this one was published as a correction, propagated into a default, and stood for a day. And a result that refuses to reproduce is a lead, not a verdict — the discrepancy was pointing at a real mechanism the whole time.
On the four files the 2026-10-08 sweep ran: n-max 3 / p-min 0 with one slot on UD-IQ4_XS, Swift IQ4_XS and Swift IQ3_XXS at xhigh; n-max 4 / p-min 0 with two slots busy (Swift IQ3_XXS) and on the Pi fine-tune at medium — measured at the cards’ own temperature-1.0 sampler, thinking on and off, shallow and after a 57,540-token prefix; the per-pick table is in the n and p subsection.C42C44 On shallow reasoning tokens against n4/p0.75 in the same session, n3/p0 read 1.111 on UD-IQ4_XS, 1.189 on Swift IQ4_XS and 1.097 on Swift IQ3_XXS; on answer tokens 1.120, 1.062 and 1.054; after the prefix 1.061, 1.052 and 1.007 on reasoning tokens, intervals that include 1. With two slots busy n4/p0 read 1.337 per slot (second pass only, one load per setting, no interval). It is also the lightest of the winning settings: n-max 3 needs about 150 MiB less than 4 and about 1 GB less than 10, and p-min costs nothing. Files, cards and slot counts the sweep did not run take n4/p0 der since 2026-10-09, keeping the n-max 4 their windows were fitted at, for the reason that follows; p-min 0 is measured on four files and derived for the rest. p-min 0.75 is not a safe default: at n-max 4, p-min 0 read higher on shallow reasoning tokens on every file (on UD-IQ4_XS 1.049, an interval of 0.995–1.121 that includes 1), and p-min saves no memory; after the prefix the gap shrank or vanished, and on UD-IQ4_XS deep answers p-min 0.75 was faster (n4/p0 0.973 of it, two prompts, both blocks).C43 n-max 10 / p-min 0.5 is the August greedy-code peak: on greedy code with thinking off the matched sweep above put it 12% over n4/p0.75, and there it is the cheapest measured setting in energy, 3.210 J per decode token against 3.743 at n4/p0.75 and 8.104 with no drafter (energy at n3/p0 and n4/p0 was not measured); even on greedy code the October probe ranked n4–n5 at p-min 0 above it (the greedy exception). At the cards’ sampler it read 0.961 and 0.935 of n4/p0.75 on shallow reasoning tokens on UD-IQ4_XS and Swift IQ4_XS, where n3/p0 read 1.111 and 1.189; after the prefix it read higher than n3/p0 on UD-IQ4_XS (1.079 against 1.061) and on the Pi file (1.021 against 1), intervals overlapping, and nothing deeper than 57,540 tokens was measured, so on the long-window picks that direction is open. The 2026-10-07 sampled run on the two fine-tunes’ IQ4_XS files (n3/p0 ×1.153 and ×1.154 over n4/p0.75 at xhigh, -c 32768) pointed the same way (Appendix B). On the August greedy English-prose probe neither flag mattered much — n-max 4 won by 10% (48.4 against 43.8) over a drafter worth only 1.16× there, so --spec-type none and the 1.8 GiB back was the better answer for that content; the 2026-10-08 sweep did not time prose on its own. One later August sweep, at the then-shipped recipe, could not reproduce the n-max gain on greedy code; it is printed immediately below.
The matched sweep above raised --spec-draft-n-max and got about 12% more throughput. A later sweep raised the same flag and could not find that gain. That sweep changed only that flag, at the recipe §03 shipped then, on a machine with nothing else running: n-max 4 measured 74.13 t/s, n-max 6 measured 73.52 (−0.8%), and n-max 8 measured 75.59 (+2.0%) meas. Read those against this rig’s run-to-run band of 0.77% — the spread the same greedy configuration shows when it is measured twice with nothing changed at all (§01). The three do not rise with the cap. Going from 4 to 6 makes it slower by 0.8%, which is the band itself; going on to 8 gains 2.0%, which is a little over twice the band and therefore real, but it is nothing like the roughly 12% the matched sweep above reported for a wider drafter. On this probe the cap is not a throughput lever — not because raising it does nothing measurable, but because what it does is small, is not monotonic, and is the wrong size to justify widening the drafter — which cost a measured 898 MiB at n-max 10 over n-max 4 on this machine (§05).
The mechanism moved and the clock did not follow, which is the part worth keeping. Mean accepted length — the average number of drafted tokens that survive each full-model check — rose from 2.71 to 3.54 across that same sweep. The drafter really did run deeper, really did get more tokens through per check, and the throughput stayed where it was. So a rising accepted length is evidence that a flag is doing something; it is not evidence that the flag is making you faster.
Two measurements disagree, and this page prints both. They are not the same experiment. The matched sweep ran seven configurations with a fresh server for each, on a 149-token novel-code prompt at -c 32768 with thinking off, and its headline pair moved --spec-draft-p-min as well as n-max — though it also holds a pair in which only n-max moves, 83.50 at n-max 4 against 90.84 at n-max 10, both at p-min 0.75. The later sweep held everything else still, changed the one flag, ran at the then-shipped recipe — but at the same -c 32768, so the window cannot account for anything here: draft-nmax-sweep.py sets -c 32768 in its base flags and only --spec-draft-n-max varies.C18 What actually differs is the prompt and the token budget — a Python priority-queue task at 400 predicted tokens against a JavaScript task at 700 — and this page’s own anti-overfit control (§06) measured speculation to be strongly content-dependent while unspeculated decode is flat to 0.8% across the same four contents. Different content is the live explanation. The later sweep stopped at n-max 8 rather than 10, so it never tested the older sweep’s winning value directly. Neither has been shown to be wrong, and the older one is not being retired. The practical reading for someone copying a recipe: the wider drafter is worth measuring on your own content and your own prompt, not worth assuming — and if you do measure it, measure your own baseline twice first, because a difference of a few percent is inside what this rig moves on its own.
Every drafting number this page measured through the 2026-08-23 round sits on a single line; two later pairs do not (below). What the 2026-08-23 round changes is what that line runs along. Draft acceptance was the obvious candidate, and within one configuration it works: same flags, different content, 0.47 → 51.5 t/s, 0.93 → 119.8, and with thinking off 0.99 → 148.7. Across configurations it inverts. The highest-acceptance configuration in the matched sweep — n2/p0.75 at 96.5% on both files — is the slowest speculating one on both; and at 91k depth two streams with statistically identical acceptance (0.895 and 0.907) run 1.69× apart. The quantity that survives both tests is the product: how many tokens the drafter dares to propose × how many of them survive = accepted tokens per verification pass. Stated as a rule: acceptance tells you whether the drafting was right; mean draft length tells you how fast you go; only their product ranks a configuration. Two later pairs break that ranking. On 2026-08-28, n10/p0.5 carried more accepted tokens per verify pass than n4/p0.00 (about 2.08 against 1.98, mean draft length times acceptance, derived) and the shorter drafts won by 8.4% (C24); on the Swift files with thinking on (one prompt, greedy, 2026-10-04), n10/p0.5 keeps more than n4/p0.75 — 1.657 against 1.479 on Swift IQ4_XS, 1.389 against 0.974 on Swift IQ3_XXS — and decodes slower, 47.51 against 59.96 and 43.95 against 50.61 t/s (§08). On 2026-10-08, at the cards’ temperature-1.0 sampler, the n/p sweep broke it on all three files that ran n10/p0: that setting kept the most reasoning tokens per check and decoded slower than n2, n3 and n4 at p-min 0 (the Pi file: 3.133 kept per check against n3/p0’s 1.932, at 0.877 of its speed), because each guess costs a head pass and a row in the check whether it survives or not (§06). The old sweep was never wrong as a measurement — it was wrong as a promise, because it reported one point on this curve as if it were the curve.
The five-arm drafter sweep was re-run with the arms in reverse order, because it starts a fresh server per arm and runs them back to back, so arm position is confounded (mixed up) with arm setting. Three of the arms that change a single flag against arm A:
| Arm | Forward | Reverse | vs A forward | vs A reverse |
|---|---|---|---|---|
| B, unquantised f16 KV | 69.56 | 61.88 | −8.9% | −8.2% |
C, -c 180224 | 73.94 | 64.26 | −3.1% | −4.7% |
| E, n-max 4 / p-min 0.00 | 82.73 | 76.32 | +8.4% | +13.2% |
Every relationship holds its sign and roughly its size. The sweep’s conclusions stand: q8_0 KV is faster than f16 for speed as well as memory, and n4/p0.00 remains a genuine open lead. Followed up 2026-10-08 at the cards’ sampler with thinking on (above).
But arm A itself read 76.32 forward and 67.41 reverse, −11.7%, and A ran first in one sweep and last in the other.
The rule this gives a reader, and it is the practically useful part: compare arms inside a single sweep; never compare a number from one sweep against a number from another. Ratios within a sweep are stable. Absolute numbers across sweeps are not.
Published speculative-decoding tables — like the DFlash2 blog's "mean acceptance length" — are speed numbers, not quality scores. Acceptance figures measure how many tokens got through per check, not whether the output text is correct. On one prompt (build 10502, greedy, thinking off, -c 32768, 2026-10-04) greedy text with the drafter on was not always identical to drafter-off text: the anchor diverged at byte 269 at n4/p0.75 and at byte 21 at n10/p0.5, and its own drafter-on runs did not always repeat; Swift IQ4_XS produced byte-identical text across all three settings. Suite scores with the drafter on are not measured for any of the three files (§08, §15).C38 For one file, speed changes with the drafter; across files it changes with the file too (§06).
Swift files against the anchor in the same sweep: speed at depth, the drafter-off floor and the drafter on one prompt
In the answer regime with the n4/p0.75 drafter the cards then shipped, Swift IQ4_XS decodes faster than the anchor at 1.5k and 28k of depth and about level at 91k, and at the same speed with the drafter off; Swift IQ3_XXS matches the anchor at 1.5k and 28k and falls below it at 91k. On one prompt with thinking on, both Swift files decode slower than the anchor with the drafter (below). Measured 2026-10-04 in one alternating sweep, greedy (temperature 0), -c 131072, q8_0 KV, the first probe after each prefill discarded, two passes n=2: with the drafter the cards shipped until 2026-10-08 (n4/p0.75, thinking off), Swift IQ4_XS ran 87.2 t/s at 1.5k depth (×1.102 against the anchor’s 79.1), 81.4 t/s at 28k (×1.116 against 73.0) and 61.1 t/s at 91k (×1.018 against 60.1) meas, ratios der — in step with its higher drafter acceptance (0.928 against 0.892 at 1.5k, 0.943 against 0.927 at 28k, 0.935 against 0.924 at 91k). Swift IQ3_XXS ran 79.3 t/s at 1.5k (×1.002), 73.2 t/s at 28k (×1.003) and 56.3 t/s at 91k (×0.938 against 60.1). With the drafter off at 1.5k depth, Swift IQ4_XS gave 42.9 t/s (×1.005 against the anchor’s 42.7) and Swift IQ3_XXS gave 44.1 t/s (×1.034). The anchor’s own speed in this sweep sat at the lower of this machine’s two levels (§03): 0.917, 0.91 and 0.928 of its August speed at 1.5k, 28k and 91k der; the absolutes behind those ratios are in the provenance chain (§17). The reasoning regime at depth is void: its two passes disagreed by more than 3% in three of its six Swift cells, and its first pass ran inside a download overlap on the test host; a re-run on 2026-10-05 failed the same agreement check in five of its six Swift cells, so no reasoning-regime ratio is published (§15). The same overlap covered the drafter-off group’s first pass and the start of the answer regime’s second pass; both groups’ passes agree within 3%. The one-prompt drafter measurements below, thinking on included, are a separate step and are not void. All speed bands in this sub-section are greedy-only; speed at the temperature-1.0 sampler the cards ship is not measured as a band (§15).
results/swift-1.5-qwen3.8-27b/arms/speed-probes.json in the measured-inference repository, passes 1,2), greedy (temperature 0), -c 131072, q8_0 KV, the first probe after each prefill discarded. Depth is the prompt the prefill probe read. Bars are means over kept probes (n per bar in the table), whiskers the min–max. Every Swift arm read the same prompt tokens as the anchor, and no thinking-off arm produced reasoning. Ratios to the anchor are within this sweep only, and printed only where the two passes agree within 3% of their mean. The reasoning-regime group is not plotted: it is void, because its two passes disagreed by more than 3%, and its first pass ran inside a download overlap on the test host; a re-run on 2026-10-05 was void as well (§15).| Regime | Depth, tokens | File | Decode t/s meas | n | min–max | vs anchor der | Acceptance | Accepted / pass |
|---|---|---|---|---|---|---|---|---|
| drafter n4/p0.75 · thinking off | 1,458 | UD-IQ4_XS (anchor) | 79.1 | 6 | 77.0–80.9 | — | 0.892 | 2.96 |
| drafter n4/p0.75 · thinking off | 1,458 | Swift IQ4_XS | 87.2 | 6 | 85.4–89.4 | ×1.102 | 0.928 | 3.21 |
| drafter n4/p0.75 · thinking off | 1,458 | Swift IQ3_XXS | 79.3 | 6 | 76.8–80.8 | ×1.002 | 0.891 | 3.05 |
| drafter n4/p0.75 · thinking off | 28,388 | UD-IQ4_XS (anchor) | 73.0 | 6 | 70.8–74.8 | — | 0.927 | 3.12 |
| drafter n4/p0.75 · thinking off | 28,388 | Swift IQ4_XS | 81.4 | 6 | 79.3–83.0 | ×1.116 | 0.943 | 3.41 |
| drafter n4/p0.75 · thinking off | 28,388 | Swift IQ3_XXS | 73.2 | 6 | 69.9–76.2 | ×1.003 | 0.908 | 3.10 |
| drafter n4/p0.75 · thinking off | 90,854 | UD-IQ4_XS (anchor) | 60.1 | 6 | 59.2–60.6 | — | 0.924 | 3.21 |
| drafter n4/p0.75 · thinking off | 90,854 | Swift IQ4_XS | 61.1 | 6 | 60.4–61.9 | ×1.018 | 0.935 | 3.24 |
| drafter n4/p0.75 · thinking off | 90,854 | Swift IQ3_XXS | 56.3 | 6 | 55.5–56.8 | ×0.938 | 0.908 | 3.04 |
| drafter off · thinking off | 1,458 | UD-IQ4_XS (anchor) | 42.7 | 6 | 42.4–42.8 | — | — | — |
| drafter off · thinking off | 1,458 | Swift IQ4_XS | 42.9 | 6 | 42.4–43.2 | ×1.005 | — | — |
| drafter off · thinking off | 1,458 | Swift IQ3_XXS | 44.1 | 6 | 43.5–44.6 | ×1.034 | — | — |
One sweep, 2026-10-04; greedy (temperature 0); -c 131072; q8_0 KV; two passes, the first probe after each prefill discarded; ratios printed only where the two passes agree within 3% of their mean. Not comparable with the August tables above in this section.
The size-derived prediction missed for both Swift files: bytes read per token alone did not predict the drafter-off decode speed of these files.
| File | Bytes read per decoded token der | Predicted ratio der | Ratio from the drafter-off probes der | Status |
|---|---|---|---|---|
| Anchor (unsloth UD-IQ4_XS) | 13,344,536,576 | — | — | baseline |
| Swift IQ4_XS | 14,550,542,336 | 0.917 | 1.005 | prediction missed |
| Swift IQ3_XXS | 11,355,027,456 | 1.175 | 1.034 | prediction missed |
Predicted ratios are decode-speed ratios against the anchor, derived before the probes ran from file size minus token_embd, header and MTP head (der; token_embd assumed host-resident). Measured ratios are the drafter-off probes at 1.5k depth, 2026-10-04, greedy, -c 131072 n=2. The data show no cause for the discrepancy.
On one prompt at -c 32768, greedy, all three files draft well with thinking off (acceptance 0.917–0.927 at n4/p0.75); with thinking on, the Swift files keep far fewer drafted tokens per verify pass than the anchor (1.479 and 0.974 against 2.684 at n4/p0.75), and for them n10/p0.5 is slower than n4/p0.75 (47.51 against 59.96 and 43.95 against 50.61 t/s). meas n=4 per cell. The board sat at its software power cap in at least 96.8% of the active samples in every window meas.
The full table for this prompt — decode, acceptance and accepted tokens per verify pass for every file, drafter setting and thinking mode — is in §08, with its conditions: one prompt (a JavaScript red-black tree) sent three times per load, the first discarded; two passes, pass 2 repeating pass 1’s text; -c 32768; greedy; 700 tokens per request; 2026-10-04; the power-cap share is over the 30 windows of that step.
With thinking off, greedy output on the one prompt was identical across drafter settings (off, n4/p0.75, n10/p0.5) for Swift IQ4_XS; it differed for the anchor (unsloth UD-IQ4_XS) and for Swift IQ3_XXS, whose drafter-on runs did not always repeat their own text (§08). Two checks on the speed probes hold: prompt_n was equal to the anchor’s at every depth (1,458 at 1.5k, 28,388 at 28k, 90,854 at 91k on the answer arms; 1,498, 28,428 and 90,894 on the reasoning arms; 1,458 on the drafter-off arms), and none of the thinking-off arms produced reasoning tokens.
Only one row of this section was measured: the RTX 3090. Every other card's number is the same arithmetic from §04 — bandwidth divided by file size, times the efficiency constant for that file's format — applied to the card's published bandwidth. That arithmetic predicted the measured 3090 row correctly, which is the only reason to trust it elsewhere, and it is still a prediction. The pattern that matters more than any individual number: how much memory a card has decides which strategy you can use; how fast that memory is decides how long you wait. A card with enormous bandwidth and too little memory loses to a slower card that fits the model, because the moment one layer lands on the processor the whole calculation changes to the processor's memory speed.
The figures below are for the reference run: about 2,000 tokens of prompt and about 100,000 tokens of thinking and output, with each card using the settings that keep quality highest. All of them exclude speculative decoding, which varies per system but helps a lot: on the reference 3090 it moved this run to roughly 34 minutes from 45. Treat the drafter as moving a card up this chart — furthest on predictable content — rather than as a new baseline.
| Card | VRAM | Bandwidth | Strategy | Decode | 100k run |
|---|---|---|---|---|---|
| RTX 5090 | 32 GB | 1792 GB/s | UD-Q4_K_XL, all on GPU, 262k window — or NVFP4 via vLLM | ~65–80 t/s der | ~25 min |
| RTX 4090 | 24 GB | 1008 GB/s | Q4_K_M, all on GPU, ~122k window | ~40–45 t/s der | ~40 min |
| RTX 3090 | 24 GB | 936 GB/s | Q4_K_M, all on GPU, 122,880 window | 40 t/s meas | 45 min meas 100k ÷ 40 t/s is 41.7 min of decode; the rest is prefill and turnaround |
| RTX 5080 / 4080 / 4070 Ti S / 5060 Ti | 16 GB | 448–960 GB/s | 4-bit partial offload (fidelity) or UD-Q2_K_XL resident at -c 65536 (speed) — 13,982 MiB meas. UD-Q3_K_XL does not fitC15 | ~8–12 / 25–50 t/s der | ~3 h / ~45 min† |
| RTX 3060 / 5070 | 12 GB | 360–672 GB/s | heavy offload — better served by a smaller model; for the Swift-1.5 files' 12 GB windows see Appendix A | ~6–8 t/s der | ~4–5 h |
| Intel Arc Pro B70 | 32 GB | 608 GB/s | Q5/Q6 GGUF (Vulkan) or OpenVINO int4; 229k with the drafter, full 262k with --spec-type none (§03) | ~18–26 t/s der | ~1.1–1.5 h |
| Intel Arc Pro B50 | 16 GB | 224 GB/s | UD-Q2_K_XL resident at -c 65536 — 13,982 MiB meas, 606 MiB spare; or OpenVINO int4. UD-Q3_K_XL does not fitC15; 4-bit does not fit either | ~8–12 t/s der | ~2.5–3.5 h† |
| Intel Arc B580 | 12 GB | 456 GB/s | even 3-bit does not fit — heavy offload (Vulkan) or 2-bit; better served by a 14B model. The 2-bit option is measured: the ladder measured every file down to 1.8 bits, and quality turns sharply below 2.91 bits per weight — a 2-bit file costs about 21% perplexity against this page's 4-bit default, five times the price per gigabyte of any step above it; the Swift-1.5 files' 12 GB windows in Appendix A are for CUDA cards and do not cover Arc | ~6–8 t/s der | ~4–4.5 h |
| DGX Spark (GB10) | 128 GB unified | 273 GB/s | Q4_K_M on llama-server (§03's recipe) · NVFP4 via vLLM for maximum speed · Q8_0 for reference quality | ~10–15 t/s der | ~2–2.8 h |
| Intel Arc B390-class iGPU (Core Ultra) | shared, ≤96 GB | 153.6 GB/s (LPDDR5X-9600, 2ch) | Vulkan Q4_K_M on llama-server (§03's recipe) or OpenVINO int4; effort ≤ medium | ~5–7 t/s der | ~4.5–5.5 h |
† the 3-bit rows trade quality for that speed — a different model for benchmark purposes (§08). They also trade window: on 16 GiB the resident 3-bit file caps context near 49,152 tokens by §05's budget arithmetic, so a single run longer than 100k tokens must either checkpoint across windows or take the offload path.
Every row here is bandwidth ÷ file size × the format constant, so the bandwidth figure carries the whole prediction. The independent re-measurement verified each one against primary vendor documents and found three ways the public numbers mislead. Precision that has no source. The RTX 3090 is 936 GB/s — NVIDIA's own GA102 whitepaper figure, 384-bit × 19.5 Gbps. The widely-copied "936.2" is a third-party derivation from an unrounded memory clock: not wrong, but false precision with nothing behind it. Marketing bandwidth that is not bandwidth. The RTX 4060 Ti's much-quoted "554 GB/s" is NVIDIA's effective-bandwidth argument about the card's 32 MB L2 cache; its peak is 288 GB/s, and only peak belongs in a bandwidth column — feed 554 into the formula and you will promise a reader roughly double the speed the card delivers. Silicon ceilings quoted as machine figures. Apple's M4 Max is 410 or 546 GB/s depending on the bin, so "up to 546" describes the part number, not the laptop in front of you. And a sourcing note for anyone re-checking this table: TechPowerUp and AnandTech both sit behind bot challenges, and a naive scrape of TechPowerUp returns a perfectly valid-looking page with zero specification content in it — which is exactly how a wrong number gets a citation.
System-memory assumption for every offload and shared-memory row: those t/s figures are derived at dual-channel bandwidth — two sticks: desktop DDR5-5600 ≈ 90 GB/s, the integrated-GPU row's LPDDR5X-9600 ≈ 153.6 GB/s. A single stick runs single-channel at half that bandwidth — halve the t/s and double the hours, and slower DDR5 grades scale down proportionally. If your machine reads slower than this table, check that assumption before blaming the table: Task Manager → Performance → Memory shows your speed and slots used, and §04's formula predicts your number — your GB/s ÷ your file's GB × 0.70 for a K-quant or 0.65 for an IQ-quant. Read the family off the filename: a name containing IQ (UD-IQ4_XS, UD-IQ3_XXS, UD-IQ2_S) is an IQ-quant and takes 0.65; a name containing K without IQ (Q4_K_M, UD-Q3_K_XL, UD-Q2_K_XL, Q6_K) is a K-quant and takes 0.70. Rows where the model is resident on a discrete card are unaffected; their bandwidth is the card's own.
On a 24 GB card, download UD-IQ4_XS from unsloth for everyday use, or Q4_K_M from lmstudio-community if you want the highest measured quality and can accept that it leaves only about 0.3 GiB of the card for your desktop. Two files of the same size from different quantizers really can differ in quality, so the choice is not arbitrary — but between these two the measured difference is 0.9% of perplexity, which is smaller than the uncertainty of the measurement itself, while the size difference is 2.1 GiB and buys you either the vision projector or about 50,000 more tokens of window. If you have a 16 GB card, take UD-Q2_K_XL — the memory each file and window needs is measured, and it is the 2.9-bit file, not the 3-bit one, that fits a 65,536-token window with its drafter still switched on (the requirement table below); if you have 32 GB, Q6_K; and if you have a Blackwell card and can run vLLM, the NVFP4 build is the one that runs natively there. If you have 12 GB, the answer is no — not a smaller quantization of this model, but a smaller model. The 24 GB answer depends on how big a window you want, and there are three files rather than one. Up to about 98,000 tokens, stay on UD-IQ4_XS. At about 131,000, UD-Q3_K_XL decodes 20% faster on this page's fill and ties the reference file on accuracy. At the full native 262,144, UD-Q2_K_XL is the faster one and the only one that can still speculate. All three are measured; the ladder that settles the quality half starts at the eight-rung table below, the memory and speed each configuration needs is in the requirement table, and the card-by-card verdicts are in the last part of this section. On the 2-bit question there is a coding-benchmark answer: on 225 programming exercises the 2-bit file finishes with exactly the same score as the 4-bit one and pays 20 to 45% more tokens, 20 to 36% more time and about 32% more energy for every exercise it solves (the coding-benchmark part below). Take the 2-bit file when you need the 4.4 GB of VRAM it saves; otherwise take the 4-bit one.
Stop at UD-Q2_K_XL — 2.912 bits per weight, 9.154 GiB. It ties the 4-bit reference on 75 paired benchmark items (one discordant answer, p=1.00), returns zero empty answers, and costs +6.07% of perplexity. 2.9 bits per weight is not the point after which quality starts to degrade — it is the last size that still performs like the full-quality file.
Four instruments now sit beside every rung, and they have different sensitivities. Perplexity falls with every step from the very top, so it never had a flat stretch that could mark a beginning. Task accuracy is flat from 4.22 bits per weight down through 2.91 — the same 75 questions score 97.30, 96.00, 93.30 and 96.00 — and one rung further down, at 2.48, it is still a statistical tie with the 4-bit file. Empty answers, the failure a reader actually notices, number exactly zero at every rung down to and including 2.91, then rise 2, 3, 5, 28 below it (the full table). Take the second and third together and 2.9 bits is where "still behaves like the big file" ends, not where quality begins to fall.
Where it stops working is lower still, and it fails in a way no quality curve predicted. The smallest file measured, at 1.835 bits per weight, does not merely answer badly: it loses the ability to stop. Twenty of its 75 answers ran into the token cap, its median answer is 932 tokens against 424 at the top of the ladder, and one arm took two and a half hours where the next longest took thirty-six minutes.
The fourth instrument settles where “stops working” actually begins, and it is two rungs higher than that. meas On 2026-08-25 the JavaScript each file had written was rebuilt and run. Everything down to 2.481 bits per weight executes and prints its answer; at 2.153 and below none of it does — one throws a TypeError, two do not even parse — while the prose around the code still reads perfectly well (the execute probe). The working floor is at 2.481, not 1.835C16 — the earlier estimate used detectors that only read output. Note what this does to the four answers: the two instruments that examine text disagree with each other, and the two that ask whether the model works — a paired accuracy test and a parser, sharing no machinery at all — break at the same rung. Four instruments, and the agreement is as informative as the disagreement.
This is the summary. The eight-rung table and the paired tests behind the word "tie" are in the part below, the empty-answer audit in full is in the one after it, what each file and window needs in memory is in the requirement table, which file is fastest at which window is in the part after that, and what to download for a 24, 16 or 12 GB card is in the last part of this section. §15 still lists what is measured and what is not — including the two things this ladder does not establish: speed and fit on any card that is not this 24 GB 3090.
| Size class | Best file | Weights (GB · GiB) | Notes |
|---|---|---|---|
| 4-bit GGUF (the default class — the comparison at the end of this section splits Q4_K_M against UD-IQ4_XS) | Q4_K_M · lmstudio-community (or UD-Q4_K_XL) | 16.5 GB 15.4 GiB | settled by measurement 2026-08-22: wikitext-2 perplexity (the table below) came out 2.3% in Q4_K_M's favour against UD-Q4_K_XL, and a 200-question GSM8K comparison agreed within noise — though that GSM8K comparison is now unaudited, see the warning below. Unsloth's claim that the XL file stays about 10% closer to the full-precision model's own next-token probabilities is measured on their own calibration text and is relative, worth 1–3 true accuracy points at most. With quality tied, the 1 GiB decides: Q4_K_M's extra gigabyte of slack holds about 23,800 more resident context tokens at the measured slope, or the vision projector (§03). bartowski’s Swift-1.5 IQ4_XS is measured in this size class below. |
| 3-bit GGUF — and, on 24 GB, the pick at about a 131,072-token window | UD-Q3_K_XL · unsloth | 13.1 GB 12.2 GiB | Measured here as of 2026-08-24: perplexity 6.7691 ± 0.047, which is +2.63% against the 4-bit UD-IQ4_XS file this page ships — the cheapest quality step down on the whole ladder, at 1.03 GiB saved — and it ties that file on 75 paired benchmark items, p=1.00. Every functional detector passes (§02). Measured 2026-08-25: at -c 131072, deep-filled with prose, it decodes 41.74 t/s against UD-IQ4_XS's 34.68 (both at n4/p0.75, where the p-min gate is part of the ordering; at p-min 0 it is not measured) and needs 19,724 MiB against 20,848, which makes it the 24 GB pick at that window (§08). Qwen's own AD-IQ3_S reportedly scores slightly better (next-token probability data) but you have to request access to download it, and it is not measured here. On a 16 GB card, what it needs is now measured and what it would do is not: at -c 65536 with the drafter it asks for 16,906 MiB, more than such a card has in total (§08) — the memory figure transfers to any card, the speed figure never can. bartowski’s Swift-1.5 IQ3_XXS is measured in this size class below. |
| 2-bit GGUF — the 16 GB pick, and on 24 GB the full-window pick | UD-Q2_K_XL · unsloth | 9.83 GB 9.154 GiB | Measured here 2026-08-24 and 2026-08-25: at 2.912 real bits per weight it ties the 4-bit reference file on 75 paired benchmark items — they answered exactly one of the 75 differently, p=1.00 (§02) — and returns zero empty answers, for +6.07% of perplexity. It is the smallest file on the ladder that still performs like the full-quality one. On a 24 GB card it is not the daily pick — the drafter makes UD-IQ4_XS 12.9% faster at a short window (n10/p0.5, greedy code) — but at the full 262,144-token window it is 34% faster (its drafter at n4/p0.75) and the only one of the two that can still speculate (§08). On a 16 GB card its requirement is measured: 13,982 MiB at -c 65536 with the drafter on, which fits inside such a card with 606 MiB left after this page's desktop reserve (§08). How fast it would run there is not measured and cannot be measured on this machine. And one limit that matters more than speed: almost everything on this page was measured on single-turn prompts of at most 16,384 tokens. Nothing here has tested a long agentic loop where a model calls tools dozens of times in a row. An independent tester (§08) reports this file completing a web app in one shot but locking into an infinite tool-calling loop on a game-engine task. A coding benchmark measures part of that gap (2026-08-26): on a 225-exercise coding benchmark whose harness shows the model its own failing unit tests and lets it try again, this file solves exactly as many exercises as the 4-bit reference — 96 of 225 each — and spends 20 to 45% more tokens, 20 to 36% more wall-clock time and about 32% more energy per solved exercise doing it (§08). That is a two-attempt repair loop and not a fifty-turn tool-calling loop, so the long-agentic caution below applies in full. |
| 6-bit GGUF (quality, 32 GB cards) | Q6_K · lmstudio-community | 22.4 GB 20.9 GiB | wins its size class; near-8-bit quality. Unmeasured here |
| 8-bit GGUF (reference) | Q8_0 · lmstudio-community | 29 GB 27.0 GiB | the reference-quality choice, for machines that can hold it (DGX Spark; 32 GB cards at short context). Unmeasured here |
| NVFP4 (native on Blackwell) | Qwen3.8-27B-NVFP4 · unsloth, via vLLM | ~15 GB ~14.0 GiB | for Blackwell cards (RTX 50 / GB10 / B200): about 1.5× the speed of the unquantized model while keeping 92–97% of its accuracy, per the vendor. Runs only in vLLM — its FP8 output layer is not supported by llama.cpp or SGLang. cited, unmeasured here |
| NVFP4 GGUF (any GPU — measured) | NVFP4-MTP GGUF tiers · community, via llama.cpp | 14.8–24.9 GB 13.8–23.2 GiB | The model card claims these need a Blackwell GPU. That claim was tested on an RTX 3090 and it is wrong — llama.cpp falls back to software decoding and the files run. The catch: on older cards you keep the small file and lose the speed. Prompt processing is about 20% slower than Q4_K_M (1058 against 1288 t/s on the 08-21 build), speculative decoding is slower too (54.6/48.2 t/s for VERY-LOW/HIGH against Q4_K_M's 57.9 with the same head), and the energy is worse: 9.293 J per decode token against UD-IQ4_XS's 8.198, +13.4%. The one reason to pick NVFP4 on an older card: the 13.8 GiB VERY-LOW file is about 1.6 GiB smaller than Q4_K_M, which buys roughly 38,000 extra tokens of window at the measured slope. One quirk: llama-bench reports these files as "Q8_0" — a metadata bug, so tell them apart by file size |
| OpenVINO int4 (Intel Arc / iGPU) | Qwen3.8-27B-int4-ov · Intel official | ~16 GB ~14.9 GiB | INT4_ASYM g128; first-party Intel runtime, OpenAI-compatible server; experimental — validate long runs. cited, unmeasured here |
| OpenVINO int8 (32 GB Arc quality play) | Qwen3.8-27B-int8-ov · Intel official | ~28 GB ~26.1 GiB | Intel's equivalent of the 6-bit/8-bit quality tier; needs a 32 GB card. cited, unmeasured here |
Units, because they bite: download pages list decimal GB while VRAM budgets are binary GiB, and the gap is about 7%. The 17.6 GB UD-Q4_K_XL is 16.39 GiB resident, which is why "17.6 GB of weights + 4.25 GiB of cache" still fits inside 24 GiB (§05). Every GiB figure in the table above is the GB figure divided by 1.073741824.
For a benchmark, use the same quantization on every machine — a different quantization is a different model and its scores are not comparable. For daily use, pick the best quantization that fits entirely in your VRAM, and remember that the file you pick is also an energy decision: the three files measured here span 8.198 to 9.293 J per decode token, a 13% range, because a bigger or slower-to-decode file spends more board-seconds per token.
Does FP4 beat INT4? A scored smoke test
A common belief says floating-point number formats always beat integer formats at the same bit width. It did not survive the test here: all three 4-bit files scored the same on one benchmark, and on the other the integer-grid file answered two more questions than either floating-point file. Twenty questions cannot rank healthy files — that is what the rest of this section and §10 are about — but it is enough to say the belief is not a rule you can apply blind.
First, what the two formats actually are. A 4-bit quantization stores each weight as one of only 16 allowed values; the formats differ in where those 16 values sit:
The three quantizations above were graded for correctness on GSM8K and MATH-500, 20 problems each, greedy decoding — always take the most likely next token, so the same prompt gives the same answer every time (§02) — identical prompts, a 4,096-token budget, with answers that ran out of budget mid-thinking counted wrong and reported as truncations. One framing rule before the numbers: at 20 problems this is a smoke test, not a ranking — each question is worth 5 points (§02). Its job is to catch a broken file: a bad conversion or a mangled chat template collapses a score by 20–40 points or more, which even 20 questions detects reliably. It cannot rank healthy files, whose true gaps are a point or three (the arithmetic is in §10). Measured on the reference 3090:
| Quant | Format | GSM8K | MATH-500 | Decode (drafter on) | Draft accept len | J / decode token |
|---|---|---|---|---|---|---|
| Q4_K_M | K-quant — integer grid; a filename with K and no IQ | 85% (1 trunc) | 75% (5 trunc) | 63–70 t/s | 4.7–4.9 | 8.853 meas |
| NVFP4-HIGH | FP4, most extras upgraded | 85% (1 trunc) | 65% (5 trunc) | ~55 t/s | 4.7–4.8 | 9.293 meas |
| NVFP4-VERY-LOW | FP4, smallest tier | 85% (0 trunc) | 65% (5 trunc) | ~49 t/s | 3.33 | — not measured |
| UD-IQ4_XS (not in the scored run) | IQ-quant — integer grid too; a filename with IQ | — | — | — | — | 8.198 meas |
The MATH-500 column is cap-taxed and its grader is unaudited. Five of 20 answers per file — 25% — ran out of the 4,096-token budget and were graded wrong. This page's own rule, stated in §10, is to raise the budget rather than shrink the test, and that rule was applied to the effort comparison's GSM8K run and never to this table. Worse, the 2026-08-23 benchmark sweep found a MATH-500 grader bug of exactly the shape that would bite here — it compared how an answer was written rather than what it was worth, marking 145 wrong against a reference of 145^\circ — and fixing it moved one arm by 32 points at n=25. Read this column as "no file collapsed", nothing finer. The decode column comes from a different llama.cpp build and workload than §06's configuration table (both are listed in §15) — compare speeds within this table only. The energy column comes from a third run, the 2026-08-23 power matrix, measured with the drafter off on a different prompt, so it ranks the files against each other and does not price the configurations in the decode column.
What the numbers say:
- "Floating point always wins" did not survive the test. All three tied on GSM8K; on MATH-500 the integer-grid K-quant answered two more problems than either FP4 file. At 20 samples that edge is weak evidence — but it is the opposite direction from the belief. What actually decides 4-bit quality is how cleverly a format spends its bits — block sizes, scale precision, which layers stay at higher precision — and mature K-quants are very good at that. Both families use fine-grained scaling; the float-versus-integer grid underneath matters far less than people assume.
- The two NVFP4 tiers were indistinguishable at this sample size — which is what noise looks like, not evidence of equivalence. The file structure does predict a tie: the family shares one byte-identical NVFP4 backbone, with tiers differing only in the surrounding layers. The surprise is where VERY-LOW's savings actually landed: on these maths prompts at n-max 10 its draft head drafted noticeably worse (accepted length 3.33 against about 4.8) — a speed cost, not an accuracy one. The penalty is configuration-dependent: re-measured 2026-08-22 at n-max 4 / p-min 0.75 on code generation, both tiers accepted identically (0.83) and VERY-LOW was the faster file (54.6 against 48.2 t/s) — smaller weights, less memory traffic per token.
- Sample-size honesty: each problem is worth 5 points. Every gap in this table is one or two questions wide. Treat it as "no meaningful quality difference, mild edge to the K-quant" — not a ranking. The same prompts were also run through a different serving stack and grader version, and individual cells moved by up to 10 points, which means measurement variance is the same size as the differences being measured. A one-run quantization ranking published without error bars is not evidence, whoever published it.
How the files rank on perplexity
Perplexity is the dotted line on the ladder chart.
Accuracy at 20 questions cannot separate healthy files, but a per-token measure can, because it feels quantization damage at every scored token instead of collecting one right-or-wrong bit per question. Perplexity is defined in §02; the run below scores 294,912 token positions — 36 chunks × 8,192, forced by the corpus, the tokenizer and -c 8192. Measured 2026-08-22 on the reference 3090 with llama-perplexity over the full wikitext-2-raw test set:
| File | Perplexity (lower is better) | vs best |
|---|---|---|
| Q4_K_M | 6.535 ± 0.044 meas | — |
| UD-IQ4_XS (13.3 GiB) | 6.596 ± 0.045 meas | +0.9% — a gap of 0.061 against a ±0.063 combined error bar, so unresolved |
| UD-Q4_K_M | 6.654 ± 0.045 meas | +1.8% |
| UD-Q4_K_XL | 6.682 ± 0.046 meas | +2.3% |
| NVFP4-HIGH | 6.826 ± 0.047 meas | +4.4% |
| NVFP4-VERY-LOW | 6.877 ± 0.047 meas | +5.2% |
Read it with three limits attached. One corpus. Wikitext rewards plain-text prediction while unsloth's Dynamic calibration targets chat-style data, so the ordering among the top files may not carry over perfectly to instruction-following work. The scored-position count is tokenizer-bound. A later run on the same corpus files recorded 297,193 tokens under Qwen's tokenizer and 295,216 under Gemma's, both at 36 chunks — so 294,912 is a count of scored windows (36 × 8,192), not of corpus tokens, and it moves with the tokenizer. The practical consequence is a hard rule: perplexity is never comparable across model families. And the reference file reproduces. A separate 2026-08-23 run re-measured the UD-IQ4_XS figure as a gate before further work and returned 6.5956 against an expected 6.5956 — bit-identical, which is what pins the error bars and the chunk count.
Two robust findings survive all three limits: both NVFP4 tiers trail the K-quants — the same direction the scored table's MATH-500 column hinted, reached by an independent measurement — and none of the unsloth Dynamic variants beat the plain lmstudio Q4_K_M on this corpus, including at identical size. Two cheap metrics agreeing beats one expensive one.
Alongside perplexity, a 200-question GSM8K run compared the same files: Q4_K_M 94.0%, UD-Q4_K_XL 94.5%, UD-IQ4_XS 93.0%, NVFP4-HIGH 92.5%, NVFP4-VERY-LOW 90.5%, all ±3–4 points and therefore a tie by §10's arithmetic. Those numbers now need re-grading before they are quoted again. The 2026-08-23 sweep found that its GSM8K grader had been comparing the whole answer line instead of the number, so #### 156 kg failed against a reference of 156 — and that single bug was worth 8 points, all of it on the arm that reasons in units. The runs above used a different grader (chinkeong/benchmark), so this is a re-grade order, not a proven error. The blast radius matters: §03's the Q4_K_M recipe was argued on "two independent metrics agreeing", and one of the two is this comparison. Until the saved transcripts are re-graded with a number-extracting grader, the Q4_K_M recipe rests on perplexity alone, and that caveat is printed at the Q4_K_M recipe as well as here.
The same corpus, the same 36 chunks and the same error bars also price the cache instead of the weights. Measured 2026-08-23 on UD-IQ4_XS, every other flag identical to the runs above (-c 8192 -fa on -ngl 99), on a corpus file byte-verified identical to the one the fp16 and q8_0 baselines used:
| K/V cache | Perplexity | vs fp16 | What it buys |
|---|---|---|---|
-ctk f16 -ctv f16 | 6.5956 ± 0.045 meas | — | the reference — and about 1.6× the window cost of q8_0 on the one measured pair (16,428 against 15,661 MiB at -c 32768, which puts the fp16 slope near 64,500 B per token against q8_0's measured 39,936). The arithmetic floors would say 1.88×; budget from a load, not from the floors (§05) |
-ctk q8_0 -ctv q8_0 (shipped) | 6.6160 ± 0.045 meas | +0.309% | halves the cache. Every context ceiling on this page is built on it |
-ctk q4_0 -ctv q4_0 | 6.6413 ± 0.045 meas | +0.693% | halves it again, for a further +0.383% on top of q8_0 |
Three readings, and the third is the one that keeps this honest. Degradation is super-linear in bits — the second halving costs slightly more than the first (0.383 against 0.309), so the cut stops paying rather than scaling. Both are small next to the model quantization: the Q4_K_M-to-UD-IQ4_XS gap on this same corpus is 0.93%, larger than q4_0's entire cache penalty, so anyone who chose the smaller weights file on quality grounds has already accepted a bigger hit than moving to a 4-bit cache would add. And the error bars overlap: ±0.045 on each of two independent estimates gives ±0.063 on their difference, so the fp16-to-q4_0 gap of 0.0457 is about 0.7 of one standard error and a single 36-chunk run does not resolve it on its own. (That combined figure, not the ±0.045 on a single estimate, is the right yardstick for every file-against-file comparison on this page, including the 0.061 gap between the two 4-bit files above.) the direction is consistent with the q8_0 point and with theory, and that is all it is. Two limits on how far to carry this. It was measured at -c 8192, so it says nothing about the long-context retrieval failure that is the actual argument against a 4-bit cache at 200k-plus (§05's mechanism footnote — still unmeasured). And q4_0 was not faster in any way that matters: 266 s against 298 s of perplexity wall-clock is a bandwidth effect on a prefill-shaped workload, not a decode result. Run -ctk q8_0 -ctv q8_0. It is what every recipe on this page ships and what every context ceiling here is built on.
The quantization ladder: nine files of one model, four readings
UD-IQ4_XS file as the reference. Conditions: the same 75 questions
throughout, one RTX 3090, driver 596.36, llama.cpp build 10502, measured
2026-08-21 to 2026-08-27. Read the three lines separately — they
disagree, and that is the point: the paired question set stays a statistical tie
all the way down to 2.48 bits, while perplexity is already climbing at every rung.
The accuracy column runs out of resolution before the files run out of quality.
The floor is 2.912 bits per weight, and it rests on perplexity —
6.9957 against 7.5481 at 2.481 bits, a 7.9% gap on a shared
tokenizer. The empty-answer line was counted under greedy decoding;
at the sampler the recipes ship
(--temp 1.0 --top-p 0.95 --top-k 20) all three low-bit files returned
zero blanks in 300 generations. A reader who runs greedy decoding
should read that line as live.
measuredOne model, squeezed by one vendor to eight sizes from 13.3 GiB down to 5.8 GiB, every one of those eight measured the same way on the same day. Three readings sit beside each file below — perplexity, task accuracy and the empty-answer count — and a fourth runs the code each file wrote (measured 2026-08-25) (further down). It answers the question people actually ask — how small can I go before it stops being the same model? — and it answers it in more than one voice, because the readings disagree about where the damage starts and each is right about something the others cannot see.
Two of those readings are not independent of each other, and this table would mislead if it did not say so. The empty-answer column is counted from the same 75 greedy generations the accuracy column is scored on, and every empty answer, on every file, is also scored wrong — checked across all nine arms, with no exceptions. The empty column is therefore a breakdown of the accuracy column’s failures rather than a second opinion about them. It earns its place for a different reason: counting one named failure is more sensitive than testing a 75-item score for significance, so it moves a full rung earlier, and it costs no GPU time at all (below).
The ninth row is not a rung of this ladder, and it is printed anyway. QAT-Q2_0 is a different vendor’s build of the same model — sdkyuan’s, made by quantisation-aware training, which trains the model with the rounding already in place instead of rounding a finished model afterwards. It is here because it is the one file that does not follow the pattern the other eight make, and a table ordered by bits per weight that quietly dropped its own counterexample would be worth less than one that prints it.
Read the columns like this. Perplexity is a per-token measure of how surprised the model is by ordinary text: it feels a little damage at every rung and is the earliest warning, but a file can lose perplexity and still do your work. Accuracy Mean is the composite of three graded benchmarks — GSM8K, HumanEval and MBPP, 25 questions each, 75 items in total — and it is the only column in the reader's own units. The empty-answer column counts the times the model returned nothing at all, and it is kept separate from truncations because the two are different failures: a truncation hit the token cap, an empty often terminated normally and simply emitted zero characters, which no truncation counter can see. Those last two columns are the functional detectors that §01 counts as the third of this campaign's instruments — the automated checks for what a score cannot see (§02). The two vocabularies name one thing.
| File (unsloth/Qwen3.8-27B-GGUF except where marked) | Weights GiB | Bits per weight | Perplexity vs the 4-bit reference file | Accuracy Mean n=25 × 3 sets | Empty (silent) | Trunc | Plain-language verdict |
|---|---|---|---|---|---|---|---|
| UD-IQ4_XS the reference file · this page's daily file | 13.274 | 4.223 | 6.5956 the reference | 97.30 | 0 (0) | 0 | Full quality. Everything below is measured against this file, on the same 75 questions |
| UD-Q3_K_XL | 12.244 | 3.895 | 6.7691 +2.63% | 96.00 | 0 (0) | 0 | The cheapest step down on the whole ladder. Ties the reference file on the paired test, no empties, no truncations |
| UD-IQ3_XXS | 10.184 | 3.240 | 6.9187 +4.90% | 93.30 | 0 (0) | 0 | Still a tie with the reference file, still clean. Saves 3.1 GiB against it |
| UD-Q2_K_XL where the perplexity curve turns | 9.154 | 2.912 | 6.9957 +6.07% | 96.00 | 0 (0) | 0 | The last rung that still performs like the reference file — ties it on the paired test, and it is the last file with zero empty answers. Below here the perplexity cost per gigabyte jumps about fivefold and the empty count leaves zero for good |
| QAT-Q2_0 sdkyuan · not a rung of this ladder: another vendor, and quantisation-aware training rather than rounding a finished model | 8.158 | 2.595 | 7.4996 +13.71% | 90.70 | 1 (0) | 1 | The counterexample — and it is a counterexample only on the instrument this table does not carry. It spends 0.114 more bits per weight than UD-IQ2_S below it, and on everything printed here the two are a tie: perplexity 0.0485 apart against a ±0.072 combined error bar, the same 90.70 Mean, two questions each way on the paired test (p=1.00), and one empty answer against two. On next-word agreement with the reference file they are nowhere near each other — 77.663% against 83.988%, which puts this 2.595-bit file below the 2.153-bit rung, between it and the 1.994-bit one (below). More bits did not buy closeness to the reference, and nothing in a column ordered by bits per weight would have predicted it |
| UD-IQ2_S | 7.797 | 2.481 | 7.5481 +14.44% | 90.70 | 2 (1) | 1 | Still a tie on accuracy, but the weakest one on the ladder and right at the edge — and the first rung that ever returns nothing at all. The empty count sees damage a full rung before the benchmark can |
| UD-IQ2_XXS | 6.767 | 2.153 | 8.0079 +21.41% | 78.70 | 3 (2) | 1 | Measurably worse than the reference file, and the first rung where the paired test says so rather than shrugging. Not recommended |
| UD-IQ1_M | 6.267 | 1.994 | 8.1418 +23.44% | 85.30 | 5 (5) | 0 | Also measurably worse than the reference file. Its higher Mean than the rung above is not a reversal — see below — and it has the worst silent-empty record short of the bottom: five questions answered with nothing, while its truncation counter read zero |
| UD-IQ1_S | 5.767 | 1.835 | 8.9265 +35.34% | 34.70 | 28 (10) | 20 | Broken, and in an unusual way: it stops being able to stop. Median answer 932 tokens against 424 at the reference file, 20 of 75 answers running into the cap, one arm taking two and a half hours where the next longest took thirty-six minutes. This is the ladder's "clearly no" rung. The 34.70 is the 16,384-token-cap arm: the automatic rerun at double the cap was stopped by the session as a priority inversion rather than allowed to finish, so this row stands as measured, with its 20 truncations reported beside it |
Conditions, and they differ between the two scored columns. Perplexity meas: frozen wikitext-2-raw test set, 36 × 8,192 = 294,912 scored positions per file — the same count and corpus as the table above it — -ngl 99 -c 8192 -fa on --load-mode mmap, fp16 KV, llama.cpp build 10502. Accuracy and the two failure columns meas: a frozen suite — fixed once and never edited again, so every file is asked the identical questions — 1cdf54f8eb9d3f8f, GSM8K + HumanEval + MBPP at n=25 each, greedy, seed 42, --max-tokens 16384, -c 32768, q8_0 KV, reasoning_effort=low, and no drafter — which turns out to matter, and is dealt with in the next part. "Silent" empties are the ones that did not hit the cap: they terminated normally and returned zero characters, so no truncation counter ever saw them. Means are printed at the harness's own precision; what they can be read to is the subject of the next two paragraphs. QAT-Q2_0 was measured later and separately, and its row says so: perplexity on 2026-08-25 over the identical corpus, chunk count and flags (7.4996 ± 0.04784), and the accuracy suite and both failure columns on 2026-08-26 over the identical frozen suite. Every condition matches; the day does not.
Twenty-five questions per set detects a collapse and nothing finer. That is §10's own first row — about 25–30 samples resolves a ~20-point gap — and the accuracy column above is built from exactly that sample size. So it is in the class this page warns readers about, and the warning applies to it in full: the Mean column may say where this model breaks; it may not rank one rung against another. Every "tie" and every "worse" below comes from a paired test on the identical 75 items, not from the difference between two Means, and no ordering is claimed anywhere that the paired test does not support.
So what can be said, arm against arm, is whatever a paired McNemar test says. It compares the two files question by question on the same 75 items and counts only the ones they disagree about — b is the number the left file got right and the right file got wrong, c the reverse, and p is how often a gap this large would turn up by chance if the two files were really equal (§02). The exact test, counting a difference in either direction:
| Pair | b | c | p | What it licenses |
|---|---|---|---|---|
| UD-IQ4_XS vs UD-Q3_K_XL | 1 | 0 | 1.0000 | Tie. One discordant item of 75 |
| UD-IQ4_XS vs UD-IQ3_XXS | 3 | 0 | 0.2500 | Tie |
| UD-IQ4_XS vs UD-Q2_K_XL | 1 | 0 | 1.0000 | Tie. One discordant item of 75 — the strongest tie on the ladder, and the basis of every 2.9-bit recommendation below |
| UD-IQ4_XS vs UD-IQ2_S | 5 | 0 | 0.0625 | Tie, and the weakest one measured — right at the edge of resolving |
| UD-IQ4_XS vs QAT-Q2_0 | 5 | 0 | 0.0625 | Tie, and exactly as weak as the row above it |
| QAT-Q2_0 vs UD-IQ2_S | 2 | 2 | 1.0000 | Tie, and the flattest on the page — two questions each way. On this instrument the two files are the same file, which is what makes the agreement column below worth reading |
| UD-IQ4_XS vs UD-IQ2_XXS | 14 | 0 | 0.0001 | Different. The 4-bit file is measurably better |
| UD-IQ4_XS vs UD-IQ1_M | 9 | 0 | 0.0039 | Different. The 4-bit file is measurably better |
| UD-IQ3_XXS vs UD-Q2_K_XL | 0 | 2 | 0.5000 | Tie — and perplexity ranks these two the other way, so no ordering claim between them exists on this page at all |
| UD-IQ2_S vs UD-IQ2_XXS | 10 | 1 | 0.0117 | Different |
| UD-IQ2_XXS vs UD-IQ1_M | 5 | 10 | 0.3018 | Tie. The Mean column appears to rise again here; that is noise, and it is never to be written as a reversal |
The one sentence that reconciles a falling perplexity curve with a flat accuracy curve. Look at the c column: it is zero in seven of the eleven rows, and in every row that compares a file with the 4-bit reference. The smaller file almost never wins a question the larger one lost, and against the reference file the losses run in one direction only. So quality is falling at every step down the ladder; it is simply below what 25 questions per set can resolve until the cliff arrives. The flat accuracy column is not evidence that the files are equal. It is evidence that this instrument cannot see a gap this small, which is what §10 said before the ladder ran.
Where the boundary is, stated as an interval because that is all the data supports. UD-IQ2_S at 2.481 bits per weight ties the reference file. UD-IQ2_XXS at 2.153 does not. The boundary therefore lies between 2.48 and 2.15 bits per weight, and 25 questions per set cannot place it more finely than that. Anyone quoting a single number for where this model breaks — including a future edition of this page — is quoting something the measurement does not contain.
A second opinion on the same ladder: how often it picks the same next word
Agreement with the full-precision model is what separates files the question set cannot; both readings sit together on the ladder chart.
Agreement — how often a shrunken file picks the same next word as the reference — separates every rung of this ladder cleanly, which neither perplexity nor the 75-item accuracy suite manages. Perplexity is one average over a whole corpus, so two files can share a perplexity while behaving differently — this page already carries the proof, in that perplexity cannot separate UD-IQ2_XXS from UD-IQ1_M (they sit 1.67 % apart at 1.7 sigma).
First, a name collision worth clearing up, because the two things are
unrelated. --top-p 0.95 in the recipes on this page is
nucleus sampling — a knob that controls randomness while the
model writes. “Same top p” here is a measurement:
an output line of llama-perplexity --kl-divergence. It reports
how often a shrunken file picks the same next word as a reference
file, over the same text. Word by word, not averaged.
| File | bits/weight | Same top p agreement with the anchor ± 1 standard error |
Mean KLD lower is closer ± 1 standard error |
PPL ratio |
|---|---|---|---|---|
| UD-Q3_K_XL | 3.895 | 94.120 % ± 0.104 | 0.0190 ± 0.0003 | 1.011 |
| UD-IQ3_XXS | 3.240 | 88.924 % ± 0.139 | 0.0639 ± 0.0007 | 1.032 |
| UD-Q2_K_XL | 2.912 | 86.594 % ± 0.151 | 0.0942 ± 0.0010 | 1.057 |
| QAT-Q2_0 sdkyuan · not a rung of this ladder | 2.595 | 77.663 % ± 0.184 | 0.2970 ± 0.0025 | 1.223 |
| UD-IQ2_S | 2.481 | 83.988 % ± 0.162 | 0.1411 ± 0.0014 | 1.105 |
| UD-IQ2_XXS | 2.153 | 79.451 % ± 0.179 | 0.2255 ± 0.0019 | 1.181 |
| UD-IQ1_M | 1.994 | 76.620 % ± 0.187 | 0.2946 ± 0.0023 | 1.259 |
| UD-IQ1_S | 1.835 | 73.071 % ± 0.196 | 0.4118 ± 0.0029 | 1.408 |
Conditions MEASURED
2026-08-26: 200 chunks (102,400 scored tokens) of the same
wikitext-2 corpus, -c 512 -ngl 99 -fa on, base logits 23.6 GiB.
The base is UD-IQ4_XS, not the unquantised model. Measuring against
FP16 properly would need ~54 GB of weights that are not on this machine,
so every figure here is agreement with this page’s own anchor file
and cannot be compared against anybody else’s KL-divergence table.
It ranks these rungs against each other, which is what a ladder is for.
Every ± figure in the table is llama-perplexity’s own
standard error for that run, printed rather than dropped: it is 0.7 to
1.4 % of the mean KLD and 0.10 to 0.20 of a point on agreement, which is
small, and small is a reason to print it rather than to leave a five-decimal
figure looking exact. QAT-Q2_0 is sdkyuan’s quantisation-aware-trained
build, not a rung of this ladder; it was measured 2026-08-26 under the
identical flags, corpus and chunk count and is placed at its own bits per
weight.
Two words before the numbers, because the rest of this part is built on them. A standard error is how far a measured figure would be expected to move if the same measurement were taken again on fresh text of the same size — the wobble in the reading. Sigma is a gap counted in those units, so a gap of three sigma is three times the wobble and is very unlikely to be wobble. Comparing two files means comparing two readings, so the error on the difference is the two errors combined — the square root of the sum of their squares, which is the same rule this page applies to perplexity in the table above derived.
It resolves what perplexity cannot. Take the pair this page says perplexity cannot rank — UD-IQ2_XXS against UD-IQ1_M, 1.67 % apart at 1.7 sigma. On agreement they are 2.83 points apart against a combined standard error of 0.26 points: 10.9 sigma derived. Every adjacent step on the eight-rung ladder is 2.3 to 5.2 points against a combined error of 0.17 to 0.27, and the weakest of those separations is still 10.9 sigma, so every rung separates cleanly — which neither perplexity nor the 75-item accuracy suite manages. C11
It is also easier to picture. “UD-Q2_K_XL costs +5.66 % of perplexity” is an abstraction — and note that the figure belongs to this run’s 512-token chunks, where the table above reads +6.07 % on 8,192-token chunks of the same corpus; the same file, two chunk lengths, and the condition has to travel with the number. “UD-Q2_K_XL picks a different next word about one time in seven” is a sentence you can act on.
And it turned up one file that does not follow the pattern.
QAT-Q2_0 spends 2.595 bits per weight, which is more than
UD-IQ2_S’s 2.481, and it agrees with the anchor on 77.663 %
of next words against UD-IQ2_S’s 83.988. That is 6.33 points the wrong
way against a combined standard error of 0.25 — 26 sigma
derived — and its mean KLD of 0.2970 cannot
be told apart from the 1.994-bit UD-IQ1_M’s 0.2946 (0.7 sigma).
A file whose size places it just above UD-IQ2_S sits, on this instrument,
with UD-IQ1_M two rungs further down.
Read that as a limit on the instrument at least as much as on the
file. This column measures distance from UD-IQ4_XS, and
UD-IQ4_XS is one vendor’s rounding of a finished model. A build made by
quantisation-aware training is not trying to imitate it, so part of that
6.33-point gap is “built differently” rather than “built
worse” — and every other reading on this page says so. The same
two files tie on perplexity (7.4996 against 7.5481, 0.0485 apart
against a ±0.072 combined error bar), tie on the 75-item paired
test (two questions each way, p=1.00), and QAT-Q2_0 returns
fewer empty answers and writes JavaScript that runs. Nothing here
says it is the worse file. What it says is that agreement with a 4-bit
anchor is not a quality score, and that a table ordered by bits per weight
stops being readable as a quality order the moment more than one vendor is in
it.
What it does NOT measure, and why the recommendation still rests on
perplexity. Agreement is fidelity to the reference, not
usefulness, and this ladder contains the demonstration rather than a
hypothetical one: QAT-Q2_0 agrees on 77.663 % of next words
and the JavaScript it wrote runs, while UD-IQ2_XXS agrees
on a higher 79.451 % and the JavaScript it wrote throws. That is
exactly what the execution probe catches and no
agreement metric can see n=1 per file. And this
column is measured against a 4-bit file rather than the real model, so it
describes distance from another quantisation. Both metrics separate
2.912 from 2.481 decisively — 7.9 % on perplexity, 2.6 points at
11.8 sigma here — so the floor does not move either way.
This is a second opinion, not a replacement.
A file called Q2_K_XL measures 2.912 bits per weight, not the 2.5625 that stock Q2_K spends, and Q3_K_XL measures 3.895 against stock Q3_K's 3.4375. That is not an error in either place: k-quants (llama.cpp PR #1684, Kawrakow, June 2023) fix a bit width for each kind of the model's internal number tables, and a vendor's "UD"/"XL" mix then spends about 0.35–0.46 bits per weight more than the label by keeping the tables that matter most at higher precision. cited for the PR and the stock widths; measured for the file sizes this column is computed from. The curve above is a function of the real bits, not of the name on the file — which is exactly why the column is printed.
Two limits to carry out of this table. Every step is resolved on perplexity, and the smallest one is not close: the 4.223-to-3.895 step moves perplexity by 0.1735, about 2.8× the ±0.063 combined error bar this section uses for any file-against-file difference derived, and every step below it is larger. And a cross-model reference point, because "how small before I should just use a smaller model?" is the real question underneath. On the identical frozen suite, gemma-4-12B-it-QAT-Q4_0 — 6.497 GiB at 4.65 bits per weight, within 4% of UD-IQ2_XXS's download size — scored 73.30 with 19 truncations, against the 2.153-bit 27B's 78.70 with one. At an equal weight budget the crushed big model won, which is worth knowing before assuming that a smaller model at higher precision is automatically the safer trade. Conditions differ between those two arms (the Qwen ran at reasoning_effort=low with a q8_0 cache; gemma ran at its own defaults, having no effort knob), and that is a disclosed asymmetry, not a fair fight.
The empty-answer table, in full
The empty-answer counts in this table are the dashed line on the ladder chart.
The ladder above prints the empty count in one narrow column. It has earned a table of its own, because it is the cheapest reading in this whole campaign and the only one of the three that draws a usable line. Cheapest, literally: it costs no GPU time at all — every number below was counted from transcripts that were already sitting on disk when the accuracy run finished. And it draws a line where the other two cannot: perplexity falls at every step from the very top, so it never has a flat stretch that could mark a beginning, while the paired accuracy test still says "tie" a full rung below the point where empty answers first appear.
One thing to hold on to while reading it, because it changes what the column is evidence of. These are the same 75 greedy generations the accuracy column is scored on, and every empty answer here is also scored wrong there — no exceptions in nine arms. So this is not a second witness agreeing with the first; it is the first witness’s failures sorted into kinds. That is still worth doing, and it is why the column moves earlier: naming one failure and counting it is more sensitive than asking whether a 75-item score has moved enough to be believed. It is not an independent confirmation; read it as a breakdown of the accuracy column’s failures, not as a second witness.
Three words in plain language before the numbers, because they name three different things. An empty answer is a reply with no characters in it — you asked, and nothing came back. At cap means that particular empty reply had run out of its token budget and was cut off; the server records that as a truncation, so it appears in a log where somebody might see it. Silent means the opposite, and it is the one to worry about: the model ended the reply by itself, in the ordinary way, and handed back nothing — an answer that finished normally and returned zero characters, so no truncation counter can see it and nothing in your logs reports a problem at all. That is why the two are separate columns rather than one. Every row is n=75: three benchmark sets — GSM8K, HumanEval and MBPP — at 25 questions each, the identical 75 questions put to every file.
| File (unsloth/Qwen3.8-27B-GGUF except where marked) | Bits per weight | Empty of n=75 | At cap cut off | Silent ended normally | Median answer tokens |
|---|---|---|---|---|---|
| UD-IQ4_XS the reference file | 4.223 | 0 | 0 | 0 | 424 |
| UD-Q3_K_XL | 3.895 | 0 | 0 | 0 | 488 |
| UD-IQ3_XXS | 3.240 | 0 | 0 | 0 | 454 |
| UD-Q2_K_XL the last clean rung | 2.912 | 0 | 0 | 0 | 416 |
| QAT-Q2_0 sdkyuan · not a rung of this ladder | 2.595 | 1 | 1 | 0 | 382 |
| UD-IQ2_S first rung to detect | 2.481 | 2 | 1 | 1 | 475 |
| UD-IQ2_XXS | 2.153 | 3 | 1 | 2 | 472 |
| UD-IQ1_M | 1.994 | 5 | 0 | 5 | 502 |
| UD-IQ1_S | 1.835 | 28 | 18 | 10 | 932 |
Conditions meas: the frozen suite 1cdf54f8eb9d3f8f, greedy at temperature 0 / seed 42, --max-tokens 16384, -c 32768, q8_0 KV, reasoning_effort=low, no drafter — the same eight ladder arms that produced the accuracy column above, plus the separate QAT-Q2_0 arm of 2026-08-26 under those same settings, all read back from their saved artifacts rather than from a counter. How this table relates to the ladder's truncation column. "At cap" counts only the empties that were cut off; the ladder's Trunc column counts every answer that was cut off, empty or not. At the bottom rung the two read 18 and 20 — so 18 of that file's 20 truncated answers came back with nothing in them.
The finding, and it is a clean one. The empty-answer rate is exactly zero at every rung down to and including 2.912 bits per weight — the same rung where the perplexity curve turns and the last one that ties the reference file on the paired test — and then, across the rungs below it, rises without going back down: 2, 3, 5, 28. It notices damage at 2.481 bits, a full rung above where the paired accuracy test can resolve anything at all. At 2.481 the paired test returns p=0.0625 and says "tie"; the first verdict of "different" does not arrive until 2.153. The empty column had already stopped reading zero one rung earlier, for no GPU time. The QAT build printed between those two rows is not part of that sequence — it is another vendor’s file, it returns one empty, and that one was cut off at the cap rather than ended silently, which is the less dangerous of the two kinds.
How to read it, and how not to. The useful reading is zero against not-zero, and the shape of the rise. The difference between 2 empties and 3 out of 75 is not a ranking and must not be used as one — the same sample-size limit that governs the accuracy column governs this one. What the column does carry is a threshold under greedy decoding: exactly zero down to 2.912 bits, then a monotone rise. The reason this page stops at 2.912 bits per weight is perplexity — 6.9957 at 2.912 against 7.5481 at 2.481, a 7.9% gap on a shared tokenizer.
The blank answers are a temperature-0 finding, and nobody runs temperature 0. Every count in this column was measured with greedy decoding — the model always taking its single likeliest next word. Qwen's own card recommends --temp 1.0 --top-p 0.95 --top-k 20, and this page's own recipes ship those settings. Re-running the fifty HumanEval and MBPP prompts — where every empty answer on this ladder lives but two — at the recommended sampler, six seeds each, 300 generations per file MEASURED 2026-08-26:
UD-Q2_K_XL 0 of 300, QAT-Q2_0 0 of 300, UD-IQ2_S 0 of 300 — not one blank answer on any file, each bounded below 1% by the rule of three. That set includes HumanEval[6], which returns empty every single time under greedy on two of those files and still returns empty when the cap is doubled to 32,768. Sampled six times each, it never failed once. The failure is greedy-specific.
The floor rests on perplexity, which separates the two rungs cleanly and is not a sampler artefact: 6.9957 at 2.912 bits against 7.5481 at 2.481, a 7.9% gap on a shared tokenizer. The accuracy column does not carry it — 72 of 75 against 68 of 75 is four questions and the paired test cannot resolve it. The recommendation is 2.912 bits per weight — UD-Q2_K_XL. A reader who runs greedy decoding should still read the empty column as written, because for them it is live.
How strong is that threshold? Weaker than it looks, and a reader deserves the number. The step that places the floor is 0 of 75 at 2.912 against 2 of 75 at 2.481. Fisher’s exact test puts that single step at p = 0.50, and the 95% interval for 2 of 75 (0.7–9.2%) overlaps nearly all of the interval for 0 of 75 (0–4.9%). Split by kind it is thinner still: one of those two empties is a truncation that reproduces at double the cap (re-run at 32,768 tokens, still empty), so the silent count at 2.481 is 1, not 2. What actually carries the recommendation is the monotone rise across four rungs — 2, 3, 5, 28 — not the first step of it. Establishing that first step rather than merely indicating it needs about 300 clean runs per rung to bound an unseen rate below 1%; this page has 75. The floor stands as the cautious reading of a real trend, and it is not a measured cliff. measured for the counts; the tests are arithmetic on them. Added 2026-08-26 after a reader challenged the boundary — the challenge was arithmetically correct.
The median-answer column is a second detector hiding inside the first, and it finds a different failure. Median answer length holds inside a band of 416–502 tokens all the way down to 1.994 bits per weight — seven of the eight rungs barely move how long an answer is — and then jumps to 932 at 1.835. That is the rung that stops terminating: it is not answering worse at the same length, it is failing to stop, which is why 20 of its 75 answers ran into the cap and why that one arm took two and a half hours where the next longest took thirty-six minutes. A quality curve would never have predicted it, and neither would the empty count on its own.
A fourth instrument: we ran the code instead of reading it
Whether the generated code ran is the row of marks along the bottom of the ladder chart.
Every test above scores text. Perplexity scores it per token, the accuracy suite scores whether the answer is right, and the repetition screens check whether it has degenerated into a loop. None of them ever executed anything. So we did meas: each file had been asked to finish a JavaScript Dijkstra implementation, and every one of those programs was rebuilt with the prompt's own opening lines and run under node v24.15.0.
Seven of the eleven programs run. Four do not. The line falls in a place worth knowing:
- Down to 2.481 bits per weight the code runs —
UD-IQ4_XS,UD-Q3_K_XL,UD-IQ3_XXS,UD-Q2_K_XLandUD-IQ2_Seach build the graph, run the search and print a shortest path with its cost. So doesQAT-Q2_0at 2.595 bits, which was added to the probe on 2026-08-26 and is the seventh program. - Below that, none of it runs.
UD-IQ2_XXSthrowsTypeError: Cannot read properties of undefined.UD-IQ1_MandUD-IQ1_Sdo not even parse.
This does not change what this page recommends, and that is the point of reporting it. The recommendation floor here is 2.912 bits per weight, which is now set by perplexity — 6.9957 against 7.5481, a 7.9% gap on a shared tokenizer — after the empty-answer column turned out to be a greedy-decoding artefact (above). The executed floor is 2.481 — one rung below the recommendation. A fourth instrument, aimed at something none of the others measure, landed inside the margin the other three had already left. UD-Q2_K_XL writes code that works, which is the most direct answer this page can give to whether a 2-bit file is a real daily driver.
The failures arrive together, which is the useful part — in one direction only. The prompt said to continue the file and to emit no markdown fences. Every one of this ladder’s files whose code failed had also restarted the program from scratch instead of continuing it, two of the three wrapping it in the fences as well. The converse does not hold. Two files break it: UD-IQ2_S continued the file as asked but wrapped it in the forbidden fences, and its program runs; and the cross-model gemma-4-12B-QAT-Q4_0 followed both instructions exactly and still produced code that does not parse. So a model that starts ignoring the shape you asked for is worth watching — it preceded every failure among this ladder’s own files — but it is a warning sign and not a test, because it also happens on files whose code is fine. n=1 per file — one task, one language — so read this as a threshold, not a rate.
And a caution about a rule of thumb you will see everywhere. The usual advice is that a smaller model at more bits beats a bigger model at fewer bits. On this one task it went the other way: gemma-4-12B-QAT-Q4_0 at 4.651 bits per weight produced code that does not run — it opened a helper function inside a method, never closed the method, and declared the next one as though it had — while Qwen3.8-27B at 2.912 bits produced a working program. One sample is not a refutation and this page does not offer it as one. It is a reason to run the substitution on your own workload rather than accept either rule of thumb: the comparison is cheap, and it does not always land where the slogan says.
UD-IQ4_XS reference, so that readings in different units can share one axis: perplexity added = (PPL − 6.5956) ÷ 6.5956, accuracy lost = (97.30 − Mean) ÷ 97.30, empty rate = empties ÷ 75. Worse is further down in all three series. The measured values behind them are printed in the tables above; the eight files drawn here are the eight rungs of this one vendor’s ladder, and sdkyuan’s QAT-Q2_0, which those tables print between 2.912 and 2.481 bits per weight, is not on this figure. The run / does-not-run strip along the bottom is meas directly. Perplexity moves at the very first step down and never stops — the most sensitive line here, and the one that says least about whether the model still works. The empty-answer rate is exactly zero at every rung down to 2.912 bits per weight and first flinches at 2.481; it costs no GPU time, because it is counted from answers already on disk. Read that flinch as the beginning of a trend rather than as a result on its own: 0 of 75 against 2 of 75 is Fisher’s exact p = 0.50, so that first step alone is not distinguishable from zero, and what carries it is the monotone rise across four rungs (the empty-answer table). And the empty series is a subset of the accuracy series — every empty answer on every arm is also scored wrong, with no exceptions in nine arms — so those two lines are one set of 75 greedy generations read two ways, not two opinions about it. The pair that genuinely rests on different work is the accuracy suite and the JavaScript execution probe, and they break at the same rung: everything down to UD-IQ2_S at 2.481 runs, and nothing below it does. Even those two are not machinery-free of each other — 50 of the 75 accuracy items are graded by running generated Python — but they use a different prompt, a different language and a different harness, which is what makes the agreement worth something. The functional boundary is between 2.481 and 2.153 bits per weight. One thing the drawing shows that no summary should hide: the accuracy line is not monotone, since UD-IQ1_M at 1.994 bits sits above UD-IQ2_XXS at 2.153 — and the paired test between those two returns p=0.3018, which is a tie and never a reversal. Execution is n=1 per file.What each file and window actually needs
Every recommendation on this page ends in the same small piece of arithmetic, so here is the table that lets you do it yourself instead of trusting a verdict. Find your file and the window you want, read what it needs, and subtract your own desktop from your card's total. What is left over is your slack (§02).
The memory column transfers to any card. The speed column is this 3090's and nobody else's. How much memory a configuration needs is a property of the model, the window and the flags — the same file at the same -c asks for the same bytes on a 5080 as it does here — so a 16 GB owner can plan against these figures directly, and that is the one thing about a smaller card this page can state as measured. How fast it then decodes is a property of this card's memory bandwidth. No speed on this page was ever measured on any other card, and none can be.
| File · weights | Drafter | -c 32768board VRAM · decode | -c 65536 | -c 131072 |
|---|---|---|---|---|
| UD-IQ4_XS 13.274 GiB | on | 16,586 MiB 41.35 t/s | 17,962 MiB 42.69 t/s | 20,848 MiB 34.68 t/s |
| UD-IQ4_XS | off | 15,376 MiB 35.92 t/s | 16,678 MiB 30.73 t/s | 19,188 MiB 24.77 t/s |
| UD-Q3_K_XL 12.244 GiB | on | 15,530 MiB 39.59 t/s | 16,906 MiB 42.98 t/s | 19,724 MiB 41.74 t/s |
| UD-Q3_K_XL | off | 14,322 MiB 36.44 t/s | 15,568 MiB 31.28 t/s | 18,064 MiB 23.91 t/s |
| UD-Q2_K_XL 9.154 GiB | on | 12,606 MiB 43.19 t/s | 13,982 MiB 40.30 t/s | 16,800 MiB 29.44 t/s |
| UD-Q2_K_XL | off | 11,396 MiB 33.48 t/s | 12,644 MiB 32.18 t/s | 15,140 MiB 24.54 t/s |
Conditions, all cells meas (2026-08-25, the reference RTX 3090, board total 24,576 MiB): -ngl 99 -fa on --parallel 1 -ctk q8_0 -ctv q8_0 --jinja, reasoning off, greedy decoding — temperature 0, top_k 1, and text-only — there is no --mmproj among the recorded flags, which matters because the projector is the fourth determinant of any ceiling on this page (§05); where the drafter is on it is --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75, the flags §03's vision recipe shipped until 2026-10-08. Every window was deep-filled to about 90% of itself with ordinary English prose from the wikitext-2-raw test set — not probed shallowly, which is the mistake §05 documents — the first probe after the prefill was discarded because the card is still raising its clock (§11), and the figure printed is from two settled probes. Memory is board VRAM, so it contains this machine's desktop as well as the server. Every speed figure in this table is a GREEDY measurement, and the recipes on this page ship --temp 1.0 --top-p 0.95 --top-k 20. Greedy holds the generated text still so the file is the only thing changing, which is the right way to compare files — but it is not what you will run. Measured 2026-08-26, one load, alternating: greedy spread 3.8% while the recommended sampler spread 25.5% MEASURED (§01). So check your own machine against these numbers with --temp 0 --top-k 1; run the recipe sampler instead and a single generation can land anywhere in a quarter-wide band with nothing wrong. The three arms at the full -c 262144 are a separate deep fill and are in the table two parts below.
Your card's total, minus what the table says, is your slack. On this 24,576 MiB card, UD-Q3_K_XL at -c 131072 with the drafter leaves 4,852 MiB and UD-IQ4_XS at the same window leaves 3,728 derived, both comfortably clear of the 1,796 MiB this page reserves. On a 16 GB card the total is 16,384 MiB, so keeping that reserve free leaves 14,588 MiB to spend derived; on a 12 GB card it is 12,288 and 10,492 derived. Hold those two numbers against the column you want and the card verdicts below follow by arithmetic rather than by assertion. And be exact about the reserve itself (§02): the desktop's own share measured 1,179–1,669 MiB in direct no-server readings depending on what was on screen, load-to-load variation measured 127 MiB, and 1,796 is this page's own derived threshold built from the worst case of the first plus the second. The desktop does not need 1,796 MiB. If you know what your own screen holds, subtract that instead and you will get a better answer than this page's worst case gives you.
What the drafter costs, and a cross-check worth having. Turning speculation on costs a consistent 1,208–1,660 MiB across these arms, and the cost grows with the window rather than with the file — which is exactly what §05's two-constant model predicts, since the drafter books 1,008 MiB fixed plus 5,120 B per window token. Subtracting the pairs above gives 1,208–1,210 MiB at -c 32768, 1,284–1,338 at 65,536 and 1,660 at 131,072 derived, against 1,168, 1,328 and 1,648 predicted derived — agreement within 44 MiB, well inside that model's own 127 MiB worst residual. That model is fitted at ordinary windows and it is not reliable at the top of the native window: at -c 262144 it under-predicts the measured requirement by about 1,213 MiB, because it does not carry the compute buffers (below).
Which file is fastest changes with the window — and with what you are writing
The speed column of the table above holds a result that is worth separating out, because it contradicts the way almost everyone reasons about quantization. With the drafter on, a different file reads fastest at different windows. Not a different file per card, or per workload — per window, on the same card, on the same content, with the same flags. Here is that column on its own, with a fourth window added, all arms drafter-on and deep-filled with prose:
| File · drafter on (n-max 4 / p-min 0.75) | -c 32768 | -c 65536 | -c 98304 | -c 131072 |
|---|---|---|---|---|
| UD-IQ4_XS 4.223 bpw | 41.35 t/s | 42.69 t/s | 36.09 t/s | 34.68 t/s |
| UD-Q3_K_XL 3.895 bpw | 39.59 t/s | 42.98 t/s | 32.90 t/s | 41.74 t/s |
| UD-Q2_K_XL 2.912 bpw | 43.19 t/s | 40.30 t/s | 34.08 t/s | 29.44 t/s |
Read down the columns and the highest reading keeps moving: UD-Q2_K_XL at 32,768, UD-Q3_K_XL at 65,536, UD-IQ4_XS at 98,304, UD-Q3_K_XL again at 131,072 — three different files across four windows. Before that becomes a claim about which file is fastest, here is how much of it the measurement actually supports. At -c 32768, -c 65536 and -c 98304 the three files sit inside a narrow band — 3.6, 2.7 and 3.2 t/s from top to bottom der — and at those windows only UD-Q3_K_XL was loaded twice. Those three columns are bands, not rankings, and this page does not order the files inside them. One column is different, and it is the one the recommendation rests on.
At -c 131072 the separation is wide and it reproduced across independent server loads. Each file was loaded twice, from scratch, with five settled probes per load: UD-Q3_K_XL read 41.44 then 42.04 (standard deviation 0.36), UD-IQ4_XS 34.66 then 34.70 (0.16), UD-Q2_K_XL 29.46 then 29.41 (1.07). Every file landed within 1.4% of its own first load. The gap between the top two is 20% — 41.74 against 34.68 der — which survives a 1.4% reproduction spread comfortably. UD-Q3_K_XL was loaded twice at -c 98304 as well — 32.86 then 32.95 — which matters for the mechanism below. On the two decimal places: they are the harness's own precision, not a claim that a cell is stable to 0.01 t/s. The band any cell can be read to is the reproduction spread printed here — 0.16 to 1.07 t/s — and at the three windows where no second load was taken, the band is unknown and the column is not ranked.
Why a file can decode faster at a deeper window. For a plain decoder this is impossible, and the impossibility is the reason the result was doubted: more context means more cache to read per token, so speed falls. A speculating decoder is not a plain decoder. What you actually measure is the model's own decode rate divided by how many tokens it commits per verification pass, and that second quantity belongs to the drafter — specifically to mean draft length, which §06 already establishes is the thing that ranks a configuration while draft acceptance is not. Watch UD-Q3_K_XL move between two neighbouring windows:
| UD-Q3_K_XL, drafter on, prose fill | Draft acceptance | Mean draft length | Decode |
|---|---|---|---|
-c 98304 | 1.00000 | 2.89 | 32.90 t/s |
-c 131072 | 0.885 | 3.30 | 41.74 t/s |
Acceptance falls, draft length rises, and throughput follows the draft length. An acceptance of exactly 1.00000 sitting beside a short draft is not a triumph — it is the confidence floor cutting the draft off early. --spec-draft-p-min 0.75 tells the drafter to stop guessing the moment it is less than 75% sure (§02), so on content it finds hard it proposes almost nothing, and everything it does propose survives checking. You get a perfect score on a nearly empty exam, and you go slowly. This is the same fact §06 measured on the reasoning stream, arriving from a completely different direction.
One more disagreement, and it is the sharpest evidence in this section rather than an embarrassment. The drafter pair below also measures UD-IQ4_XS against UD-Q2_K_XL at -c 32768, and it ranks them the other way round: 86.91 t/s against 77.01, the 4-bit file 12.9% ahead. The table above, at the same -c, reads 41.35 against 43.19 with the 2.9-bit file ahead. Neither number is wrong and neither is the general answer, because two conditions differ between them: that pair generated novel code into a nearly empty window, while every row here is prose filling about 90% of the window. Change the content and change the depth and a speculating decoder reorders itself. If you take one thing from this part, take that — the ordering belongs to the workload, not to the file — and then go and measure your own.
The confidence floor gates on the drafter's confidence, and confidence is a property of the content, not of the file and not of the window. Every number in this part was measured with the windows filled with ordinary English prose from the wikitext-2 test set. Your own work is not that. Fill a window with your own code, or your own codebase, and the drafter is confident in different places, the gate cuts in different places, and the ordering can move with it.
So this table is published as "measured on prose fill" and never as a fixed property of these three files. The proof that it is content-sensitive is inside the table itself: UD-Q3_K_XL reverses its own reading between two neighbouring windows of the same prose. If a change of 32,768 tokens of the same text can do that, a change of subject matter certainly can.
The check costs about ten minutes and this page recommends you spend them. Load your two candidate files at the window you actually use, fill each one with your own content, throw away the first probe after the prefill, and take two settled ones. That is exactly the protocol behind this table, and running it on your own material is worth more than any ordering published here.
First: do not retract a speculative-decoding measurement using a plain-decoder rule. The 131,072 ordering was initially challenged with the argument that decode cannot rise with depth. That is true of a plain decoder and false of a speculating one, for the reason set out above — the observed rate is decode divided by tokens committed per verification pass, and the draft length that sets the divisor is gated by content confidence, not by depth. A retraction is a claim, and it needs its conditions stated as carefully as the claim it retires. The retraction here was built on a mechanism that did not apply to the measurement it challenged — the same error documented in (§06's 81.7 t/s case study), which is the reason it is recorded again rather than quietly fixed.
And the replication that mattered was not the one that had already been done. The suspect number had passed every internal check, twice: five settled probes, a tight spread, a repeat load. None of that could help, because replicating within a condition cannot test whether the condition itself is anomalous. Only comparison across conditions — here, a neighbouring window — can test that, and re-running both windows across independent loads is what settled this ordering.
Which file for which card — and, on 24 GB, for which window
Two further measurements change which file leads. With the drafter on — which is how every recipe in §03 runs — UD-IQ4_XS reads 86.91 t/s against UD-Q2_K_XL’s 77.01 on 700-token novel-code generations at -c 32768, 12.9% faster despite being 45% larger, the reverse of the drafter-off ordering. At -c 262144 filled to 218,233 real tokens, UD-Q2_K_XL is 34% faster at depth (21.33 against 15.96 t/s), keeps 962 MiB more slack, and is the only one of the two that can still speculate at all.
With the drafter on at n-max 10 / p-min 0.5, the ranking inverts: UD-IQ4_XS reads 86.91 t/s against UD-Q2_K_XL’s 77.01 — 12.9% faster despite being 45% larger. With no drafter the smaller file wins (45.66 against 42.34, by 7.8%), which is the obvious result and, for anyone following this page’s recipes, the wrong one. §03’s August launcher carried two drafter pairs: n-max 4 / p-min 0.75 (the conservative pair, on four picks) and n-max 10 / p-min 0.5 (the faster pair, on one); the table below was measured at the faster pair, which was faster for greedy, thinking-off code like this table’s. Since 2026-10-08 the UD-IQ4_XS and Swift recipes run n3/p0 with one slot and n4/p0 with two, and since 2026-10-09 the UD-Q2_K_XL recipes run n4/p0 der (§06). Measured on 700-token novel-code generations at -c 32768, q8_0 KV, thinking off, temperature 0, three settled probes with the first post-prefill probe discarded:
| File | GiB | Drafter off | Drafter on (n-max 10 / p-min 0.5 — the faster of the two August settings) | Speculation is worth | Acceptance | Mean draft length |
|---|---|---|---|---|---|---|
| UD-IQ4_XS | 13.274 | 42.34 t/s | 86.91 t/s | 2.05× | 0.611 | 5.70 |
| UD-Q2_K_XL | 9.154 | 45.66 t/s | 77.01 t/s | 1.69× | 0.551 | 5.08 |
Drafter off, the 2.9-bit file is 7.8% faster — in the direction fewer gigabytes per token predicts, but far below what §04’s arithmetic gives for the size difference (1.45× at one constant, 1.56× at the format-specific pair, against the measured 1.078×).C39 Drafter on — which is how every recipe in §03 runs — the 4-bit file is 12.9% faster despite being 45% larger. The mechanism is the draft head: it is quantized along with everything else, so it degrades with bit width. Acceptance falls 0.611 → 0.551 and mean draft length falls 5.70 → 5.08, a 10.9% drop against an 11.4% drop in throughput — draft length predicting throughput again, to within half a point, as §06's measurement says it does. What the bigger file keeps in speculation — 2.05× against 1.69× — is worth more than everything the smaller file wins by being smaller. Across drafter settings, neither acceptance nor mean draft length predicts speed. A setting with acceptance 0.897 ran 33% slower than one with acceptance 0.611. Draft length does not rank them either. A direct test at -c 32768 put n-max 4 / p-min 0.00 at 82.73 t/s with mean draft length 2.803 against n-max 10 / p-min 0.5 at 76.32 with length 4.035 — the shorter drafts won by 8.4%. What both rules miss is waste: a drafter capped at ten speculates ten and discards about half, and every discarded token was still computed. The quantity that ranks them is accepted tokens per unit of drafting work, not accepted length and not acceptance.
Replicated 2026-08-25, after a reader challenged it. meas The argument against this table is a good one: UD-Q2_K_XL reads 31 % fewer bytes per token, so bandwidth alone says it must win. It was therefore re-measured from scratch by a second implementation in a different language, against the original script’s conditions taken verbatim — same prompt, same flags, same probe protocol. All four arms landed inside 1.3 % (86.91 → 86.58 and 77.01 → 76.97), and both acceptance figures reproduced exactly to three decimals (0.611 and 0.551). The bandwidth argument is right about bytes and still loses: what speculation returns to the less-damaged file is worth more than what the smaller file saves by being small.
Every speed figure on this page is measured with reasoning off; benchmark your own card with reasoning on and you will read roughly 15 % lower, and the gap will look like hardware when it is a regime difference. --reasoning off changes what the model emits and therefore what the drafter has to guess — acceptance moved 0.523 → 0.611 with it. The replication attempts that first exposed this read 73.71 t/s for the 4-bit file instead of 86.9; the entire gap was the reasoning regime, not Flash Attention.C17
Both files were probed at -c 32768, so this pair ranks them at shallow context — which is the range the recommendation table's first row covers, and no further. That row's boundary is now about 98,000 tokens, and it comes from the depth series above rather than from this pair: at -c 131072 a third file overtakes both of these. Do not confuse it with UD-IQ4_XS's 180,224-token text-only ceiling (below), which is a statement about how much window that file can hold with a drafter aboard, not about which file is fastest inside it.
The general lesson is worth more than this pair: sweep at the recipe you ship, not at a clean-room default. The ladder was run drafter-off for clean determinism and, for the configuration anyone actually runs, it produced the wrong ordering. Carrying the shipped flags from the first arm would have cost nothing and would have had the right answer in hand from the start.
Then the second measurement, at the other end of the window. All three arms below were deep-filled to 218,233 real tokens — 83% of the full 262,144-token window, not a shallow probe — on the card's 24,576 MiB of board VRAM, with the first post-prefill probe discarded. Each is judged against the desktop reserve defined in §02, and it is worth restating which half of it is measured and which is a choice: the 1,669 MiB worst case of the desktop's own share is measured,C12 the 127 MiB of load-to-load variation is measured, and the 1,796 MiB reserve built by adding them is this page's own derived threshold — the amount of slack it wants a configuration to keep before calling it usable with a screen attached, not an amount the operating system demands.
Arm at -c 262144, filled to 218,233 tokens | Board VRAM at depth | Slack | Decode at depth | Verdict against the 1,796 MiB desktop reserve |
|---|---|---|---|---|
| UD-Q2_K_XL, drafter on (n-max 4 / p-min 0.75) | 22,859 MiB | 1,717 MiB | 21.33 t/s | 79 MiB SHORT of the corrected 1,796 MiB reserve.C12 Run it headless, or turn the drafter off — that drops the requirement to 20,567 MiB and leaves 4,009, which clears comfortably. Tight, and published as tight — but the whole native window is resident and speculating |
| UD-Q2_K_XL, drafter off | 20,567 MiB | 4,009 MiB | 17.96 t/s | Keeps 4,009 MiB — comfortably more than the reserve. The drafter is still worth +18.8% here even though mean draft length has collapsed from 5.70 to 2.43 at this depth |
| UD-IQ4_XS, drafter off | 23,821 MiB | 755 MiB | 15.96 t/s | Keeps only 755 MiB — 1,041 MiB short of the reserve. It will run with no graphical session using the card (§02), and there is no room left to turn speculation on at all |
At the full window the 2.9-bit file beats the 4-bit file on every axis at once: 34% faster at depth, 962 MiB more slack, and it can run speculation at all, which the 4-bit file cannot. That last point is the one that decides it — UD-IQ4_XS reaches 23,821 MiB with the drafter already off, so there is nothing to trade. This is now measured directly rather than derived from arithmetic alone, and the number says it: full native context on the 4-bit file needs the card free of any graphical session.
Two honest costs ship with that recommendation. Filling 218,233 tokens took 424 seconds at 514.7 t/s of prefill — seven minutes before the first token of the answer. This is a long-document configuration, not an interactive one. And the two-constant VRAM model in §05 under-predicts this window by about 1,213 MiB — it predicts 19,353 and 21,647 MiB against measured 20,567 and 22,859 — because it does not count the compute buffers — the scratch memory the server needs while it is actually working, on top of the weights and the cache. §05 already says its arithmetic is a floor rather than a budget; this measures how big the gap between floor and reality gets at the top of the window, and it is big enough to matter. A reader budgeting from the formula alone would think the drafter-on arm kept 2.9 GiB when it keeps 1.7.
Putting all of it together — the ladder, the empty-answer audit, the requirement table, the depth series, the drafter pair and the full-window trial — gives this. Every "fits" below means fits and still leaves this page's 1,796 MiB desktop reserve free, which is a stricter test than fitting on the card:
| Your card, and what you want from it | Take | Because | Tier |
|---|---|---|---|
| 24 GB · everyday windows, up to about 98k | UD-IQ4_XS — stay where you are | On novel code into a nearly empty window it is 12.9% faster than the 2.9-bit file with the drafter on (n10/p0.5, greedy code), and that beats the size advantage outright. On prose filling about 90% of the window the picture is not an ordering at all: across -c 32768, 65536 and 98304 the three candidate files sit inside a band of 3.6 t/s or less, and this page does not rank files inside a band that narrow (§08). Either way there is no speed argument for moving, and of the three the 4-bit file has the lowest measured perplexity | measured |
| 24 GB · a window of about 131k | UD-Q3_K_XL — switch | 20% faster at that window — 41.74 t/s against UD-IQ4_XS's 34.68, both at n4/p0.75, where the p-min gate is part of the ordering (§08); at p-min 0, which UD-IQ4_XS’s cards and the unswept cards now carry, the ordering is not measured — and this is the one window where the separation is wide and reproduced: each file was loaded twice from scratch and every one landed within 1.4% of its own first load (§08). It needs 19,724 MiB with the drafter on, leaving 4,852 MiB free where the 4-bit file leaves 3,728 der, and it ties the reference file on accuracy — the two answered exactly one of 75 questions differently, p=1.00. The condition travels with the recommendation: the ordering was measured on prose fill, and the drafter's confidence gate is a property of your content. Spend ten minutes checking it on your own material before you commit | measured on prose fill |
| 24 GB · you need close to the full 262,144 window | UD-Q2_K_XL — switch | 34% faster at depth (its drafter at n4/p0.75), keeps speculation, comes 79 MiB SHORT of the 1,796 MiB this page reserves for a desktop — so run it headless or with the drafter off (§08), which drops it to 20,567 MiB and leaves 4,009, and ties the 4-bit file on accuracy (the two answered exactly one of 75 questions differently, p=1.00). The 4-bit file cannot hold this window with a desktop running at any speed. Price: a seven-minute prefill | measured |
| 16 GB | UD-Q2_K_XL at -c 65536, drafter on | Settled by arithmetic on measured requirements. A 16 GB card holds 16,384 MiB; keep the 1,796 MiB reserve free and 14,588 MiB is what you have to spend der. UD-Q2_K_XL at -c 65536 with the drafter needs 13,982 MiB — it fits, with 606 MiB spare der. UD-Q3_K_XL does not fit at that window at all: 16,906 MiB with the drafter, more than the whole card. The largest configuration of it that fits inside 14,588 is -c 32768 with the drafter off, at 14,322 MiB. So the smaller file gives you twice the window, a drafter you can keep, and 2,924 MiB less memory at the same window and drafter setting (13,982 against 16,906). It also ties the reference file on accuracy, p=1.00, for +6.07% of perplexity. On this 3090 those two configurations decode at 40.30 (drafter at n4/p0.75) and 36.44 t/s — but that is this card's bandwidth, not a 16 GB card's, and no measurement here can tell you what yours would do. §03’s 16 GB recipe ships this same configuration: UD-Q2_K_XL at -c 65536C15 | VRAM requirement measured · the subtraction derived · speed on a 16 GB card not measured |
| 12 GB | Borderline. Do the sum with your own screen | Measured directly 2026-08-26 — peak board VRAM with the file loaded and answering four short requests, minus the idle board reading taken just before the load — UD-Q2_K_XL at -c 32768 allocates 10,497 MiB of server memory MEASURED.C30 A 12 GB card holds 12,288 MiB, and what you can spend is that minus whatever is drawing your desktop. Do not subtract a desktop reserve from the 11,396 MiB figure in the requirement table above — that is a board figure which already contains this machine’s desktop; subtracting a desktop reserve from the card as well counts it twice. 11,396 minus the ~899 MiB of desktop on the measuring machine is 10,497, so the two readings never disagreed. Do the sum with your own screen: a light desktop of 1,179 MiB leaves 612 MiB spare, a heavy one of 1,669 leaves 122, and this page's deliberately pessimistic 1,796 MiB worst case comes up 5 MiB short. It is borderline, and which side of the line you land on is decided by your own desktop, not by this page. Run your screen off the motherboard or a second card and it fits outright. Two smaller files fit comfortably, by more than a gigabyte: QAT-Q2_0 at 9,437 MiB and UD-IQ2_S at 9,507 MiB MEASURED, both at the same window and flags (§08). Nothing below -c 32768 was measured, and a window smaller than that leaves this model very little room to think (§09). The answer on this card is a smaller model, not a smaller quantization of this one. §03's existing 12 GB recipe — Q4_K_M with most layers on the processor, 6–8 t/s — remains the only way this page can defend running this model there, and it is slow on purpose. For the Swift-1.5 files' 12 GB windows see Appendix A | VRAM requirement measured · the subtraction derived |
| 8 GB | No | The smallest file that still performs like the full-quality one is 9.154 GiB. It does not fit, and the files that do fit are on the wrong side of the boundary | derived |
| Anything below 2.48 bits per weight | Not recommended, on any card | Not because the score looks worse — because the model starts failing to answer. The empty count is exactly zero at every rung down to 2.912 and then rises 2, 3, 5, 28 (the empty-answer table). Empty replies first appear at the 2.481-bit rung itself, which is why this page recommends nothing below 2.912 either. Many are silent: of those four counts, 1, 2, 5 and 10 ended normally and returned zero characters, so nothing in your logs reports a problem at all. And at 1.835 bits the model loses the ability to stop — 20 of 75 answers ran into the cap, and the median answer jumps from a 416–502-token band held all the way down the ladder to 932 tokens | measured |
The boundary of what this machine can say. This machine owns one 24 GB card and has never loaded any of these files on anything else, so the 16 and 12 GB rows have to be split into three kinds of statement rather than dismissed as one. Quality is measured and it transfers — a file's perplexity and its benchmark answers do not depend on which card decodes them, so the accuracy tie between UD-Q2_K_XL and the 4-bit reference file is a measurement a 16 GB owner can rely on directly. The VRAM requirement is also measured, and it also transfers, because how many bytes a configuration asks for is a property of the model, the window and the flags rather than of the card: the 13,982 and 16,906 MiB figures above were read off this 3090 and are what those same configurations will ask for on any card (§08). The arithmetic that turns a requirement into a verdict is derived and its steps are printed in the cells: card total, minus the reserve, against the requirement. The speed is not measured, and on a card this machine does not own it never can be — every t/s in this section is this 3090's memory bandwidth. How fast UD-Q2_K_XL runs on a 5080 or a 4080 is not known here at all, and §15 lists it as a gap rather than filling it with an estimate.
UD-IQ4_XS against Q4_K_M — the two files this page recommends
These two files are the whole decision on a 24 GB card, and the honest summary is that they are tied on quality and not tied on anything else. At 13.3 GiB, UD-IQ4_XS is 2.1 GiB smaller than Q4_K_M, matches it on perplexity to within the measurement's own error bar, and — because decode is limited by memory bandwidth — is faster everywhere it was measured: prefill 1365 against 1303 t/s, plain decode 42.97 against 39.99 with no drafter, the same maths benchmark with speculation 73 against 63, the matched code sweep 93.9 against 81.7 — the first and last pairs measured on the same prompt in the same sweep. It also costs about 8% less energy per token. The saved memory becomes context or vision.
The measured ceilings for the smaller file, each row naming all four determinants (file · drafter flags · projector · desktop). Apart from the 218,233-token fill of the 262,144-token window and the figure derived for that window at load, these are memory readings at load — board VRAM, and the server's shared GPU memory where it is printed — each with one short request and no fill:C34 with vision and the n-max 4 drafter, 163,840 is fully resident at load (996 MiB slack); with vision and the n-max 10 drafter the same window keeps only 537 MiB, which is less than a busy desktop holds, so 131,072 is the largest window that still leaves a desktop comfortable (2,426 MiB slack) and 122,880 is what §03 ships. Text-only with the n-max 10 drafter the ceiling at load is 180,224 (847 MiB slack) — not the ~196k the projector arithmetic implied: 196,608 leaves 415 MiB and 212,992 leaves 280, with decode already sagging 56.6 → 51.9 → 49.2 across those three on a short probe. Do not try to read the budget model out of those three slack figures: the board VRAM total is saturating, so the growth goes somewhere the board VRAM figure cannot show — the server's shared memory use across the same three windows is 478, 638 and 744 MiB, which is the spill beginning before the board VRAM slack runs out. The full native 262,144 fits only text-only and drafter-off, at 23,216 MiB with 1,360 MiB of board VRAM left der — and that figure is derived for the moment of load (§05): filled to 218,233 real tokens on 2026-08-25 the same configuration reaches 23,821 MiB and keeps only 755, which is 1,041 MiB short of this page's 1,796 MiB desktop reserve, so the window now needs the card free of any graphical session by measurement rather than only by arithmetic (§08). With the projector loaded and the n-max 4 / p-min 0.75 drafter, shallow probes read 59–64 t/s from -c 32768 to the full native 262,144 with no collapse (2026-08-23) — that is precisely the probe artefact §05 documents: fill that configuration at -c 262144 to 90,886 tokens, greedy, and it delivers 8.0 t/s.C29
So why does Q4_K_M keep being named a default anywhere? Because there are two defaults with two scopes. Q4_K_M is this page's measurement base and its maximum-quality pick — perplexity leans slightly its way, it is the file most of the 08-21 and 08-22 numbers were taken on, and IQ-format decode support varies more across backends (everything here was measured on CUDA). UD-IQ4_XS with vision at -c 122880 is the reference machine's shipped daily default (§03, menu row 1). Short version: context-hungry or speed-hungry, take UD-IQ4_XS; maximum measured quality at 122k with little on screen, stay on Q4_K_M. Either is defensible, and the caveat on the Q4_K_M recipe's quality argument above applies to this paragraph too.
Swift IQ4_XS for one slot on a 24 GB card, Swift IQ3_XXS for a 16 GB card or two slots — ranked by perplexity on a shared tokenizer
On a 24 GB card take Swift IQ4_XS, the lower-perplexity of the two measured Swift-1.5 files (15.48 GB, 14.41 GiB, 4.5316 bpw der). Swift IQ3_XXS (12.32 GB, 11.47 GiB, 3.6076 bpw der) is the pick for a 16 GB card, where Swift IQ4_XS fits in no configuration at this page’s 1,796 MiB desktop reserve. On a 24 GB card it is the file for the longest window (§05) or for two slots, where with vision it holds 93,184 tokens per slot against Swift IQ4_XS’s 57,344 der (Appendix A). Neither file fits a 12 GB card (Appendix A). All three files — including the anchor (unsloth UD-IQ4_XS of the base model) — tokenize the wikitext-2 test corpus to the same 297,193 token ids meas, so raw perplexity across these files is a legal comparison: Swift IQ4_XS scores 6.7742 ± 0.04728 against Swift IQ3_XXS’s 7.3810 ± 0.05387 meas, a gap of 0.6068 (+8.96%, 8.47× the combined error of ±0.0717) der. That ranking is a quant rank — both files carry the same Swift-1.5 weights at different bit widths. Perplexity cannot rank either Swift file against the anchor, because the two models have different weights and any wikitext difference is drift from the corpus, not a quality rank. The quality comparison against the base model is in §09.
bartowski’s IQ4_XS of bytkim’s Qwen3.8-27B-pi, a fine-tune of Qwen3.8-27B for the Pi coding agent, is the same recipe at the same size: 15,475,952,128 bytes, 192 more than Swift IQ4_XS, with the tensor names, shapes and types, tokenizer arrays, MTP head (227.91 MiB, F32 + Q4_0) and chat template identical to Swift IQ4_XS’s (header read 2026-10-06) meas, so the memory and window figures for Swift IQ4_XS apply to it der. Its perplexity was not measured, so it is not ranked here; its measurements are in Appendix B.
Every file holds the same 866 tensors, named and shaped alike; their quantization types differ throughout each file, and the table names three of them: the token embedding, the output layer and the MTP head.
Files hashed and headers read 2026-10-04.
| File | Bytes meas | GB der | GiB der | bpw der bytes × 8 ÷ 27,320,697,856 | token_embd meas | output meas | MTP head meas |
|---|---|---|---|---|---|---|---|
| Anchor (unsloth UD-IQ4_XS, base model) | 14,252,845,984 | 14.25 | 13.27 | 4.1735 | Q3_K | Q5_K | F32, Q6_K, Q8_0 |
| Swift IQ4_XS (bartowski) | 15,475,951,936 | 15.48 | 14.41 | 4.5316 | IQ4_XS | Q6_K | F32, Q4_0 |
| Swift IQ3_XXS (bartowski) | 12,320,168,256 | 12.32 | 11.47 | 3.6076 | Q4_K | Q5_K | F32, Q4_0 |
The bpw column divides each whole file, header included, by its 27,320,697,856 parameters meas, the MTP layer among them; the ladder above divides by a fixed 27,000,000,000, which puts the same anchor file at 4.223 der, and Appendix A counts tensor bytes only, which puts Swift IQ3_XXS at 3.60 der. Every other Swift-1.5 tier, with its bits per weight on that tensor basis and the weight it puts on the card, is in Appendix A.
Each file’s perplexity on wikitext-2, with the error the tool prints; the shared tokenizer (above) makes the rows comparable. The anchor ran first and reproduces this page’s August 6.5956 (above) to the printed digit, drift 0.0% meas: the October runs sit on the August scale. Conditions: f16 KV unless the row says q8_0, -c 8192, 36 chunks = 294,912 positions, -fa on, --load-mode mmap, wikitext-2 test.
One sweep, 2026-10-04; its anchor row matches the August reference above, and a Swift row against any base-model row, August or October, is drift, not a rank.
| File | KV | Perplexity | Reading der |
|---|---|---|---|
| Anchor (UD-IQ4_XS, base model) | f16 | 6.5956 ± 0.04453 meas | reproduces this page’s August figure to the printed digit (0.0% drift) — the reference for this sweep |
| Swift IQ4_XS | f16 | 6.7742 ± 0.04728 meas | +0.1786 (+2.71%) against the anchor, 2.75× the combined error of ±0.0649 — different weights: drift from wikitext, not a quality rank |
| Swift IQ3_XXS | f16 | 7.3810 ± 0.05387 meas | +0.6068 (+8.96%) against Swift IQ4_XS — quant rank, 8.47× the combined error of ±0.0717. +0.7854 (+11.91%) against the anchor, 11.24× ±0.0699 — different weights: drift from wikitext, not a quality rank |
| Swift IQ4_XS | q8_0 | 6.7708 ± 0.04724 meas | −0.0034 (−0.05%) against f16 of the same file, inside the combined error of ±0.0668: a tie. The anchor’s own q8_0 cost, measured 2026-08-23 on the same corpus and chunks, is +0.309% (above; §03), also inside its error: neither file shows a measurable q8_0 cost. |
With reasoning off, the MTP drafter lifts decode from the 42.88–44.85 t/s drafter-off range to 84.93–100.81 t/s meas, a speculation gain of 1.89× to 2.35× der across the three files and two drafter settings. With reasoning on, both Swift files kept fewer drafted tokens per verify pass than the anchor on this one prompt, greedy: at n4/p0.75 Swift IQ4_XS decodes 59.96 t/s against the anchor’s 76.23 meas, keeping 1.479 drafted tokens per pass against 2.684 der, and for both Swift files n10/p0.5 is slower than n4/p0.75 (47.51 against 59.96, 43.95 against 50.61 meas), the reverse of the anchor.
One sweep, 2026-10-04; not comparable with the August rows above in this section.
| File | Regime | Drafter | Decode (t/s) meas | Gain over drafter off der | Acceptance meas | Accepted per verify pass der accepted ÷ (predicted − accepted) |
|---|---|---|---|---|---|---|
| Anchor (UD-IQ4_XS) | think off | off | 42.88 | — | — | — |
| Anchor | think off | n4/p0.75 | 87.42 | 2.04× | 0.923 | 3.281 |
| Anchor | think off | n10/p0.5 | 100.81 | 2.35× | 0.666 | 5.965 |
| Swift IQ4_XS | think off | off | 43.38 | — | — | — |
| Swift IQ4_XS | think off | n4/p0.75 | 94.33 | 2.17× | 0.917 | 3.348 |
| Swift IQ4_XS | think off | n10/p0.5 | 98.1 | 2.26× | 0.641 | 5.667 |
| Swift IQ3_XXS | think off | off | 44.85 | — | — | — |
| Swift IQ3_XXS | think off | n4/p0.75 | 84.93 | 1.89× | 0.927 | 3.142 |
| Swift IQ3_XXS | think off | n10/p0.5 | 99.53 | 2.22× | 0.686 | 6.071 |
| Anchor | think on | n4/p0.75 | 76.23 | — | 0.901 | 2.684 |
| Anchor | think on | n10/p0.5 | 83.53 | — | 0.590 | 4.582 |
| Swift IQ4_XS | think on | n4/p0.75 | 59.96 two outputs: about 67 and 53 | — | 0.837 | 1.479 |
| Swift IQ4_XS | think on | n10/p0.5 | 47.51 | — | 0.404 | 1.657 |
| Swift IQ3_XXS | think on | n4/p0.75 | 50.61 | — | 0.772 | 0.974 |
| Swift IQ3_XXS | think on | n10/p0.5 | 43.95 | — | 0.424 | 1.389 |
Conditions: -c 32768, -ngl 99, --parallel 1, -ctk q8_0 -ctv q8_0, greedy (temperature 0, top_k 1), 700 predicted tokens per probe. One prompt (a self-contained JavaScript red-black tree, code only; 43 tokens with reasoning off, 83 with it on meas) was sent three times per load, the first output discarded, in two passes in alternating order n=4 per cell (two probes per pass, two passes). With reasoning on, every probe spent all 700 tokens reasoning and printed no answer, so those rows are the decode speed of reasoning tokens. Reasoning off is --chat-template-kwargs {"enable_thinking":false}; reasoning on is the chat template’s default, which opens a <think> block, run with --reasoning-preserve; n4/p0.75 is --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75, and n10/p0.5 the same with 10 and 0.5. At temperature 0 the second pass repeated the first pass’s text probe for probe, so each cell’s acceptance rests on one or two distinct outputs for that one prompt. These speeds are greedy-only; speed at the temperature-1.0 sampler the cards ship is not measured as a band (§15). Decode here is at -c 32768 on a short prompt, not at depth (§06 covers depth). Greedy-identity check (reasoning off): Swift IQ4_XS produced byte-identical text with the drafter on and off; the anchor and Swift IQ3_XXS did not, and their own drafter-on runs did not always repeat: in six runs per setting the anchor gave two different texts at n4/p0.75, and Swift IQ3_XXS three at n4/p0.75 and two at n10/p0.5 meas, so their differences cannot be laid on the drafter alone. Identity on one prompt is evidence, not proof, that accuracy scored with the drafter off (§09) carries over to drafter-on use for Swift IQ4_XS; for the anchor and Swift IQ3_XXS this check cannot say either way.
A second sweep on 2026-10-07 used the same prompt, window and token count, two passes in reversed order, with Swift IQ4_XS alternated with the Pi fine-tune’s IQ4_XS (Appendix B). It reproduced this table’s kept-per-pass figures for Swift IQ4_XS at n4/p0.75 and n10/p0.5 (1.479 and 1.657 with reasoning on, 3.348 and 5.667 with it off) and added n-max 3 / p-min 0. With reasoning on, Swift IQ4_XS decoded 72.70 / 72.47 t/s at n3/p0 against 66.93 / 66.66 at n4/p0.75 in that sweep (×1.087), keeping 1.682 drafted tokens per verify pass against 1.479 at an acceptance of 0.565 against 0.837. With reasoning off the two were level (102.59 / 102.58 against 103.68 / 103.54). Under the cards’ temperature-1.0 sampler (4 prompts × 2 fixed seeds, reasoning on) n3/p0 ran ×1.153 and was faster on 8 of 8 probes, and it read 152 MiB less board VRAM at -c 32768. Absolute t/s from that sweep are not comparable with the rows above (§06); only its ratios are. Since 2026-10-08 the Swift cards run n3/p0 with one slot and n4/p0 with two, the n/p sweep’s best score per file and slot count (Swift IQ4_XS shallow reasoning tokens 1.189 of n4/p0.75 at xhigh, §06); their windows were filled at n4/p0.75, and neither change needs more memory (n-max 3 about 150 MiB less, p-min none).
The smallest Swift-1.5 file that performs like the full Swift-1.5 weights — the counterpart of the floor the ladder above sets for the base model — is not located here (§15). The script that locates it from each file’s divergence against the unquantised weights reports “Not enough rungs carrying fidelity data (0). A knee needs three.”: no Swift file was measured here for that divergence, and with two Swift files it would be one file short of three, because the anchor is a different model.
2-bit against 4-bit on a 225-task coding benchmark: the same score, and not the same cost
The tie this section measures is the same tie the paired question set finds, and both are on the ladder chart.
Everything above ranks these files by perplexity, by a 75-item question set, and by whether the code they wrote will actually run. None of that says what happens when a model has to repair its own work, which is most of what coding with a model consists of. On 2026-08-26 the two candidate files were each run through aider’s official polyglot benchmark — 225 programming exercises across several languages, each one shipped with its own unit tests — one full arm per file. It is the closest this page comes to measuring these two files at the work most readers will actually give them.
They finish level. They do not spend the same amount getting there. Both files solved 96 of the 225 exercises, a pass rate — the share of exercises whose every unit test passed — of 42.7% each meas. The verdict is therefore not in the score. It is in the bill: the 2-bit file needed about 20 to 45% more tokens, ran 20 to 36% slower, and cost about 32% more energy for every exercise it solved. Those ranges are not vagueness. They come from counting the same thing more than one way — over every exercise an arm solved, and paired over only the exercises both arms solved — and each way is printed on its own line in the table below.
The mechanism is what makes this worth remembering. The benchmark gives every exercise two attempts. The model writes an answer, the harness runs the unit tests, and if they fail the model is shown the failure and asked to fix its own mistake. Most of the successes on both files arrive on that second attempt — the 4-bit file took 69% of its passes that way and the 2-bit file 80%. That is the regime this whole comparison lives in, and it is where the two files separate. On the first attempt the 4-bit file solved 13.3% of the 225 exercises against the 2-bit file’s 8.4% meas. Because 2-bit quantisation introduces more noise, it misses clean first answers more often, so it leans harder on the repair step to catch up — and a repair costs a whole extra round of reading a failure and writing a fix. That extra fixing is the overhead, and the overhead is the entire finding, because the score is a tie.
| What was counted | UD-IQ4_XS 4-bit · 14.25 GB | UD-Q2_K_XL 2-bit · 9.83 GB | What it means |
|---|---|---|---|
| Exercises solved, of 225 — the score | 96 · 42.7% | 96 · 42.7% | a tie, and the reason the rest of this table is the story |
| Solved on the first attempt, before any test failure was shown | 13.3% | 8.4% | the 2-bit file misses more clean first answers — this is the cause of every row below |
| Completion tokens per solved exercise, each arm counted over its own solved exercises | 2,102 | 3,248 | +55% |
| Completion tokens, median exercise, counted only on the exercises both files solved | reference | 1.198 times | +20% |
| Completion tokens, total over those same jointly solved exercises | reference | 1.421 times | +42% |
| Wall-clock seconds per exercise, whole run | 53.1 | 72.2 | +36% |
| Wall-clock time, median exercise, jointly solved only | reference | 1.197 times | +20% |
| Board energy per solved exercise, paired over the 36 exercises both files solved and for which board power was logged | reference | 1.319 times | +32% |
| Board energy per completion token | 6.915 J | 6.657 J | the 2-bit file is 3.7% cheaper per token — and needs far more of them, which is why tokens rather than speed are the thing to watch |
| Exercises solved by exactly one of the two files | 30 | 30 | 60 disagreements — an upper bound, for the reason in the box below |
Conditions, and they travel with every figure above meas: one RTX 3090 at its stock 350 W board limit, llama.cpp, aider’s official polyglot benchmark run unmodified in its own container, the whole edit format, reasoning off, a 32,768-token window, greedy sampling (temperature 0), and board power only — read from the card’s own sensor, which excludes the power supply’s conversion loss, the processor, system memory, drives and the display (§15). The 4-bit file is UD-IQ4_XS at 14.25 GB and the 2-bit file is UD-Q2_K_XL at 9.83 GB; the difference between them is 4.4 GB of VRAM. One machine, one benchmark, one arm per file. And the sampler caveat this page attaches everywhere else applies here too: greedy decoding holds the text still so that the file is the only thing changing, but the recipes in §03 ship --temp 1.0 --top-p 0.95 --top-k 20, which is not what was run here.
The two files do not solve the same 96 exercises. Sixty of the 225 were solved by exactly one of them, thirty by each, so the tie in the score is a tie in the total rather than agreement about which work is within reach. Read 60 as a ceiling on what the quantisation explains, and not as the amount it explains. The reason is that this benchmark’s own flip rate — how many exercises change verdict when the same file is run twice with nothing changed at all — is not zero, and it has now been measured meas. Measured 2026-08-27: the same file, the same flags and the same machine, run twice — UD-IQ4_XS, whole edit format, reasoning off, -c 32768, n-max 4 / p-min 0.75, greedy, 113 exercises paired by case and language. Both runs scored 43.4%, identical to one decimal place, while 18 of the 113 exercises flipped — 15.9%, 95% interval 10.3 to 23.8, and the split was exactly symmetric at nine each way. A sixth of the individual verdicts moved underneath a pass rate that did not move at all, which is why two single runs cannot be read exercise by exercise.
What that does to the 60. If the two files were identical in capability, comparing one run of each is exactly comparing two runs of one file, so the expected discordance is the noise rate. Both files have now been measured, and they do not have the same one:
| comparison | flips | rate | 95% interval |
|---|---|---|---|
| cross-quantisation, one run of each file | 60 of 225 | 26.7% | 21.3 – 32.8 |
| UD-IQ4_XS against itself | 18 of 113 | 15.9% | 10.3 – 23.8 |
| UD-Q2_K_XL against itself | 25 of 113 | 22.1% | 15.5 – 30.6 |
Which floor you test against decides the answer, and that is the whole finding. The question is whether the cross-file disagreement rate of 26.7% is too high to be noise alone. A p-value measures that probability — below 0.05 is the conventional line for “probably real” — and a z-score is how many standard errors the gap spans. Against the 4-bit file’s floor of 15.9%, the gap is wide enough to clear the line (z = 2.21, p = 0.027). Against the 2-bit file’s noisier floor of 22.1%, the same 26.7% is easily explained by chance (z = 0.91, p = 0.364). Pooling both files’ retests — 43 flips out of the 226 exercises both retests covered, a combined floor of 19.0% (interval 14.4 to 24.6) — lands just above the threshold (z = 1.93, p = 0.053). So no quantisation effect is demonstrated: the verdict depends entirely on which floor, and two of the three say no. The 60 exercises these two files disagree on are consistent with what this benchmark does to a single file run twice.
A noise floor taken from one arm cannot validate a two-arm comparison meas. The single-floor test that appeared to show a quantisation effect at p = 0.027 compared the 60 against the floor of the quieter file alone, silently assuming both arms are equally noisy — and that assumption was doing all the work. What also cannot be claimed: the two floors themselves — 15.9% against 22.1% — do not differ significantly from each other (z = -1.19, p = 0.236), so “the 2-bit file is noisier” is a point estimate, not a finding, however tempting its direction.
Nor does it license subtracting a floor from 60. Both floors are measured now, and neither subtraction is valid — a noise floor is not a constant to deduct but a rate against which a gap is tested, and neither gap clears the test. Two things are known about the mechanism. First, greedy decoding is byte-reproducible on this rig only when the work is shaped the same way twice. §10 measured that directly and it held: 139 of 139 prompts reproduced byte-for-byte when the cap and the window were raised. An agentic run is the case that does NOT hold it fixed — context length, cache reuse and how many speculated tokens are accepted all differ from request to request, and a change in the shape of a batch changes the order of a reduction, which flips a near-tie between two candidate tokens — the same effect §06 measured as about 1% of answer length between configurations that should have produced identical text. Second, 69% of the 4-bit file’s passes and 80% of the 2-bit file’s arrive as a repair after a failure, and that is exactly the regime in which a small perturbation flips an outcome most easily.
2-bit works, and it reaches the same final accuracy. It is less stable getting there and it is more expensive. Take UD-Q2_K_XL when you need the 4.4 GB of VRAM it saves — on a 16 GB card, where it is this page’s pick because the larger files do not fit (§08), or on a 24 GB card at the full native window, where it is the only one of the two that can still speculate (§08). Otherwise take the 4-bit file. Paying 20 to 45% more tokens and about 32% more energy per solved exercise to arrive at the same score is a poor trade whenever the memory is there to avoid it. And the shape of the result is worth carrying away on its own: the 2-bit file is genuinely the cheaper of the two per token and still the more expensive one per finished job, because it needs so many more tokens. A cheaper token is not a cheaper answer.
The limit that remains: nothing here has watched a long agentic loop
Almost every number on this page comes from single-turn prompts of at most 16,384 tokens. You ask once, the model answers once, and the answer is scored. The coding benchmark above is the one exception, and it is a modest one: it gives the model a second attempt with its own test failure in front of it, which is two turns rather than one. That is still not how a coding agent works. An agent calls a tool, reads the result, calls another, and keeps going for dozens of turns — and a model that goes wrong once in fifty turns will fail an agent while scoring perfectly here. So the smaller files on this page are measured over one turn and over two, and nothing here has ever watched one run for fifty.
A video published as "Qwen 3.8 27B Quantizations Q1 - Q8 compared" (channel reported as Luke's Dev Lab; the title was confirmed but the video was not watched from this machine, so everything below is cited from a viewer's written breakdown rather than measured here) ran the same unsloth files at high reasoning effort against three multi-turn agentic tasks: a Kanban web app, a Blender 3D asset job over MCP, and a Godot 4 game.
Where it agrees with this page. It found UD-Q2_K_XL one-shotting a full web app with no intervention and recommends it for 16 GB cards — which is this page's 16 GB pick, reached independently from perplexity, paired accuracy and a memory requirement. It found the 4-bit class the overall sweet spot, which is what this page ships. And it found IQ1_M unusable on every task, failing by locking into infinite loops — the same failure this page detected from short prompts at no GPU cost, as five empty answers out of 75, every one of them terminating normally and returning nothing (§08). That agreement matters in a specific way: this page's accuracy score for that file was 85.30, which looks survivable. The empty-answer count was the number that saw the problem.
Where it finds something this page cannot see. It reports UD-Q2_K_XL and UD-Q3_K_XL both failing the Godot task — one locking into a tool-calling loop immediately, the other generating crashing scripts until the project was unusable. Both of those files return zero empty answers here and tie the 4-bit reference on paired accuracy. There is no contradiction: those are long multi-turn runs, and nothing on this page tested one. Error compounding across turns is precisely the failure mode a single-turn benchmark is blind to.
What to take from it. If you are running chat, single-file coding, or web work, the recommendations above stand and now have independent support. If you are running an agent, the 2.9-bit file is measured over two turns and unmeasured beyond them. On aider’s 225-exercise polyglot benchmark, one arm per file, reasoning off, the 2.9-bit file matches the 4-bit file’s 96 of 225 (42.7%) while spending 20 to 45% more tokens and about 32% more energy per solved exercise (the part above). What that does not establish is a loop of dozens of tool calls, and the caution that survives is a specific one rather than a blanket one: one machine, one benchmark, one arm per file, reasoning off. The files below 4 bits that were not in that run remain untested for this kind of work. Try any of them on your own workflow before trusting it, and watch specifically for a loop that never terminates rather than for wrong answers. That tester also went above 4 bits, to Q5, Q6 and Q8, and reports those needed for the hardest visual and game-logic work; this page measured nothing above 4.4 bits per weight and has no opinion there at all.
This model has a dial that decides how long it thinks before it answers, and the honest finding is that the dial moves hours, not scores. On a 175-prompt benchmark suite run on this machine — all seven sets scored, the two open-ended ones by a blind three-seat judge panel (2026-08-24, §09) — the three levels scored 80.2, 80.3 and 79.7 out of 100, a spread of 0.6 points at a sample size where a single benchmark cell is worth about ±16 points, which is a tie — while the wall clock went from 1.0 to 1.5 to 2.7 hours and the average answer grew from 830 to 2,217 tokens. So: use medium when you are waiting at the keyboard, and xhigh when you are not. This page ships xhigh as the default in every recipe whose window can hold it and whose speed makes the wait sane — the 12 GB recipe's window holds it and still ships medium, because at 6–8 t/s an answer takes three hours — for one reason that the benchmark suite cannot measure and the smaller judged tests could: on complex single-file tasks, xhigh was the only level that delivered 100% of the specification with no fatal defect, twice out of two. Buy it as insurance against an incomplete answer, not as a higher score. And not for prose: the two open-ended sets, once judged, are the only place in this whole campaign where the levels separate at all, and they separate against xhigh — it lost to medium on 14 of 25 MT-Bench prompts and won on 3 (§09). For writing, medium.
The strongest evidence: a 175-prompt suite, three levels, a tie
Measured 2026-08-23 on the reference 3090: five scored benchmark sets at 25 questions each, per effort level, greedy decoding, seed 42, no drafter (to keep the scores clean), 525 generations in 8.48 hours of GPU time. Two further sets — ALPACA and MT-Bench — needed a judge and were scored on 2026-08-24 by a blind three-seat panel (§09 below). Both indices are ties: the seven-set index (80.2 / 80.3 / 79.7) is what the benchmark protocol specifies; the five-set mechanically-scored index (82.1 / 80.5 / 81.3) is kept because every earlier comparison on this page is against it. They are comparable only with another run using the same sets, and the seven-set index only when both runs’ judged sets were rated in one judge session (§09).
| Benchmark (n=25 each) | Scorer | low | medium | xhigh |
|---|---|---|---|---|
| GSM8K | exact match | 100 | 100 | 100 |
| MATH-500 | exact match | 96 | 100 | 100 |
| HumanEval | execution pass@1 | 100 | 96 | 92 (2 trunc) |
| MBPP | execution pass@1 | 92 | 84 | 92 (1 trunc) |
| MeetingBank | ROUGE-L F1 | 22.6 | 22.4 | 22.3 |
| ALPACA | judge 1–10 (3-seat panel) | 70.2 | 74.1 | 72.1 (1 empty) |
| MT-Bench (turn 1) | judge 1–10 (3-seat panel) | 80.7 | 85.3 | 79.7 |
| Composite index over the 5 mechanically scored sets | mean | 82.1 | 80.5 | 81.3 |
| Composite index over all 7 sets (what rule 21 actually specifies) | mean | 80.2 | 80.3 | 79.7 |
Conditions: UD-IQ4_XS, llama.cpp build 10502, q8_0 KV, -ngl 99, --parallel 1, temperature 0 / top-k 1 / seed 42, no speculative decoding. This table is a merge of two caps, and it matters which cell came from which. Six cells were re-run at --max-tokens 32768 with -c 65536 — those are the ones that had truncated: MATH-500 at low and at xhigh, and HumanEval and MBPP at medium and at xhigh. The other nine scored cells are the original run at --max-tokens 16384 with -c 32768, carried across unchanged. What licenses carrying them is not an argument but the byte-comparison in §10: 139 of 139 untruncated prompts reproduced identically when the cap and the window were both raised. Decode speed, and which instrument read it: the benchmark harness reports 42.2 / 42.0 / 41.9 t/s as an unweighted mean over requests, while the server's own token-weighted timings read 38.4–40.2 t/s for the same arms. The harness figure is pulled up by short, shallow answers. The two are never averaged together on this page (§15). The three scorer names in that table are defined in §02: execution pass@1 is the share of problems whose first generated program runs and passes its tests, and ROUGE-L F1 measures how much of a reference summary's wording a generated summary reproduces. MeetingBank's ~22 is a property of the metric pairing rather than a failure: the model writes the 60–120-word summary the prompt asks for — measured across all 75 MeetingBank generations at 79–129 words, mean 105 — while the reference is a ~40-word resolution title, so ROUGE-L F1 is structurally capped. It is flat across the three arms and contributes no signal. Note also that mixing a ROUGE score into a mean of percentages is what makes this an index rather than an accuracy.
Every individual cell above is 25 questions, so one flipped question moves that cell by 4 points and a single cell carries roughly ±16 points of confidence interval. The composite pools about 125 scored samples per arm, which is why it is the interpretable number — and it still cannot resolve a 1.6-point difference. Two of the five sets are also saturated: GSM8K returned 100 for all three levels, which means it discriminates nothing at any sample size, and MATH-500 returned 100 for two of them. The correct reading is that this suite could not detect a quality difference between the three effort levels, not that there is none. §10 puts numbers on how many questions detecting one would take.
The two sets a machine cannot score, and what a judge found in them
Five of the seven benchmark sets can be marked by a computer: there is a right answer, or the code runs or it does not. The other two — ALPACA and MT-Bench — ask for open-ended writing, and the only way to score writing is to have something read it. On 2026-08-24 the transcripts were scored by a blind panel.
The reader was a panel of three independent Claude Opus 5 judge seats, working blind: every answer carried an opaque identifier, the mapping from identifier to effort level was sealed in a file no seat could open, and each seat received the answers in its own shuffled order so that ordering effects could not line up across seats. All three seats rated all 150 answers, on the standard 1-to-10 single-answer rubric, giving 450 ratings with none missing. Ratings convert to the 0–100 scale as (rating − 1) ÷ 9 × 100.
The answers were written by Qwen and read by Claude, so nothing here is a model grading its own work. But this is not independence in the strongest sense: the judge and the author of this page are both Claude models, which makes them a correlated instrument. That is why the panel is three seats rather than one, why the spread between seats is published beside every mean, and why the conclusions below lean on the paired test rather than on the raw scores. Seat-to-seat spread averaged 0.28–0.92 rating points out of ten, so the seats largely agreed; agreement is not the same as being right.
| Set (n=25 each) | Measure | low | medium | xhigh |
|---|---|---|---|---|
| ALPACA | mean rating, 1–10 | 7.32 | 7.67 | 7.49 |
| score, 0–100 | 70.2 | 74.1 | 72.1 | |
| MT-Bench (turn 1) | mean rating, 1–10 | 8.27 | 8.68 | 8.17 |
| score, 0–100 | 80.7 | 85.3 | 79.7 | |
| Mean spread between the three seats (rating points) | 0.28–0.92 | 0.56–0.64 | 0.44–0.56 | |
Conditions. MT-Bench's generations are the 16,384-cap ones described above; it never truncated. ALPACA's xhigh answers come from the raised-cap re-run — one of its 25 answers spent its whole budget inside a reasoning block and returned nothing, so the budget rule required a re-run at 32,768. The re-run reproduced it exactly: the same answer consumed all 32,768 tokens and again returned nothing, while the other 24 came back token-for-token identical. That is not a budget shortfall, it is a generation that does not terminate, so the remedy is exhausted and the score is final, carrying one disclosed non-terminating answer that all three seats rated 1. Because the re-run replaced one arm's answers, all three ALPACA arms were re-judged together, so the three numbers come from one session rather than two. Judging protocol, packet builder and scorer: scripts/bench/judge-panel.py; every rating and the sealed key are kept with the run data.
The finding, and it is the only thing in this campaign that separates the effort levels at all. Because the same 25 prompts were put to every level, the arms can be compared prompt-by-prompt rather than only mean-against-mean, which is a far more sensitive test. Doing that (a bootstrap: the same set of per-prompt differences is re-scored 20,000 times with the prompts drawn at random and with replacement, and the range the answer falls in 95% of the time is reported): on MT-Bench, medium beats xhigh by 0.51 rating points, interval +0.21 to +0.80, and the count is lopsided — xhigh wrote the better answer on 3 prompts, the worse on 14, and tied on 8. All five other pairings are ties, ALPACA included.
medium against low on ALPACA, and the judge’s own repeatabilitymedium against low on ALPACA is a tie: +0.35 rating points, interval −0.04 to +0.73. An earlier judging pass had put the gap at +0.44 (interval +0.013 to +0.907), clearing zero by thirteen thousandths — marginal, not a finding. The re-judge confirmed it was noise.
The re-judge also measured how repeatable the judge itself is. Seventy-four of the seventy-five answers were byte-identical between the two passes, so any change in their ratings is the instrument, not the model. Across those 74: mean absolute change 0.33 rating points, largest change 1.33, and 25 answers scored exactly the same both times. Averaged over 25 items an arm’s score moves by at most 1.0 point on the 0–100 scale — which is why a 0.51-point paired gap on MT-Bench is a result and a 0.44-point one on ALPACA was not. Measured on ALPACA, between two sessions that judged the same packets; MT-Bench’s packets were never re-judged, so applying this band to the MT-Bench gap is an assumption, stated as one. Measured 2026-10-04: the 25 MT-Bench answers of the xhigh arm above, re-run as the anchor of the Swift comparison (§09), are byte-identical to August’s, yet the same 25 items scored 79.7 in the August session and 87.0 in the 2026-10-04 session — 20 items rated higher and 1 lower, 7.3 points on the 0–100 scale. In a later session, with different answers beside them, the same answers moved well past that band. The cause of the move is unmeasured (§15). Judged scores are read only as paired comparisons inside one session.C37
xhighThis does not overturn the recommendation in §09; it sharpens it. The case for xhigh was never a higher score — it was completeness on complex code tasks, where it was the only level to deliver 100% of the specification with no fatal defect, twice out of two. That finding stands and is categorical. What the judged pair adds is the other half: on MT-Bench, more thinking made the answers measurably worse, not better — while on ALPACA the same pairing is a tie. Longer deliberation produced more places to go wrong: invented specifics, padded lists, and in the worst cases output that stopped being an answer at all. One of six comparisons survived a 95% test, so this is a set-specific result, and overthinking is a candidate explanation for it rather than an established one. So: xhigh for hard code you need finished, medium for writing. Five of the nine recipes ship medium and four ship xhigh (§03); the 24 GB default is one of the four xhigh recipes. And a reasoning_effort field in the request body does not switch the level: llama-server ignores it, which is why every recipe sets it at launch. A per-request chat_template_kwargs object is a different field: on build 10502 the server’s /apply-template rendered the low and xhigh effort lines from it over a launch setting of medium, and no line without it (2026-10-06, on the Pi fine-tune’s IQ4_XS, whose template is the Swift files’; Appendix B); generation through /v1/chat/completions with it was not tested.C41 The verified way to change level is relaunching with a different --chat-template-kwargs, and raising max_tokens to about 120,000 for xhigh, which fits one turn per 131,072-token window.C20
What the seats caught that no mechanical scorer could. The value of a judge is not only the score; it is the failure modes a right-answer test is blind to. Three kinds turned up, and all three matter more than the point differences above.
Confident invention. The travel-writing prompts drew fluent, well-structured answers containing things that are simply not true — a “Hawaiian $2 bill”, a claim that the first surfers ever to ride a wave did so at Sunset Beach, a misdated battle, a hula centre that does not exist, and mistranslated Hawaiian words. A Chicago-transit prompt produced invented Metra line names, expressways matched to the wrong interstate numbers, and a count of “14 L lines”. A finance prompt turned a 1-for-4 bonus issue on 1,000 shares into 1,500 shares. None of this is detectable by perplexity, by an exact-match scorer, or by running code.
Degeneration well short of any limit. One low MT-Bench answer, asked for a short story, began spelling out numbers and never stopped — “one hundred and one, one hundred and two…” — and ended without a story. It used 1,682 tokens, nowhere near a cap, so nothing in the harness flagged it; all three seats rated it 1. One xhigh ALPACA answer expanded into 100 near-duplicate list entries that collapsed into nonsense, and another repeated the same two words three times each to pad a list to the requested length.
A failure that belongs to the model, not the setting. On one ALPACA prompt demanding a full sentence, low and medium independently returned the same bare noun phrase, and every seat rated both a 3. When two settings fail a prompt the same way, the effort dial is not the variable.
| Per arm (full 7-set run at the 16,384 cap) | low | medium | xhigh |
|---|---|---|---|
| Wall clock | 1.00 h | 1.47 h | 2.70 h |
| Mean output tokens (all 7 sets) | 830 | 1,228 | 2,217 |
| Decode, mean (harness, unweighted over requests — the server's token-weighted figure for the same arms is 38.4–40.2) | 42.2 t/s | 42.0 t/s | 41.9 t/s |
| Truncated at the 16,384 cap (scored sets) | 1 | 2 | 8 |
xhigh costs 1.84× medium's wall clock on this suite and produces 1.81× the tokens — the same ratio, because decode speed is flat across the three levels. On a single long authoring task the multiplier is higher: xhigh cost about 4× medium’s wall clock. Both are measured and neither is wrong; the multiplier is task-shaped, not a constant, and it grows with how much room the task gives the model to keep thinking.
A truncated answer scores zero, and xhigh truncates most. At the 16,384-token cap the three arms scored 81.3 / 80.5 / 77.3 — quality falling as effort rises, an attractive, quotable, wrong headline. The entire xhigh penalty was the cap: it truncated 8 times in the scored sets against 1 and 2 for the other levels. Raise the cap to 32,768, re-run only the affected sets, and the ranking becomes the 82.1 / 80.5 / 81.3 above. One prompt needed 18,273 tokens and was correct; another needed 17,025 and was correct; both had scored zero at 16,384. Three xhigh prompts still exceed 32,768 and are genuine runaway loops rather than long solutions, and they are counted in the table.
The other evidence: judged deliverables, at n=2
The suite above measures whether the model gets short, self-contained questions right. It does not measure whether it produces a complete working deliverable, and that is where the effort dial visibly earns its keep. Measured 2026-08-22: the same ~1,700-token coding task — build an animated aquarium page — run twice per level at temperature 1.0, with the six resulting pages graded blind against the specification by independent reviewers.
| Effort | Tokens (2 runs) | Wall clock | Blind quality score /100 n=2 | GSM8K n=200 |
|---|---|---|---|---|
low | 16.7k–29.8k | 4:34–8:34 | 40, 47 — both runs produced pages whose JavaScript crashes (one never draws a frame, one freezes on its first) | 96.0% (0 truncated) unaudited grader |
medium | 20.7k–23.4k | 5:54–7:04 | 86, 93 | 97.5% (0 truncated) unaudited grader |
xhigh | 73.1k–75.8k | 24:58–26:48 | 92, 95 · 100% of the specification, both runs | 95.0% (94.0 under a 4,096 cap†) unaudited grader |
† the first run's 4,096-token cap cut 5 of xhigh's 200 answers mid-thinking; re-run at a 16,384 cap it scored 95.0% with zero truncations, so the budget had cost about one point. Only the xhigh arm needed re-running, because greedy decoding is deterministic and the never-truncating arms are byte-identical under any larger cap — a licence established by measurement: 139 of 139 prompts reproduced byte-for-byte (§10). Size your token cap against your hardest benchmark set, not against GSM8K: "16,384 proved sufficient" was true on GSM8K. On the seven-benchmark suite above, the same 16,384 cap truncated xhigh eight times in the scored sets, and three prompts exceeded even 32,768. Size your cap against your hardest set, not against GSM8K. The GSM8K percentages in this table also carry the grader caveat from §08 and should be re-graded before they are quoted again.
The findings from the judged runs, stated plainly. The dial changes the answer sharply at one end and hardly at all at the other, and neither is where you would guess. The sharp change is between low and medium, and it moves with task difficulty: low matched the others on short maths, then shipped fatally broken JavaScript on the complex page twice out of two, with one broken run burning more tokens than medium. The evidence base is two runs at temperature 1.0, on one task — and both directions have been observed. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) ran the same task once per level and got the opposite defect direction: its low page rendered correctly while its medium page shipped a real initialisation bug — the resize routine ran after every creature array was built, so every mobile creature spawned in the top-left corner, violating the specification's "distribute them across the tank". Three runs across two campaigns, both directions observed. What is solid is that a fatal defect is a live risk at every effort level at n=1 — not that low specifically manufactures them. Above that point the change is small: medium and xhigh overlap on score, and medium's own run-to-run spread beats the between-level gap. Only xhigh delivered 100% of the specification with zero fatal defects on both runs, and it won both blind head-to-head rankings. Its edge is completeness insurance, not average score — and the 175-prompt suite above says the same thing from the other side: on short scored questions, there is no score to win.
Both bodies of evidence above are built from greenfield tasks: write this page from scratch, solve this self-contained problem, implement this function. Exactly one read-then-edit task against existing code was ever run on this machine, and it failed. An agentic pipeline was validated end to end on one real sandboxed repair task in 2026-08-23: 22 model calls, 9 minutes 36 seconds of wall clock, and 0 of 96 target tests turned green, though all 561 previously-passing tests kept passing — so the harness worked and the fix did not. That is one sample of plumbing, not a quality measurement, and no scored read-then-edit suite was run at any effort level. Open a file, understand it, change one thing without breaking the rest is the shape of work an agent actually does, and it is the use this page recommends the model for in §13 and §14. So the quality verdict here is scoped to greenfield single-file work, and the measurement that would close the gap is named in §15's register of what was not measured.
The independent re-measurement read the model's chat template and then proved the knob reaches it. Four things worth knowing before you set it: only low, medium and xhigh are accepted (anything else raises an error) — and high is silently aliased to xhigh, so a client sending the familiar "high" gets the most expensive setting, not a middle one. That aliasing belongs to the unsloth UD-IQ4_XS file’s template (9,993 characters, §14); lmstudio-community’s Q4_K_M of the base model, the repository this page’s Q4_K_M recipe downloads from, carries the 8,952-character template that the Swift-1.5 and Qwen3.8-27B-pi files also carry byte for byte, and that template raises on high instead of aliasing it (jinja2 render; headers read 2026-10-06/07).C40 Only low and xhigh inject a reasoning-instruction system line; medium injects nothing at all — it is the template's unmodified default path, not a third instruction. The proof that the setting is not being swallowed is the prompt-token differential: identical task, prompt_n of 1,689 / 1,659 / 1,701 for low / medium / xhigh — medium's is the short one, exactly as the template predicts. And separately from effort, --chat-template-kwargs "{\"enable_thinking\":false}", which works per request and not only at load, is the switch between the two token regimes §06 keeps separating.
Your window sets an effort ceiling
Effort has a second price besides time: thinking tokens live inside the context window (§02), so a window that cannot hold a level's thinking cannot offer that level at all. Measured appetite on one hard task: xhigh’s appetite is a distribution, not a number. Four samples of the same hard task give 61,500, 73,100, 75,800 and one run that hit a 65,536-token cap and returned nothing usable. Read it as 61,500–75,800 straddling 65,536 — which is the operationally important part: the re-run at a 120,000-token cap then wanted only 61,476 tokens, fewer than the cap the previous sample had blown through. Those four are August’s. Two more runs on the same file and task, measured 2026-10-04 in a different sweep (temperature 1.0, -c 131072, uncapped), generated 22,718 and 80,213 tokens (below), so the spread is wider than the August four show at both ends, and the longer run fits every window of 100k or more. A cap sitting near the middle of that spread looks generous and is a truncation machine. Size caps and windows against the distribution's upper tail, never its middle.
Your -c window | Levels offered | Level not offered, and why |
|---|---|---|
| ≥ 100k (24 GB+, DGX Spark, Arc Pro B70, 12 GB offload) | low · medium · xhigh | none — the 61,500–75,800 appetite distribution fits with room for a prompt and an answer, including its upper tail |
| < 76k (16 GB cards and Arc Pro B50, at 65,536) | low · medium | xhigh not offered — a WINDOW limit. Its measured appetite alone can overflow the window mid-thought, and a truncated xhigh run scores worse than a completed medium one and often returns nothing at all. At -c 65536 the window reaches into the 61,500–75,800 appetite band rather than clearing it — so xhigh would fit on its shortest runs and truncate on its longest, and you would not know which you got. Not offeredC15 |
| 12 GB offload at ~112k | low · medium shipped; xhigh for unattended runs | a WALL-CLOCK limit, not a window one. The window holds the appetite fine; at 6–8 t/s an xhigh answer takes about three hours against about one for medium. The fix is patience, not memory |
| integrated GPUs at 65,536 | low · medium | both limits at once, and they have different fixes. 61,500–75,800 of appetite against a 65,536-token window (fix: raise -c — an integrated GPU borrows system memory, so it can), and 3–4 hours per answer at 5–7 t/s (fix: nothing but patience) |
Set it server-side in llama.cpp with --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"; graphical front-ends usually expose it as a reasoning-effort setting. llama-server ignores a per-request reasoning_effort field, which is why every recipe in §03 sets it at launch and why the multi-agent setup in §14 needs two servers to route between two levels. (That is the request-body reasoning_effort field. A per-request chat_template_kwargs did reach the template at the server’s /apply-template on build 10502, 2026-10-06; generation with it was not tested, so two servers remain the verified way to serve two levels; Appendix B.)
Every Swift card in §03 offers xhigh on the 24 GB RTX 3090, judged on one task at two runs per file. On a 16 GB card, Appendix A gives Swift IQ3_XXS at most 43,008 tokens der (at a 1,272 MiB desktop reserve; 30,720 at this page’s 1,796 MiB) and Swift IQ4_XS no fit, so neither holds one of the xhigh runs below, and the window rule in the table above applies. On the aquarium task at temperature 1.0 (2026-10-04), one xhigh run generated 55,621–60,991 tokens on Swift IQ4_XS, 52,929–67,455 on Swift IQ3_XXS and 22,718–80,213 on the anchor (unsloth UD-IQ4_XS of the base model) meas n=2 — each count is everything the model generated, the thinking and the finished HTML page together; the smallest per-slot window across the six cards is 93,184, so the 1,701-token prompt and the longest Swift run fit in every one of them. With --reasoning-preserve, each turn’s reasoning stays in the window, and the table below shows how many xhigh turns each card holds before it fills. The overflow event — what the server does on the turn after the window fills — is not measured for the Swift files (§15).
| Configuration | Window per slot the card’s setting; fully resident windows in §05 | Longer of two xhigh runs meas | Fits one turn der | Turns per window der | Minutes per xhigh turn der |
|---|---|---|---|---|---|
| Swift IQ4_XS, text, 1 slot | 159,744 | 60,991 | yes | 2 | 16.6 |
| Swift IQ4_XS, vision, 1 slot | 128,000 | 60,991 | yes | 2 | 16.6 |
| Swift IQ3_XXS, text, 1 slot | 216,064 | 67,455 | yes | 3 | 17.6 |
| Swift IQ3_XXS, text, 2 slots | 123,904 | 67,455 | yes | 1 | 17.6 |
| Swift IQ3_XXS, vision, 1 slot | 195,584 | 67,455 | yes | 2 | 17.6 |
| Swift IQ3_XXS, vision, 2 slots | 93,184 | 67,455 | yes | 1 | 17.6 |
One sweep, 2026-10-04; not comparable with the August tables above. Turns per window: floor(window per slot ÷ (prompt 1,701 + the file’s longer xhigh run)), with --reasoning-preserve keeping each turn’s reasoning in the window, derived from single-turn runs. Minutes per turn: that run’s length at its own decode speed (one slot decoding, -c 131072, temperature 1.0, speculative decoding on at n4/p0.75, the cards’ drafter before 2026-10-08, the machine not held idle), a cross-reference rather than a speed band; the two-slot rows assume only one slot is decoding.
Measured 2026-10-04: the xhigh ranges above come from one server setting for all three files (-c 131072, one slot, q8_0 KV, n4/p0.75, --reasoning-preserve, temperature 1.0), two runs per file in alternating order, every run stopping on its own and none truncated meas n=2; the anchor’s August run that hit the 65,536-token cap (above) counts as a lower bound, ≥ 65,536. On this task at this temperature, two runs per file show neither Swift file generating fewer tokens than the anchor — both files’ ranges sit inside the anchor’s same-sweep spread. At low, Swift IQ4_XS generated 12,350–16,704 tokens and Swift IQ3_XXS 14,002–14,614; at medium, Swift IQ4_XS generated 16,278–20,478 and Swift IQ3_XXS 10,584–20,072 meas n=2, most of it the HTML page itself at these two levels. No anchor run at these levels is in this sweep; the anchor’s single August 2026 runs on the same task generated 17,321 at low and 24,973 at medium meas n=1 (temperature 1.0), a different sweep, so no Swift figure is compared with them. The benchmark comparison of Swift IQ4_XS against the anchor is below.
What each level costs in electricity
A local card's tokens are free but its electricity is not quite. Measured with a power log integrated over each generation window (2026-08-23, one run per level, with xhigh measured twice — once at a 64k cap that truncated and once at 120k that completed), an answer to the same task costs 20.6 Wh at low, 36.0 at medium and 104.8 at xhigh. Those are board joules only, so they are not a household electricity figure and must not be treated as one (§15). The number worth noticing is not the total but the rate: energy per token rises with effort, because the thinking stream decodes more slowly than the answer stream and the board draws the same watts either way.
| Effort level | Tier | Mean load W | J / decode token, gross | J / decode token, net of 34.1 W idle | J / prompt token | tokens / kWh | EDP (J·s) | Wh per complete answer, gross | Wh, net |
|---|---|---|---|---|---|---|---|---|---|
low | in-band NVML | 344.6 | 4.26 meas | 3.84 der | 0.166 meas | 844,778 der | 1.57e7 meas | 20.58 meas | 18.54 der |
medium | in-band NVML | 345.0 | 5.18 meas | 4.67 der | 0.120 meas | 694,661 der | 4.84e7 meas | 36.01 meas | 32.45 der |
xhigh · 64k cap — truncated, returned no file | in-band NVML | 338.7 | 6.60 meas | 5.94 der | 0.162 meas | 545,326 der | 5.52e8 meas | 120.25 meas | 108.15 der |
xhigh · 120k cap — completed, delivered the file | in-band NVML | 341.8 | 6.13 meas | 5.52 der | 0.178 meas | 586,830 der | 4.16e8 meas | 104.84 meas | 94.38 der |
| E_comm (interconnect energy) | N/A — single GPU, no interconnect | ||||||||
Conditions: RTX 3090, UD-IQ4_XS, -c 131072, no projector, drafter at n-max 4 / p-min 0.75 on, temperature 1.0 / top-p 0.95 / top-k 20, n=1 per level, one HTML-authoring task, mixed regime (thinking and answer tokens both timed and both counted), 1 Hz power sampling with no clock or power-state columns, idle not subtracted in the gross columns. Instrumentation tier: in-band GPU board power (NVML); the power supply, processor and the rest of the machine are excluded and unmeasured. All four gross Wh figures may be cited without hedging. They were independently re-integrated from the raw logs on 2026-08-23 and reproduced by two methods: the original rectangle-rule method returned the originally published 20.55 / 35.96 / 120.21 / 104.9 Wh to within 0.05%, and a trapezoid integration with interpolated edges — which is what the four values printed above are — agrees to within 0.15%. The net columns subtract the dated 34.1 W loaded idle and are derived. The first-minute clock ramp makes each of these runs look 0.01–0.63% better than its settled rate, which is the right sign and a negligible size, so no ramp correction is applied — but it is a 17–20% artefact on 10-second probes, which is why short probes on this page are read differently (§01).
Three things hide inside that table. First, energy per token rises with effort — 4.26 → 5.18 → 6.13 J — but effort is not the cause. The mechanism is arithmetic: joules per token equals watts divided by tokens per second, and it predicts these four numbers as 4.24 / 5.16 / 6.60 / 6.13 against the measured 4.26 / 5.18 / 6.60 / 6.13, with no residual at all. The thinking stream simply decodes more slowly. The control that proves it: on the 25-question arms, where all three levels decoded at the same drafter-off speed, energy per token was flat to 0.3% (7.921 at xhigh against 7.947 at medium) while answers differed 3.1× in length. Effort changes joules per answer, not joules per token. Second, energy-delay product spans 1.57e7 to 5.52e8 — a factor of 35 across three levels, where energy alone spans 5.8×, because a slow setting is punished twice by that metric. Third, the 120 Wh row returned nothing. That xhigh run hit its 65,536-token cap mid-document: 5.8× low's energy for zero deliverable. The completed run beneath it cost less — 104.84 Wh — because a cap that fits lets the model stop when it is finished. Wall clock is still the bill that matters; truncation is the one that wastes it entirely.
Turning the drafter on cuts energy per token by roughly half, and it is free in quality: measured 7.884 ± 0.307 J per decode token with the drafter off across 115 requests, against 4.26–6.60 with it on. On the controlled comparison — one server, one prompt, only the flags changing — it is 8.104 → 3.210 J, a factor of 2.52 (§03). The drafter-off figure and the drafter-on figures come from different tasks with different sampling, so treat that pair as a reference comparison rather than a controlled experiment; the controlled version is the one in §03, and both point the same way. Anyone optimizing electricity here should reach for the drafter and for cached prefixes at depth — starting follow-up requests with exactly the same opening text, so the server does not read it twice (§02, §05) — long before touching the effort dial.
Which level to use for which kind of work
Within your window's ceiling, the trade is time. Pay it knowingly, and drop one level when you are sitting there waiting.
| Effort | Use for | Cost |
|---|---|---|
xhigh | the quality-first default this page ships — 100% specification coverage and zero fatal defects in the two judged runs that completed; a third xhigh run on the same task hit a 65,536-token cap and returned no file at all, which is the failure this level is most exposed to; worth it whenever a shipped bug or a missed requirement costs more than about twenty minutes of GPU time. Buy it as completeness insurance: on the 125 scored prompts it won nothing | 1.8× medium's wall clock across a benchmark suite, ~4× on one long authoring task. Raise max_tokens or thinking eats the budget |
medium | the value setting — most of the judged quality in a fraction of the wall clock, and statistically tied with the other two on the scored suite; the right choice for interactive sessions where the wait matters | minutes, not hours |
low / none | short, self-contained tasks. On the scored suite it posted the highest composite (82.1) and the fastest wall clock — though the suite cannot resolve a 1.6-point spread, so read that as "tied, and cheapest", not as a win. Either way it breaks any simple effort-quality ladder | not reliably cheaper on hard tasks, and the deliverable may not run: one campaign saw fatal defects at low in 2 of 2 runs while another saw the fatal defect at medium instead. At n=1 a fatal defect is a live risk at every level |
With greedy decoding, Swift IQ4_XS generates fewer tokens on the same 175 prompts at no detectable quality cost
Swift IQ4_XS generated 45% of the total tokens der that the anchor (unsloth UD-IQ4_XS of the base model) produced on the same 175 single-turn prompts at xhigh (greedy, drafter off, 32,768-token cap, 2026-10-04) — 52% over the 171 items both finished, with a median per-item ratio of 0.87, and fewer tokens on 139 of 175 items. Quality is indistinguishable at this sample size, 25 questions per set n=25, which is a smoke-test scale (§10): the 5-set composite is 82.7 against 81.3 meas, a difference of +1.49 (from the unrounded 82.75 and 81.26) der whose paired 95% bootstrap interval (−0.2 to +3.9; 10,000 resamples, seed 42) includes zero. The judged pair beside it (below) is level on MT-Bench; on ALPACA, Swift IQ4_XS is marginally ahead, which is not a finding: the interval clears zero only because of the anchor’s one empty answer at the cap. With the drafter off, derived decode time over the whole suite is 0.44 of the anchor’s der (token ratio 0.446 divided by the drafter-off decode ratio of 1.005); like the token total, it leans on the anchor’s capped items, and for the typical item the figure is about 0.86 der. Measured GPU energy over the suite, drafter off, is 0.47 of the anchor’s der, and it follows the sweep’s own cell time, also 0.47 of the anchor’s, almost exactly (r2 = 0.9996).
The pooled total-token ratio leans on the anchor’s 4 items at the 32,768-token cap, none of which produced an answer: 3 are degenerate greedy loops and 1 circular to the cap. Per item the typical saving is about 13% (median 0.87), and over the 171 items both finished the totals ratio is 0.52 rather than 0.45. The saving is measured on greedy single-turn prompts only (scope below).
| Set (n = 25) | File | Tokens meas | Reasoning meas | Answer meas | Score meas | Truncated | Empty | Ratio of means der | Median item der |
|---|---|---|---|---|---|---|---|---|---|
| GSM8K | anchor | 314 | 210 | 101 | 100 | 0 | 0 | ||
| Swift IQ4_XS | 269 | 152 | 113 | 100 | 0 | 0 | 0.85 | 0.96 | |
| MATH-500 | anchor | 3,055 | 2,761 | 290 | 100 | 0 | 0 | ||
| Swift IQ4_XS | 1,370 | 1,086 | 281 | 100 | 0 | 0 | 0.45 | 0.84 | |
| HumanEval | anchor | 5,192 | 5,011 | 178 | 92 | 2 | 2 | ||
| Swift IQ4_XS | 2,316 | 2,081 | 232 | 100 | 0 | 0 | 0.45 (0.44 without capped items) | 0.83 | |
| MBPP | anchor | 5,084 | 4,990 | 91 | 92 | 1 | 1 | ||
| Swift IQ4_XS | 1,808 | 1,717 | 88 | 92 | 0 | 0 | 0.36 (0.47 without capped items) | 0.59 | |
| ALPACA | anchor | 1,887 | 1,544 | 340 | judged | 1 | 1 | ||
| Swift IQ4_XS | 440 | 218 | 219 | judged | 0 | 0 | 0.23 (0.61 without capped items) | 0.86 | |
| MeetingBank | anchor | 1,597 | 1,457 | 137 | 22.3 | 0 | 0 | ||
| Swift IQ4_XS | 999 | 861 | 136 | 21.7 | 0 | 0 | 0.63 | 0.91 | |
| MT-Bench | anchor | 1,954 | 1,171 | 780 | judged | 0 | 0 | ||
| Swift IQ4_XS | 1,300 | 461 | 836 | judged | 0 | 0 | 0.67 | 0.88 | |
| 175 prompts | anchor | — | — | — | 81.3 (5-set) | 4 | 4 | ||
| Swift IQ4_XS | — | — | — | 82.7 (5-set) | 0 | 0 | 0.45 | 0.87 |
One sweep, 2026-10-04, xhigh, greedy, drafter off, q8_0 KV, seed 42, 32,768-token cap, suite 1cdf54f8eb9d3f8f; on every item it finished below the cap, the anchor’s response and token count equal its August runs byte for byte (scope below); August kept no reasoning text, so the reasoning column has no August counterpart. Scores: exact match (GSM8K, MATH-500), execution pass@1 (HumanEval, MBPP), ROUGE-L F1 (MeetingBank). ALPACA and MT-Bench are judged by a blind three-seat Claude Opus 5 panel. Tokens, reasoning and answer columns are mean per item. Ratio of means: total Swift IQ4_XS tokens divided by total anchor tokens for that set; median item: median of the 25 per-item ratios (all-items view); the hint beside a ratio of means leaves out the items either file left at the cap.
A spot-read of 24 transcripts, 12 per arm (every truncated one, the longest of each file, and the longest the automatic loop detector flagged), classified the anchor’s long-tail items as 3 degenerate greedy loops, 5 circular and 4 productive-long; Swift IQ4_XS had 0 degenerate, 3 circular, 8 productive-long and 1 other. The automatic loop detector over-flags code and arithmetic; its raw counts are not loop counts (§15).
llama-tokenize on the GGUF’s own vocabulary (thinking = the tokenized reasoning text, measured; the thesis table above prints the total beside it). Cap per item: 32,768 tokens. Per file: the quantizer and bit-width differ, so this is not a claim about the Swift-1.5 fine-tune. Tokenizer failures: 0.The time ratio derives from two measured quantities: Swift IQ4_XS generated 0.446 of the anchor’s tokens over the suite der and decoded at 1.005 times the anchor’s speed der (probe with the drafter and reasoning off on a 1,458-token prompt: 42.9 against 42.7 t/s meas, ratios 1.0016 and 1.0087 in its two passes; the first pass ran while two downloads were running on the host, the second after they stopped), giving a derived drafter-off decode time over the suite of 0.44 der (load, prefill and scoring excluded). The same derivation gives about 0.86 for the typical item and 0.52 over the totals of the 171 items both files finished der, because the suite total leans on the anchor’s four capped items (above). File-size arithmetic — Swift IQ4_XS reads 14,550,542,336 bytes per decode step against the anchor’s 13,344,536,576 (×1.0904; each file minus its token embeddings, header and drafter head, with the embeddings assumed to stay in system memory) — predicted a decode ratio of 0.917 der; the measured ratio is 1.005, not 0.917. The sweep’s own wall time, measured while other work was using the machine, is not used in this derivation; the Swift speed bands are greedy-only (temperature 0); speed at the temperature-1.0 sampler the cards ship is not measured as a band (§15). The 0.44 is a drafter-off figure. The Swift IQ4_XS cards ship with the drafter on, and with the drafter and reasoning on, Swift IQ4_XS decoded slower than the anchor at n4/p0.75 on the one prompt measured in §08, so this time ratio is not a measurement of the cards.
No quality difference larger than this suite can detect: across 8 comparisons on the same 175 prompts (the five scored sets, their 5-set composite and the two judged sets; the 16k and 7-set composites are further views of the same items; 2026-10-04, greedy, drafter off, 32,768-token cap), the 5-set composite for Swift IQ4_XS is 82.7 against 81.3 for the anchor (unsloth UD-IQ4_XS of the base model), +1.49 points before rounding (82.75 against 81.26), with a 95% paired bootstrap interval of −0.2 to +3.9 (10,000 resamples, seed 42) n=25 per set — the interval includes zero. No cell gap reaches 20 points, and GSM8K and MATH-500 sit at 100 for both files, so those two sets discriminate nothing at any sample size (§10). HumanEval is the widest: Swift IQ4_XS scored 100.0 (0 truncations) against the anchor’s 92.0 (2 truncations), +8.0 points; the anchor’s two misses are those two capped, empty answers. At the derived 16k cap der, licensed by greedy determinism, the composites are 81.9 against 77.3 — the anchor’s truncations at the lower cap widen the spread. Beside the 5-set composite, the judged pair (3 blind Claude Opus 5 seats, a correlated instrument): on ALPACA, Swift IQ4_XS scored 79.7 against 73.5, +0.56 rating points, 95% interval +0.04 to +1.21 — the interval barely clears zero, so this is marginal, one of 8 comparisons, and not a finding. It rests on the anchor’s one ALPACA answer cut at the 32,768 cap, which was empty and was rated 1 by all three seats (final at this cap): that item adds +0.20 of the +0.56, and without it the difference is +0.38 with an interval of −0.04 to +0.92, which includes zero. On MT-Bench the two files tied: 86.7 against 87.0, −0.03 rating points, interval −0.35 to +0.21. The 7-set composite (the five scored sets plus the judged pair, printed separately from the 5-set one) is 82.9 against 81.0, +1.9 der.
| Benchmark (n=25 each) | Scorer | Anchor meas | Swift IQ4_XS meas | W | L | T | Gap (pts) der |
|---|---|---|---|---|---|---|---|
| GSM8K | exact match | 100.0 | 100.0 | 0 | 0 | 25 | 0.0 |
| MATH-500 | exact match | 100.0 | 100.0 | 0 | 0 | 25 | 0.0 |
| HumanEval | execution pass@1 | 92.0 (2 trunc) | 100.0 | 2 | 0 | 23 | +8.0 |
| MBPP | execution pass@1 | 92.0 (1 trunc) | 92.0 | 0 | 0 | 25 | 0.0 |
| MeetingBank | ROUGE-L F1 | 22.3 | 21.7 | 10 | 15 | 0 | −0.54 (21.74 − 22.28) |
| 5-set composite | mean | 81.3 | 82.7 | 95% CI: −0.2 to +3.9 | +1.49 | ||
| Derived 16k composite der | mean | 77.3 | 81.9 | ||||
| ALPACA | judge 1–10 (3-seat panel) | 73.5 (1 trunc, empty) | 79.7 | 9 | 4 | 12 | +0.56 (rating) |
| MT-Bench | judge 1–10 (3-seat panel) | 87.0 | 86.7 | 10 | 5 | 10 | −0.027 (rating) |
| 7-set composite (incl. judged pair) | mean | 81.0 | 82.9 | +1.9 | |||
One sweep, 2026-10-04; the Swift IQ4_XS rows exist only in this sweep. The anchor’s five scored rows and both composites of them equal its August runs exactly (scope below); its judged rows, and so its 7-set composite, do not, because this panel session is not the August one: the same 25 MT-Bench answers, byte-identical to August’s, score 87.0 here against 79.7 in the August session (above), so the judged pair is read only as a paired comparison inside this one session. Both files at xhigh, q8_0 KV, greedy (temperature 0, top-k 1, seed 42), drafter off, 32,768-token cap, n=25 per set. W = items where Swift IQ4_XS scored higher; L = where the anchor scored higher; T = same. For the judged pair: W/L/T are per-prompt comparisons across 3 averaged seats, the interval is a paired bootstrap of 20,000 resamples (seed 42), and the mean spread between seats is 0.60 and 0.68 rating points on ALPACA and 0.52 and 0.32 on MT-Bench (anchor, Swift IQ4_XS); scores on the 0–100 scale, converted as (r − 1) / 9 × 100; gaps in the 1–10 rating scale. The 5-set bootstrap: 10,000 resamples within each set, same indices both arms, seed 42; composite = mean of the 5 set means.
Swift IQ4_XS used 0.47 of the anchor’s measured GPU energy over the suite der (the same 2026-10-04 sweep, drafter off, measured while other work was using the machine: the cell windows whose wall time is not used for the time ratio) — 3.375 Wh per answer (gross) against the anchor’s 7.235 meas; idle-subtracted, 3.242 against 6.955 der. That ratio is not independent evidence of anything the token count did not say: energy against cell time across the 14 cells gives r² = 0.9996 der, with mean board power between 312 and 344 W in every cell, so energy is a near-perfect restatement of time. The windows are coarse — each cell window holds load, prefill, decode and scoring together; prefill and decode are not separated.
The quantization effect — Swift IQ4_XS against Swift IQ3_XXS on quality — is not measured in this sweep: Swift IQ3_XXS was not run on the scored suite (it is ranked against Swift IQ4_XS on perplexity, speed and appetite only; §15).
Scope of this comparison: 175 single-turn prompts with one answer each, greedy (temperature 0, top-k 1, seed 42), drafter off, q8_0 KV, xhigh, 32,768-token cap, n=25 per set. Whether these drafter-off scores carry over to Swift IQ4_XS with the drafter on rests on one prompt (§08). The anchor (unsloth UD-IQ4_XS of the base model) reproduced its August 2026 runs exactly: all 171 items that finished below the cap in both the August run and this one have byte-identical responses and identical token counts, and its perplexity is 6.5956 meas, exactly the August value (the same file and build). At temperature 1.0 on one task, two runs per file did not show Swift IQ4_XS generating fewer tokens than the anchor (above), so the token saving measured here is shown for greedy single-turn prompts only (§17).
The Pi fine-tune’s IQ4_XS ran the same 175 prompts on 2026-10-06/07 (greedy, drafter off, 32,768-token cap), paired item for item with this sweep’s anchor and Swift IQ4_XS cells. At xhigh its 5-set composite is 80.6 meas: −2.15 against Swift IQ4_XS (paired 95% interval −5.1 to +0.2) and −0.66 against the anchor (−3.5 to +1.9) der, no difference the suite can detect n=25. It generated 1.46× Swift IQ4_XS’s total tokens (median item 1.03, fewer on 70 of 175) and 0.65× the anchor’s der. ALPACA and MT-Bench were not judged for it. Its medium run and the details are in Appendix B.
Enough for what? That is the whole question, and it has an arithmetic answer. A 25-question benchmark can tell you a model file is broken and nothing finer. A 25-question benchmark can tell you a setting has not collapsed. Telling two healthy 4-bit files apart, where the real difference is one to three points, needs thousands of questions and days of machine time — which is why nobody ranks healthy files by accuracy scores, and why this page ranks them by perplexity instead. Before you believe any benchmark number, including the ones on this page, ask three things: how many questions, what was the token budget, and how many answers were cut off. The rest of this section is those three questions with numbers on them.
A worked example from this page's own data. A GSM8K score of 85% at 20 questions carries a 95% confidence interval of roughly 64–95% — a thirty-point window around a table where every score sits within ten points of every other. One flipped question moves that score 5 points. Healthy 4-bit files differ by perhaps 1–3 true points; inside that window, such a gap is invisible.
The arithmetic is unforgiving because sampling noise shrinks with the square root of the sample count: halve the gap you want to detect and you need four times the samples. What that costs on this 3090 — paired evaluation on the same questions, temperature 0, about 80% base accuracy, 95% confidence at 80% power, about 1,200 tokens per sample at about 55 t/s with the drafter on:
| True gap to detect | Samples needed | Wall clock per model | Board energy per model |
|---|---|---|---|
| ~20 pts — a broken file | ~25–30 der | ~9–11 min der | ~51–61 Wh der |
| ~10 pts | ~100–150 der | ~35–55 min der | ~0.20–0.31 kWh der |
| ~5 pts | ~300–500 der | ~2–3 h der | ~0.61–1.0 kWh der |
| 1–3 pts — typical 4-bit against 4-bit | ~2,000–10,000 der | ~12–60 h der | ~4.1–20.4 kWh der |
The energy column is derived at the measured 6.13 J per decode token for a drafter-on run at 55.8 t/s (§09), in-band GPU board power only. Two measured reference points for calibration, both from the 2026-08-23 sweep with the drafter off and real answer lengths rather than a 1,200-token assumption: 25 MATH-500 questions at low cost 126.2 Wh, and 50 questions at medium cost 165.9 Wh. Real benchmark work costs more than the table's assumption because real answers are longer than 1,200 tokens and because a drafter-off run spends about 7.9 J per token instead of 6.1.
The first row is the one job a local accuracy run does well, and it is the only claim §08's scored smoke test makes. The last row is why nobody ranks healthy files by accuracy: the run that could would take days per file and about twenty kilowatt-hours. And one failure mode the table cannot show: a saturated benchmark discriminates nothing at any sample size. On the 2026-08-23 suite, GSM8K returned 100% for all three effort levels. No amount of extra questions from that set would have separated them, because the set had stopped measuring anything about this model. Check for a ceiling before you buy more samples.
The other way a benchmark lies: the token budget
Sample size is not the only silent corrupter. A thinking model spends tokens reasoning before it answers, and a fixed max_tokens cap silently converts its hardest questions into wrong answers. The run does not error; the score just drops. This page has the cleanest possible demonstration of it, because the same measurement was run twice and the two runs disagree about which setting is best.
| Token cap | low | medium | xhigh | Truncations (scored sets) | What the table appears to say |
|---|---|---|---|---|---|
| 16,384 | 81.3 | 80.5 | 77.3 | 1 / 2 / 8 | "quality falls as effort rises" — a clean, quotable, wrong headline |
| 32,768 | 82.1 | 80.5 | 81.3 | 0 / 0 / 3 | all three levels within 1.6 points — indistinguishable at n=25 |
Same model, same prompts, same grader, same machine, same day. The only thing that changed is how many tokens each answer was allowed. The entire xhigh penalty was the cap, because by the grading rule a truncated answer scores zero, and xhigh truncated eight times against one and two for the other levels. Two individual prompts make it concrete: one MATH-500 problem needed 18,273 tokens and was correct; one HumanEval problem needed 17,025 and was correct. Both scored zero at 16,384. And the failure is worse than a cut-off answer: 11 of the 12 truncated generations came back with an empty answer field (eleven of the twelve are the scored-set truncations counted above; the twelfth is an ALPACA prompt — the judge panel rated that empty answer 1 out of 10 on all three seats, and the raised-cap re-run reproduced it exactly — the same answer consumed all 32,768 tokens and again returned nothing, §09), because the runaway happens inside the reasoning block, so the model never emits its end-of-thinking marker and never reaches an answer at all. Three xhigh prompts still exceed 32,768 and are genuine non-terminating loops rather than long solutions; they are reported as truncations rather than hidden.
The tempting fix is the wrong one: "re-run using only the questions that did not truncate." That selects the question set based on one arm's behaviour — dropping precisely the hard items where a quality difference could live — and quietly changes the question from "which setting scores better?" to "which setting scores better on easy questions?". It also breaks comparability with every published number on the same benchmark. The correct fix is to raise the budget, not to shrink the test, and to re-run only the arms and datasets that actually truncated.
Raising --max-tokens and -c changes nothing about any generation that had room to finish: every prompt that did not hit the old cap was compared byte-for-byte between the 16,384 run and the 32,768 re-run — low 24/24, medium 48/48, xhigh 67/67, a total of 139 of 139 with zero drift — because greedy decoding with the drafter off is deterministic, and the re-run doubled the serving window as well as the cap. That is the evidence that re-running only the truncating datasets loses no information relative to a full re-run, and it is the cheapest way to fix a capped benchmark honestly.
Before trusting any thinking-model benchmark score — yours or a leaderboard's — ask two questions: what was max_tokens, and how many answers truncated? A score published without its cap cannot be checked, for the same reason a speculative-decoding speedup published without its acceptance rate cannot (§06). And a third question, exposed by the 2026-08-23 campaign: which token regime was being timed? On a model that thinks by default, a "code generation" probe with a 700-token cap can spend every timed token reasoning about the task and return an empty answer field — the label describes the task, the number describes the thinking. A speed number without its token regime is exactly as uncheckable as one without its acceptance rate. Pin enable_thinking or report the regime beside the number, and name probes after the token stream rather than the task.
Tokens ÷ wall-clock is not decode speed at depth — it silently averages the prefill in. The same measured run reads 47.1 t/s of decode and 9.2 t/s naive at a 28,000-token-deep prompt, and 35.8 against 2.4 at 92,000. Both are arithmetically correct; only one describes the model, because the other is mostly a statement about how long the prompt was. Quote the server's own timings, and quote prefill as its own number (§05, §06).
GPQA-Diamond: 79.8% on 198 graduate-level science questions
GPQA-Diamond is a graduate-level science benchmark — 198 questions, harder than the scored arms in this section and nearly eight times the sample. It tested two things at once: whether a 30,000-token cap still turns hard questions into wrong answers at this difficulty, and what happens when the question file you are reading is not in random order.
GPQA-Diamond, all 198 questions (2026-08-27 and 2026-08-28, UD-IQ4_XS, reasoning effort xhigh, vendor thinking sampler at temperature 1.0 / top-p 0.95 / top-k 20, max_tokens 30,000, -c 32768, seed 42, drafter at n-max 4 / p-min 0.75, one RTX 3090 at its stock 350 W board cap, llama.cpp b10502):
| Reading | Score | 95% interval | What it means |
|---|---|---|---|
| All 198 questions | 158 / 198 = 79.8% | 73.7 – 84.8 | the headline — every question in the file, answered once |
| Excluding answers that hit the cap | 158 / 166 = 95.2% | 90.8 – 97.5 | an upper bound — assumes every truncated answer would have been right |
| Answers that hit the 30,000-token cap | 32 / 198 = 16.2% | — | scored wrong — a harness condition, not a fact about the model |
The gap between 79.8% and 95.2% is the token budget, not the model’s knowledge. The truth is somewhere between them, and nothing here can say where: a truncated answer might have been right or wrong, and scoring it wrong is the only defensible choice.
The first run stopped at question 100 and scored 82.0%. That number looked publishable and was not, because the frozen question file is sorted by subject. Counting how often two neighbouring questions share a subject gives 106 such pairs against 48.1 expected if the file were in random order; shuffling the labels 20,000 times never once produced 106 or more (permutation p < 0.00005).
So the first hundred questions were not a sample of the benchmark. They were every one of its 16 biology questions and 3 of its 21 quantum ones. Publishing 82.0% as “GPQA-Diamond” would have published a subject mix.
The fix was to run the exact complement — questions 101 to 198, which answered biology 0 of 16 and quantum 18 of 21 — and combine. Two complementary halves cover every row exactly once, and the ordering stops mattering. That is the only reason the 79.8% above may be quoted plainly.
An early count looked significant but was not: the second half’s answers appeared to hit the cap far more often than the first half’s — 7 of the first 10 against 13 of 100, testing at p = 0.0117, with the physics-heavy half’s longer derivations as a plausible mechanism.
Over the finished halves: 19 of 98 against 13 of 100, z = −1.22, p = 0.222. Not significant. The early reading was an artefact of testing a count while it was still moving. The two halves also do not differ in accuracy (82 of 100 against 76 of 98, z = 0.78, p = 0.436).
The lesson is not that the mechanism is wrong. It is that a test applied to a running total is a test you can stop at a moment that flatters it, and the honest procedure is to fix the sample size first and look once.
Throughput is deliberately not pooled across the two halves: they differ in how much they generated (mean 10,166 against 10,860 tokens per answer) and in the board temperature they met, so a single average would describe neither. First half 52.1 t/s, second half 54.2 t/s.
Perplexity and how close a file stays to the full-precision model's own next-token probabilities rank model files; a small-sample accuracy pass smoke-tests everything those metrics cannot see. Per-token metrics compare probabilities against the full-precision model over 294,912 scored positions on this page's corpus (36 chunks × 8,192 at -c 8192 — and note the tokenizer caveat in §08: that is a count of scored windows, not of corpus tokens, and it moves with the tokenizer), in an estimated 15–30 minutes per file on a 3090-class card. That estimate is not a benchmark from this page: it covers the per-file pass against saved full-precision baseline probabilities, and generating that baseline needs the unquantized model, which exceeds a 3090's memory and is a slower one-time cost. Those metrics resolve differences no feasible accuracy run can, which is why quantization authors compare files this way. But they run below the sampling and template layer, so a broken chat template or a mangled tool-call format sails through them unnoticed while collapsing scored answers by 20 points or more. Run both: per-token metrics to pick between files, one cheap accuracy pass to confirm the file you picked actually works end to end. §08's perplexity table is this advice in practice.
Four things on this machine produce the same complaint — the command that did 40 tokens per second yesterday does 25 today — and in all four cases the server starts without an error; on the spill loads it prints one start-up memory warning, easy to miss.C31 Two are real slowdowns you can fix: the card quietly ran out of memory and spilled part of the model into system memory, or the command left one layer on the processor. One is not a slowdown at all but a client that never delivered your request. And the fourth does not slow anything down at all: it makes your own measurement read up to 25% low, so the server is fine and the number is wrong. Before any of them, though, comes a class of problem whose fix is never in the server command at all, so it is stated first.
Failure class 0 · the client contract
These are the failures a reader most often blames on the model, and every one of them is fixed in a different file from the server command. They belong together because they share one cause: the model has no awareness of its own limits and cannot budget its own tokens. Your configuration is the only thing enforcing them.
- One token pool.
prompt + thinking + output ≤ context. A 100,000-token thinking run inside a 32,768-token window fails by arithmetic, not by bug (§02). - Cut-offs are client-side. The server generates until it is told to stop. A truncation comes from your client's
max_tokens, its request timeout, or Ctrl+C. Sizemax_tokensso the whole run fits one response, and check that your client's context setting matches the server's-c— those two numbers drifting apart is the classic silent failure (§14 shows the setting per agent). - A cut generation cannot resume. With OpenAI-compatible interfaces there is no seamless continuation: the thinking is gone and the next request starts over. This is why a cap placed near the middle of the thinking-appetite distribution is a truncation machine rather than a saving (§09).
- Timeouts outlive defaults. A long
xhighrun outlasts most default client timeouts; the server log shows one asClient disconnected. Stopping generation. - A signature that separates client from server in one look. Check the llama-server log for the task. If the server logged no task at all, the request never left your machine and nothing about the model is implicated. If the server logged the task and returned something empty, that is the token-budget trap above, not a transport problem.
Failure 1 · the VRAM spill that raises no error
The NVIDIA driver does not refuse an allocation that exceeds VRAM — by default it quietly overflows into shared system memory across the PCIe bus. On the reference machine, two open browsers held 2–3 GB of VRAM while the 122,880-token configuration needed about 22 GiB (§05), and about 1.9 GB spilled. The server loaded without an error and decoded at 20–35 t/s — and stayed there no matter how far the context was shrunk, because the weights stay spilled either way. This is §05's cliff in its most treacherous form: the server does not refuse the load. Nothing errors, and everything runs at half its speed or worse. In the spill loads measured on 2026-08-23, the one memory line logged is the start-up warning described in the callout below (§11), which reads like noise.C31
The signature lives in Task Manager: dedicated GPU memory pinned at 23.5–24.0 / 24.0, shared GPU memory growing, processor use around 70% from driver paging. The fix is freeing the VRAM. The prevention is making the failure loud: NVIDIA Control Panel → Manage 3D Settings → Program Settings → llama-server.exe → CUDA — Sysmem Fallback Policy → Prefer No Sysmem Fallback. With that set, an over-budget load fails with an out-of-memory error at start-up instead of running slow all day.
Failure 2 · the -ngl off-by-one
llama.cpp counts the output layer as one more "layer" than the model has transformer layers. This model has 64, so -ngl 64 reads as complete — and leaves the output layer, a 5120 × ~151k-vocabulary matrix multiplication that runs for every generated token, on the processor. Measured cost: decode 25.7 against 39.7 t/s, prefill 106 against 290 t/s (server-log prefill on the short test prompt — a different probe from llama-bench's, which reads about 1300 on this card), GPU utilisation 53–67% instead of about 90%, and a stack of processor threads pinned doing the vocabulary multiplication. Independently replicated on a second file: the independent re-measurement (2026-08-23, same machine) measured UD-IQ4_XS at 29.84 t/s with -ngl 64 against 42.31 with -ngl 99 — −29.5% — on a different quantization. The Q4_K_M pair above is a steeper −35.3%; call the penalty 30–35% of your decode speed. The detail that makes it dangerous: the load succeeded and board VRAM read 15,256 against 15,661 MiB, near enough identical. There is no memory signature. The only signature is the speed — which is why §04's bandwidth figure for your file is the number to check against. Every recipe in §03 says -ngl 99 for exactly this reason: any number past the real count means "everything", with no off-by-one to get wrong.
High processor use alone proves nothing. llama.cpp worker threads spin while waiting on the GPU, so 50–70% processor use during perfectly healthy all-GPU decode is normal on a 20-thread machine. The trouble signal is the combination — high processor use and low t/s. Both failures above show both. And one start-up warning that is easy to miss: failed to fit params to free device memory: n_gpu_layers already set by user means that llama.cpp's automatic fit projected this load would leave less free memory on the card than its 1,024 MiB target (the --fit-target default in build 10502, set in common/common.h and tested in common/fit.cpp at commit 0adcc3bb5), and because -ngl was given explicitly it abandoned fitting and loaded exactly what was asked. It is not triggered by -ngl 99 alone — in the 2026-08-23 ceiling sweep (UD-IQ4_XS, BF16 projector, n-max 4 / p-min 0.75), with -ngl 99 on every load, the line is present at -c 163840 and above and absent at -c 131072 and below. The load that collapsed to 8.0 t/s meas at 90,886 tokens (-c 262144, same file, projector and drafter, greedy) printed it; its fully resident -c 131072 control did not. But at -c 163840 the same warning printed while the server's shared GPU memory read 446 MiB meas at load, in line with the resident loads below it — so the warning can fire on a load that has not spilled. When it prints, read it together with the signals that separate the two failures in the table below, and with shared GPU memory at load: in the cache spill measured on 2026-08-23 the server’s own shared-memory reading was 3,502 MiB at load and 3,546 MiB from the first request through 90,886 tokens: high from the start, then flat.C31
| Signal | VRAM spill | -ngl off-by-one |
|---|---|---|
| Shared GPU memory (Task Manager) | grows during the run | stays flat |
| Decode speed | 20–35 t/s, flat at every context size | 25.7 t/s, fixed (against 39.7) |
| Dedicated VRAM | pinned at 23.5–24.0 of 24.0 | normal — 15,256 against 15,661 MiB, no signature at all |
| Persistence | vanishes when other applications release VRAM | survives a reboot — it lives in your command line |
| Fix | free the VRAM · set Prefer No Sysmem Fallback | -ngl 99 |
The diagnostic commands that actually work on Windows
One tooling trap cost this page's own debugging session real time: on Windows, nvidia-smi's per-process listing shows [N/A] or "Insufficient Permissions" for memory, and its memory.used counts only dedicated VRAM — it cannot see the spill. The tools that exposed both failure modes:
# totals + who is on the GPU (dedicated only - the spill is NOT visible here): nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv # the spill itself - total shared GPU memory in use (PowerShell): (Get-Counter '\GPU Process Memory(*)\Shared Usage').CounterSamples | Measure-Object CookedValue -Sum # divide Sum by 1MB; growing = spilling # per-process dedicated commitment - this is what caught the compositor (dwm) # holding 3.6 GiB during a browser-UI session: (Get-Counter '\GPU Process Memory(*)\Dedicated Usage').CounterSamples | Where-Object { $_.CookedValue -gt 200MB } # or simply Task Manager - Performance - GPU: the dedicated/shared split and, # under Performance - Memory, your RAM speed and "Slots used" (the section 07 channel check)
Failure 3 · PowerShell 5.1 silently drops your screenshot
One more Windows trap, and it is the kind that looks like a model problem. The independent re-measurement (2026-08-23, same machine) posted a 1440p screenshot to the vision endpoint from Windows PowerShell 5.1 and got nothing back — no answer, no error worth reading. Invoke-RestMethod cannot POST the roughly 261 KB body that a 1440p PNG data-URI produces. The identical request from Python succeeded in 3.8 s.
The diagnostic signature is the one from Failure class 0, and it separates client bugs from server bugs in one look: check the llama-server log for the task. If the server logged no task at all, the request never left your machine — nothing about the model, the projector, or --mmproj is implicated, and re-tuning any of them is wasted time. The fix is the transport: use Python's requests, curl, or PowerShell 7 or later for anything carrying a base64 image. Keep PowerShell 5.1 for the counters above, where it is excellent.
Failure 4 · the first probe after a long prefill — a 45% measurement swing
This one does not slow your server down; it makes your stopwatch lie (it affected the projector measurement in §06). Measured 2026-08-23: four runs of one identical configuration read 18.27, 18.82, 19.21 and 26.60 t/s — the highest 45.6% above the lowest, with nothing changing but what the card had been doing beforehand. Every one of those was the probe fired immediately after a roughly 105-second prefill. This is the measurement that produces the ±25% band declared once in §01, and the two figures are the same measurement stated two ways: half of a 45.6% end-to-end spread is ±22.8%, rounded to ±25%.
The nvidia-smi samples taken before and after each prefill say why, and it is not heat. The first load of a session starts on a genuinely idle card — 56 °C, 225 MHz, 34 W — and its prefill drives the SM clock to a full 1,455 MHz boost. Every later load starts while the card is still winding down from the previous server's teardown and model load — 63–66 °C, 315–435 MHz, 62–136 W drawn at 1% utilisation — and its prefill only reaches 900–990 MHz. The decode probe that follows inherits whatever clock state that was. With the drafter on the same effect is milder but present: post-prefill probes read 54.3–60.0 against a settled 62.7.
The obvious suspect is thermal throttling, and it is wrong. Within a single cooled load, five consecutive probes ran the card from 57 °C to 82 °C and decode moved from 27.18 to 26.90 t/s — 1.0%. Once settled, this card's decode is nearly temperature-insensitive. All of the damage lives in the boost-clock ramp during and just after a long prefill, which is exactly where a benchmark script naturally puts its first measurement.
The fix is a protocol — this page's cooled protocol (§02) — and it costs about a minute per data point. Fire the prompt, discard the decode timing that comes back with it, wait 30–45 s, then send several short repeat requests that re-use the cached prefix and report the median of those. Under that protocol repeat probes on this machine agree to 0.4% — which is the only reason a 0.04% difference between two configurations was measurable at all. The same ramp is a power caveat: a prefill fired at a cold board draws unrepresentative watts over an unrepresentative duration, which is why 10-second power probes read 277–287 W against a sustained 306–341 (§03).
Two traps that only affect people taking measurements
Both were found while producing this page, and both are the kind of mistake that yields a plausible number rather than an obvious error.
- "The server is down" is not "the GPU is idle." A settled idle window measured after the last job read a mean of 37.8 W against a median of 31.1, because five short excursions to 121–124 W punched through it — each one the desktop rendering plot images, at 0–14% GPU utilisation, minutes after llama-server had exited. If you measure idle power, measure it with the desktop quiet and discard the first 60 seconds while the board is still cooling. The settled figures on this page (29.9 W with no server, 34.1 W with the model resident) were taken that way; the earlier 33.2 W and 34.6 W readings were not, which is how they came out above the loaded figure.
- Never charge a whole token count against a partly covered power log. Joining request timings to a power log that started mid-run, and crediting each request its full token count, produced a physically impossible 5.32 J per token and 59.8 t/s on a card that cannot do either with the drafter off. The fix is to count only requests whose entire window lies inside the log — 115 of 150 in that case. Any per-arm energy join across a log boundary needs the same filter.
This model can look at images if you start the server with the optional projector file, and on a 24 GB card that is a real option rather than a compromise: it costs 1,138 MiB of memory and, measured carefully, nothing at all in speed. What it does cost is context. A 1440p screenshot occupies about 3,600 tokens of your window and a 4K one about 8,244, so a feedback loop that keeps ten screenshots around has spent 36,000 tokens before the model writes anything. Use 1440p, set --image-min-tokens 1024 and --image-max-tokens 10580, send each picture as base64 data inside an ordinary request (that is the standard way to put an image into a text request), budget max_tokens generously, and clear old screenshots between iterations. The model reads the picture — 7 of 7 with the image, 0 of 7 with the image withheld on a generated 1440p image with seven questions whose answers are known exactly (2026-08-25, greedy, meas n=1 per question; §15). The aquarium critique below is one demonstration with no such test of its own, so its quality is unverified.C32
7 of 7 with the image and 0 of 7 with the image withheld — five refusals and two wrong guesses on UD-IQ4_XS — on a generated 1440p detail target with seven questions whose answers are known exactly, scored by string equality (2026-08-25, greedy, n=1 per question). But that control was a different task from the aquarium critique below (2026-08-22), which had no withheld-image arm of its own, so the end-to-end run is not evidence of a critique loop: a model that has been told "this is an animated aquarium page" can produce plausible criticism of an aquarium page without seeing one. Treat the quality of the critique as unverified. The control is in §15's register.C32
What was demonstrated end to end on the reference 3090, 2026-08-22: Chrome captured a 4K screenshot (3840×2160, a 2.3 MB PNG) of a test page — an animated canvas aquarium — and the model, served with the configuration below, named the visible creatures individually, identified the weakest visual element (kelp drawn as wire-thin lines), and wrote working replacement drawing code with tapered, phase-offset swaying blades. Round trip: about 39 s, of which the 4K image cost 8,244 prompt tokens. n=1
Image cost is resolution, not file size — and it obeys an exact law. The vision tower uses patch 16 with a spatial merge of 2, so one visual token covers a 32 × 32-pixel block: tokens = pixels ÷ 1024 + 2, the +2 being the start and end markers. The independent re-measurement (2026-08-23, same machine) confirmed it to within 2 tokens at two resolutions — 720p measured 922 against 920 predicted, 1440p measured 3,602 against 3,600. The 4K row needs the rounding stated to match: 3840×2160 divided flat gives 8,102, while rounding each dimension up to whole 32-pixel blocks (120 × 68) gives 8,162 against the measured 8,244 — 1.0% high on the block form, 1.8% on the flat one, so quote the block form and expect about a percent of slack at 4K. The 1080p row is derived arithmetic and unmeasured; the correction history is in §15:
| Screenshot | Context cost | Source | Feedback rounds alongside a full xhigh cycle at -c 122880 |
|---|---|---|---|
| 1080p (1920×1080) | ~2,040 tokens | der arithmetic — unmeasured; the height rounds up to 34 blocks | ~16–21 der |
| 1440p — recommended (2560×1440) | 3,602 | meas | ~9–12 der |
| 4K (3840×2160) | 8,244 | meas | ~4–5 der |
1440p under --image-max-tokens 1024 | 1,010 | meas — 3.6× cheaper; meas 5 of 7 against 7 of 7 on a detail target (UD-IQ4_XS, greedy, n=1 per question; box below)C32 | ~33–43 der |
The rounds column is your own window’s arithmetic: a full xhigh cycle wants 80,000–90,000 of the 122,880-token window (§09), leaving 33,000–43,000 for images, and the column is that remaining window divided by the row’s per-image cost.
--image-max-tokens 1024 takes a 1440p screenshot from 3,602 tokens to 1,010, measured — a genuine 3.6× saving and the knob to reach for when a feedback loop is eating the window. What it does to the answer is measured too. On the same seven-question test image (2026-08-25, UD-IQ4_XS, greedy, n=1 per question): full budget (--image-max-tokens 10580) scored 7 of 7; at --image-max-tokens 1024 it scored 5 of 7 meas — both layout questions right, the misses on fine 12–15 px type, with confident wrong values rather than refusals (a latency of 207 read as 287). The cheaper budget costs fine detail specifically. Use it to fit a window, not to read numbers or small text off a screenshot.C32
The two budget flags behave differently, and it is worth knowing which one actually saves you anything. --image-max-tokens is a real cap that bites, as above. --image-min-tokens is only a floor: it lifted a 720p shot from 922 to 1,077 tokens and left the 1440p shot at 3,602, untouched. It raises small images and does nothing to anything above the floor — so --image-min-tokens 1024 costs you nothing on the screenshots you care about, and is not a lever for shrinking them. The default cap is 4,096 tokens, which by the law above is about 4.2 megapixels: that is exactly why a 1440p shot (3,602) passes through untouched at default flags while a 4K one does not, and why this page's vision configuration raises the cap to 10,580 for 4K-class detail. (4K handling upstream still has open quirks, so 1440p remains the recommended resolution.)
The serving configuration, and the capture command that feeds it:
llama-server -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b -c 122880 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 # n3/p0, matching section 03's default card since 2026-10-08: the n/p sweep's best # score for this file at this sampler (section 06), about 150 MiB lighter than n4/p0.75. # Measured live at n4/p0.75: ~3 GiB of board VRAM left over with a desktop up. # Plan on 83-86 t/s of answer tokens (measured at n4/p0.75) # (sections 03 and 06). An older live check of this configuration # read 71.3 t/s short-context, but its token regime was never # recorded, so it is not used as a speed band. # The projector itself books 1,138 MiB of that (section 05) --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --jinja --host 127.0.0.1 --port 1234 # capture the page your code just rendered, using Chrome's no-window mode. The flag # below is spelled --headless, but it has NOTHING to do with the graphics card or # with "headless" as section 02 defines it: it only means the browser draws no # window on screen. It does not free any VRAM. chrome --headless --disable-gpu --window-size=2560,1440 ^ --screenshot=shot.png --virtual-time-budget=6000 file:///path/to/page.html
Copy-paste version — the same command with the comments removed
llama-server -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b -c 122880 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --jinja --host 127.0.0.1 --port 1234 chrome --headless --disable-gpu --window-size=2560,1440 ^ --screenshot=shot.png --virtual-time-budget=6000 file:///path/to/page.html
Multi-image works, and it can compare versions. Measured: three screenshots of three different aquarium implementations in one request — 4K, 1440p, 1080p — cost 13,913 prompt tokens with a 74-second round trip, and the model ranked them by visual richness correctly and flagged the third as "broken / nearly empty… looks like the animation failed to spawn its fish". That page is the same low-effort output whose code was blind-scored 40/100 in §09 for freezing on its first frame — two independent reviews reaching the same verdict, though again without a withheld-image control. That 13,913 is also a third confirmation of the token law: 8,244 + 3,602 + ~2,040 = about 13,890, some twenty tokens from the measurement once the text prompt is counted. Under the old interpolated figures the same three shots would have cost about 15,500.
Which coding agents can carry a screenshot to the model
All five were driven from a script, with no interactive terminal session, against this server with the same 2.3 MB 4K PNG and a question only answerable by seeing it (2026-08-22, one attempt each). The encouraging result: no agent hallucinated sight. Every failure was an honest "I see no image", which is the failure mode you want, because the alternative — an agent that describes a picture it never received — is undetectable from the transcript. Two of the five needed configuration before they passed, and in both cases the default behaviour was to silently send the image as text.
| Agent | Verdict | How — and the trap |
|---|---|---|
| OpenCode | PASS with configuration · FAIL-honest without it | a bare model entry silently reads the PNG as bytes — tested, and the model honestly reported seeing nothing. Declare "attachment": true and "modalities": {"input": ["text","image"]} on the model (the configuration in §14 includes them), and, run from a script, opencode run "@shot.png …" passes — verified on v1.18.21 after tracing issue #15728 |
| aider | PASS | image path as a plain file argument (aider … --message "…" shot.png); requires supports_vision: true in the model metadata file |
| Qwen Code | PASS with configuration · FAIL-honest without it | an environment-variables-only setup silently drops images. Declare the model in settings.json under modelProviders with capabilities: {vision: true}, then qwen -p "@shot.png …" works |
| Pi | PASS | image path as a positional argument after -p; the provider's models.json must list "image" in input. Swift IQ3_XXS: see below. |
| DeepSeek Harness | PASS | no attach flag exists or is needed — reference the path in the task text; settings.yaml must declare input: [text, image]. Gave the richest correct answer of the five |
One attempt per agent, one image, one question — so this table says "the plumbing works", not "this agent is good at vision". The image-attachment matrix at scale, across resolutions and multi-image requests, was never run.
For a scripted screenshot loop, the OpenAI-compatible interface itself is the zero-dependency path: send the image as a data:image/png;base64,… entry in an image_url content part alongside the text — exactly what the demonstration above used. Three rules regardless of route. Budget max_tokens generously: a thinking model reasons about the image before answering, and a tight cap returns an empty answer with the thinking complete (measured; §11's client contract applies to vision too). Clear old screenshots between iterations; each one keeps its thousands of tokens in the window. And on Windows, do not send image data-URIs from PowerShell 5.1 — Invoke-RestMethod silently fails to POST the roughly 261 KB body a 1440p PNG produces, and the signature is that the server log shows no task at all (§11). Python, curl, or PowerShell 7 and later all work.
Both Swift files score by sight, at the anchor’s image cost; Pi carries the screenshot
Swift IQ4_XS and Swift IQ3_XXS both scored 7 of 7 at the full image budget (--image-max-tokens 10580) and 0 of 7 with the image withheld (2026-10-04, greedy, meas n=1 per question, 7 questions per arm), so the score is sight, not guessing. Both files produced a median of 3,635 prompt tokens at the full budget, identical to the anchor’s (unsloth UD-IQ4_XS of the base model) run of the same instrument; the image costs the same prompt tokens under both chat templates, which differ elsewhere (§14); this is an instrument check, not a speed or quality comparison. In Pi, on the October launcher’s Swift IQ3_XXS two-slot vision configuration (123,904 per slot), the model read an attached screenshot correctly and declined to answer when it was withheld (below).
| File | Arm | Coarse meas | Fine meas | All meas | Blank meas | Prompt tokens (median) meas |
|---|---|---|---|---|---|---|
| Swift IQ3_XXS | FULL 10580 | 2 / 2 | 5 / 5 | 7 / 7 | 0 | 3,635 |
| REDUCED 1024 | 2 / 2 | 2 / 5 | 4 / 7 | 0 | 1,043 | |
| BLIND | 0 / 2 | 0 / 5 | 0 / 7 | 0 | 33 | |
| Swift IQ4_XS | FULL 10580 | 2 / 2 | 5 / 5 | 7 / 7 | 0 | 3,635 |
| REDUCED 1024 | 2 / 2 | 2 / 5 | 4 / 7 | 0 | 1,043 | |
| BLIND | 0 / 2 | 0 / 5 | 0 / 7 | 0 | 33 | |
| Anchor (UD-IQ4_XS), 2026-08-25 meas | FULL 10580 | 2 / 2 | 5 / 5 | 7 / 7 | 0 | 3,635 |
| REDUCED 1024 | 2 / 2 | 3 / 5 | 5 / 7 | 0 | 1,043 | |
| BLIND | 0 / 2 | 0 / 5 | 0 / 7 | 0 | 33 |
Two sweeps in one table: the Swift rows are one sweep, 2026-10-04; the anchor rows are this page’s 2026-08-25 run of the same instrument, the run §12 above reports. Only the prompt-token column is read across the two sweeps, as a check that the image costs the same; the scores are not compared across them. Greedy, n=1 per question, 7 questions per arm — smoke-test scale. Both Swift files loaded the projector mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf.
In Pi, on the October launcher’s Swift IQ3_XXS two-slot vision configuration (123,904 per slot, n4/p0.75, xhigh, 2026-10-04) meas n=1, the screenshot attach passed: the model answered “171”, the number printed on the screenshot (40.6 s); with the screenshot withheld it said no image was attached and declined to guess (14.5 s). The Pi configuration and the one-shot coding task are in §14. The first run of this probe failed in the harness, not the model: the card sets --api-key and the probe sent no key, so the server answered 401; the probe was fixed and re-run.
The Pi fine-tune’s IQ4_XS, with bartowski’s f16 projector for it, answered “171” from the same screenshot target (one question at the full image budget, 3,636 prompt tokens) and declined when the image was withheld (2026-10-06) meas n=1. Through Pi at medium, the attach probe passed (provisional) in 9.4 s and the withheld control held in 5.9 s. The seven-question instrument was not run on it (Appendix B).
A well-tuned local 27B is genuinely good at bounded work: answering factual questions, writing and explaining ordinary code, summarizing, and completing a self-contained task you can check when it finishes. It gets less reliable as a task gets longer, as the knowledge gets rarer, as the number of simultaneous requirements grows, and as the number of unchecked steps between you and the result increases. The decline is gradual rather than sudden, and much larger models are further along the same decline rather than free of it. The practical rule: give this model work you will look at, and escalate work that has to be right without you looking. This section is the evidence for that rule — and unlike the rest of this page, most of it is argued from published literature rather than measured here.
What this page did measure about capability: a five-set, 175-prompt benchmark suite at three effort levels, composite index 82.1 / 80.5 / 81.3 with GSM8K saturated at 100 and code execution pass rates of 92–100% on HumanEval and 84–92% on MBPP (§09). That is a real capability measurement on this file and this machine, and it is not comparable to the vendor's published figures for this model: different precision (4-bit against full), different harness, different sampling, different prompt formatting. Nor does it speak to the five territories below, all of which are about behaviour over long horizons that a 25-question set cannot reach. One agentic pipeline was also built and validated end to end on this machine — 22 model calls, 9 minutes 36 seconds of wall clock on one sandboxed repair task against real code — and it is worth reporting its outcome rather than only its existence: 0 of 96 target tests turned green, while all 561 previously-passing tests kept passing. The harness worked; the fix did not. That is one sample, and it proved the plumbing and the cost rather than the capability; the 10-task, three-effort sweep that would have measured capability was cut against a four-hour time gate on an eight-and-a-half-hour projection.
A model does not need trillions of parameters to act as an encyclopedia. High-frequency knowledge — geography, mainstream history, standard biology, common coding syntax — compresses comfortably into 27B (a model stores on the order of 2 bits of factual knowledge per parameter), which is why local factual question-answering feels frontier-grade. What frontier models buy with their scale looks like a difference in kind but is built from differences in degree: per-step reliability improves smoothly with capability, and because a long task succeeds only if every step does, those smooth gains compound into cliff-like gaps on exactly the work below — as reasoning engines, software architects, and autonomous agents. The boundary runs through five territories:
- Long-horizon reasoning. Direct questions are effortless at 27B; what degrades is the 30-step deduction, the formal proof, the architecture analysis — work where one early flaw silently poisons everything downstream: agent success decays exponentially with task length, and transformers solve deep compositional problems by pattern-matching that collapses as the reasoning graph grows — a collapse that reaches even frontier models at sufficient depth, so the frontier extends the horizon rather than escaping the regime. This page's judged runs located the local edge precisely: at
xhigh, the 27B delivered 100% of a complex single-file specification, twice. Multi-file systems, where the flaws hide in the seams between files, are where the footing gives — and note from §09 that no read-then-edit task against existing code was ever run here, so that boundary is argued rather than measured. - The long tail of knowledge. 27B keeps what training data repeated often — fact recall tracks how many times a fact appeared in pre-training — and frontier models retain far more of the rare: the niche legal precedent, the legacy framework, the unusual drug interaction (the studies measure trivia-style entity facts; these are the practical analogue). But the tail deficit shrinks rather than disappears — scaling fails to appreciably improve memorization of tail facts, and closing the gap by scale alone would take many orders of magnitude more model. The danger is the failure mode: on obscure ground no model reliably says "I don't know" — training and benchmarks reward confident guessing over abstention — and a small model simply has more gaps for the fluent guess to fill. Treat confident answers about rare domains as unverified drafts, or ground them with retrieval: retrieval-augmented small models beat plain models orders of magnitude larger on exactly these facts.
- Ultra-long recall. A whole codebase or a 500-page audit in one prompt is two problems wearing one coat: holding the text, which this card's 122k window does (§05), and keeping every cross-document dependency straight, which improves with model scale — in RULER's within-family test, a 34B degrades far less than a 6B trained on the same context — but scale mitigates rather than cures: only about half of models claiming 32K or more actually handle 32K, information buried mid-context is used far worse than at the edges, and once retrieval requires inference instead of literal matching, most long-context models fall below half their short-context accuracy by 32K.
- Multi-constraint instruction following. "Five custom interface rules, strict typing, an exact schema, and ten edge-case tests — simultaneously." Rule stacks like that degrade small models first, and quietly: the dominant failure at high instruction density is silent omission — the output stays a coherent, compliant-looking document — and violations are hard enough to spot that benchmarks need model judges to find them. Every model declines as constraints stack, so fidelity under heavy constraints is one of the clearest capability dividends — frontier reasoning models hold near-perfect through about 150 simultaneous instructions where small models decay exponentially — but it buys later, slower degradation, not immunity: even the best fall to about 68% at 500 instructions.
- Autonomous agentic reliability. An agent that calls tools, reads its own error logs, and self-corrects across dozens of steps multiplies its per-step error rate through every step — measured agent success declines exponentially with task length, like a half-life — and worse than independent errors predict, because models self-condition on their own earlier mistakes, amplifying drift into failure. Frontier agents push the horizon from minutes toward hours, but at 50% success — at their quoted horizon they still fail half the time. This is why the setups in §14 shine for supervised sessions and wobble on multi-hour autonomous loops, and why §09's defect risk matters most in exactly that mode, where nobody is watching the output.
| Task type | 27B local | Frontier model |
|---|---|---|
| Factual question-answering and summarization | excellent | excellent |
| Standard code and prose | great | exceptional |
| Multi-file system architecture | moderate — struggles | exceptional |
| 50-step autonomous tool loops | high drift / failure | much longer horizons — still ~50% at their own limit |
| Obscure or niche domain expertise | prone to confident guessing | better, still unreliable — verify or retrieve |
This table is a summary of the literature above plus this page's own judged runs; no cell in it is a measurement of a frontier model on this machine.
A well-tuned local 27B is a fast, private, tireless colleague for the bounded work that fills most of a day — and this page exists to make it as good at that as the hardware allows. Frontier models exist for the multi-step, high-stakes cognitive work where accuracy must decay as slowly as possible — and even they decay. The two compose: route the bounded work locally and escalate the architecture, which the planner-and-executor split in §14's DeepSeek Harness setup expresses as a configuration file. Knowing which task you are holding is itself a setting — the one no flag can fix.
Every recipe in §03 ends with an OpenAI-compatible server, so any OpenAI-compatible agent can connect with the same three facts: the address http://localhost:1234/v1, the model name the server reports (qwen/qwen3.8-27b, set by --alias), and the context and output limits from your -c setting. Two things go wrong often enough to state once for all five agents below. Tell the client the model's real limits — if its context setting and the server's -c drift apart, generations get cut off silently. And override the temperature — several clients send 0 by default, which puts this model into repetition loops; it wants 1.0 with top-p 0.95. The five setups below are ordered by popularity, and each was installed, configured exactly as printed, and run through a live request against the reference server on 2026-08-22.
OpenCode · the most-starred coding agent
# install: npm i -g opencode-ai (or the curl installer / brew) # config: ~/.config/opencode/opencode.json (or opencode.json in the project) { "$schema": "https://opencode.ai/config.json", "provider": { "llamacpp": { "npm": "@ai-sdk/openai-compatible", "name": "llama-server (local)", "options": { "baseURL": "http://localhost:1234/v1", "apiKey": "dummy" }, "models": { "qwen/qwen3.8-27b": { "name": "Qwen3.8 27B (local)", "attachment": true, "modalities": { "input": ["text", "image"], "output": ["text"] }, "limit": { "context": 122880, "output": 106496 } } } } } } # run: opencode --model llamacpp/qwen/qwen3.8-27b (or pick via /models) # the attachment/modalities pair is REQUIRED for images: without it OpenCode # silently reads a referenced PNG as text bytes (section 12's agent table) # # THE TWO NUMBERS ABOVE ARE TIED TO YOUR SERVER'S -c. Derive them, do not copy: # context = the server's -c value exactly # output = -c minus your longest expected prompt; 106496 = 122880 - 16384. # For recipe [2] at -c 180224 that is context 180224, output 163840; # for recipe [3] at -c 262144, context 262144, output 245760.
Two gotchas from the documentation: baseURL must sit inside options, not at the provider root, and the provider identifier must be a custom name ("llamacpp", "local") — reusing a built-in provider name breaks resolution. (llama.cpp example)
aider · the git-native pair-programmer
The most detailed walkthrough of the five — its metadata-and-limits pattern is the template every other client repeats in its own configuration format.
Step 1 · Install aider
# needs Python 3.9+ — any one of these: python -m pip install aider-install && aider-install # official installer (recommended) pip install aider-chat # plain pip uv tool install --force --python python3.12 aider-chat # uv, isolated environment
Step 2 · Point aider at the server
Aider's openai/ model prefix means "any OpenAI-compatible server" — llama-server and OpenVINO Model Server both qualify. Set two environment variables (the key is required by aider; make it match the server's --api-key if you set one, otherwise any placeholder works):
# Windows (cmd / batch) # Linux / macOS
set OPENAI_API_BASE=http://localhost:1234/v1 export OPENAI_API_BASE=http://localhost:1234/v1
set OPENAI_API_KEY=dummy export OPENAI_API_KEY=dummy
Step 3 · Tell aider the model's limits
Aider knows nothing about a local model until you describe it. Create .aider.model.metadata.json in the directory you launch aider from, or in your home directory. Match max_input_tokens to the server's -c value — these two numbers drifting apart is the classic silent failure:
{
"openai/qwen/qwen3.8-27b": {
"max_input_tokens": 122880, // = the server's -c, exactly
"max_tokens": 106496, // = -c minus your longest prompt (122880 - 16384)
"input_cost_per_token": 0,
"output_cost_per_token": 0,
"supports_vision": true
}
}
The model name after openai/ must match what the server reports at /v1/models — llama-server sets it with --alias, so every recipe in §03 serves as qwen/qwen3.8-27b.
Step 4 · Fix aider's defaults for a thinking model
Aider sends temperature 0 by default, which the model card advises against for thinking mode and which is a known repetition risk on reasoning models — this page's own audit of ten long greedy transcripts found 10 of 10 clean, so treat it as a risk worth avoiding rather than a certainty. It also needs to know to strip thinking blocks before parsing edits. Create .aider.model.settings.yml next to the metadata file:
- name: openai/qwen/qwen3.8-27b
use_temperature: 1.0
reasoning_tag: think
extra_params:
top_p: 0.95
Step 5 · Launch
aider --model openai/qwen/qwen3.8-27b --edit-format diff --timeout 3600
--edit-format diff— the model writes only changed sections instead of whole files. At local speeds this is the biggest quality-of-life setting there is: a 60 KB file edit drops from minutes to seconds. Diff mode also degrades gracefully on cut-offs, because complete sections still apply.--timeout 3600— long thinking runs outlive default client timeouts; the server log shows a timeout asClient disconnected. Stopping generation.- First-run sanity checks: the model-warning banner should be gone (the metadata file was found), and the banner should read diff edit format. If edits arrive mangled, the fallback is removing
--edit-format diffto use whole-file mode. - Vision: with
supports_vision: trueand the server launched with--mmproj,/add screenshot.pngworks and the local model will see it (§12).
Qwen Code CLI · Qwen's own terminal agent
# install (Node.js 20+) npm install -g @qwen-code/qwen-code # configure via environment variables — put them in ~/.qwen/.env to persist: OPENAI_BASE_URL=http://localhost:1234/v1 OPENAI_API_KEY=dummy OPENAI_MODEL=qwen/qwen3.8-27b # then just run: qwen
The key must be present even for a local server ("Missing credentials" otherwise), and set the context window explicitly in its settings if offered — Qwen Code's built-in defaults assume the cloud model's limits, not your -c value. For images, declare the model in settings.json under modelProviders with capabilities: {vision: true}; an environment-variables-only setup silently drops them (§12). (setup notes)
Pi coding agent
# custom providers live in ~/.pi/agent/models.json: { "providers": { "llamacpp": { "baseUrl": "http://localhost:1234/v1", "api": "openai-completions", "apiKey": "dummy", "compat": { "supportsDeveloperRole": false, "supportsReasoningEffort": false }, "models": [ { "id": "qwen/qwen3.8-27b", "name": "Qwen3.8 27B (local)", "contextWindow": 122880, "maxTokens": 106496, "reasoning": true, "input": ["text", "image"] } ] } } } # one-shot from the command line, no interactive session: # pi --model llamacpp/qwen/qwen3.8-27b --no-session -p "task" # contextWindow = the server's -c; maxTokens = -c minus your longest prompt.
The compat flags matter: they tell Pi to send a plain system role and to skip the reasoning_effort request parameter, which is the dialect llama-server expects — effort is a server-side setting here (§09). Pi enforces maxTokens on the client side, so size it like aider's; this is the client whose default cap causes the mid-generation cut-offs described in §11. (documentation)
bytkim’s Qwen3.8-27B-pi, a fine-tune of this model for the Pi harness, runs in Pi through the same provider block. Because supportsReasoningEffort is false, Pi sends no effort and the server’s launch setting is what the model gets (a reading of the provider setting; Pi’s request bodies were not captured). Its smoke test is below, and its measurements and launch command are in Appendix B.
Pi passes end to end on Swift IQ3_XXS with vision at two slots
Pi works end to end with Swift IQ3_XXS on the October launcher’s two-slot vision configuration n=1 (123,904 per slot, n4/p0.75, xhigh, 2026-10-04): it wrote a script that prints FizzBuzz for 1 to 30, ran it and reported its last line correctly (PASS, 584.4 s) meas, and it carried a screenshot to the model, with the withheld-image control holding (§12). One setting is known to differ from the anchor (unsloth UD-IQ4_XS of the base model): the Swift chat template rejects high effort when it is rendered, where the anchor’s treats high as xhigh, so set xhigh (below). aider, OpenCode, Qwen Code and DeepSeek Harness were not tested on the Swift files (§15).
# as tested with Swift IQ3_XXS, two slots (-c 247808 --parallel 2): # contextWindow = per-slot window (-c / --parallel); maxTokens as tested; # /slots read n_predict 108853-110099 on Pi's requests, below the window { "providers": { "llamacpp": { "baseUrl": "http://localhost:1234/v1", "api": "openai-completions", "apiKey": "dummy", "compat": { "supportsDeveloperRole": false, "supportsReasoningEffort": false }, "models": [ { "id": "qwen/qwen3.8-27b", "name": "Qwen3.8 27B 124k per slot ctx [8] default", "reasoning": true, "input": ["text", "image"], "contextWindow": 123904, "maxTokens": 123904, "thinking": { "mode": "effort", "efforts": ["low", "medium", "xhigh"], "defaultLevel": "xhigh", "requiresEffort": false } } ] } } }
The file is reproduced as it was used. Its thinking block is inert: Pi 1.0.2 reads no thinking, defaultLevel or requiresEffort field, and with supportsReasoningEffort false it sends no effort, so the server’s --chat-template-kwargs decides (checked in Pi’s installed code, 2026-10-08).C45 The server ran the card’s flags unchanged, plus five the probe supplied for this machine and for Pi’s routing: --model and --mmproj (the local file paths), --alias (the model id Pi asks for), and --host and --port (to match Pi’s baseUrl); Pi’s provider file was read, never edited. The server applied the card’s own sampler to each of the seven Pi requests that llama-server’s /slots endpoint showed during the runs — temperature 1.0, top_k 20, top_p 0.95, min_p 0.0 meas, polled once a second, so a shorter request can be missed; /slots cannot show whether Pi also sent those values, and the reasoning_format deepseek and chat_format peg-native it lists are the server’s own parsing settings.
Set xhigh, not high, in a Swift launcher’s --chat-template-kwargs. Swift IQ3_XXS and Swift IQ4_XS share a chat template (8,952 characters meas) that accepts xhigh (its default), medium and low, and raises on high: “Unexpected reasoning effort high. Supported types are xhigh (default), medium, and low.” The anchor’s template (9,993 characters meas) accepts high and renders it as xhigh. The same 8,952-character template, byte for byte, is in bartowski’s IQ4_XS and IQ3_XXS of the Pi fine-tune and in lmstudio-community’s Q4_K_M of the base model (headers read 2026-10-06, and 2026-10-07 for the Pi IQ3_XXS), so high raises on those files too; the template that accepts it is the unsloth file’s.C40 This is the template’s own behaviour under a jinja2 render; llama-server renders templates with its own engine and was not started with high on a Swift file (§15), and it ignores a per-request reasoning_effort field (§09). The server’s own renderer was checked on this template for the three accepted levels: on build 10502, /apply-template on the Pi fine-tune’s IQ4_XS gave no effort line at the launch setting medium, and the low and xhigh lines when a request carried them in chat_template_kwargs, as the jinja2 render predicts (2026-10-06, Appendix B). high was not sent.
The setups above for aider, OpenCode, Qwen Code and DeepSeek Harness were verified with the anchor only. Rendered with jinja2, the Swift template produced the same bytes as the anchor’s for one test prompt at low, medium and xhigh and for the 175 benchmark prompts of §09 meas; no tool definitions or tool calls were rendered, so treat those four setups as untested with the Swift files, not as known to fail (§15).
Pi also ran end to end on the Pi fine-tune’s IQ4_XS (2026-10-06) meas n=1. It used the October launcher’s Swift IQ4_XS vision settings with the Pi file and its projector (139,264 × 1, n4/p0.75), with medium set at launch, no effort sent by Pi, and Pi’s existing 123,904-token provider entry above. The one-shot FizzBuzz task passed in 48.1 s, the screenshot attach passed (provisional) in 9.4 s, and the withheld control held in 5.9 s, with the server applying temperature 1.0 to every task. A model download was running at the time, so these wall times are single readings under that load. This is a smoke test of the harness, not a measure of agent-task quality, and its wall times are not comparable with the Swift IQ3_XXS run above (different file, effort and slots; Appendix B).
DeepSeek Harness · plugin-based agent framework (developer preview)
# install globally, then run the web UI (an npx cold cache can hang for minutes): npm i -g @deepseek-ai/dsh dsh web # custom OpenAI-compatible providers live in $DSH_HOME/settings.yaml. # BOTH blocks are required: registering the provider is not enough — without the # routing block the harness silently keeps its cloud default and fails with # MISSING_CREDENTIAL (verified end-to-end on the reference machine, 2026-08-22): agent-default-model: provider: llamacpp model: qwen/qwen3.8-27b llm-pi-ai: providers: llamacpp: api: openai-completions baseURL: http://localhost:1234/v1 apiKeyEnv: LLAMA_API_KEY # set LLAMA_API_KEY=dummy to match --api-key compat: supportsDeveloperRole: false # llama-server wants a plain system role maxTokensField: max_tokens models: - id: qwen/qwen3.8-27b input: [text, image] # vision works when the server runs with --mmproj contextWindow: 122880 # match the server's -c value exactly maxTokens: 106496 # REQUIRED: dsh enforces this client-side (same # schema as Pi). Without it the small default cap # cuts long thinking replies mid-run — the symptom # is "Output token limit reached ... send continue" # on any xhigh answer (section 11)
The harness's headline trick is routing: it plans with heavy reasoning, then fans the plan out to sub-agents that execute with light reasoning. Two things to know when backing it with llama-server. First, llama-server ignores a per-request reasoning_effort field (the same dialect issue as Pi above) — effort is fixed server-side with --chat-template-kwargs, so either run one server at xhigh (this page's quality-first default), or run two llama-server instances — one xhigh planner on one port, one low executor on another — and register both as separate models in settings.yaml so the harness can route between them. A per-request chat_template_kwargs carrying reasoning_effort did reach the template at /apply-template on build 10502 (2026-10-06, §09); generation with it is untested, so the two-server setup is still the verified one. Second, its parallel sub-agents each hold their own context: with --parallel 1 (every recipe here) requests queue rather than run concurrently, which is the right default on a single 24 GB card for window reasons — with --parallel 2 each concurrent sub-agent gets half of the window -c sets, and the second slot itself added only 134 MiB of board VRAM meas n=1 at -c 32768 (2026-08-23, UD-IQ4_XS, q8_0 KV, drafter off, kv_unified off as logged).C28 It is also the right default for latency reasons: a matched pair (2026-08-25, UD-IQ4_XS, drafter on at n10/p0.5) measured two concurrent slots at +22.0% aggregate throughput but −35.4% per slot (§03), so a harness whose sub-agents you are waiting on finishes each one faster with a single slot. For non-thinking "fast mode" runs, Qwen's official instruct sampling is temperature 0.7 · top-p 0.8 · presence penalty 1.5.
Whatever the agent, its context and maximum-output settings must match the server's -c value — the client-side limits are the only thing preventing silent mid-generation cut-offs, because the model itself cannot budget its own tokens (§11).
Everything below is from one instrumented run of the aider polyglot agentic coding benchmark on one RTX 3090 (GA102, 24 GB GDDR6X, 350 W enforced board limit), running Qwen3.8-27B UD-IQ4_XS at 14.25 GB resident with MTP speculative decoding on (mean accepted length 3.55). Board power and clocks were logged by NVML at roughly 0.5 s; the server’s own per-request split was polled at 1 Hz; host CPU counters were sampled at 3 s. Run tag iq4xs-agentic, 2026-08-26.
Wall and system power — every watt here is the GPU board rail read in-band by NVML, so power-supply conversion loss, CPU, RAM, drives and fans are all excluded. Per-process VRAM and power attribution is impossible under Windows WDDM, where nvidia-smi pmon reports a dash for every process. Memory-junction temperature reads NULL on this part (the NVML field is absent in every sample), so memory thermal headroom is unknown. And there is no time-series of SM clock, board power and throughput against wall clock: the sustained drift — 1,453 to 1,606 MHz with board power rising 305.5 to 341.1 W at constant throughput — is published as two endpoints, which cannot separate monotone drift from stepped or oscillatory behaviour. The telemetry to draw it was collected; the plot has not been made.
What limits this part
The first question about any accelerator workload: is it starving for memory bandwidth, or is it sitting on a power limit?
Speculation is a workload transform, not a speed-up: it moves the operating point off the memory roof and onto the power limit. What a better draft head would buy is the next question.
What the clocks actually did
The roofline shows what the power cap costs in principle. These two figures show what it did to the clocks in practice, over 128 minutes of real traffic.
Which limit was binding
The operating point says where the part sits. These two figures say which clock-event limit NVML reported, moment by moment and in aggregate.
What capping the board costs
The limit is binding 93% of the time. What happens if you lower it?
Saturating power-cap sweep (measured 2026-08-28), with every arm actually reaching its cap — the earlier sweep’s stock arm sat 32.1 W under its own cap:
| Cap | t/s | Mean W | J/token | SM MHz | Temp |
|---|---|---|---|---|---|
| 350 W | 76.76 | 337.5 | 4.392 | 1,688 | 80.8 °C |
| 300 W | 73.60 | 293.2 | 3.978 | 1,580 | 77.8 °C |
| 250 W | 65.58 | 246.1 | 3.748 | 1,281 | 76.9 °C |
Relative cost of each cap, against each run’s own 350 W baseline (bold column governs):
| Superseded (non-saturating, 2026-08-25) | Saturating | |
|---|---|---|
| 300 W | −5.0% t/s, −11.2% W, −6.5% J | −4.1% t/s, −13.1% W, −9.4% J |
| 250 W | −12.8% t/s, −24.8% W, −13.7% J | −14.6% t/s, −27.1% W, −14.7% J |
The mechanism: throughput falls less than the core clock does (1,688 → 1,580 → 1,281 MHz) because decode is partly memory-bandwidth-bound and -pl does not touch the memory clock. Capping to 300 W costs 4.1% of throughput for 13.1% less board power and 9.4% less energy per token; 250 W costs 14.6% for 27.1% less power and 14.7% less energy.C22
Real agentic traffic sits pinned on the power limit, not near it. The next figure shows the regime the workload actually occupies.
Four benchmark sets, two passes each, second pass in reversed order so no set is confounded with the board temperature it met. Every set reproduces within 3.8% (three of four within 2.4%). But sort each set by which of its two passes was hotter and the hotter, more-throttled pass is cheaper per token in all four cases:
| Set | Throttle residency | J/token |
|---|---|---|
| GSM8K | 0.0% → 67.4% | 8.7420 → 8.4104 (−3.79%) |
| ALPACA | 29.5% → 70.9% | 8.4711 → 8.2681 (−2.40%) |
| MeetingBank | 0.2% → 1.3% | 11.6387 → 11.5873 (−0.44%) |
| MT-Bench | 63.0% → 73.4% | 8.2886 → 8.2542 (−0.42%) |
Thermal throttling is the power-cap trade taken involuntarily. Two independent routes now say the stock 350 W limit sits past the efficiency knee for this workload. MeetingBank’s 11.6 J/token outlier is not noise: long prompts, short answers — 3,876 generated tokens against MT-Bench’s 18,609 for the same 25 items — so whole-window energy over generated tokens charges all the prefill to a small denominator, and prefill never heats the board the way decode does. measured
Where the time and energy go
The power limit constrains the part, but the workload alternates between prefill, decode and idle. These figures show how the time, energy and tokens split between phases, and what drives the per-exercise cost.
The spread looks large, but the next figure shows it is not an independent degree of freedom.
What the host contributes
Everything above is the GPU board. The host — an i5-13600KF running the llama-server, a Docker-based scoring container, and the WSL2 boundary — contributes its own load, and the question is whether that load reaches the GPU.
Which lever is worth pulling
Of the nine levers measured in this campaign, few moved anything: speculative decoding is by far the largest (2.2×), then the quantisation format at equal file size (1.30×), then the power cap (350 → 250 W). The rig’s noise floor is 1.1% total range on throughput and 2.9% on J/token; six of nine levers were never paired for energy; no interaction between levers was measured. The tornado below ranks all nine on one chart.
The workload that drove it
The figures above are only as representative as the workload that produced them. These show what the agentic coding loop actually asked of the card.
What this page got wrong, and what killed each claim
A guide that publishes what it got wrong — with the date, the old number, and the measurement that replaced it — is more trustworthy than one that quietly overwrites its claims. The corrections below were scattered through the body as inline interruptions: “Earlier editions printed…”, “SUPERSEDED”, “was wrong”. Individually each was honest. Together they broke the reading of nearly every paragraph to tell you about a number you never saw. They are collected here so that what you act on is separated from what you might study, and so that the body prose reads as a guide rather than as a changelog.
The retractions from the coding benchmark arm (2026-08-26, aider’s polyglot benchmark, 225 exercises, two full arms on this same RTX 3090) are folded in here as well. Two findings from that run set the stage for the corrections. Speculative decoding moved the workload off the memory roofline and onto the power limit — the single largest lever measured anywhere in the campaign. Drafter off, the card ran at 69% of its bandwidth ceiling and was bandwidth-bound. Drafter on, measured throughput exceeded the one-weight-pass-per-token ceiling of 65.7 t/s: real traffic fell to 43% of bandwidth while 97% of busy samples hit the 350 W board limit. And the levers that actually moved were few: speculation itself (2.2×), the quantisation format at equal file size (1.30×), and a power cap from 350 to 250 W (−14.6% throughput for −14.7% J/token). KV cache quantisation to q4_0 cost nothing measurable in retrieval out to 119,435 tokens. Host CPU contention cost −5.4% while the GPU clock rose — invisible to the logs. Everything else was either not a lever or was retracted after measurement.
Relocated corrections
Each note says what the page used to claim, what it claims now, and the measurement that changed it. Superscript references in the body point here.
- 2026-08-25 · drafter trade priced on the wrong token regime. The vision recipe’s drafter trade was priced at 5.6% against a reasoning stream when the reader was writing code. Measured on code at 112,735 tokens of depth: +17.1% (62.47 against 53.33 t/s). §03
- 2026-08-25 · window shrink called “pure loss”. Shrinking the window from 122,880 to 98,304 was called “pure loss”. Measured: 87.04 against 87.49 t/s — inside both arms’ own probe spread. What the 24,576 tokens buy is the desktop reserve, not throughput. §03
- 2026-08-25 · effort constraint blamed on prefill. The big-window pick shipped
medium“because of the 278-second prefill”. Effort changes how much the model writes, not how long it takes to read a prompt. The actual constraint is the window pot: at 163,124 tokens of depth, only 33,484 remain, and a complete-buildxhighanswer wants 61,500–75,800. §03 - 2026-08-25 ·
xhighpriced at “about three hours”. That is the cost of a complete build — one task, four samples. An ordinaryxhighanswer averages 2,217 thinking tokens across 175 benchmark prompts and takes about five minutes at 6–8 t/s. The old figure overstates ordinary work by roughly thirty times. §03 - 2026-08-23 · idle power printed backwards. Published as 33.2 W with no server against 30.7–31.1 W loaded — physically backwards. The board was still cooling from the previous job. Settled: 29.9 W with no server, 34.1 W with the model resident. §03
- 2026-08-23 · board power printed as a flat “~344 W”. Sustained draw drifts 305.5 → 341.1 W (+11.7%) at constant throughput and constant temperature, tracking its own SM clock from 1,453 to 1,606 MHz. §03
- 2026-08-25 · power capping listed as “unmeasured”. It read “unmeasured, requires elevated shell”. Measured: 300 W costs −5.0% throughput for −6.5% J/token; 250 W costs −12.8% for −13.7%. The stock cap sits past the efficiency knee. §03
- 2026-08-23 · VRAM budgeted by arithmetic floors. Earlier budgets used 64 KiB per token at fp16, 34,816 B/token at
q8_0and “1–1.5 GiB of compute buffers”. Each of those is a floor, not a figure. The fitted model from 17 server loads reads 39,936 B/token drafter-off and 45,056 drafter-on (+15% and +29% over arithmetic), reproducing all 17 loads to within 127 MiB. §05 - 2026-08-23 · big window “verified” by short probes. The 262,144-token window was declared safe on the strength of short probes that never filled it. Filled to 90,886 tokens, speed collapsed from 30.6 to 8.0 t/s (3.82×). Short probes test a ceiling that does not exist. §05
- 2026-08-23 · speed ceiling timed with thinking on. The copying-text ceiling was printed as 119.8 t/s at 0.93 acceptance. The model was reasoning about copying rather than copying. With thinking off: 148.7 t/s at 0.99. §06
- 2026-08-27 · sigma separations divided by the wrong denominator. Every sigma in the agreement table was divided by one file’s standard error instead of by the combined error of the difference — the square root of the sum of the two squares. Values were printed about a third too large: 15 sigma where the figure is 10.9, 17 where it is 11.8. No conclusion changes; every rung still separates cleanly. §08
- 2026-08-26 · desktop worst case was a spill reading. Published as 1,181 MiB, taken while the card was in a spill state rather than at rest. Resting worst case: 1,669 MiB. The VRAM reserve moved from 1,308 to 1,796 MiB, and one verdict flipped with it: UD-Q2_K_XL at the full window keeps 1,717 MiB, published as 409 MiB MORE than the reserve when that reserve was 1,308, and 79 MiB SHORT against the corrected 1,796. §08
- 2026-08-25 · slack threshold was derived from a spill-state reading. The glossary’s “slack” entry used 1,308 MiB, built on a desktop maximum of 1,181 MiB that an audit showed was not the maximum: the campaign’s own power-matrix round carries a direct reading with no server loaded at all of 1,669 MiB, and an earlier bare-idle reading the same day of 1,179 — so this desktop moved 490 MiB between two idle measurements. The old 133 MiB floor was worse than stale: it came from a board at 24,296 of 24,576 MiB that was already spilling, so it recorded a desktop being evicted rather than a desktop at rest. Corrected to 1,796 MiB. §02
- 2026-08-27 · drafter trade verdict stated as “no window to trade”. The vision recipe’s drafter callout said “there is no window to trade” and presented 1,776 MiB as room left over. The measurement file marks that arm
"clears_reserve": true— but computed it against the superseded 1,308 MiB reserve, and against the corrected 1,796 MiB threshold it does not clear. §03 - 2026-08-26 · two effort-level claims in §09. (a) “The recipes ship medium” was wrong after the 24 GB default was changed to
xhigh: five of nine shipmedium, four shipxhigh. (b) “Override per request” was wrong: llama-server ignores areasoning_effortfield in the request body, so a reader who followed the advice gotmediumwhile believing they hadxhigh. §09 - 2026-08-24 · Q4_K_M labelled “maximum measured quality”. The perplexity gap between Q4_K_M and UD-IQ4_XS is 0.061 — inside this page’s own combined error of ±0.063. The label overstated what the measurement resolved. §08
- 2026-08-27 · drafter-sweep discrepancy blamed on the window. The later sweep’s lower absolute figures were attributed to its larger
-cwindow. Wrong:draft-nmax-sweep.pysets-c 32768in its base flags. What actually differs is the prompt and the token budget. §06 - 2026-08-25 · replication gap blamed on Flash Attention. The paragraph said
-fa ontogether with--reasoning offexplained the gap. Wrong:-fadefaults toauto, auto resolves to on, and those runs already had Flash Attention. The whole gap is--reasoning off, which changes acceptance 0.523 → 0.611. §08 - 2026-08-24 · “the one file that fails” was written before the ladder. The 1.835-bit-per-weight file was the only known failure. The eight-rung ladder found three more: 2, 3, 5 and 28 empty replies at 2.481, 2.153, 1.994 and 1.835 bits per weight. §08
- 2026-08-25 · 16 GB recipe shipped a file that does not fit. Both the NVIDIA 16 GB and Intel B50 recipes shipped
UD-Q3_K_XLat-c 49152, priced at “about 15.3 GiB” from a derived bandwidth estimate. Measured file by file (§08):UD-Q3_K_XLneeds 15,530 MiB at-c 32768and 16,906 MiB at-c 65536, so-c 49152lands near 16,218 MiB — 1,142 MiB over the budget. A derived number that was never checked against a measurement. The B50 copy was missed for a day after the NVIDIA recipe was corrected. §03 - 2026-08-28 · coding benchmark flip rate called “never measured”, then a p = 0.027 published and retracted the same day. The page claimed this benchmark’s own run-to-run flip rate “has never been measured here”. Measured 2026-08-27: UD-IQ4_XS against itself, whole edit format, reasoning off,
-c 32768, n4/p0.75, greedy, 113 exercises paired by case and language. Both runs scored 43.4%; 18 of 113 flipped (15.9%, 95% interval 10.3–23.8), split exactly nine each way. A second floor followed overnight: UD-Q2_K_XL against itself, 25 of 113 (22.1%, interval 15.5–30.6). On 2026-08-28 the page briefly claimed “a real quantisation effect does survive” at z = 2.21, two-sided p = 0.027, testing the 60 exercises the two quantisations disagree on (60 of 225 = 26.7%, interval 21.3–32.8) against the 4-bit floor only. That test silently assumed the two arms are equally noisy. Against the 2-bit floor: z = 0.91, p = 0.364. Against both pooled: z = 1.93, p = 0.053 — does not clear 0.05. No quantisation effect is demonstrated; the 60 exercises are consistent with what this benchmark does to a single file run twice. The claim that the 2-bit file is noisier (15.9% against 22.1%) is a point estimate, not a finding: z = −1.19, p = 0.236. §08 - 2026-08-28 · power-cap savings understated (supersedes corr-7). The 2026-08-25 sweep ran on synthetic decode averaging 305 W against a 350 W limit — the stock arm sat 32.1 W under its cap and never reached it, so the saving looked smaller than it is. Re-measured 2026-08-28 with every arm saturating its cap (busy-sample cap-flag share 99.0–100.0%). Corrected: 300 W costs −4.1% throughput for −9.4% J/token (was −5.0%/−6.5%); 250 W costs −14.6% throughput for −14.7% J/token (was −12.8%/−13.7%). Capping to 300 W is a better deal than the page claimed. §11
- 2026-08-28 · headline speed measured at a window the recipe does not ship. Recipe pick [2], UD-IQ4_XS text-only at
-c 180224, n10/p0.5, printed 93.9 t/s and 847 MiB of slack, measured. The speed was measured at-c 32768, not at the-c 180224the recipe ships. The same drafter at-c 180224measures −3.1% — 73.94 t/s, three settled probes, 1.91% spread — larger than either arm’s own scatter. Slack at-c 180224is 582 MiB, which does not clear the 1,796 MiB desktop reserve; the 847 MiB figure is the same window measured days earlier, and it does not clear the reserve either. §03 - 2026-08-28 · drafter rule wrong, and the rule it replaced also wrong. The glossary read “the second is the one that predicts speed” and “Read draft length first” — mean draft length as the predictor of speculative-decoding throughput. That rule had itself replaced an earlier one: that acceptance predicts throughput (contradicted by a setting at acceptance 0.897 running 33% slower than one at 0.611). Measured 2026-08-28, UD-IQ4_XS,
-c 32768,q8_0KV, three settled probes each after a discarded warmup: n4/p0.00 at 82.73 t/s, acceptance 70.5%, mean draft length 2.803; n10/p0.5 at 76.32 t/s, acceptance 51.6%, mean draft length 4.035. The shorter drafts win by 8.4%. Neither quantity alone ranks configurations when n-max differs; the ranking quantity is accepted tokens per unit of drafting work. §06 - 2026-08-28 · 9–13% energy-arm speed gap called “unexplained”. Two three-probe samples of the same speculative configuration differed by 9–13% and the page logged this as an open item. Twenty back-to-back probes (UD-IQ4_XS, n10/p0.5,
-c 32768,-ctk/-ctv q8_0, greedy, 700 tokens, one warmup discarded, RTX 3090 at 350 W, build 10502) measured mean 88.76 t/s, standard deviation 2.86, range 86.22–94.00, coefficient of variation 3.22%. The probes fall into two clusters — fourteen near 87.0 and six near 92.9 — with nothing between and no trend over time, so a three-probe sample lands in one cluster or the other. The previously published 86.91 sits inside the range (z = −0.65). Plain-decode arms reproduced within 1–3% on the same re-run, so this spread is a property of speculative decoding, not of the machine. One discarded warmup is not enough for a speculative arm: a fresh-server three-probe re-run gave 77.59, 78.12, 85.79 (mean 80.50, about 8% below the settled distribution), still climbing through all three probes. Retracted 2026-08-28 at 16:30. That closure was published at 16:00 from one twenty-probe session and presented it as the configuration. Two further twenty-probe runs of the identical configuration, at 16:15 and 16:22, gave session means of 77.35 and 77.15 t/s (standard deviations 3.22 and 2.70, ranges 73.44–84.21 and 74.70–83.34). Runs 1 and 2 do not overlap at all; their means differ by 14.8%. The previously published 86.91 sits inside run 1 (z = −0.65) and outside runs 2 and 3 (z = +2.97 and +3.61). Session-to-session variation is larger than the spread within any session, so one session could not measure this band and the closure was wrong. Within a session, throughput correlates with mean draft length (r = +0.683 to +0.894 across six greedy sessions, every one positive). Between sessions, draft length cannot be the mechanism: six greedy sessions produced bit-identical acceptance and draft-length sequences while session means spanned 75.71–78.65 t/s. Across all nine sessions the means span 75.71 to 88.76 t/s, a 16.6% range, and the drafter does precisely the same work in each. The between-session cause is unmeasured.C26 §03 - 2026-08-28 · session-to-session gap attributed to the drafter’s draft length. This is the third correction to this passage today. The mechanism claim was published at 16:30 as part of the C25 retraction and withdrawn at 17:00. Six greedy sessions (six probes each, identical configuration, one server per session torn down between, all 2026-08-28) produced bit-identical acceptance and mean-draft-length sequences probe for probe while session means spanned 75.71–78.65 t/s. A within-session correlation was extended to a between-session gap it cannot explain. The within-session correlation is real (r = +0.683 to +0.894, every session positive); the between-session cause is unmeasured. §03
- 2026-08-28 · speed band printed as 83–86 t/s. Twenty-three sessions of the same speculative configuration (UD-IQ4_XS, n10/p0.5,
-c 32768,-ctk/-ctv q8_0, greedy, 700 tokens, fresh server per session, one warmup discarded, RTX 3090 at 350 W, build 10502, all 2026-08-28) landed on two distinct levels: 20 sessions at 75.71–78.65 t/s and 3 sessions at 88.18–88.76 t/s, about 13% apart, with nothing in between. The earlier 83–86 band mixed sessions from both levels and was narrower than the real spread. A designed paired test (four alternating COLD/HOT pairs, five probes each) refuted the prior-load hypothesis: difference −0.43%, ranges overlapping. Seven factors have been ruled out; the cause is unmeasured. §03, §06 - 2026-10-04 ·
--parallel 2said to double the KV cache. Two passages claimed that raising--paralleldoubles or multiplies the KV cache. Board VRAM at load measured 16,050 MiB with--parallel 1and 16,184 MiB with--parallel 2— 134 MiB more (2026-08-23, UD-IQ4_XS,-c 32768,q8_0KV, drafter off,kv_unifiedoff as logged, build 10502, n=1 per setting). A second full 32,768-token cache at this page’s measured slope of 39,936 B per window token would be 1,248 MiB der. The cache is sized by-cand split between the slots; it is not allocated twice. §06, §14 - 2026-10-04 · 8.0 t/s set against the wrong comparator, spill figure and configuration. The glossary, recipe launcher and menu table paired 8.0 t/s with 2,364 MiB spilled. 2,364 MiB belongs to the text-only configuration at load (
-c 262144, drafter n-max 4 / p-min 0.75, projector off, 2026-08-23); the 8.0 t/s run loaded the BF16 projector as well and spilled 3,502 MiB at load (UD-IQ4_XS,-c 262144,q8_0KV, drafter n-max 4 / p-min 0.75, greedy, prompt filled to 90,886 tokens, 2026-08-23). The glossary and §05 set the 8.0 t/s against 50 t/s, the level of the shallow 2026-08-22 probes, while the 26% printed beside it is 8.0 against the 30.6 t/s the fully resident-c 131072window read at the same 90,886-token depth (same file, projector and drafter, greedy, 2026-08-23). §05’s 262,144 budget row attributed the 2,364 MiB spill and the 25,504 MiB (24.9 GiB) requirement to the n-max 10 coding drafter, and the launchers printed them beside their n-max 10 variants with no drafter setting named; both are n-max 4 / p-min 0.75 figures (the n-max 10 coding drafter costs 898 MiB more). §08 attributed the 8.0 t/s to the text-only drafter-off window (“fill that same configuration”), which measured 15.96 t/s filled to 218,233 tokens (2026-08-25), and printed the shallow probes of the 8.0 t/s configuration at 65–70 t/s where the 2026-08-23 sweep read 59.1–64.1 t/s. §02, §03, §05, §08 - 2026-10-04 · 12 GB called “a plain no, 416 MiB over”. The footer compared the 11,396 MiB board figure — which contains the measuring machine’s desktop — against 12,288 minus a 1,308 MiB desktop reserve (10,980), counting the desktop twice: 11,396 minus 10,980 is 416. Measured 2026-08-26 — peak board VRAM with the file loaded and answering four short requests, minus the idle board reading taken just before the load — UD-Q2_K_XL at
-c 32768(-ngl 99 --parallel 1 -fa on -ctk q8_0 -ctv q8_0 --load-mode none --spec-type none): 10,497 MiB of server memory. On a 12,288 MiB card that leaves 612 MiB with a light 1,179 MiB desktop, 122 MiB with a heavy 1,669 MiB one, and 5 MiB short of the 1,796 MiB worst-case reserve: borderline, decided by the reader’s own desktop. §08 - 2026-10-04 · start-up no-fit warning called a red herring. The page described
failed to fit params to free device memoryas a harmless line triggered by any explicit-nglvalue. Measured 2026-08-23 (UD-IQ4_XS, BF16 projector, n-max 4 / p-min 0.75,q8_0KV,-ngl 99on every load, greedy probes): the warning is present at-c 163840,196608,229376and262144and absent at-c 32768,65536,98304and131072. The-c 262144load that collapsed to 8.0 t/s at 90,886 tokens printed it; its fully resident-c 131072control did not. At-c 163840it printed while the server’s shared GPU memory at load was 446 MiB, in line with the resident loads — an early warning, not proof of a spill. The-c 122880n-max 10 configuration logged the same warning on every surviving log (2026-08-25) and did not at-c 98304. §11, §05, §03 - 2026-10-04 · vision controls described as never run, and three closed items listed as not measured. §12 stated that the withheld-image control “was never run” and that the cost of
--image-max-tokens 1024on the answer was “never measured”. Both were measured on 2026-08-25: a generated 1440p target, seven questions scored by string equality, greedy, n=1 per question — full budget 7 of 7, reduced budget 5 of 7 on UD-IQ4_XS and 4 of 7 on UD-Q2_K_XL, image withheld 0 of 7. The §15 register row printed the reduced-budget figures of UD-Q2_K_XL (4 of 7; fine type 2 of 5) without naming the file; the full-budget 7 of 7 holds for both files. §05’s budget row and §15’s register row printed the whole-prompt medians (3,635 and 1,043 tokens, 3.49×) as the image’s cost; the image itself is 3,602 and 1,010 tokens, 3.57×. The front matter listed power capping, system RAM under either load mode, and board power for files below 4 bits as not measured; all three are closed in §15’s register. §05, §12, §15 - 2026-10-04 · cliff chart labelled as measuring the cliff, and probe conditions unstated. The chart annotation read “the cliff, measured” and the lead-in called it “the cliff chart”, but every solid plotted point is a short temperature-0 wall-clock probe (2026-08-22, Q4_K_M, a reconstruction — the record does not name the file;
q8_0KV, drafter on, projector loaded; completion tokens divided by the stopwatch time of the whole request, prefill included, first request after each load) — 55.4 t/s at 122,880, 51.8 at 212,992, 41.2 at 217,088, 19.5 at 221,184. The dashed 256k point is an older, different run. §05 and the chart’s description said speed held at 50–55 / between 51 and 55 t/s through 212,992; the same sweep read 45.1 at 147,456 and 47.1 at 180,224, all above its 41.6 t/s pass line. A short probe cannot locate the collapse: on-c 262144(UD-IQ4_XS, projector loaded, greedy, 2026-08-23) a 1,531-token prompt read 42.9 t/s while the same window filled to 90,886 tokens read 8.0 t/s. The chart shows where an unfilled window stops being fast, not where a filled one collapses; the collapse table below it shows that. The caption also called every solid segment a probe; the solid run-in from the chart’s left edge to 122,880 has no probe under it (the 2026-08-22 sweep started at-c 122880), and the segment from 122,880 to 212,992 joins its end points without plotting the ten readings between them. Relabelled, not redrawn. §05 - 2026-10-04 · ceilings labelled “deep-fill measurements” were short probes. The §08 ceiling paragraph claimed “These are deep-fill measurements, not shallow probes”. The 163,840-token ceiling (vision, n-max 4, 996 MiB board-VRAM slack) was a 2026-08-23 sweep read at load with one short probe (200 generated tokens). The 180,224, 196,608 and 212,992 text-only windows (n-max 10, 847 / 415 / 280 MiB board-VRAM slack) and the 131,072 vision ceiling (n-max 10, 2,426 MiB board-VRAM slack) were read at load with one short request and no fill (2026-08-23); slack is 24,576 MiB minus board VRAM, so the measuring machine’s desktop is inside it. The 262,144 at-load figure (23,216 MiB, 1,360 MiB of board VRAM left) is derived. Only the 218,233-token measurement (2026-08-25, 755 MiB of board VRAM remaining) is a deep fill. §08, §15
- 2026-10-04 · MTP head called “always 8-bit”. The page claimed the draft head “always stays at 8-bit whatever the main quantization is”. Measured 2026-10-04 from the GGUF file headers: in UD-IQ4_XS the 15 head tensors are Q8_0, Q6_K and F32 (334.75 MiB); in bartowski’s IQ4_XS and IQ3_XXS quantizations of Swift-1.5, a fine-tune of this model with the same architecture, they are Q4_0 and F32 (227.91 MiB). The head’s precision is the quantizer’s choice, not a property of the format. §06
- 2026-10-04 · 262,144 budget row printed 179 MiB of margin against a retired desktop figure. The row and the
:pick_3launcher gave the 1,360 MiB left at load as “179 MiB once a desktop takes its measured worst case of 1,181 MiB”; C12 retired 1,181 on 2026-08-26, and against the resting worst case of 1,669 MiB the same 1,360 MiB is 309 MiB short. The launcher also printed the 25,504 MiB requirement as “~25.5 GiB”; 25,504 ÷ 1,024 = 24.9 GiB (25.5 is ÷ 1,000), the same slip §15 lists for three other budgets. §03, §05 - 2026-10-04 · an arm’s judged score said to move by “at most 1.0 point on the 0–100 scale”. The §09 callout measured that band on ALPACA, averaged over 25 items, between two sessions that judged the same packets, and applied it to MT-Bench as a stated assumption. The anchor’s 25 MT-Bench answers are byte-identical to August’s (171 of 171 comparable answers identical across the two sweeps), yet the same 25 items scored 79.7 in the August session and 87.0 in the 2026-10-04 session — 20 items rated higher, 1 lower, 7.3 points on the 0–100 scale. The band holds for a re-judge of the same packets; in a later session, with different answers beside them, the same answers moved well past it. The cause of the move is unmeasured. Judged scores are read only as paired comparisons inside one session, which every August comparison was. §09
- 2026-10-05 · the drafter’s verification step said to guarantee every configuration writes exactly what the base model would have written, and quality said to change only when the weights change. On one prompt (a JavaScript red-black tree, thinking off, greedy, 700 tokens,
-c 32768, build 10502, 2026-10-04) with the drafter off each file repeated its own text exactly across six runs. With the drafter on, the anchor’s text differed from its drafter-off text — at byte 21 at n10/p0.5 and at byte 269 at n4/p0.75 — and its n4/p0.75 runs produced two distinct texts in six; Swift IQ3_XXS likewise diverged, with three distinct texts at n4/p0.75 and two at n10/p0.5. Swift IQ4_XS produced byte-identical text with the drafter on and off across all three settings. The Swift comparison’s quality scores (§09) were measured with the drafter off; drafter-on quality is not measured there. §02, §06, §08 - 2026-10-05 · the 2.9-bit file’s 7.8% drafter-off lead called “exactly” what fewer gigabytes per token predicts. The size difference predicts 1.45× at a single constant (14.25 GB ÷ 9.83 GB) and 1.56× at §04’s format-specific constants (K-quant 0.70, IQ-quant 0.65); the measured ratio is 1.078× (45.66 against 42.34 t/s). In October the same size arithmetic predicted 0.917 and 1.175 of the anchor’s drafter-off speed for the two Swift files and measured 1.005 and 1.034 (greedy, drafter off, 1.5k of depth,
-c 131072). §08, §04, §06 - 2026-10-07 · the Swift chat template said to differ from the base model’s. It differs from the anchor’s: unsloth’s UD-IQ4_XS carries a 9,993-character template that accepts
highand renders it asxhigh. lmstudio-community’s Q4_K_M of the base model carries the same 8,952-character template as the Swift-1.5 files and bartowski’s Qwen3.8-27B-pi files, byte for byte, and that template raises onhigh(jinja2 render; headers read 2026-10-06/07). The difference belongs to the unsloth file, not to the base model. §14, §17, §15 - 2026-10-07 · “You cannot switch level per request.” What llama-server ignores is a
reasoning_effortfield in the request body (C20). A per-requestchat_template_kwargsis a different field: on build 10502 the server’s/apply-templaterendered thelowandxhigheffort lines from it over a launch setting ofmedium, and no line without it (2026-10-06, on the Pi fine-tune’s IQ4_XS, whose template is the Swift files’). Generation with it through/v1/chat/completionswas not tested, so relaunching remains the verified way to change level. §09, §14 - 2026-10-08 · “n-max 10 / p-min 0.5 is the speed pick wherever the VRAM is there to spend”; “Put n10/p0.5 on the text-only rows”; “THE SPEED PICK”. Those rested on greedy probes, mostly code with thinking off. At the cards’ own temperature-1.0 sampler with thinking on, the n/p sweep measured n10/p0.5 on UD-IQ4_XS at 0.961 of n4/p0.75 on shallow reasoning tokens and 1.049 on answer tokens, against n3/p0’s 1.111 and 1.120 (
xhigh,-c 81920, text only, two passes, paired intervals); after a 57,540-token prefix n10/p0.5 read higher, 1.079 against n3/p0’s 1.061, intervals overlapping, and nothing deeper was measured. Greedy code still favours longer drafts. The recipes for the files the sweep ran, and the October launcher, now run n3/p0 with one slot and n4/p0 with two slots and for the Pi fine-tune. §06, §03 - 2026-10-08 · n4/p0.75 as “a safe starting point, not the peak”, and “the card keeps n4/p0.75, the setting its window was filled at”. At n-max 4, p-min 0 read higher than 0.75 on shallow reasoning tokens on every file the n/p sweep ran: 1.049 (UD-IQ4_XS, interval 0.995–1.121, which includes 1), 1.110 (Swift IQ4_XS) and 1.034 (Swift IQ3_XXS, one slot) of n4/p0.75, about 1.15 on the Pi file (a ratio of two ratios, no interval) and 1.337 per slot with two slots busy (second pass only, one load per setting). After a 57,540-token prefix the gap shrank or vanished (UD-IQ4_XS 1.000, Swift IQ3_XXS 0.996), and on UD-IQ4_XS deep answers p-min 0.75 was faster (n4/p0 0.973 of it, two prompts, both blocks). p-min costs no memory (Pi file at n-max 4: 18,872 MiB at p-min 0, 18,871 at 0.75), and n-max 3 needs about 150 MiB less than 4, so a window filled at n4/p0.75 only gains margin. On 2026-10-09, for that reason, the commands for files, cards and slot counts the sweep did not run moved from n4/p0.75 to n4/p0, keeping n-max 4 der. §06, §03
- 2026-10-08 · “the Pi fine-tune’s launch command uses n-max 3”. It was chosen on 2026-10-07 from two settings at
xhighand-c 32768. At the command’s ownmedium, with the launcher’s arguments and sampler, n4/p0 decoded shallow reasoning tokens 1.038 of n3/p0 (paired 95% interval 1.000–1.092) and 1.073 after a 57,540-token prefix (range 1.041–1.116 over 3 prompts); a repeat block read 1.038 and 1.082, and n5/p0 and longer were slower. The command now uses n-max 4 / p-min 0, the n-max its window was deep-filled at. Appendix B, §06 - 2026-10-08 · “in Pi’s provider file set …
defaultLeveltomedium, so that Pi’s display matches the server”. Pi 1.0.2, the installed version, reads nothinking,defaultLevelorrequiresEffortfield: none appears in its installed code, which selects effort throughthinkingLevelMapand, withcompat.supportsReasoningEffortfalse, sends none to llama-server. The effort the model gets is the server’s--chat-template-kwargsalone. The §14 provider file shows the block as it was used; it is inert. Appendix B, §14
Retractions from the coding benchmark arm
Each of the following was measured, published, and then refuted by a better measurement — in every case one this campaign ran against its own work. They come from the paired benchmark arm (2026-08-26, 225 exercises, UD-IQ4_XS against UD-Q2_K_XL, identical recipe, same RTX 3090 at 350 W).
Four deliberate choices — reasoning off, the whole edit format, quantisation and a 32k context — all depress the absolute score, so it is not comparable with published leaderboards.
Checked against aider’s own published data, three of the four go the other way or do nothing. whole beat diff on every published Qwen3 pair by 2.2 to 4.5 points. For Qwen, aider recommends thinking off and its no-think arm scored 11.5 points higher. The 32k window shows no evidence of harm — 100% well-formed is the signature of a window that is not truncating. Only quantisation clearly costs. The conclusion survives, but for a simpler reason: no comparable published score exists at all.
The two files succeed on materially different tasks: 60 of 225 exercises went to exactly one of them.
The count is right; the attribution is not. Nothing here separates a quantisation effect from run-to-run variation, because the benchmark has never been run twice with the same weights on this rig. Since 69% and 80% respectively of passes arrive only on a second attempt, and decoding here is not bit-reproducible run to run, the baseline flip rate is certainly above zero and its size is unknown.
The flip rate has now been measured (2026-08-27, UD-IQ4_XS, aider polyglot benchmark, whole edit format, reasoning off, -c 32768, n-max 4 / p-min 0.75, greedy, 113 exercises paired by case and language). Both runs scored 43.4%, identical, while 18 of 113 exercises flipped verdict — a flip rate of 15.9% (95% interval 10.3–23.8), split exactly 9 pass→fail and 9 fail→pass. At that noise rate, comparing one run of each file across 225 exercises would produce about 36 disagreements (23–54) even if the files were identical in capability. The actual cross-file count is 60 of 225 = 26.7% (95% interval 21.3–32.8), 10.7 points above that floor. Whether that gap is real depends on which file’s noise it is tested against — a p-value below 0.05 would mean noise alone is an unlikely explanation. Against the 4-bit file’s floor of 15.9% the gap clears the line (z = 2.21, p = 0.027). That test was against the wrong floor, and the second one has since been measured. The UD-Q2_K_XL retest (2026-08-28, same benchmark, same flags, 113 exercises paired by case and language) flipped 25 of 113 = 22.1% (95% interval 15.5–30.6). Against that noisier floor the same 60 is unremarkable (z = 0.91, p = 0.364). Pooling both files’ retests — 43 flips out of the 226 exercises both retests covered, a combined floor of 19.0% (interval 14.4–24.6) — gives z = 1.93, p = 0.053, which does not clear the line either. No quantisation effect is demonstrated. The two floors do not themselves differ significantly (z = −1.19, p = 0.236), so “the 2-bit file is noisier” is a point estimate, not a finding. Most of the 60 — possibly all of it — is the benchmark talking to itself.C21
A 2-bit quantisation solved the same number of coding tasks as a 4-bit one.
True as arithmetic, misleading as a claim. Both finished 96 of 225 — but only 66 were solved by both, and 59 by exactly one, split 29 to 30. The equal totals conceal disagreement on a quarter of the benchmark.
The 2-bit file costs 1.52× the energy per solved task: 23.2 kJ against 15.3 kJ.
The first arm’s power sampler was started after its benchmark, so it has telemetry for 132 of 225 exercises while the second has 222. The cost table averaged each arm over whatever it happened to have — different exercises. Recomputed over only the exercises where both arms have power: 1.32×. Direction survived; magnitude overstated by about 15%.
Acceptance sits at 89% of the configured cap and stays 59% at the last permitted position, so the flag limits speculation; raising it should be worth about +24%.
Measured: −0.8% at n-max 6, +2.0% at n-max 8. Mean accepted length rose from 2.71 to 3.54 and throughput did not follow. Deeper drafts cost more per token and are accepted less often — 99/73/55/44%, decaying to 13% by position 7 — so wasted work grows faster than accepted work.
Energy per completion token varies 30× across exercises.
Measured correctly: 3.2×. The alignment window billed test-time idle to the model. The cheapest exercise was published at 0.417 J/token against a true 5.830.
A poll-derived phase split overstates the long phase by 6.3 points.
Measured over the full arm: −1.5 points — wrong in magnitude and in sign. The four-minute window was dominated by start-up transients.
Thermal throttling never fires on this part.
It fires on 3.0% of busy samples. The earlier reading counted idle samples, which are 14% of the trace and swamped the histogram. Idle is a state, not a limit.
With 59 discordant pairs, a difference of up to 1.3 points remains consistent.
The true interval is about five times wider. The rule of three is the bound for zero discordant pairs and does not apply once there are any. Understating uncertainty is the one direction this campaign must never round.
This section exists so that a reader who distrusts one number can find out exactly what produced it. It has four parts: which instrument read each class of number; where every external fact came from and how strongly it is sourced; a numbered list of the places this page corrected itself; and a register of what it never measured at all, so that nobody mistakes an omission for a result. It ends with one command you can run to check whether your own setup matches this machine.
Which instrument produced which number
Paths in this table are in the measured-inference repository: scripts/ from its root, and work/ and data/ from results/swift-1.5-qwen3.8-27b/; the Pi fine-tune study’s rows give their paths from the repository root.
| Class of number | Instrument | Notes and known biases |
|---|---|---|
| Decode and prefill speed | llama-server's own timings block — prompt eval time … / N tokens and eval time … / M tokens | Token-weighted. Never tokens ÷ wall-clock, which averages prefill into decode and reads 5–20× low at depth. A benchmark harness's own mean over requests is unweighted and read up to 3.8 t/s higher than the server's on these arms (42.2 against 38.4); the two are never mixed in one table here |
| Draft acceptance, mean draft length | the same timings block: draft_n and draft_n_accepted | Accepted per target pass is derived as draft_n_accepted ÷ (predicted_n − draft_n_accepted) |
| VRAM, server process | llama-server's own dedicated-VRAM report at load | This is what the budget model in §05 is fitted to, and it is stable to 127 MiB across 17 loads |
| VRAM, board total and spill | nvidia-smi --query-gpu=memory.used for dedicated; Windows performance counters \GPU Process Memory(*)\Shared Usage and \Dedicated Usage for the spill and per-process split | nvidia-smi on Windows cannot see the spill and shows [N/A] per process. The board VRAM total includes the desktop, which measured 1,179–1,669 MiB. The three counters and what each of them can and cannot see are in §02 |
| Board power and all energy figures | NVML through nvidia-smi --query-gpu=power.draw, logged at 1 Hz (the 08-22 effort arms) or 500 ms (the 08-23 sweep and matrix), integrated by scripts/power/attribute-power.py (trapezoid rule, edges linearly interpolated, gaps over 2 s excluded) | In-band tier — read from the card's own sensor, not from a meter at the wall. Inside the number: the graphics chip, its memory, board regulators and fans. Excluded and never measured: the power supply's conversion loss, the processor, system memory, drives, the display, and any facility overhead. The integrator's self-test passed 7 groups and 27 assertions on 2026-08-23. Every joule is joined to the server's own per-request timings, so prefill and decode are attributed separately |
| Open-ended writing quality (ALPACA, MT-Bench) | a blind panel of three Claude Opus 5 seats over the kept transcripts, standard 1–10 single-answer rubric, normalized (r−1)/9×100; scripts/bench/judge-panel.py | The most subjective instrument on this page, and the one with a named conflict. Qwen wrote the answers and Claude read them, so no model graded its own output — but the judge and this page's author are both Claude models, which is a correlated instrument, not an independent one. Mitigations, all of them partial: three seats rather than one; blinding by opaque salted identifier with the arm mapping sealed; a different shuffle seed per seat; seat spread published beside every mean (0.28–0.92 rating points); and arm-against-arm conclusions drawn from a paired bootstrap over the same prompts inside one judge session rather than from raw means, because byte-identical MT-Bench answers scored 7.3 points apart in two sessions (§09). Known residual bias: judges tend to reward length, and these seats were instructed against it and reported correcting for it — which is self-report, not proof. A second-vendor or human judge is an open item in §15 |
| Perplexity | llama-perplexity over the full wikitext-2-raw test set at -c 8192 -fa on -ngl 99 | 36 chunks × 8,192 = 294,912 scored positions; the corpus file was hash-verified identical across the cache-precision comparison. The count is tokenizer-bound (§08) |
| Benchmark scores, 2026-08-23 and 2026-10-04 | measured-inference/scripts/bench, suite hash 1cdf54f8eb9d3f8f, 175 prompts, greedy, seed 42 | Three grader bugs were found and fixed symmetrically during this run — applied to prediction and reference alike, so a fix can only make one value written two ways compare equal, never make two different values match. All arms were re-graded offline from kept transcripts with the final grader; the self-test stayed at 78 passed, 0 failed. The 2026-10-04 suite ran scripts/bench/bench.py on the same suite hash with the same build, and the anchor reproduced its August scores on the five automatically graded sets exactly (§09) |
| Benchmark scores, 2026-08-21 and 2026-08-22 | chinkeong/benchmark — final-number / boxed-answer exact match, truncations counted wrong | A different grader from the one above, and unaudited against the three bugs it found. Every score from these two days carries that caveat: §08's 20-question table and its GSM8K n=200 comparison |
| Whether the emitted code runs | measured-inference/scripts/bench/execute-probe.py — each file's probe-A output rebuilt with the prompt's own prefix and executed under node v24.15.0, 15 s timeout, greedy (temperature 0, top_k 1), n=1 per file | One task, one language, one sample per file — a threshold, not a rate, and it must not be read as a pass rate. Probe A is a continuation task, so an output beginning mid-expression is correct rather than truncated; the prefix is prepended per output according to whether the file continued the program, restarted it, or wrapped it in the markdown fences the prompt forbade. The demo graph is model-generated, so there is no fixed correct answer to check — the test is that the program parses, runs to completion and prints the path and cost it promised. Zero GPU time: it reads files already on disk. Added 2026-08-25 after the judge panel and this probe both found degeneration that the four lexical detectors (§02) cannot see — a missing closing brace contains no repeated n‑grams, so no threshold on them could ever catch it |
| Image token counts | the server's own prompt_n differential with and without the image | Confirmed the arithmetic law to within 2 tokens at two resolutions |
| Clock, temperature and power state | nvidia-smi --query-gpu=clocks.sm,clocks.mem,temperature.gpu,pstate,utilization.gpu | Present in the 08-23 logs and absent from the 08-22 ones, which is why the earlier logs cannot prove whether a low power sample was a ramping board or an efficient one |
| Tokenizer identity across files | work/ppl.py tokenizer gate | Fails closed: rc 0 SHARED TOKENIZER (the corpus token count, a hash of the corpus token ids and the ids of a short probe word all agree across files), rc 1 NOT SHOWN IDENTICAL (all three produced, and they differ), rc 2 ERROR (the queue stops). Every file must produce all three values, so a tokenizer that never ran cannot pass as shared. Corpus 1,290,590 bytes. This sweep: rc 0, SHARED TOKENIZER (2026-10-04) — raw perplexity comparison across these files is legal |
| Speed ratios (Swift file against anchor) | The anchor-ratio method (data/arms/speed-anchor-check.json) | Each Swift speed is read against the anchor in the same sweep, same probe group and depth, pass by pass; a group whose two passes disagree by more than 3% of their mean is void. Separately, the anchor’s own speed is read against its August value: on this sweep the four readings with an August value sat at 0.910–0.928 of it (answer regime at 1.5k 0.917, 28k 0.910, 91k 0.928; reasoning regime at 91k 0.915; the drafter-off group has no August value), the lower of the two levels this machine shows (§03); the Swift cards in §03 print each answer-regime speed also on the higher level, derived: the October figure divided by the anchor’s ratio at the same depth. Groups: answer regime OK, drafter off OK, reasoning regime void — its first pass overlapped downloads on the test host and its second ran after they stopped, so that disagreement is not clean evidence about the models; a re-run on 2026-10-05 failed the same check in five of its six Swift cells (data/arms/speed-anchor-check-thinkdepth-rerun.json, the register below). The drafter-off group’s first pass also ran inside those downloads (12:53–12:58, 2026-10-04) and agreed with its second within the 3% |
| Window safety (fit, spill, knee) | work/g5-gate.py on the deep fills and the knee sweep | Fit = max dedicated depth over the step’s loads plus desktop max plus desktop sd ≤ board MiB. Spill = shared load > 100 MiB over the next-lower step. Knee = decode ≥ 15% below the next-lower step, confirmed by a second low load at the same window or a confirmed failure above it. Desktop allowance from 73 desktop readings (72 loads and one direct reading): max 855 MiB, mean 358.2 MiB, sd 121.0 MiB (sample, n−1). Fit limit 23,600.0 MiB dedicated depth. Comparator: the nearest lower step within 32,768 tokens measured before this one in ledger order, else the first measured after it (every compared step here had one measured before). A step with no lower step within 32,768 tokens has spill and knee untested and is a baseline if it fits. A vision step is also judged against its text twin at the same window, by the same spill and decode rules. Deep fills: 1 baseline and 2 passes. Knee sweep: 64 steps over its 69 loads, 7 baselines, 22 passes and 35 failures. In five picks each failure is a step above the window the gate allows; pick 7 (vision at one slot) fails at both of its windows, so the gate allows it none; picks 3 and 5 have no failure and pass at the full 262,144 tokens |
| Drafter-regime invariance (greedy text) | work/d7-check.py | Byte identity of greedy thinking-off bodies across drafter regimes on the one drafter prompt, matched by (rep, probe_index, probe); last ledger line per key. Identical for Swift IQ4_XS; different for the anchor and Swift IQ3_XXS. On one prompt that is evidence, not proof, that drafter-off scores carry over to drafter-on use for Swift IQ4_XS; the anchor’s and Swift IQ3_XXS’s drafter-on runs did not always repeat their own text, so for them the check cannot say either way (§08) |
| Repetition and loop detection | scripts/bench/loop-detect.py through work/loop-scan.py | Four signals at WARN and STRONG thresholds: n1_loop_frac (0.1 / 0.2), n2_skeleton (0.32 / 0.45), n3_compress (0.3 / 0.26), n4_worst_ttr (0.45 / 0.3). MIN_WORDS 15; texts under that are NO-DATA. Per-field verdict on each text; an item is flagged if any field is. The flag totals are not a loop count and are not published as one: on 24 transcripts read (12 per arm, chosen from every truncation and the longest outputs), 3 of 22 flagged items were degenerate loops (11 of 22 counting circular) and no degenerate item was missed; every false positive carries the vocabulary-collapse signal (n4_worst_ttr) in code or arithmetic. The thesis in §09 cites the spot-read classes instead |
| Answer and reasoning token counts | llama-tokenize through rule21-sweep.py split, exact counts via the GGUF’s own vocab (--no-escape, --ids, --show-count) | Answer is exact (the GGUF’s tokenizer); reasoning_direct is measured; reasoning_derived (tokens minus answer) is DERIVED and includes 2–4 template tokens. 696 texts, 667 counted in this run, 4 concurrent jobs, 96.9 s wall, per-call preflight 0.623 s. Zero tokenizer failures, no unpublishable datasets. The tokenizer is identical across the three files (header gate), so one GGUF serves all arms. Resumable via sha256-keyed cache; a preflight pair (literal backslash-n against real newline, 6 ids each) confirms the escape mode before any text is counted |
| Board power and throttle state (October sweep) | NVML through nvidia-smi, logged at 500 ms; scripts/power/sample-power.ps1 for power, detached nvidia-smi for throttle | One power CSV (from 2026-10-04, 500 ms) and two throttle CSVs (09:38 to 17:21 and from 17:35, both 500 ms). Throttle state is unlogged from 17:21:28 to 17:35:37, which covers about the first five minutes of the suite re-run; power was logged throughout. Throttle query: clocks_event_reasons.active (the throttle mask), clocks.sm, clocks.mem, utilisation, temperature, power.draw. Known bias: the first samples after an idle board read low because the SM clock is ramping (~900–990 MHz against 1,455 settled); warm the GPU with a throwaway request before any arm to be published |
| Host CPU, memory and I/O during the sweep | Windows performance counters, data/telemetry/swift-campaign-host.csv | Columns: processor, privileged and user time (%); context switches and system calls per second; run-queue length; available memory; page faults and pages read in per second; committed bytes; disk throughput and queue length; interrupts per second. One row about every 3 s from 06:58 on 2026-10-04, unbroken through the sweep: 14,363 rows between 09:38 and 21:47, no gap over 6 s. The logger kept running after the sweep, so the file holds more rows than these |
| File identity and chat-template correctness | work/header-diff.py (GGUF header parse), jinja2 template byte comparison; Appendix A: hf_headers.py (HTTP Range), tensor_cmp.py (sha256 of the first 1 MiB of sampled tensors), repo_diff.py, q8_hdr.py, make_table.py | Suite control: 175 prompts, all rendered identically for the anchor and Swift IQ4_XS (0 different, 0 render errors). Effort probe: default equals xhigh on both files; Swift IQ4_XS rejects reasoning_effort high with a template exception (the Swift template accepts only xhigh, medium and low; the anchor’s renders high the same as xhigh). The comparison uses jinja2, not the server’s own minja engine — it confirms that the two templates produce the same bytes on identical input, not that the server will. For the effort lines it was later checked on the server: on 2026-10-06, build 10502’s /apply-template rendered the same 8,952-character template (in the Pi fine-tune’s IQ4_XS) with no line at medium and the low and xhigh lines on request; high was not sent to the server. Appendix A’s tools read file headers and the first 1 MiB of sampled tensors over HTTP Range, not whole files |
| Pi fine-tune: file identity | results/qwen-pi-study/work/headers.py (GGUF header parse and jinja2 template render) | Compares tensor names, shapes and types, tokenizer arrays and the chat template with the Swift twins and with lmstudio-community’s base Q4_K_M; not weight values. Only the IQ4_XS and IQ3_XXS files were read |
| Pi fine-tune: server checks and Pi smoke test | results/qwen-pi-study/work/pi-stage0.py; results/qwen-pi-study/work/pi-agent.py, which runs the 2026-10-04 agent-probe.py unedited | Effort lines read from the server’s /apply-template, not from Pi’s rendered prompts; one request or probe per check, n=1 |
| Pi fine-tune: suite comparison | results/qwen-pi-study/work/pi-suite.py (the 2026-10-04 suite runner rule21-sweep.py, imported unedited with its paths repointed) and results/qwen-pi-study/work/a1-compare.py | Pairs each Pi item with the same item of the anchor’s and Swift IQ4_XS’s 2026-10-04 cells, which are reused, not re-run; composite differences by paired bootstrap (10,000 resamples, seed 42); token totals, median per-item ratios and fewer-token counts over all 175 items, ALPACA and MT-Bench included as tokens only (they were not judged) |
| Pi fine-tune: speed against Swift IQ4_XS, and the drafter choice | results/qwen-pi-study/work/a2-speed.py over the arm logs in results/qwen-pi-study/data/arms/ | Two passes in alternating order, the first probe after each prefill discarded; a ratio is void when its two passes differ by more than 3%. The sampled follow-up fixes each probe’s seed, so its passes repeat each probe and its comparison is across 8 probes, not across passes |
| Pi fine-tune: window at depth | scripts/bench/window-knee.py’s measure(), imported unedited by results/qwen-pi-study/work/run-phase-a23.py | Order Pi, Swift, Swift, Pi per configuration, n=2 per file; the desktops were not matched across loads (500–1,247 MiB); one window per configuration, so no knee can be located |
| n/p drafter sweep: plan and driver | results/qwen-pi-study/work/np-make-plan.py (writes results/qwen-pi-study/work/np-plan.json), results/qwen-pi-study/work/np-sweep.py; the feasibility check results/qwen-pi-study/work/nk-probe-test.py | One fresh server load per setting, because llama-server ignores per-request n-max and p-min (checked on one load by nk-probe-test.py); each pick’s launcher arguments plus --cache-ram 0 --metrics, -c 81920 at one slot and 123,904 × 2 with the projector at two; a gate before every load (commit room, a quiet card, no other server) and, from 2026-10-08 00:24, at least 180 s without keyboard or mouse input; in the one-slot blocks, reference loads at the start, middle and end of each pass; pass 2 in the reverse order of pass 1, with different seeds |
| n/p drafter sweep: analysis and facts | results/qwen-pi-study/work/np-analyze.py, results/qwen-pi-study/work/np-facts.py (writes results/qwen-pi-study/data/np-sweep/np-facts.json) | Rate = Σ generated tokens ÷ Σ decode time per setting, pass, token regime and depth, divided by its own session’s reference (a pass splits into sessions at gaps over 30 min). 95% intervals by a paired bootstrap over prompts within a session (B 2,000, seed 42), on 3 and 2 prompts after the prefix. Survival per guess position from llama-server’s /metrics counter spec_decode_num_accepted_tokens_per_pos_total; server VRAM = board peak minus the desktop read before load, median of loads. Known biases: the passes use different seeds by design, so the pre-registered 3% pass rule flags text-path differences as level changes; 13 disturbed loads were excluded and re-run; the two-slot block has one load per setting and pass (§06) |
| n/p semantics | results/qwen-pi-study/work/np-design-review/sem.json | llama.cpp source read at this build’s commit, 0adcc3bb5, from raw files at that commit rather than from master; these are source semantics, not measurements, except where the page cites a check |
External sources, graded
Each bullet is marked [P] primary document, [A] arithmetic on published specifications, [C] carried from a dated prior fact-check on this page, or [S] secondary or user-reported.
- [P] §13 task boundaries — long-tail knowledge: Kandpal et al., ICML 2023, PopQA / Mallen et al., ACL 2023, knowledge-capacity scaling laws, Why Language Models Hallucinate · long-horizon reasoning: Faith and Fate (NeurIPS 2023), emergent-abilities mirage (NeurIPS 2023), METR time horizons, agent half-life (Ord, 2025) · long context: RULER (COLM 2024), Lost in the Middle (TACL 2024), NoLiMa (ICML 2025) · constraint stacking: FollowBench (ACL 2024), IFScale, ComplexBench (NeurIPS 2024) · agentic reliability: METR blog, self-conditioning in long-horizon execution
- [P] Arc Pro B70 — 32 GB, 256-bit, 608 GB/s: Intel datasheet, corroborated by Tom's Hardware and Puget Systems
- [P] RTX 50-series bandwidth (5090 1792, 5080 960, 5060 Ti 448 GB/s) — GamersNexus, TechSpot; the RTX 3090's 936 GB/s is NVIDIA's own GA102 whitepaper figure (384-bit × 19.5 Gbps)
- [P] DGX Spark GB10: 128 GB unified LPDDR5X, 273 GB/s — Tom's Hardware review, LMSYS
- [P] Arc B390-class iGPU (Core Ultra): dual-channel LPDDR5X-9600 mandate, 153.6 GB/s, ≤96 GB shared, shipping in Core Ultra X9 388H laptops — Intel ARK, TechPowerUp, Tom's Hardware, HotHardware, shipping laptop list
- [C] NVFP4 on Blackwell: vLLM-only, about 1.5× BF16, 92–97% accuracy, FP8 KV — Unsloth documentation, NVIDIA forums, Kaitchup. Vendor claims, unmeasured here
- [C] Quantization quality claims (unsloth Dynamic v3, Qwen's AD-IQ3_S, the lmstudio Q6_K comparison) — HF discussion #65, kingy.ai
- [C] The
xhighoverthinking anecdote (22,276 thinking tokens, 21 minutes, for one simple drawing) — Simon Willison - [P] Intel software stack: IPEX-LLM archived January 2026 — the repository itself; Vulkan beats SYCL on Battlemage — llama.cpp #22413, B70 benchmark notes, Phoronix
- [P] Model card, OpenVINO builds, drafting support — Qwen/Qwen3.8-27B, int4-ov, int8-ov, qwen38-mtp recipe and benchmarks, llama.cpp discussion 27164, NVFP4-MTP-GGUF card (whose Blackwell-only claim this page's Ampere measurements contradict)
- [P] MTP drafting semantics in llama.cpp at commit
0adcc3bb5(build 10502):common/speculative.cpp(the draft loop, the p-min check and the n-max stop),common/common.h(defaults n-max 3, p-min 0.00),common/sampling.cpp(sample-and-match verification),tools/server/server-schema.cpp(per-request speculative fields disabled),src/llama-memory-recurrent.cpp(one rollback row per n-max per slot); PR 22838, PR 22673 and PR 23287. Read from raw files at that commit, not from master (§06) - [P] DFlash2 block drafting and path selection — inco.ai announcement, llama.cpp PR 27342, drafter GGUFs
- [A] Every non-3090 row of §07: bandwidth ÷ file size × the format constant (0.70 K-quant, 0.65 IQ-quant), using the bandwidth figures above
- [P] Swift-1.5 GGUF file provenance, recipes, imatrix chain and chat-template behaviour — ukisai/Swift-1.5-Qwen3.8-27B-GGUF, bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF, ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF (ukisai’s 8–12 GB set). Cited: bartowski’s imatrix (calibration-v6, 583 chunks of 512 tokens); ukisai’s card stating that its tiers reuse the importance matrix and per-tensor layouts bartowski computed for Swift-1.0; the Swift-1.5 template’s restriction to xhigh, medium and low effort, from the base repository ukisai/Swift-1.5-Qwen3.8-27b. Read from every file’s header here rather than cited:
general.samplingtemp 1.0 / top_p 0.95 / top_k 20. ukisai’s cards: mean KLD against Swift-1.5 BF16 — standard IQ2_XXS 0.2866 (wikitext-2 test, 100 × 512 tokens), GSQ-RCO IQ2_XS 0.189979 (development set; held-out C4 0.161594), GSQ-RCO IQ3_XXS 0.097774 (development set; for comparison, standard IQ2_M 0.1523, Q2_K 0.1655). GSQ-RCO ships no vision projector; its MTP files “require a runtime with support for this model’s MTP implementation” (cited, not reproduced) - [P] Qwen3.8-27B-pi provenance and its author’s claims — bytkim/Qwen3.8-27B-pi (model card, Apache-2.0), bytkim/Qwen3.8-27B-pi-GGUF, the author’s blog post, and bartowski/bytkim_Qwen3.8-27B-pi-GGUF (the files measured in Appendix B). Cited, not reproduced: a llama.cpp quickstart at
--spec-draft-n-max 3that sets no effort level, so the template’s defaultxhighapplies; Terminal-Bench 2.1, GPQA Diamond and SciCode results, including “Pi’s medium setting matches Base’s xhigh completion rate with about 41% fewer output tokens” (“in these selected results”); that all of the author’s Pi evaluations used FP8 weights; and a GGUF agent smoke test of eight quants atmediumwith a 16,384-token output cap on subsets, which the author calls a smoke test, not a quant ranking. The research notes written before this page’s measurements read the Terminal-Bench 2.1 scores as best observed across at most 3 attempts, on a task set also used to select the training checkpoint, with every Pi-against-Base gap within about one standard error; the 41% comparesmediumwith the base model’sxhigh, and at the same level the author’s own figures give about 20%; and they read the GGUF smoke test as 15 to 20 points below FP8 - [S] Card VRAM totals for non-3090 boards — nvidia-smi readings pasted by users in GitHub issues: 5090 (32,607 MiB), 5080 (16,303), 5070 (12,227). Appendix A’s 16 GB column uses the lowest cited total; the 4080, 4070 Ti SUPER, 5060 Ti 16 GB and 4060 Ti 16 GB read 16,311–16,380 (cited or assumed); the 4090’s 24,564 is assumed. The RTX 3060’s 12,288 MiB is assumed, from a user statement
This page's own measurement runs
- 2026-08-21 — llama.cpp build 5ecbe1ac1 (PR 27342), CUDA 12.8, RTX 3090: the drafting comparison (baseline / built-in head / DFlash2 / NVFP4 on Ampere) and the 18-configuration n-max and p-min sweep, all temperature-0 600-token code generation, plus llama-bench prompt-processing and generation figures. Also the FP4-against-INT4 scored comparison: GSM8K and MATH-500, 20 each, greedy, 4,096-token budget, llama-server build 10502, graded with chinkeong/benchmark
- 2026-08-22 — RTX 3090: the context-ceiling probe (stepped
-cwith short temperature-0 probes), the reasoning-effort comparison at temperature 1.0 with blind judging, the acceptance demonstration at n-max 10 / p-min 0.5, the wikitext-2 perplexity table over six files, the GSM8K n=200 runs per effort level and per file, the 10-configuration drafting re-sweep, the two half-speed failure modes, and the vision demonstration. Reference recipe: Q4_K_M,q8_0KV, drafter on, projector loaded - 2026-08-23, independent re-measurement round (00:19–02:19, ~2 h GPU) — the source of most corrections dated 08-23 in §04–§12: the drafter's VRAM bill, the two-constant VRAM model, the measured window slope, the deep-fill collapse, the answer-against-reasoning token regimes, depth on the shipped recipe, the vision token law, per-effort energy, appetite as a distribution, chat-template forensics, the perplexity position count, the PowerShell image-POST trap, and the bandwidth citation fixes. Same RTX 3090, driver 596.36, Windows 11 Pro 26200, i5-13600KF, 31.8 GiB DDR4-3200 dual-channel; llama.cpp build 10502 (version 0.1.2-dev, commit
0adcc3bb5, Clang 20.1.8, binary dated 2026-08-19); Qwen3.8-27B-UD-IQ4_XS.gguf (14,252,845,984 B) withmmproj-Qwen3.8-27B-BF16.gguf;--load-mode mmap, port 1235,--parallel 1,q8_0KV,-ngl 99. Run without sight of this page, so where its numbers differ from an 08-21 or 08-22 row, the difference is either a genuine condition change or a correction, and each is labelled as one or the other where it lands - 2026-08-23, follow-up round (02:48–04:29, ~1.7 h GPU) — same machine and build. Four measurements: (1) the matched drafter-flag sweep — 7 configurations × 2 files × 2 probes,
-c 32768, thinking off, temperature 0 / top-k 1, 700 tokens, fresh server per configuration, on a novel 149-token rate-limiter prompt chosen to avoid the acceptance inflation a textbook algorithm causes (§06). (2) the projector-at-depth pair — byte-identical 90,862-token prompts, loads ordered A-B-B-A, post-prefill probe discarded, 45 s cooldown, 5 probes per load with the drafter off and 4 with it on, plus the cooled depth ladder and the thinking-on-against-off isolation that produced the mean-draft-length result. (3)llama-perplexityat-ctk q4_0 -ctv q4_0, flags identical to the campaign's cache series, on a corpus file hash-verified identical to the fp16 andq8_0baselines'. (4) a no-GPU repetition audit of 10 long greedy transcripts: 10 of 10 clean, the single detector flag a false positive, and the one incomplete file cut by its 65,536-token budget rather than looping - 2026-08-23, effort benchmark sweep (04:40–13:19, 8.48 h GPU) — the evidence behind §09 and §10. Suite hash
1cdf54f8eb9d3f8f, 175 prompts (7 sets × 25) × 3 effort arms = 525 generations, greedy at temperature 0 / top-k 1 / seed 42,--max-tokens 16384with a rule-driven re-run of only the truncating datasets at 32,768 and-c 65536, no speculative decoding,q8_0KV. ALPACA and MT-Bench ran unscored at the time — transcripts kept, and scored on 2026-08-24 by the judge panel below. Includes the determinism check (139 of 139 byte-identical) and the three grader fixes - 2026-08-24, judge panel (zero GPU) — the evidence behind §09. The kept ALPACA and MT-Bench transcripts from the three effort arms, 150 answers, read blind by three independent Claude Opus 5 seats for 450 ratings with none missing or partial. Opaque salted identifiers; the identifier-to-arm mapping sealed in a key file no seat reads; a separate shuffle seed per seat; the standard 1–10 single-answer rubric; ratings normalized (r−1)/9×100. Arm comparison by 20,000-resample paired bootstrap over the per-prompt differences, seed 42, six comparisons. Builder, scorer and paired test:
scripts/bench/judge-panel.py. Instrument limit, stated with the result: the answers are Qwen's and the judge is Claude, so no model grades its own output — but the judge and this page's author are both Claude models, a correlated instrument, which is why three seats run and the seat spread is published - 2026-08-23, energy joins (zero GPU) — two analyses computed entirely from logs already on disk, using
scripts/power/attribute-power.py. The first joined a 500 ms power log (17,715 samples, 10:48–13:19) to the sweep's per-request timings, counting only the 115 of 150 requests whose whole window lay inside the log, and produced the drafter-off reference of 7.884 ± 0.307 J per decode token plus the sustained-power drift anomaly. The second re-integrated the 08-22 effort arms' own 1 Hz logs and reproduced all four published Wh figures to within 0.05%, adding joules per token, tokens per kWh and energy-delay product per level - 2026-08-23, power matrix (13:21:49–14:05:20, 43.5 min of logged span, 19 arms, 100% sample coverage) — thirteen of those arms carry full per-arm energy tables (the drafter ladder, three quantizations, two cache precisions, both token regimes, three depths); the other six are the two idle arms, the two concurrency arms, and the two power-cap arms that were skipped. The source of every figure in §03 and §03: settled idle with and without a server, the three-point drafter ladder, three quantizations, two cache precisions, both token regimes, three depths, and the concurrency pair. In-band NVML board power at 500 ms, three requests per arm, 700 generated tokens each. The two power-cap arms were skipped for lack of an elevated shell; the runner's own net columns used a 31.0 W idle constant rather than the measured 34.1 W loaded idle, a uniform shift of about 1% that changes no ranking, and the net columns published here are re-derived at 34.1 W
- 2026-08-23 to 08-24, quantization ladder (14:13 → 12:35, ~6 h GPU) — complete. Nine files of this model ranked on perplexity over 294,912 scored token positions each, under one identical condition set, with functional detectors beside every rung. Its gate re-ran the UD-IQ4_XS reference file and reproduced 6.5956 bit-identically twice at 0.000% drift before starting, which is what pins §08's error bars, and a same-size pair resolved a 4.27% quality gap against ±0.045 error bars — so the rig both reproduces and still discriminates. It also recorded that the same corpus tokenizes to 297,193 tokens under Qwen's tokenizer and 295,216 under Gemma's, which is the source of the tokenizer caveat. The full treatment is now in §08: the eight-rung table, the point where the perplexity curve turns, the accuracy ladder, the empty-answer audit and the paired tests. This entry's earlier "not yet on this page" caveat is retired by the two rounds below
- 2026-08-24 into 08-25, accuracy ladder over the same files — eight arms run across the day; most took under thirty minutes each (1,705 s and 1,766 s on two of them) and the bottom rung took 8,965.8 s — two and a half hours, more than four times the next-longest arm, because at 1.835 bits per weight the model stops terminating. The evidence behind §08's Mean, empty and truncation columns and behind every paired McNemar row. Frozen suite
1cdf54f8eb9d3f8f— GSM8K + HumanEval + MBPP, 25 questions each, 75 items per file — greedy at temperature 0 / seed 42,--max-tokens 16384,-c 32768,q8_0KV,reasoning_effort=low, no speculative decoding, eight files scored on the identical items so that every arm-against-arm claim is paired. Two instrument failures are recorded rather than quietly fixed: a grader crashed on a bare####with an empty tail, a shape no file above 2.15 bits had ever produced (fixed to grade it wrong; no recorded score could change, because any earlier run reaching that path would have crashed and none did); and an empty-answer count was first published from the truncation counter and was wrong — UD-IQ1_M reported zero truncations while returning nothing to five questions. Both columns are now read from the saved artifacts bywork/ladder-repcheck.py, which is why they are separate columns. A cross-model arm on the identical suite — gemma-4-12B-it-QAT-Q4_0 at 6.497 GiB — ran the same day under disclosed condition asymmetry - 2026-08-25, the four rounds that turned the ladder into a recommendation — (1) the drafter on/off pair on both candidate files: 700-token novel-code generations,
-c 32768,q8_0KV, thinking off, temperature 0, three settled probes per arm with the first post-prefill probe discarded, which found the ranking inverting between drafter-off and the shipped drafter-on recipe. (2) the full-native-window trial: three arms at-c 262144, each loaded, VRAM read as a drafter on/off pair, then deep-filled to 218,233 real tokens — 83% of the window, 424 s of prefill at 514.7 t/s — and probed with the first post-prefill probe discarded. A self-caught instrument error belongs with it: the probe script averagedprompt_nacross settled probes, which read 4 because probes 2 and 3 hit the server's prompt cache; the real fill is in the server's ownprompt eval timeline and the decode figures are at depth as published. (3) the requirement sweep behind §08's memory-and-speed table and §08's depth series: three files — UD-IQ4_XS, UD-Q3_K_XL, UD-Q2_K_XL — × three windows (-c 32768,65536,131072) × the drafter on and off, plus a fourth window at-c 98304drafter-on, each window deep-filled to about 90% of itself with wikitext-2-raw prose, board VRAM read at that depth, first post-prefill probe discarded and two settled probes taken. At-c 131072every file was loaded twice from scratch with five settled probes per load, and at-c 98304UD-Q3_K_XL was loaded twice, which is what settled the ordering after it had been published, withdrawn and reinstated (the episode is recorded at §08 rather than only in a log). The condition on that round is stated wherever its numbers appear: the fill is prose, the drafter's confidence gate is content-dependent, and the ordering is therefore published as a measurement on prose rather than as a property of the files. (4) the execute probe (§08), which used no GPU at all: every file's probe-A JavaScript, already on disk, rebuilt with the prompt's own prefix and run undernode v24.15.0at a 15-second timeout. Seven of eleven programs run — the eleventh, sdkyuan’s QAT-Q2_0, was added to the probe on 2026-08-26; the three files below 2.481 bits per weight do not run, and neither does the cross-model gemma arm at 4.651. It moved a published number — the “functional floor” had been read off four detectors that never executed anything — and it is the cheapest round in the whole campaign. The same day's matched--parallelpair closed the concurrency question (§03) - 2026-08-26, the coding benchmark on both candidate files (two full arms) — the evidence behind §08. aider’s official polyglot benchmark, 225 programming exercises each carrying its own unit tests and each given two attempts, run unmodified in its own container against a local llama-server on the same RTX 3090 at its stock 350 W board limit:
wholeedit format, reasoning off,-c 32768, greedy at temperature 0, one arm per file. UD-IQ4_XS and UD-Q2_K_XL both scored 96 of 225. Board power was logged for part of the run, which is why the energy comparison is paired over the 36 exercises that both files solved and that fall inside the logged window, rather than over all 66 that both files solved. (96 is what each file solved on its own, and no paired comparison can be taken over it: the two files did not solve the same 96.) Each arm ran once, so this page’s statement about the 60 exercises the two files disagree on is an upper bound, and the benchmark’s own run-to-run flip rate was an open gap until it was measured on 2026-08-28 at 15.9% (18 of 113 exercises, 95% CI 10.3–23.8) on the UD-IQ4_XS arm (§15) - 2026-08-28, the two-level finding and the reverse-order sweep — (1) twenty-three sessions of one identical configuration (UD-IQ4_XS, n10/p0.5,
-c 32768,-ctk/-ctv q8_0,-ngl 99,--parallel 1,-fa on, greedy, 700 tokens, fresh server per session, one warmup discarded, RTX 3090 at 350 W, build 10502), nine measured earlier in the day (three runs of twenty probes each plus six runs of six probes each) and fourteen measured in the paired test and subsequent sessions. The finding: 20 sessions at 75.71–78.65 t/s and 3 at 88.18–88.76, two distinct levels about 13% apart with nothing between. (2) the prior-load paired test: four alternating COLD/HOT pairs (COLD = 150 s idle, HOT = 180 s sustained burn on a different prompt), five probes each, arms alternated so drift falls on both equally. Result: COLD mean 80.48, HOT mean 80.13, difference −0.43%, ranges overlapping. Prior load refuted as the cause. (3) the five-arm drafter sweep in reverse order: same arms as the forward sweep, reversed, to test whether arm position was confounded with arm setting. Arm A read 76.32 forward and 67.41 reverse (−11.7%); all three varied arms held their sign and roughly their size against A (B −8.9%/−8.2%, C −3.1%/−4.7%, E +8.4%/+13.2%). Ratios within a sweep are stable; absolute numbers across sweeps are not. (4) the saturating power-cap re-measurement (SwPowerCap residency 99–100%): 300 W −4.1% t/s, −13.1% W, −9.4% J; 250 W −14.6%, −27.1%, −14.7%. Capping to 300 W is a better deal than the non-saturating sweep claimed - 2026-10-03 to 10-04, knee sweep (69 loads, 32,431.9 s of fill) — the window-and-VRAM data behind Appendix A and the Swift IQ3_XXS cards in §03. Eight picks across two files (the anchor and Swift IQ3_XXS), each loaded at stepped context windows with deep fills, VRAM and decode speed read at depth. The window gate (
work/g5-gate.py) judged its 64 steps (Appendix A, §05) - 2026-10-04, the Swift-1.5 comparison sweep (09:38–21:47) — the evidence behind the Swift rows in §03, §05, §08, §09, §12 and §14, and of Appendix A’s two deep-filled cells. Same RTX 3090 and driver; llama.cpp build 10502 (commit
0adcc3bb5), the same binary as August. In order, with each step’s wall: machine detection, the tokenizer gate and a loaded-idle reading (434.4 s for the idle reading); appetite atxhigh(2,074.1 s, then 3,494.6 s for its second pass) and at low and medium (1,878.7 s); the speed probes (1,261.6 s, then 4,029.2 s); the drafter regimes (2,629.8 s); the Swift IQ4_XS deep fills (777.5 s); perplexity (803.7 s) and theq8_0-cache perplexity (262.2 s); the 175-prompt suite for the anchor and Swift IQ4_XS (6,691.6 s, then 13,392.4 s); the answer/reasoning split (98.5 s); the Pi agent probe (656.9 s); vision on both Swift files (171.4 s and 187.4 s). Four steps needed a second run, and each second run completed: this host’s memory-commit guard refused part of the appetite and speed runs and stopped the suite at a cell boundary, so those second runs did only the work the first runs had not done; the agent probe’s first run sent no API key, so it ran again in full. The harness self-test reported one failing check both times it ran: its runner started that check’s script under WSL’sbash, which could not read the Windows path, so the check itself never ran. The needle tests did not run: the host never had the commit headroom they need - 2026-10-04, judge panel (zero GPU) — the evidence behind the judged pair in §09. The ALPACA and MT-Bench answers of the anchor and Swift IQ4_XS, 100 answers, read blind by three Claude Opus 5 seats in one session with none missing: the same three-seat design as the August panel, each seat run as a headless
claude -pprocess rather than as one of August’s subagents (§09) - 2026-10-04, anchor reproduction — the anchor’s published-view scores (GSM8K 100, MATH-500 100, HumanEval 92, MBPP 92, MeetingBank 22.3, Mean 81.3) and the derived narrower view (100, 92, 84, 88, 22.3, Mean 77.3) equal the August references exactly. Transcript identity: 171 of 171 comparable (non-truncated) anchor items byte-identical in response text and generated token count to the August transcripts. Perplexity reproduces at 6.5956 exact (§09)
- 2026-10-05, the reasoning-regime re-run — pass 1 of seven arms at 02:28–02:39 (800.9 s, until the memory-commit guard refused the rest), one arm at 11:16 in an attempt that was stopped, and the last arm’s pass 1 and every pass 2 at 19:58–20:18 (1,378.4 s); void again (§15).
- 2026-10-06 to 10-07, the Pi fine-tune study (23:07 on 10-06 to 10:34 on 10-07) — the evidence behind Appendix B and the notes dated 2026-10-06 and 2026-10-07 elsewhere on this page. Same RTX 3090 (24,576 MiB); llama.cpp build 10502 (commit
0adcc3bb5). Files: bartowski’s IQ4_XS and IQ3_XXS of bytkim’s Qwen3.8-27B-pi and its f16 projector, each checked by size and Hugging Face LFS sha256 before use, with Swift IQ4_XS and Swift IQ3_XXS beside them; no anchor arm. Pre-registered on 2026-10-06 at 22:16, before any measurement: server checks and a Pi smoke test; the 175-prompt suite atxhighandmedium, paired with the 2026-10-04 anchor and Swift IQ4_XS cells, reused read-only; speed probes (1,780 s) and drafter regimes (2,388 s) against Swift IQ4_XS; and deep fills at three windows in the order Pi, Swift, Swift, Pi (08:58–10:08). Not pre-registered: a sampled drafter follow-up (1,300 s from 10:12). Job memory cap per step, replacing the pre-registered fixed 36 GB after the first start was refused: 30 GB for the server checks, the suite and the follow-up, 31 GB for the speed probes and fills. The speed probes, fills and follow-up ran on a quiet host. The server checks, the Pi smoke test and about the first hour of thexhighsuite overlapped the IQ3_XXS download (23:06 on 10-06 to 00:13 on 10-07): the checks’ t/s and the smoke test’s wall times are single readings under that load and are not compared with anything else, and the suite’s wall times are not claimed. The IQ3_XXS header was read after that download, at 00:13 on 10-07. The first start of the server checks was refused by the memory-commit guard before it measured anything, and the suite waited from 01:06 to 06:19 for commit room. Every paired speed cell passed the 3% agreement between passes; none is void - 2026-10-07 17:05 to 2026-10-08 10:04, the n/p drafter sweep — the evidence behind §06’s n and p subsection and the notes dated 2026-10-08 elsewhere on this page, run on request, for clarity on n and p and for the best pair in the launcher. Same RTX 3090, driver 596.36, build 10502. Pre-registered at 17:05 before any result, with a two-load pilot kept as part of the run; 138 server loads ok. Blocks: the Pi fine-tune’s IQ4_XS at
medium; Swift IQ4_XS and the anchor (UD-IQ4_XS) atxhigh; repeat blocks for those three in fresh sessions, added at 18:12 when the launcher’s best pair was requested; Swift IQ3_XXS at one slot and at the two-slot vision shape. Paused 18:43 to 00:10 for another GPU task. Thirteen loads run while the reference machine was in use were excluded at 00:24, kept on disk and re-run behind an idle gate. Dated deviations: survival read directly from the per-position counter (17:12); the paired bootstrap (01:24); like-for-like scoring and n2/p0 with depth (01:57); and the launcher taking each file’s highest score where the UD-IQ4_XS tie-break would pick n2/p0 (08:15). The launcher’s drafters were changed at 10:20 - Methodology version. This page follows the campaign methodology as of 2026-10-05: 31 numbered rules (two more drafted, not yet law) plus four review gates, a two-voice writing law, a recipes-first structure, and a standardized-energy-metrics mandate. Rule 30 records that throughput on this rig has two levels and that nothing recorded predicts which one a session reaches, so ratios travel and absolute numbers do not (§03). Rule 21 requires that judged sets compare inside one panel session only: the same 25 byte-identical MT-Bench answers scored 79.7 in one session and 87.0 in another (C37). Rule 20 now requires empty answers and truncations as separate columns, because the truncation counter is structurally blind to an answer that terminates normally and returns nothing — a blindness that produced a false statement in an earlier draft of that very table (amendment of the 2026-08-25 revision). Rule 25 now requires a sweep to be run at the recipe that ships: the whole quantization ladder was measured with the drafter off and gave the wrong file ordering for the configuration every recipe on this page uses (amendment of the 2026-08-25 revision). The methodology is pinned here for the same reason the build is pinned to commit
0adcc3bb5: a page that names a rule number without naming the ruleset cannot be checked later
The fifty-four places this page corrected itself
Each entry is a claim an earlier edition of this page made and this edition retires, with the measurement that forced it. They are kept visible rather than quietly deleted, because a page that hides its corrections teaches its readers to trust the wrong things.
- "ALPACA and MT-Bench are unscored by design." They were unscored by circumstance — no judge was available — and calling that a design choice dressed a gap up as a decision. A blind three-seat panel scored all 150 kept answers on 2026-08-24, the headline index became a seven-set index, and the pair turned out to be the only instrument in the campaign that separates the effort levels at all:
xhighlost tomediumon 14 of 25 MT-Bench prompts. The gap that remains is real and now stated as a gap — the judge and this page's author are both Claude models. §09, §15 - "The built-in draft head costs no VRAM." It costs 1,008 MiB fixed, plus 5,120 B per window token, plus another 898 MiB at n-max 10 — about 1.8 GiB at a 163,840-token window. §05, §06
- The vision projector was budgeted at its file size (~0.9 GB). It occupies 1,138 MiB resident, measured four times. §04, §05
- Large windows were declared safe on the strength of short probes — 196,608 shipped with a claimed ~3 GiB of slack, and 262,144 shipped as verified. The 415 MiB at 196,608 and the 180,224 text-only ceiling with the n-max 10 drafter were both read at load with a short probe; 262,144 text-only with the n-max 4 / p-min 0.75 drafter on spills 2,364 MiB at load (3,502 MiB with the projector), so it needs the drafter off; filled to 218,233 tokens drafter-off it keeps 755 MiB of board VRAM, 1,041 MiB short of the 1,796 MiB desktop reserve, so it needs the card free of any graphical session (§08).C34 §05, §08
- "Past the resident ceiling, speed degrades progressively." It collapses, and only when real tokens land there: 3.82× at a 91,000-token fill, from 30.6 to 8.0 t/s. §05
- The projector was thought to cost about 15% of decode at depth. Paired probes measured 0.04–0.09%: it costs memory and nothing else. §06
- One universal 0.7 bandwidth-efficiency constant. It is format-specific: 0.70 for K-quants, 0.649 measured for UD-IQ4_XS. §04, §07
- Perplexity was said to score "~330k token positions". It scores 294,912 — 36 chunks × 8,192 — and that count is tokenizer-bound, so it is a count of scored windows rather than of corpus tokens. §08, §10
- 1080p and 1440p image costs were interpolated and ran 28–30% high. The measured law is tokens = pixels ÷ 1024 + 2; a 1440p shot costs 3,602, not "~4.7k". §12
- 81.7 t/s was demoted to "best case, not reproducible on real work". It reproduces at 81.71 on a deliberately novel code prompt; the demotion was an unlabelled token regime. §06
- The drafter flags were split per file. A matched sweep across both files, on greedy thinking-off code, ranks n-max 10 / p-min 0.5 first on each (at the cards’ temperature-1.0 sampler n3/p0 led on UD-IQ4_XSC42); acceptance lands within 1.6 points at six of the seven configurations and 3.7 at the seventh, so it is a property of the draft head and not of the quantization. §06
- Configurations were ranked by draft acceptance. Across configurations acceptance inverts: the 96.5%-acceptance setting is the slowest speculating one. Mean draft length — strictly, accepted tokens per verification pass — is what ranks them. §06
- "Acceptance never moves with depth." It rises, 0.80 → 0.92 across a 60× span, in two independent series. §06
- The agent-depth band was published as 34–44 t/s. That was the reasoning stream; the tokens an agent receives run 65–70 t/s at the same depth. §06, §03
- "The board pulls a flat ~344 W." Sustained draw is a range — 306–341 W drafter-off — drifting 11.7% with its own SM clock at constant throughput, constant temperature and constant memory clock. §03
- Idle was published as 33.2 W with no server against 30.7–31.1 W loaded — physically backwards, and the reason was a board still cooling. Settled with little on screen: 29.9 W with no server, 34.1 W with the model resident. §03
- "Never use a
q4_0KV cache." It costs +0.693% perplexity against fp16 — real, super-linear in bits, and smaller than the gap between two respectable 4-bit weight files. A knowing trade, not a prohibition. §06, §08 - "16,384 proved sufficient for every
xhighthought." True on GSM8K only. On a seven-set suite the same cap truncatedxhigheight times, and three prompts exceeded even 32,768. §09, §10 - An effort sweep read as "quality falls as effort rises" (81.3 / 80.5 / 77.3). The entire penalty was the token cap; at 32,768 the arms read 82.1 / 80.5 / 81.3, a tie. §09, §10
- "
lowships broken code, 2 of 2 runs" was printed as a flat claim. A second campaign saw the fatal defect atmediuminstead, and on the 175-prompt suitelowwas the highest-scoring arm. What is solid is that a fatal defect is a live risk at every level at n=1. §09 - "Two concurrent requests gain +60.3% aggregate throughput." That arm ran with the drafter off, and this page quarantined it against an older ~+11% drafter-on figure. The matched drafter-on pair ran on 2026-08-25 and the quarantine is lifted: +22.0% aggregate and −35.4% per slot (82.98 → 101.25 t/s aggregate, 85.79 → 55.41 per slot), acceptance unmoved at 0.618 → 0.620. The older figure was much closer to right, for the reason the quarantine named — the drafter has already saved most of the repeated reading of the model that batching would otherwise save. Every recipe still ships
--parallel 1, now on evidence rather than on caution. §03, §06 - "Quality falls about 1% of perplexity per gigabyte down to 2.9 bits, then five times that — so stop at about 9 GiB." The advice stands and the reasoning behind it did not. The per-gigabyte figures were right, but 2.9 bits per weight is not the point after which quality degrades — it is the last rung that still performs like the full-quality file. Accuracy on 75 paired items is flat from 4.22 bits through 2.91 and still a tie at 2.48, and empty answers are exactly zero down to 2.91. §08
- The whole quantization ladder was swept with the drafter off, and it gave the wrong ordering for the configuration every recipe ships. Drafter off, UD-Q2_K_XL is 7.8% faster than UD-IQ4_XS and the ladder reads "swap your daily file". Drafter on, UD-IQ4_XS is 12.9% faster despite being 45% larger, because the draft head degrades with bit width. Sweep at the recipe you ship. §08
- Full native context on UD-IQ4_XS was said, from arithmetic alone, to need the card free of any graphical session. It is now measured: deep-filled to 218,233 real tokens the 4-bit file reaches 23,821 MiB with speculation already off, leaving 755 MiB — 1,041 MiB short of the 1,796 MiB this page reserves for a desktop. UD-Q2_K_XL holds the same window with the drafter at 1,717 MiB of slack and decodes 34% faster there. §08, §05
- "Board watts fell as the drafter got wider — the win compounds." They did not. That claim came from the whole-window mean watts (325.2 → 308.4 → 302.4), which fall because each request also contains a short low-power prefill and setup segment drawing 101–139 W, and that segment is a larger share of a shorter run. Power while actually decoding is flat: 344.6 → 341.7 → 341.0 W, a 1.0% change. The 2.52× energy saving is the 2.50× throughput gain and nothing else — a duller mechanism and the true one, and the same distinction is why joules per token divide by the decode watts, not by the column headed mean load. §03
- The word "cliff" named two unrelated things. It is this page's term for the collapse that happens when a window overflows the card — a real discontinuity, 3.8× at a 91k fill — and it was also being used for the quality step between
lowandmedium, which is a step in a small sample and not a discontinuity in anything. The second use is gone. §05, §09 - The short-context speed band was published as 71–84 t/s. Its lower endpoint came from one live check of the vision recipe whose token regime was never recorded — the one thing this page says no speed number may lack. Rebuilt from two measurements that do carry their regime: 83.5 t/s on a novel code prompt at
-c 32768and 86.3 at a 1,458-token fill under the cooled protocol, both answer tokens at n4/p0.75, the flags shipped then. The band is now 83–86 t/s, and the 71.3 figure is kept only as a labelled loose end. §03, §12 - The drafter-off floor was printed as 39.7–43.8 t/s, and prose speculation was credited with 1.16×. Both were inherited from a summary bullet in the campaign log that this page checked against the log's own table and found wrong. 43.80 t/s is a drafter-ON prose figure, so it cannot be a floor endpoint; the measured drafter-off floor for this file is 41.46–42.97 across five contents and both token regimes. And 43.80 ÷ 41.55 is 1.05×, not 1.16× — the 1.16× belongs to
n4/p0.75's 48.35 t/s on the same content, which is what the log's own "best vs floor" column says. §06 - Three memory budgets converted MiB to GiB by dividing by 1000. The 262,144-token window costs 9,984 MiB (9.75 GiB, not "9.98"), that configuration totals 23,216 MiB (22.7 GiB, not "23.2"), and the worked vision example totals 21,556 MiB (21.05 GiB, not "21.6"). The slack figures those pages quoted were right; the GiB labels were not. §05
- “Treat every file below 4 bits as untested for agentic work.” That was true when it was written and it is no longer true of the file this page recommends. Two full 225-exercise arms of aider’s polyglot coding benchmark — one per file, every exercise carrying its own unit tests and a second attempt after a failure is shown — were run on 2026-08-26. UD-Q2_K_XL and UD-IQ4_XS finish level at 96 of 225 each, and the 2-bit file pays 20 to 45% more tokens, 20 to 36% more wall-clock time and about 32% more energy per solved exercise to get there. What is still untested is a loop of dozens of tool calls, which is a narrower claim and is still on the page. §08
- “bits per weight → smaller file · worse is higher” on the ladder figure. Worse is drawn further down. That figure’s vertical scale runs 0% at the top to 60% near the bottom, so a larger loss is a lower point, and the note contradicted the drawing it labelled. The drawing was right: every plotted point was re-checked against the tables beside it and all twenty-four agree. Only the note was wrong, and it is corrected. §08
- “On agreement, UD-IQ2_XXS and UD-IQ1_M are 15 sigma apart.” They are 10.9 sigma apart. The gap had been divided by one file’s standard error instead of by the combined error of the difference — the square root of the sum of the two squares, which is the rule this page already applies to perplexity two parts earlier. Every separation in that part was printed about a third too large: the 17-sigma figure is 11.8, and the “error near 0.13” is 0.17 to 0.27. No conclusion in that part changes — every rung still separates cleanly, the narrowest at 11 sigma — but the numbers do. §08
- “Every file whose code ran had followed both of the execution probe’s instructions.” Two files break that.
UD-IQ2_Scontinued the file as asked but wrapped it in the markdown fences the prompt forbade, and its program runs; the cross-modelgemma-4-12B-QAT-Q4_0followed both instructions exactly and produced code that does not parse. Ignoring the shape you asked for came before every failure in this probe, but it does not predict one, and the paragraph now says which direction the rule runs in. §08 - “Four instruments read the ladder.” Three, plus a breakdown of one of them. The empty-answer column is counted from the same 75 greedy generations the accuracy column is scored on, and every empty answer on every arm is also scored wrong — checked across nine arms, with no exceptions — so it cannot disagree with the accuracy column and is not independent evidence beside it. It stays in the table because counting one named failure is more sensitive than testing a 75-item score for significance, which is a real reason and a different one. §08
- “The energy comparison is paired over the 36 exercises both files solved inside the logged window rather than over all 96.” A paired comparison cannot be taken over 96 at all: 96 is what each file solved separately, and the two did not solve the same 96. They jointly solved 66, and 36 of those fall inside the window where board power was logged. The 36 was always right; the thing it was being contrasted with was not. §15
- “Six of the ten programs run.” Seven of eleven. The execution probe gained an eleventh file on 2026-08-26 — sdkyuan’s QAT-Q2_0, whose program runs — and this page went on printing the old count. The line the probe draws has not moved: the three files below 2.481 bits per weight still fail, and so does the cross-model gemma file at 4.651. §08
- "Doubles the KV cache" / "multiplies the KV cache." The cache is sized by
-cand split between the slots. A second slot at-c 32768added 134 MiB of board VRAM (2026-08-23, UD-IQ4_XS,q8_0KV, drafter off,kv_unifiedoff as logged, n=1 per setting), not the 1,248 MiB a second full cache would cost. §06, §14 - "8.0 t/s on a 91k-token document with 2,364 MiB spilled." 8.0 t/s belongs to the run with the BF16 projector loaded and the n-max 4 / p-min 0.75 drafter, which spilled 3,502 MiB at load (UD-IQ4_XS,
-c 262144,q8_0KV, greedy, prompt filled to 90,886 tokens, 2026-08-23); the 2,364 MiB is the text-only configuration with the projector off, at load. The glossary and §05 set 8.0 t/s against 50 t/s, the level of the shallow 2026-08-22 probes; the 26% printed beside it is 8.0 against the 30.6 t/s the fully resident-c 131072window read at the same depth. §05’s 262,144 budget row attributed the 2,364 MiB spill and the 25,504 MiB requirement to the n-max 10 coding drafter, and the launchers printed them beside their n-max 10 variants with no drafter setting named; both are n-max 4 / p-min 0.75 figures (the n-max 10 coding drafter costs 898 MiB more). The ceiling paragraph printed the shallow probes of the 8.0 t/s configuration at 65–70 t/s where the 2026-08-23 sweep read 59–64 t/s, and attributed the 8.0 t/s to the text-only drafter-off window, which measured 15.96 t/s filled to 218,233 tokens (2026-08-25). §02, §03, §05, §08 - “12 GB is a plain no, 416 MiB over.” The 416 came from holding a board VRAM figure that contains the measuring machine’s desktop against the card size minus a desktop reserve, counting the desktop twice. UD-Q2_K_XL at
-c 32768with the drafter off allocates 10,497 MiB of server memory (measured 2026-08-26); on a 12,288 MiB card that is borderline, decided by the reader’s own desktop. §08 - "
failed to fit params to free device memoryonly means that an explicit-ngloverrode the automatic fit. Any value triggers it." The warning is window-dependent: present at-c 163840and above, absent at-c 131072and below, with-ngl 99on every load (2026-08-23 ceiling sweep: UD-IQ4_XS, projector, n-max 4 / p-min 0.75). The load that collapsed to 8.0 t/s (greedy) at 90,886 tokens printed it; its resident control did not. At-c 163840the warning fired while the server’s shared GPU memory read 446 MiB at load — an early warning, not proof of a spill. §11, §05, §03 - “The control that would prove the model is actually using the image was never run” and “what it does to the answer was never measured.” Both were measured on 2026-08-25: a generated 1440p target, seven questions scored by string equality, greedy, n=1 per question. Full image budget: 7 of 7. Image withheld: 0 of 7 — the model reads the picture. At
--image-max-tokens 1024: 5 of 7 on UD-IQ4_XS, 4 of 7 on UD-Q2_K_XL — fine 12–15 px type is lost, large type survives. §05’s budget row and §15’s register row printed the whole-prompt medians (3,635 and 1,043 tokens, 3.49×) as the image’s cost; the image itself is 3,602 and 1,010 tokens, 3.57×. Three items in the front matter (power capping, system RAM, sub-4-bit board power) were listed as not measured; all three are closed in the register. §05, §12, §15 - The chart annotation read “the cliff, measured” and the prose called it “the cliff chart”. Every solid plotted point is a short wall-clock probe (prefill included, first request after each load; the record does not name the file, Q4_K_M is a reconstruction); the sweep read 45.1 t/s at 147,456 and 47.1 at 180,224, not 50–55 throughout, though all above its 41.6 t/s pass line. The caption called every solid segment a probe; the solid run-in from the chart’s left edge to 122,880 has no probe under it (the sweep started at
-c 122880). A short probe cannot locate the collapse. On-c 262144a 1,531-token prompt read 42.9 t/s while the same window filled to 90,886 tokens read 8.0 t/s (greedy) — the table below the chart, not the chart itself, measures that. §05 - “These are deep-fill measurements, not shallow probes.” Apart from the 218,233-token fill (2026-08-25, 755 MiB of board VRAM remaining) and the figure derived for 262,144 at load, the paragraph’s ceiling and slack figures were read at load with a short request and no fill (2026-08-23); the label was backwards. §08
- “The head itself always stays at 8-bit whatever the main quantization is.” The head’s precision is the quantizer’s choice: in UD-IQ4_XS the 15 head tensors are Q8_0, Q6_K and F32 (334.75 MiB); bartowski’s quantizations of Swift-1.5, a fine-tune of this model with the same architecture, carry them at Q4_0 and F32 (227.91 MiB). §06
- "179 MiB once a desktop takes its measured worst case of 1,181 MiB" and "~25.5 GiB". C12 retired 1,181 MiB on 2026-08-26; against the resting worst case of 1,669 MiB the 1,360 MiB left at load is 309 MiB short, not 179 MiB over. 25,504 MiB is 24.9 GiB (25,504 ÷ 1,024); dividing by 1,000 gives the printed 25.5. §03, §05
- “An arm’s score moves by at most 1.0 point on the 0–100 scale.” The band holds between two sessions that judged the same ALPACA packets. The anchor’s 25 MT-Bench answers are byte-identical across the two sweeps, yet scored 79.7 in the August session and 87.0 in the October one — 7.3 points apart, 20 items higher and 1 lower. In a later session, with different answers beside them, judged scores are not bound by that band. §09
- “The verification step guarantees every configuration writes exactly what the base model would have written” and “quality only changes when the weights change”. On one prompt (build 10502, greedy, thinking off,
-c 32768, 700 tokens, 2026-10-04) the anchor’s drafter-on text diverged from its drafter-off text at byte 21 (n10/p0.5) and byte 269 (n4/p0.75), and its own drafter-on runs did not always repeat; Swift IQ3_XXS likewise. Swift IQ4_XS was byte-identical across all three drafter settings. §06, §08 - “Exactly as fewer gigabytes per token predicts” — the 2.9-bit file’s 7.8% drafter-off lead matched to the size arithmetic. The arithmetic predicts 1.45× (one constant) or 1.56× (format-specific constants); the measured ratio is 1.078×. In October the same size arithmetic missed for both Swift files (predicted 0.917 and 1.175, measured 1.005 and 1.034). §08
- “The Swift chat template differs from the base model’s.” It differs from the anchor’s (unsloth UD-IQ4_XS, 9,993 characters, accepts
high); lmstudio-community’s Q4_K_M of the base model carries the Swift files’ 8,952-character template byte for byte, and it raises onhigh. §14, §17 - “You cannot switch level per request.” A
reasoning_effortfield in the request body is ignored; a per-requestchat_template_kwargsreached the template at the server’s/apply-templateon build 10502 (2026-10-06). Generation with it is untested, so relaunching remains the verified route. §09 - “Set n-max 10 / p-min 0.5 where the VRAM is spare and the work is code”; “n4/p0.75 wherever the projector is loaded”. That split was priced on greedy probes, mostly code. At the cards’ sampler with thinking on, n-max 3 / p-min 0 beat both on UD-IQ4_XS on shallow reasoning tokens (1.111 against 0.961 and 1) and on answer tokens; after a 57,540-token prefix n10/p0.5 read higher, intervals overlapping. n10/p0.5 remains the August grid’s greedy-code peak. §06
- “A safe starting point, not the peak” (n-max 4 / p-min 0.75). At n-max 4, p-min 0 read higher than 0.75 on shallow reasoning tokens on every file the n/p sweep ran (on UD-IQ4_XS the interval includes 1), and p-min costs no memory; on 2026-10-09 the commands for files, cards and slot counts the sweep did not run moved to n4/p0 der. §06
- “The Pi fine-tune’s launch command uses n-max 3.” At its own
medium, n4/p0 read 1.038 of n3/p0 on shallow reasoning tokens and 1.073–1.082 after the prefix in two blocks; the command now uses n-max 4 / p-min 0. Appendix B - “Set
defaultLeveltomediumso that Pi’s display matches the server.” Pi 1.0.2 reads no such field and sends no effort here; the server’s launch setting is the only one. Appendix B
What was not measured
Stated plainly so that nobody mistakes an omission for a result. Each entry names the measurement that would close it and roughly what it would cost.
| Gap | Why it is missing | What would close it, and its price |
|---|---|---|
| Power capping — CLOSED 2026-08-25 | The earlier probe returned exit 4 Insufficient Permissions and the cap was never changed. It has since been changed from an elevated shell, and the default was read first, restored in a finally, and verified by reading it back — the setting is persistent and a card left capped would corrupt every later measurement on this machine. | Measured twice. The 2026-08-25 sweep (305.4 W mean against a 350 W limit, never saturating) read −5.0% t/s, −6.5% J at 300 W and −12.8%, −13.7% at 250 W. Re-measured 2026-08-28 with every arm saturating its cap (SwPowerCap residency 99–100%) MEASURED: 300 W −4.1% t/s, −13.1% W, −9.4% J; 250 W −14.6%, −27.1%, −14.7%. Capping to 300 W is a better deal than the first sweep claimed. |
| Session-to-session speculative throughput band — two levels, cause unmeasured | Twenty-three sessions of one configuration (UD-IQ4_XS, n10/p0.5, -c 32768, -ctk/-ctv q8_0, greedy, 700 tokens, fresh server per session, one warmup discarded, RTX 3090 at 350 W, build 10502, all 2026-08-28) landed on two distinct levels: 20 sessions at 75.71–78.65 t/s and 3 sessions at 88.18–88.76 t/s, about 13% apart, with nothing in between. Ruled out: the workload (bit-identical), the drafter (bit-identical acceptance and draft length), the llama.cpp build (unchanged), the core clock (a high session ran 49 MHz slower), board temperature (75–83 °C alike), the memory clock (pinned at 9501 MHz), and prior load (refuted by a designed paired test: four alternating COLD/HOT pairs, difference −0.43%, ranges overlapping). Within a session, throughput correlates with mean draft length (r = +0.683 to +0.894, every session positive). Between sessions, draft length cannot be the mechanism: greedy sessions produced bit-identical drafter sequences while session means still moved. The cause is unmeasured.C27C26C25 | Identify why speculative throughput lands on one of two levels at the same flags. This may require llama.cpp instrumentation or a controlled multi-session campaign varying one hardware or software state at a time. Until then, every speculative throughput figure on this page is a point drawn from a band at least 15% wide |
| Two concurrent requests with the drafter on — CLOSED 2026-08-25 | The one concurrency arm that had ever run used the drafter off, and it contradicted an earlier measurement, so both were quarantined | The matched pair ran: UD-IQ4_XS at n10/p0.5, three reps each, +22.0% aggregate and −35.4% per slot with acceptance unmoved (§06, §03). What is still open is the energy half — the pair measured throughput, not board power, so the retired 5.19 J/token figure has no drafter-on replacement. One re-run of the two arms under the power logger would close it. Minutes |
| Speed of any sub-4-bit file on a 16 or 12 GB card — narrowed 2026-08-25 | Quality was measured on every rung and does not depend on the card. The VRAM requirement is now measured too, at three files × three windows × the drafter on and off, deep-filled (§08) — and because a requirement is a property of the model, the window and the flags, it transfers to any card. So the fit half of this gap is closed and only the arithmetic of card-total-minus-reserve is derived. What remains open is speed, which belongs to a card's memory bandwidth and to nothing else | One 16 GB card and one 12 GB card, loading UD-Q3_K_XL and UD-Q2_K_XL at a stated window and timing decode at depth. Nothing on this machine can close it — it owns one 24 GB card (§08) |
| Board power for any file below 4 bits — CLOSED 2026-08-26 | The 19-arm power matrix predated the ladder and all three of its quantisation arms were 4-bit, so nothing here priced the energy of the file this page recommends. | Measured, -c 32768, q8_0 KV, drafter n4/p0.75, three settled probes MEASURED: UD-IQ4_XS 4.199, UD-Q3_K_XL 4.389, UD-IQ3_XXS 3.903, UD-Q2_K_XL 4.085 J per decode token. Smaller files are NOT more energy-efficient — those span 12 % with no trend and the most efficient is the middle rung. Board power barely moves (311–314 W); throughput sets J/token and drafting sets throughput. UD-IQ2_S reads 6.829, but it contains no MTP layers so it ran drafter-off — that is the price of losing speculation, not of quantisation, and it is not comparable to the four above it. |
| This coding benchmark’s own run-to-run flip rate — CLOSED 2026-08-28 | The two 225-exercise arms in §08 were each run once, so the 60 exercises the two files disagree about contained both the effect of the quantisation and whatever this benchmark flips on its own. | Measured on both files (UD-IQ4_XS 2026-08-27, UD-Q2_K_XL 2026-08-28; aider polyglot benchmark, whole edit format, reasoning off, -c 32768, n-max 4 / p-min 0.75, greedy, 113 exercises paired by case and language) MEASURED. The 4-bit file flipped 18 of 113 exercises = 15.9% (95% interval 10.3–23.8), split 9 each way, both runs scoring 43.4%. The 2-bit file flipped 25 of 113 = 22.1% (95% interval 15.5–30.6). The cross-file disagreement of 60 out of 225 (26.7%) is tested against each floor — a p-value below 0.05 would mean the gap is too wide to blame on noise. Against the 4-bit floor, p = 0.027 — below the line, but this was the quieter of the two files. Against the 2-bit floor, p = 0.364. Pooling both retests — 43 flips out of the 226 exercises both retests covered, a combined floor of 19.0% — gives p = 0.053. No quantisation effect is demonstrated. The two floors do not differ significantly from each other (p = 0.236) |
| A read-then-edit task against existing code | Every scored generative task here writes new code. The coding benchmark in §08 adds one repair round, but what it repairs is the model’s own answer from a minute earlier rather than an existing codebase somebody else wrote. One real repair task did run, inside an agentic pipeline validation: 22 model calls, 9 m 36 s, 0 of 96 target tests fixed, all 561 regression tests still passing. The 10-task × 3-effort sweep that would have made it a measurement was cut on a ~4-hour time gate against a ~8.5-hour projection | That cut sweep, or a smaller suite of repair tasks against real files judged blind at n≥10 per effort level. Several hours, and it is the measurement that would let this page speak about the work it recommends the model for |
| Blind, repeated quality judging | The 08-22 pages were graded blind but only n=2 per level; the 08-23 single runs were n=1 and not blind. Neither is enough | n≥10 per level on one task with blind judging. Separating a real effort-to-defect effect from sampling noise needs that and nothing less |
| The vision critique loop's control — CLOSED 2026-08-25 | The same question had never been asked with the image withheld, so no claim on this page about the model “seeing” anything was more than an assumption | Measured. A generated 1440p target whose every answer is known exactly, scored by string equality, asked three ways. Image at full budget: 7 of 7. Image withheld: 0 of 7 MEASURED. The model is reading the picture rather than inferring from the question, and the control now exists to say so. |
The quality cost of --image-max-tokens 1024 — CLOSED 2026-08-25 | The saving was measured (3.57×) and the consequence was not, so the flag shipped in a recipe on a token count alone.C32 | Measured, and it is not free. Same image, same seven questions, two budgets: full 7 of 7, reduced to 1,024 tokens 5 of 7 on UD-IQ4_XS and 4 of 7 on UD-Q2_K_XL (greedy, n=1 per question) MEASURED. The split is the whole finding — questions about large type survive 2 of 2 at both budgets on both files, while questions about 12–15 px type fall from 5 of 5 to 3 of 5 on UD-IQ4_XS and to 2 of 5 on UD-Q2_K_XL.C32 The cheaper budget costs fine detail specifically. Use it to fit a window, not to read small text. |
Long-context retrieval with a q4_0 cache — CLOSED 2026-08-26 | Perplexity at -c 8192 cannot see a retrieval failure deep in a window, and that was the actual argument against a 4-bit cache. | Measured, and harder than a needle test: hundreds to thousands of near-identical numeric records with one specific value as the answer, one server load per depth. UD-Q2_K_XL scored 5 of 5 at 61,727, 89,589 and 119,435 tokens on f16, q8_0 AND q4_0 alike MEASURED — KV precision made no difference at any depth. A q4_0 cache costs nothing measurable for retrieval on this file. |
| System RAM under either load mode — CLOSED 2026-08-26 | The 15 GB saving attributed to --load-mode none was an inherited figure, never a reading taken here. | Measured on UD-IQ4_XS: none holds a 1,275 MiB working set against mmap’s 13,480 — a 12,205 MiB saving, confirmed by two measures agreeing to 3 MiB MEASURED. The inherited claim was directionally right and about 20 % too large. But private bytes are the same within 500 MiB (16,471 against 15,973), so what mmap holds is evictable page cache, not private allocation. none costs 4 seconds of load time and no throughput at all (43.0 against 43.4 t/s). auto is byte-identical to mmap. |
| Any drafter but the built-in head — CLOSED 2026-08-26 | DFlash2 was measured and lost; no other mechanism had been tried, so an unmeasured alternative read as a nonexistent one. | Measured on UD-Q2_K_XL at -c 32768 MEASURED: none 45.7, ngram-simple 45.0 (acceptance 0), ngram-cache 42.0 (0.148), draft-mtp 82.6 (0.943), and ngram-mod 353 t/s (acceptance 1.000). That last one was checked before it was believed — all three drafters produce identical distinct-trigram ratios (0.9515) and a normal stop, as speculation that leaves the text unchanged would, so it is real. But that 353 was one unusually favourable prompt, and a wider sweep corrected it the same day. Six content types, full-length generations, same load per drafter MEASURED: ngram-mod returns 1.10× on a novel algorithm, 1.32× on boilerplate and on JSON, 1.36× on prose, 1.48× on a refactor — a range of 1.10 to 1.48×, not 3.9. And draft-mtp beats it on every content type except prose (1.81× to 2.39×). An n-gram drafter only wins where the output repeats text already in the window, so a single figure for it is meaningless — but on this model the built-in head is the better default almost everywhere, and ngram-mod matters mainly for files that have no MTP layers and therefore cannot use it. |
| Any machine but this one | Every measured row on this page is the same RTX 3090 on the same driver | Nothing here can close it. The non-3090 rows in §07 are arithmetic and are labelled as such |
| ALPACA and MT-Bench scores — CLOSED 2026-08-24 | Both sat unscored for want of a judge. A blind three-seat Claude Opus 5 panel read all 150 kept answers on 2026-08-24; the scores are in §09 and the suite now reports a seven-set index | What remains open is stronger independence: the judge and this page's author are both Claude models. Closing that needs a judge from a different vendor, or human raters. The transcripts, the sealed key and all 450 ratings are kept, so any other judge can be run over exactly the same answers and compared |
| A second-vendor or human judge | Opened by the line above. One judge family has read these answers; nothing tests whether a different family would rank the arms the same way | Re-run judge-panel.py's packets through another model, or a small human panel. The packets are built and blinded already, so the cost is the judging itself |
| Wall power, cost per answer, and carbon | Only in-band board power was ever measured. The power supply, processor, memory, drives and display are excluded | A plug meter. Until then, no figure on this page may be called system power or divided into an electricity bill |
| Energy for four of the seven benchmark sets — CLOSED 2026-08-26 | GSM8K, ALPACA, MeetingBank and MT-Bench ran entirely outside the window the power logger covered, so the per-benchmark energy story was 43 % missing and nothing marked which numbers were absent. | Measured, each set in its own server load and its own power window MEASURED: GSM8K 19.99, ALPACA 36.77, MeetingBank 12.19, MT-Bench 41.17 Wh. A hundred answers cost 110.12 Wh of board energy, about 1.10 Wh each, with a 3.4× spread between question types. Board power is near-constant across all four (325.6–331.5 W). MeetingBank has the highest J/token and the lowest energy per answer at once — long prompts, short answers — which is why both columns are printed. |
| An image-attachment matrix at scale | Each agent got one image and one question — enough to say the plumbing works, not enough to rank them | Several resolutions and multi-image requests per agent. An afternoon |
| Quality ranking between quantizations at usable sample size | §10's own arithmetic says it needs thousands of questions per file | About 12–60 hours and 4–20 kWh per file. This is why the page ranks files by perplexity instead |
| Quality at depth for every Swift card | The needle test did not run on either Swift file: commit headroom on the test host never reached the 34 GB job cap within the 1 h wait (2026-10-04). Every Swift card says its quality at depth is not verified | Two needle runs, one per file at the window it was set for (159,744 tokens on Swift IQ4_XS, 262,144 on Swift IQ3_XXS), with 34 GB of commit free. About 25 minutes of GPU time by the campaign plan’s estimate (10 minutes for Swift IQ4_XS, 15 for Swift IQ3_XXS); neither run started, so this sweep has no wall for it |
| Swift IQ3_XXS on the 175-prompt suite | Cut for time when the recipes were fixed: with Swift IQ3_XXS as a third arm the suite was projected at 11.6 h, so only the anchor and Swift IQ4_XS ran. Swift IQ3_XXS is ranked against Swift IQ4_XS on perplexity, speed and appetite only, and the quantization effect on quality and generation length between the two Swift files is unmeasured (2026-10-04) | One arm on the 175-prompt suite: about the Swift IQ4_XS arm’s own 6,452 s (1.8 h) of cell time (2026-10-04) |
| Swift IQ4_XS: decode-verified window and collapse point | No knee sweep was run on Swift IQ4_XS, so its decode-verified window and collapse point above 159,744 tokens (text) and 139,264 (vision) at one slot are unmeasured, and no deep fill was run with the n10 drafter (2026-10-04) | A knee sweep at the same loaded steps. The knee sweep’s comparable anchor picks ran 7,429.4 and 4,872.2 s of fill time across 20 and 9 loads (2026-10-03/04) |
| Accuracy of the 12 and 16 GB Swift files in Appendix A | No file in Appendix A’s 12 and 16 GB brackets was run on the 175-prompt suite. Swift IQ3_XXS (bartowski’s IQ3_XXS, in the 16 GB bracket) was cut from the suite and has only perplexity, the seven-question vision probe and one Pi task; no other file in the brackets was run on any quality instrument (2026-10-04) | A suite run per file, each about one Swift arm (6,452 s of cell time for Swift IQ4_XS, 2026-10-04) |
| Divergence of any Swift file against the unquantised weights | No file’s KL divergence against the unquantised Swift weights was measured, so the smallest Swift file that performs like the full weights is not located. Only two Swift files were measured, too few for a ladder (2026-10-04) | One reference pass on the BF16 or Q8_0 file (about 55 or 29 GB to download; neither fits this card’s 24 GB, so that pass needs CPU offload, at a speed not measured here), then one KLD pass per file at about the 266–269 s each perplexity pass took on this sweep (2026-10-04) |
| ukisai’s Swift files against bartowski’s | The two repos’ files are not compared on quality at any tier: Appendix A compares their recipes, headers and tensor samples, not their output. Their tokenizer values and embedded chat templates are not compared, and the two repos’ files are not compared by whole-file hash (2026-10-04) | A KLD pass per tier for each repo’s file against the same reference, at the per-file price in the row above. The template and tokenizer comparison and the hashes need no GPU |
| Agents on the Swift files: aider, OpenCode, Qwen Code, DeepSeek Harness | The Swift chat template differs from the anchor’s (unsloth UD-IQ4_XS) template; it is byte-identical to lmstudio-community’s base Q4_K_M template (header read 2026-10-06).C40 The base-model setups in §14 are untested with the Swift files. The one agent measured on a Swift file is Pi, on Swift IQ3_XXS (2026-10-04); Pi also passed a smoke test on the Pi fine-tune’s IQ4_XS (2026-10-06, Appendix B) | One run per agent per file. The Pi probe took 656.9 s for one agent on one file (2026-10-04); 4 agents at comparable depth cost about 4 times that per file |
reasoning_effort: high sent to llama-server with a Swift file | The Swift template raises under a jinja2 render at effort high (TemplateError: Unexpected reasoning effort high); the server’s own minja engine was not tested (2026-10-04) | One server load with high set at launch in --chat-template-kwargs (llama-server ignores a per-request reasoning_effort, §09), and one request. Under the loaded-idle step’s 434.4 s (2026-10-04) |
| The turn after an xhigh session fills a Swift window | What the server does when the window is full (error, context shift or re-prefill) and its wall-clock cost are unmeasured; per-turn times in a real multi-turn session are unmeasured. The turns-per-window table is derived from n=2 single-turn appetite probes at temperature 1.0 (2026-10-04) | A multi-turn session to window fill and one turn beyond, per file. Several xhigh turns per file, each about one appetite answer long: the second xhigh appetite pass took 3,494.6 s for one answer from each of 3 files, about 19 minutes per answer (2026-10-04) |
| Appetite beyond n=2 per file and level, and Swift speed and thinking length at the temperature-1.0 sampler at usable sample size | Appetite is n=2 per file at temperature 1.0, on one task. The thesis’s token saving is measured greedy-only; at the cards’ temperature-1.0 sampler, two runs per file on that one task did not show Swift IQ4_XS generating fewer tokens than the anchor (§09, 2026-10-04). Narrowed 2026-10-08 for speed ratios only: the n/p sweep ran both Swift files at the cards’ sampler (6 prompts × 768 tokens with thinking on, 4 × 384 off, two seeds), as within-session drafter ratios, not a band or token counts (§06) | Further reps per file and level. Each xhigh rep across the 3 files took one pass: 2,074.1 s for the first and 3,494.6 s for the second; the low and medium reps on the two Swift files took 1,878.7 s (2026-10-04) |
| A drafter grid (n-max, p-min) on Swift’s Q4_0 head — narrowed 2026-10-08; n-gram and draft-model drafting on Swift | The n/p sweep ran Swift IQ4_XS with no drafter, at n2–n6 and n10 with p-min 0, at n4 with p-min 0.25, 0.5 and 0.75, and at n10/p0.5; Swift IQ3_XXS at n3, n4 and n5 with p-min 0 and n4/p0.75 with one slot, and n3/p0, n4/p0 and n4/p0.75 with two slots busy (cards’ sampler, xhigh, -c 81920, shallow and after a 57,540-token prefix; §06). Not run: n2 on Swift IQ3_XXS, p-min above 0 at n-max 3, the cards’ own windows; n-gram and draft-model drafting were not tested | n-gram and draft-model drafters on each Swift file, and the grid at each card’s own window. The two-setting drafter sweep took 2,629.8 s (2026-10-04) |
| Suite scores with the drafter on, as the cards ship | Every suite arm ran with the drafter off. On the one drafter prompt, greedy text with reasoning off was byte-identical with the drafter on and off for Swift IQ4_XS, which is evidence, not proof, that its scores carry over; the anchor’s and Swift IQ3_XXS’s drafter-on runs did not always repeat their own text, so for them the check cannot say either way (§08, 2026-10-04) | One suite arm per file with the drafter on: about 6,452 s of cell time for Swift IQ4_XS and 13,619 s for the anchor, their drafter-off arms’ own times (2026-10-04) |
| The reasoning-regime speed group | Void twice: its first pass ran inside a download overlap on the test host (2026-10-04), and a re-run of its 9 arms (2026-10-05) failed the same 3% agreement between passes in five of its six Swift cells (data/arms/speed-anchor-check-thinkdepth-rerun.json). Its passes ran in separate sessions: pass 1 at 02:28–02:39 except the two Swift arms at 91k (11:16 and 19:58, so those two cells divide by an anchor reading from another session), and pass 2 at 20:02–20:18; of the four cells whose first-pass arms shared a session, three failed. Light CPU checks ran on the host at 19:57–20:01, during the Swift IQ3_XXS 91k arm | More passes per cell, in one session, so the pass-to-pass spread is measured rather than tested against the 3% band. About 20 minutes of GPU time per pass of the 9 arms (2026-10-04: 12:34–12:53 and 13:18–13:36) |
Swift energy by KV type, --parallel, effort and power cap | The October sweep measured the 175-prompt suite’s energy and joined energy to the speed probes and the drafter regimes at their fixed settings on the Swift files, which gives decode energy at the three probe depths; nothing varied KV type, --parallel, effort or the power cap. Prefill energy at depth is the row on long prefills below (2026-10-04) | A probe sweep with the power logger running, per varied setting, at about the drafter-regime sweep’s scale (2,629.8 s, 2026-10-04) |
| What moved the judge between sessions: the same 25 MT-Bench answers scored 79.7 and 87.0 | The anchor’s 25 MT-Bench answers are byte-identical across the August and October sweeps; the difference (0.653 rating points per item, 20 items higher, 1 lower) comes from the judging, since the answers did not change. Candidates: the seat harness (subagents in August against headless claude -p in October), the date, and the number of arms per packet. The 7.3-point gap is the MT-Bench figure; the ALPACA gap is about 0.6 to 1.3 points (2026-10-04) | Re-judge the August packets in one new session under both harnesses and compare. Zero GPU; the blinded packets and ratings are kept |
| A loop count over the suite | The automatic loop flags are not a loop count: of 22 flagged items read, 3 were degenerate loops (11 counting circular), and the false positives carry the vocabulary-collapse hit in code and arithmetic, so the raw flag count is not published (2026-10-04) | Read and classify every flagged item. Zero GPU |
| The 12 GB baseline behind correction C30 | The baseline behind C30’s 12 GB allocation reads 1,988 MiB in the stored record and 1,665 MiB in the August log; only a re-measure can tell which holds (2026-10-04) | A re-run of the 12 GB fit measurement with its baseline recorded. About 10 minutes of GPU time (the campaign log’s estimate; the loaded-idle step took 434.4 s on this sweep, 2026-10-04) |
| A token_embd placement print from the loader | llama.cpp source says token_embd stays in host RAM (llama-model.cpp lines 1605–1607), and the measured delta matched CPU placement (2,278 MiB against 2,278.34 predicted, where GPU placement predicts 2,439.38), but no loader print from this build confirms it (Appendix A’s flagged assumption) | One load with verbose logging. Under the loaded-idle step’s 434.4 s (2026-10-04) |
| Prefill energy per prompt token at the 28k and 91k probe depths on the Swift files | The long prefills sit in the discarded probe rows, which the energy join (events-from-arms.py) drops by design (2026-10-04) | A second join over the speed-probe logs on disk. Zero GPU |
| Swift IQ3_XXS with vision at one slot, deep-filled at a window that fits | The knee sweep loaded this configuration (drafter n4/p0.75) only at 253,952 and 262,144 tokens, and both loads spilled into shared memory at load, so the one-slot vision card’s window and Appendix A’s V1 cells for Swift IQ3_XXS rest on the per-card arithmetic alone (2026-10-04) | One deep fill with vision at the card’s window. At most about 10 minutes of fill: the knee’s two fills, at larger windows, took 578 and 602 s (2026-10-04) |
| Host RAM of the Swift cards’ prompt cache at depth | The cards keep llama-server’s host-RAM prompt cache on. The window fills ran with it off (--cache-ram 0), because during the knee sweep it tried to save slot snapshots of 5–6 GB at depth; its RAM use in a real deep session is unmeasured (2026-10-04) | One deep session per card with the cache on, reading host RAM. About one deep fill each: the knee’s fills took 578–602 s at one slot (2026-10-04) |
| Agent-task success of the Pi fine-tune | Only a three-probe headless smoke test ran in Pi (one-shot FizzBuzz, screenshot attach, withheld-image control; n=1 each, 2026-10-06). No pass rate, turn count, tool use or wall time was compared with Swift IQ4_XS or the base model, and the author’s agent results are cited, not reproduced (Appendix B) | The study’s pre-registered agent-task benchmark (its Phase B), with repeats per task and alternated order. The plan prices it at about 26 GPU hours or more, plus new harness code (2026-10-06) |
| ALPACA and MT-Bench for the Pi fine-tune | Not judged: in the Pi study the two sets contribute token counts only (2026-10-07) | The blind three-seat panel over the kept answers, run in one session with the anchor’s and Swift IQ4_XS’s answers, because judged sets compare only inside one session. Zero GPU |
| Quality of any Pi file but IQ4_XS, and header identity beyond two quants | Only the IQ4_XS ran the suite. Pi IQ3_XXS has one two-slot deep fill (speed and fit) and nothing else, and only the IQ4_XS and IQ3_XXS headers were compared with the Swift twins (2026-10-06/07) | A suite run per file: on 2026-10-04 the anchor’s xhigh suite took 3.8 h and Swift IQ4_XS’s 1.8 h. A header read per file needs no GPU |
| The Pi fine-tune’s launch command as shipped: n-max 4 / p-min 0 at 139,264 with vision — narrowed 2026-10-08 | n4/p0 at medium was measured on 2026-10-08 at -c 81920, text only, one slot, shallow and after a 57,540-token prefix (§06). The 139,264 fill, the server checks and the Pi smoke test ran n4/p0.75: the same n-max and the same memory, since p-min costs none. Decode at that window with p-min 0 and an image at depth was not filled (2026-10-08) | One deep fill of the shipped command at medium, with decode probes for both drafters at depth. The study’s twelve fills at three windows took about 70 minutes (2026-10-07) |
| The Pi fine-tune’s tokens and speed at the cards’ temperature-1.0 sampler, and in multi-turn or agent use | Every Pi token figure is greedy and single-turn on the 175-prompt suite; the sampled measurements are reasoning-token decode on 8 probes of 700 tokens at -c 32768 (2026-10-07) and the n/p sweep’s drafter ratios at medium (6 prompts × 768 tokens with thinking on, 4 × 384 off, two seeds, 2026-10-08) — ratios, not a speed band or token counts | Appetite runs at the cards’ sampler with repeats per file and level, and the agent-task benchmark above |
Generation with a per-request effort, and high on the server, for the 8,952-character template | On build 10502 the server’s /apply-template rendered per-request low and xhigh lines over a medium launch; no request generated text through /v1/chat/completions with them, and high was not sent (2026-10-06, Appendix B) | One request per level through /v1/chat/completions, comparing prompt_n as in §09, and one load with high. The Pi file loaded in 10.6 s (2026-10-06) |
| The 2026-10-08 drafters at each card’s own window, deep-filled with an image | The n/p sweep ran at -c 81920, one slot without the projector (two slots: 123,904 × 2 with the projector, no image sent), and its depth was a 57,540-token prefix; no window was re-filled at n3/p0 or n4/p0, so the windows rest on the memory argument: the same or a smaller n-max, and p-min free. After the prefix, longer settings such as n10/p0.5 read as high as or higher than n3/p0 on two files, so the best n-max at 139,264–262,144 tokens is open (2026-10-08) | One deep fill per card at its window with an image in flight, with decode probes at the old and new drafter. The Pi study’s twelve fills at three windows took about 70 minutes (2026-10-07) |
| n-max and p-min on the files, efforts and cards the sweep did not run | Q4_K_M, UD-Q2_K_XL, every non-3090 card, the August listing’s picks 4–6 and Appendix A’s other drafter-on cells were not run; since 2026-10-09 they carry n4/p0 in place of n4/p0.75 der: p-min costs no memory, so each window fitted at n-max 4 still fits, and on the four files the sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot, reasoning and answers at 0.67–0.75 per slot with two (Swift IQ3_XXS, the one file run with two slots), and 0.90–1.03 after a 57,540-token prefix, where on UD-IQ4_XS deep answers it was slightly faster (n4/p0 read 0.973 of it). Their best n-max is not measured (Pi IQ3_XXS has no drafter card: Appendix B runs it drafter off). Not run: n2 on Swift IQ3_XXS, n2 or n5 with two slots, p-min between 0 and 0.25 (n4/p0.25 tied n4/p0 on the Pi file and sat in Swift IQ4_XS’s tie band), the Pi file at xhigh, the others at any effort but xhigh (2026-10-08) | The same grid per file, effort and card. This sweep’s 138 loads ran from 2026-10-07 17:05 to 2026-10-08 10:04, with a pause from 18:43 to 00:10 |
| Quality, agent wall time and energy at n3/p0 and n4/p0 | No suite arm and no agent session ran with these drafters, and no power logger ran in the sweep (the page’s 3.210 J per token is n10/p0.5 on greedy code). On the sweep’s one greedy probe, 256 tokens of code, every setting gave the same text per file; that is one prompt, and C38 records one where drafter-on text diverged (2026-10-08) | One suite arm per file with the drafter on, as above; an agent benchmark per drafter; a probe sweep with the power logger running |
| A greedy-code ranking with an interval, and a clean two-slot reading at depth | The greedy probe is one prompt with no interval. The two-slot block has one load per setting and pass, no repeated reference and no interval; its pass 1 desktop changed between loads (1,094, 602 and 125 MiB), so only pass 2 is quoted, its n4/p0-over-n3/p0 gap is about the size of a level change between server starts, and its deep row is confounded: the two slots’ deep probes decoded at very different speeds and neither reused the cached prefix (2026-10-08) | Several greedy code prompts per file with two passes; a two-slot repeat with matched desktops, repeated references and staggered deep probes |
The reproduction check
One command, one number, one pass band. If your machine returns something inside the band, your build and your card behave like the one every measurement on this page came from. If it returns something below the band, the diagnostic table in §11 tells you which of the two half-speed failures you have.
:: 1. start the server with the drafter OFF — this measures the bandwidth floor, :: which is the one number that does not depend on your content: llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf -c 32768 -ngl 99 --parallel 1 ^ -ctk q8_0 -ctv q8_0 --spec-type none --jinja --host 127.0.0.1 --port 1235 :: 2. send TWO identical requests and discard the first (it inherits the cold-clock :: state that section 11 measures at up to 25%). Each request: 700 tokens, :: temperature 0, top-k 1, thinking OFF, a short novel code prompt. :: ONE LINE - a line break inside the -d body will not survive cmd: curl -s http://127.0.0.1:1235/v1/chat/completions -H "Content-Type: application/json" -d "{\"model\":\"qwen/qwen3.8-27b\",\"max_tokens\":700,\"temperature\":0,\"top_k\":1,\"chat_template_kwargs\":{\"enable_thinking\":false},\"messages\":[{\"role\":\"user\",\"content\":\"Write a sliding-window rate limiter class in Python.\"}]}" :: 3. read the SECOND run's decode speed from the server's own log line: :: eval time = ..... ms / 700 tokens ( ... ms per token, NN.NN tokens per second) :: Use that number. Do NOT use tokens divided by wall-clock (section 10).
Copy-paste version — the same command with the comments removed
llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf -c 32768 -ngl 99 --parallel 1 ^ -ctk q8_0 -ctv q8_0 --spec-type none --jinja --host 127.0.0.1 --port 1235 curl -s http://127.0.0.1:1235/v1/chat/completions -H "Content-Type: application/json" -d "{\"model\":\"qwen/qwen3.8-27b\",\"max_tokens\":700,\"temperature\":0,\"top_k\":1,\"chat_template_kwargs\":{\"enable_thinking\":false},\"messages\":[{\"role\":\"user\",\"content\":\"Write a sliding-window rate limiter class in Python.\"}]}"Expected: 42.97 t/s (measured 2026-08-23, RTX 3090, driver 596.36, llama.cpp build 10502, UD-IQ4_XS). Pass band: 41.7 to 44.3 t/s — that is ±3%, and the 3% is not arbitrary: it is this campaign's measured spread of the drafter-off decode floor across five different contents and both token regimes (§01). One clarification, because the two numbers look like they disagree: the floor across contents runs 41.46–42.97 t/s, and the low end of that belongs to a different prompt. This check pins the prompt, so the band around 42.97 is the right one to judge your own run by. Inside it, your setup matches. Between about 34 and 41 t/s, suspect a different file, a different build, or fp16 rather than 8-bit cache. Below about 35 t/s something is wrong, and there are only two common causes: the -ngl off-by-one reads 29.8 t/s on this file, and a memory spill reads 20–35 (§11). Above the band, check that the drafter really is off — --spec-type none makes the server log unused tensor blk.64.*, and if that line is missing you are measuring a speculating server, which starts at about 73 t/s.
Everything above is drawn from two complete agentic runs on this card, and both are published whole rather than summarised. Each report carries 19 figures with their conditions on the figure itself: the roofline and its operating points, the voltage/frequency cloud and clock residency, which limit was active at each sampled moment, the power-cap sweep against the operating cloud, per-request workload shape and KV reuse, the phase timeline, and two deliberate null results. Captions name what was not measured as well as what was — per-process VRAM attribution is impossible under Windows WDDM, and memory-junction temperature reads NULL on this part.
UD-IQ4_XS agentic run — 19 figures · UD-Q2_K_XL agentic run — 19 figures
Self-contained pages, about 6.9 MB each; every figure is embedded, nothing is fetched.
Still missing, and named because a reader asked for it: there is no time-series of SM clock, board power and throughput against wall clock. The sustained drift — 1,453 to 1,606 MHz with board power rising 305.5 to 341.1 W at constant throughput — is reported as two endpoints, and two endpoints cannot tell monotone drift from stepped or oscillatory drift. The telemetry to draw it was collected; the plot has not been made.
Swift IQ4_XS holds a text window of 159,744 tokens per slot fully resident on a 24 GB card, and a vision window of 139,264 with a desktop of up to 1,330 MiB — both deep-filled on the 24,576 MiB RTX 3090 on 2026-10-04, with desktops of 841–846 MiB at load meas n=1. The text fill left 1,880 MiB for a desktop, more than this page’s 1,796 MiB threshold (§02); the vision fill left 1,330 MiB, so at the threshold the vision window is 128,000 der. The per-card tables below give the largest fully resident window for each Swift-1.5 file on cards from 12 to 32 GB. Every cell is arithmetic on one measured machine: the 1-slot slope rests on this RTX 3090’s August 2026 runs of unsloth files of the base model plus one October load; the constants rest on the knee sweep of 2026-10-03/04 (the vision constant on pick 8, from which the conservative text constant follows; the expected text constant on pick 5), with the drafter-off cost from August on/off pairs; the 2-slot slope rests on pick 8 (below); and the deep fills of 2026-10-04 checked the 159,744 and 139,264 cells directly.
On a 12 GB card, at a 1,272 MiB desktop reserve (a 555 MiB desktop plus 0.7 GiB), a Q2 file at about 2.5 bits per weight fits for text and one conversation at a time; a standard Q3 at 3.60 bits per weight or larger does not fit in any configuration. The only file labelled IQ3_XXS that fits is GSQ-RCO’s 3.00-bit standard file, and only with the drafter off; at this page’s 1,796 MiB reserve it fits only at the expected constant (10,240 tokens; below). This page recommends nothing below 2.912 bits per weight for the base model (§08); that result is for a different model and quantiser, so it does not carry over to these files as a measurement, and Swift-1.5’s KL divergence figures are its maker’s, cited, not reproduced here (below).
Swift-1.5 is a fine-tune of Qwen3.8-27B with the same architecture, layer count and KV layout, so the same KV arithmetic applies (the 38 KiB-per-token floor, below), with the same 262,144-token native window. Nothing in this appendix is a quality claim: no file in its 12 and 16 GB brackets was measured for accuracy (§15). The derived cells may not be compared with the measured requirement rows of §05 and §08: those rows measured unsloth files of the base model in August 2026, with the same llama.cpp build (10502), and some are board figures that include the desktop; the cells mix August 1-slot data with 2-slot data from 2026-10-03/04.
The decode knee: where speed collapses when the window is filled
The knee sweep tested the eight picks of an October launcher, one sweep on 2026-10-03/04, not comparable with the August tables in §06. Picks 1–3 use the file and projector of §03’s menu rows 1–3 and the drafter those rows carried before 2026-10-08 (n4/p0.75, n10/p0.5, off; rows 1 and 2 now carry n3/p0), at larger windows than §03 ships (219,136 and 196,608 per slot against 122,880 and 180,224; 262,144 for both); picks 4–8 are other configurations, named in each row. A knee is a window step whose decode falls at least 15% below the next-lower window of the same configuration, confirmed by a second load or by a confirmed collapse above it. Each pick carries two window numbers. The decode-verified window is the last step whose decode held with the window filled (to about 91%, 98% for picks 3, 5 and 7), minus 0.7 GiB at the configuration’s per-token slope, floored to 1,024 der, at the knee loads’ 258–539 MiB desktops: it is not judged against this page’s 1,796 MiB threshold, it is fully resident only where the table says so, and that October launcher ships these windows. The fully resident window is the highest loaded step at which three tests pass at every loaded step up to it: dedicated memory at depth plus a desktop allowance of 855 MiB and its 121.0 MiB spread fits the card, shared memory at load rises by no more than 100 MiB over the next-lower step, and decode holds. They differ for two reasons. The desktop allowance (the 855 MiB maximum, from the deep fills of 2026-10-04, plus the 121.0 MiB spread of 73 desktop readings, 72 loads and one direct reading) is heavier than any knee load saw (258–539 MiB), and that alone sets the fully resident window of picks 1, 2, 4 and 8; and a small spill of buffers that decode does not touch costs no speed, so a decode-verified window can sit above the fully resident one. Steps between two loaded steps were not tested. Conditions: thinking off, temperature 0, 400 predicted tokens, NVIDIA GeForce RTX 3090.
window-knee.py (q8_0 KV, -ngl 99, prompt cache off), one pick per panel. Regime: answer tokens, --reasoning off, temp 0, 400 predicted. Each window was filled to 90–98% of its length before the decode probes; desktop VRAM before each load 258–539 MiB (n=69 loads); job memory cap per row: 32 GB, 38 GB, default 0.75 x RAM (23.9). Knees and the last window whose decode held are judged by a 15% drop, re-run to confirm; spill onsets by a rise of more than 100 MiB in shared memory at load over the next-lower step. “Held” is about decode only: for picks 1, 2, 4, 6, 7, 8 the fully resident window (fit + spill + knee, in the table) is lower. Picks ran in a fixed order, so decode is not compared across panels. 3 fill(s) killed under the job memory cap are set aside in knee-capped-failures.jsonl and not drawn. Not drawn (off the axis): pick 8 adhoc 16,384 (39.1 t/s).| Pick | File · mode · drafter · slots | Last window whose decode held (t/s) meas | Decode knee (t/s) meas | Decode-verified /slot der last held − 0.7 GiB, or the native 262,144 where decode held to it; the October launcher’s | Spill onset | Fully resident fit + spill + knee, loaded steps only | Loads | Desktop before, MiB |
|---|---|---|---|---|---|---|---|---|
| 1 | UD-IQ4_XS anchor · vision · n4/p0.75 · 1 | 235,520 (35.8) n=1 | 237,568 (13.3) re-run | 219,136 | 204,800 | 180,224 next step loaded: 188,416, fails | 20 | 281–379 |
| 2 | UD-IQ4_XS anchor · text · n10/p0.5 · 1 | 212,992 (44.5) n=1 | 215,040 (15.7) n=1 | 196,608 | 212,992 | 188,416 next step loaded: 196,608, fails | 9 | 258–286 |
| 3 | UD-IQ4_XS anchor · text · off · 1 | 262,144 (16.0) n=1 | none to 262,144 | 262,144 | – | 262,144 | 2 | 288–288 |
| 4 | UD-IQ4_XS anchor · vision · n10/p0.5 · 1 | 212,992 (37.0) n=1 | 215,040 (14.5) n=1 | 196,608 | 180,224 | 155,648 next step loaded: 163,840, fails | 19 | 258–299 |
| 5 | Swift IQ3_XXS · text · n4/p0.75 · 1 | 262,144 (36.2) n=1 | none to 262,144 | 262,144 | – | 262,144 | 2 | 415–415 |
| 6 | Swift IQ3_XXS · text · n4/p0.75 · 2 | 133,120 (29.2) n=1 | 135,168 (14.0) n=1 | 123,904 | 133,120 | 124,928 next step loaded: 133,120, fails | 7 | 356–479 |
| 7 | Swift IQ3_XXS · vision · n4/p0.75 · 1 | 262,144 (30.8) n=1 | none to 262,144 | 262,144 | 253,952 (against the text-only twin) | none: the lowest step loaded, 253,952, fails | 2 | 366–415 |
| 8 | Swift IQ3_XXS · vision · n4/p0.75 · 2 | 133,120 (23.1) n=1 | 135,168 (11.6) n=1 | 123,904 | 122,880 | 16,384 meas fit only next step loaded: 114,688, fails | 8 | 338–539 |
One sweep, 2026-10-03 to 2026-10-04; not comparable with the August tables in §06. Conditions: as the knee figure (NVIDIA GeForce RTX 3090). Windows are tokens per slot. “Decode-verified” is the last step whose decode held with the window filled, minus 0.7 GiB at the configuration’s per-token slope, floored to 1,024, or the native 262,144 where decode held to it (picks 3, 5 and 7), at the knee loads’ desktops (last column); it is not judged against this page’s 1,796 MiB threshold, and it is fully resident only where the next columns say so. The October launcher the sweep tested ships these windows. “Fully resident” is the highest loaded step with every loaded step at or below it passing three tests: fit (server dedicated memory at depth + a desktop of 855 MiB + its spread of 121.0 MiB ≤ 24,576 MiB), spill (shared memory at load up by no more than 100 MiB over the next-lower step) and knee (decode within 15% of the next-lower step’s); a step with no loaded step of its pick within 32,768 tokens below it is tested for fit only, unless it is a vision step with a text-only twin at the same window, which it is then judged against for spill and decode. Steps between two loaded steps were not tested. Each step is one server load unless marked re-run; a knee marked n=1 is confirmed by the collapse above it. Each row’s title carries the server command.
Picks 7 and 8 show how far the two kinds can sit apart. Pick 7 (Swift IQ3_XXS, vision, one slot) held 30.8 t/s n=1 at the native 262,144 tokens, so its decode-verified window is the full native window; its fully resident check found no passing step, because the lowest step loaded for it, 253,952, fails both the fit test (23,837 MiB dedicated at depth plus 855 + 121 MiB is more than the card) and the spill test (478 MiB more shared memory at load than its text-only twin, pick 5, at the same 415 MiB desktop). Pick 8 (Swift IQ3_XXS, vision, two slots) held 23.1 t/s n=1 at 133,120 per slot and collapsed to 11.6 t/s at 135,168 n=1, so its decode-verified window is 123,904 per slot; its fully resident window is 16,384, the only step loaded below 114,688 (fit tested only), and from 114,688 up every loaded step fails the fit test against the 855 + 121 MiB desktop allowance; the knee table’s spill onset, the first step judged against a near neighbour, is 122,880 meas. The per-card tables below put the same configuration fully resident at 93,184 per slot at a 1,796 MiB reserve der.
The largest window per card, per file and configuration
Each cell in the tables below is the largest per-slot -c that keeps the server fully resident: the large figure at this page’s 1,796 MiB desktop reserve (§02) and the small one at 1,272 MiB (a 555 MiB October desktop plus 0.7 GiB). Four desktop assumptions meet in this appendix: the tables use 1,796 and 1,272; the knee’s fully resident check uses 855 + 121 MiB; and the anchor check below uses the 555 MiB measured desktop with no buffer. A dagger (†) marks a cell beyond the deepest vision two-slot window measured (knee pick 8: 139,264 tokens per slot, 2026-10-03). Two cells are measured: Swift IQ4_XS text at 159,744, which the fill showed fully resident with 1,880 MiB left for a desktop, and Swift IQ4_XS vision at 139,264, which left 1,330 MiB and so appears as the small figure (both deep-filled 2026-10-04) meas; every other cell is der. Card totals: RTX 3090 24,576 MiB meas; RTX 5090 32,607, RTX 5080 16,303 and RTX 5070 12,227 MiB cited from user-posted nvidia-smi output. An RTX 3060 is assumed to have 61 MiB more than the 12 GB column and an RTX 4090 12 MiB less than the 24 GB column, enough to move a cell that sits at a floor by 1,024 tokens or more. Intel Arc cards (B580, Pro B70 and B50) and the B390 iGPU are not covered, because the constants were measured on CUDA; DGX Spark is not covered either, because its unified memory and operating system change the constant.
Text, 1 slot, drafter on (T1)
UNVERIFIED — DERIVED CONFIG. T1 (text, 1 slot, drafter on). Largest per-slot -c by card, fully resident: C = 2,044 MiB, 45.867 KiB per token, reserve 1,796 MiB for the desktop (this page’s threshold, §02); the small figure under each cell is the same window at a 1,272 MiB reserve (a 555 MiB October desktop + 0.7 GiB). Flags: -ngl 99 -ctk q8_0 -ctv q8_0 --parallel 1 -c N --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 (p-min costs no memory, so the windows are unchanged from n4/p0.75 der); the highlighted rows’ 24 GB cells run n-max 3, as the note under the tables explains.
| File | bpw | VRAM weights, MiB | 32 GB RTX 5090 32,607 MiB cited | 24 GB RTX 3090 24,576 MiB meas | 16 GB RTX 5080 16,303 MiB cited | 12 GB RTX 5070 12,227 MiB cited |
|---|---|---|---|---|---|---|
| bartowski IQ2_XXS | 2.60 | 7,938 | 262,144 max 262,144 max | 262,144 max 262,144 max | 100,352 112,640 | 9,216 21,504 |
| bartowski IQ2_XS | 2.66 | 8,136 | 262,144 max 262,144 max | 262,144 max 262,144 max | 96,256 107,520 | no fit (5,120) 16,384 |
| bartowski IQ2_S | 2.83 | 8,704 | 262,144 max 262,144 max | 262,144 max 262,144 max | 82,944 95,232 | no fit (−317 MiB) no fit (4,096) |
| bartowski IQ2_M | 3.08 | 9,503 | 262,144 max 262,144 max | 249,856 262,144 max | 65,536 76,800 | no fit (−1,117 MiB) no fit (−593 MiB) |
| bartowski Q2_K | 3.17 | 9,789 | 262,144 max 262,144 max | 243,712 256,000 | 59,392 70,656 | no fit (−1,402 MiB) no fit (−878 MiB) |
| bartowski IQ3_XXS | 3.60 | 11,057 | 262,144 max 262,144 max | 216,064 ‡ 227,328 | 30,720 43,008 | no fit (−2,670 MiB) no fit (−2,146 MiB) |
| bartowski Q3_K_S | 3.73 | 11,457 | 262,144 max 262,144 max | 206,848 218,112 | 21,504 33,792 | no fit (−3,070 MiB) no fit (−2,547 MiB) |
| bartowski IQ3_XS | 3.75 | 11,514 | 262,144 max 262,144 max | 205,824 217,088 | 20,480 32,768 | no fit (−3,128 MiB) no fit (−2,603 MiB) |
| bartowski Q3_K_M | 3.92 | 12,091 | 262,144 max 262,144 max | 192,512 203,776 | 8,192 19,456 | no fit (−3,704 MiB) no fit (−3,180 MiB) |
| bartowski Q3_K_L | 4.13 | 12,776 | 262,144 max 262,144 max | 177,152 188,416 | no fit (−313 MiB) no fit (4,096) | no fit (−4,389 MiB) no fit (−3,865 MiB) |
| bartowski IQ3_M | 4.35 | 13,480 | 262,144 max 262,144 max | 161,792 173,056 | no fit (−1,017 MiB) no fit (−493 MiB) | no fit (−5,093 MiB) no fit (−4,569 MiB) |
| bartowski IQ4_XS | 4.53 | 14,104 | 262,144 max 262,144 max | 159,744 meas filled 2026-10-04 at n4/p0.75, fully resident, leaving 1,880 MiB for a desktop; derived 147,456 / 159,744 | no fit (−1,642 MiB) no fit (−1,118 MiB) | no fit (−5,718 MiB) no fit (−5,194 MiB) |
| bartowski Q4_0 | 4.78 | 14,899 | 262,144 max 262,144 max | 130,048 141,312 | no fit (−2,436 MiB) no fit (−1,912 MiB) | no fit (−6,512 MiB) no fit (−5,988 MiB) |
| bartowski Q4_K_S | 4.79 | 14,913 | 262,144 max 262,144 max | 129,024 141,312 | no fit (−2,450 MiB) no fit (−1,926 MiB) | no fit (−6,526 MiB) no fit (−6,002 MiB) |
| bartowski IQ4_NL | 5.10 | 15,942 | 262,144 max 262,144 max | 106,496 117,760 | no fit (−3,479 MiB) no fit (−2,955 MiB) | no fit (−7,555 MiB) no fit (−7,031 MiB) |
| bartowski Q4_K_M | 5.10 | 15,942 | 262,144 max 262,144 max | 106,496 117,760 | no fit (−3,479 MiB) no fit (−2,955 MiB) | no fit (−7,555 MiB) no fit (−7,031 MiB) |
| bartowski Q4_1 | 5.22 | 16,231 | 262,144 max 262,144 max | 100,352 111,616 | no fit (−3,769 MiB) no fit (−3,245 MiB) | no fit (−7,845 MiB) no fit (−7,321 MiB) |
| bartowski Q4_K_L | 5.51 | 17,256 | 256,000 262,144 max | 76,800 89,088 | no fit (−4,794 MiB) no fit (−4,270 MiB) | no fit (−8,870 MiB) no fit (−8,346 MiB) |
| bartowski Q5_K_S | 5.73 | 17,816 | 243,712 256,000 | 64,512 76,800 | no fit (−5,354 MiB) no fit (−4,830 MiB) | no fit (−9,430 MiB) no fit (−8,906 MiB) |
| bartowski Q5_K_M | 6.12 | 19,110 | 215,040 226,304 | 35,840 47,104 | no fit (−6,648 MiB) no fit (−6,124 MiB) | no fit (−10,724 MiB) no fit (−10,200 MiB) |
| bartowski Q6_K_S | 6.69 | 20,793 | 177,152 189,440 | no fit (−58 MiB) 10,240 | no fit (−8,331 MiB) no fit (−7,807 MiB) | no fit (−12,407 MiB) no fit (−11,883 MiB) |
| bartowski Q6_K | 6.98 | 21,750 | 155,648 167,936 | no fit (−1,014 MiB) no fit (−490 MiB) | no fit (−9,287 MiB) no fit (−8,763 MiB) | no fit (−13,363 MiB) no fit (−12,839 MiB) |
| bartowski Q6_K_L | 7.30 | 22,798 | 133,120 144,384 | no fit (−2,062 MiB) no fit (−1,538 MiB) | no fit (−10,335 MiB) no fit (−9,811 MiB) | no fit (−14,411 MiB) no fit (−13,887 MiB) |
| bartowski Q8_0 | 8.52 | 26,469 | 51,200 62,464 | no fit (−5,733 MiB) no fit (−5,209 MiB) | no fit (−14,006 MiB) no fit (−13,482 MiB) | no fit (−18,082 MiB) no fit (−17,558 MiB) |
| GSQ-RCO IQ2_XS-mtp | 2.56 | 8,089 | 262,144 max 262,144 max | 262,144 max 262,144 max | 97,280 108,544 | no fit (6,144) 17,408 |
| GSQ-RCO IQ2_S-mtp | 2.81 | 8,764 | 262,144 max 262,144 max | 262,144 max 262,144 max | 81,920 94,208 | no fit (−377 MiB) no fit (3,072) |
| GSQ-RCO IQ3_XXS-mtp | 3.06 | 9,560 | 262,144 max 262,144 max | 248,832 261,120 | 64,512 75,776 | no fit (−1,174 MiB) no fit (−650 MiB) |
| GSQ-RCO IQ3_S-mtp | 3.55 | 11,160 | 262,144 max 262,144 max | 212,992 225,280 | 28,672 39,936 | no fit (−2,773 MiB) no fit (−2,249 MiB) |
‡ knee pick 5 (the same file, at n4/p0.75, which holds about 150 MiB more than this cell’s n-max 3) passed fully resident at 262,144 with a desktop allowance of 976 MiB (855 + 121); this cell assumes 1,796 MiB (large) or 1,272 MiB (small).
Vision, 1 slot (V1)
UNVERIFIED — DERIVED CONFIG. V1 (vision, 1 slot). Largest per-slot -c by card, fully resident: C = 2,929 MiB, 45.867 KiB per token, reserves 1,796 and 1,272 MiB as above. Flags: -ngl 99 -ctk q8_0 -ctv q8_0 --parallel 1 -c N --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf (p-min costs no memory, so the windows are unchanged from n4/p0.75 der).
| File | bpw | VRAM weights, MiB | 32 GB RTX 5090 32,607 MiB cited | 24 GB RTX 3090 24,576 MiB meas | 16 GB RTX 5080 16,303 MiB cited | 12 GB RTX 5070 12,227 MiB cited |
|---|---|---|---|---|---|---|
| bartowski IQ2_XXS | 2.60 | 7,938 | 262,144 max 262,144 max | 262,144 max 262,144 max | 80,896 92,160 | no fit (−436 MiB) no fit (1,024) |
| bartowski IQ2_XS | 2.66 | 8,136 | 262,144 max 262,144 max | 261,120 262,144 max | 76,800 88,064 | no fit (−634 MiB) no fit (−110 MiB) |
| bartowski IQ2_S | 2.83 | 8,704 | 262,144 max 262,144 max | 248,832 260,096 | 63,488 75,776 | no fit (−1,202 MiB) no fit (−678 MiB) |
| bartowski IQ2_M | 3.08 | 9,503 | 262,144 max 262,144 max | 230,400 242,688 | 46,080 57,344 | no fit (−2,001 MiB) no fit (−1,477 MiB) |
| bartowski Q2_K | 3.17 | 9,789 | 262,144 max 262,144 max | 224,256 235,520 | 39,936 51,200 | no fit (−2,287 MiB) no fit (−1,763 MiB) |
| bartowski IQ3_XXS | 3.60 | 11,057 | 262,144 max 262,144 max | 195,584 207,872 | 11,264 22,528 | no fit (−3,555 MiB) no fit (−3,031 MiB) |
| bartowski Q3_K_S | 3.73 | 11,457 | 262,144 max 262,144 max | 187,392 198,656 | no fit (2,048) 14,336 | no fit (−3,955 MiB) no fit (−3,431 MiB) |
| bartowski IQ3_XS | 3.75 | 11,514 | 262,144 max 262,144 max | 185,344 197,632 | no fit (1,024) 12,288 | no fit (−4,012 MiB) no fit (−3,488 MiB) |
| bartowski Q3_K_M | 3.92 | 12,091 | 262,144 max 262,144 max | 173,056 184,320 | no fit (−512 MiB) no fit (0) | no fit (−4,588 MiB) no fit (−4,065 MiB) |
| bartowski Q3_K_L | 4.13 | 12,776 | 262,144 max 262,144 max | 157,696 168,960 | no fit (−1,198 MiB) no fit (−674 MiB) | no fit (−5,274 MiB) no fit (−4,750 MiB) |
| bartowski IQ3_M | 4.35 | 13,480 | 262,144 max 262,144 max | 141,312 153,600 | no fit (−1,902 MiB) no fit (−1,378 MiB) | no fit (−5,978 MiB) no fit (−5,454 MiB) |
| bartowski IQ4_XS | 4.53 | 14,104 | 262,144 max 262,144 max | 128,000 139,264 meas (filled 2026-10-04 at n4/p0.75, leaving 1,330 MiB for a desktop) | no fit (−2,526 MiB) no fit (−2,002 MiB) | no fit (−6,602 MiB) no fit (−6,078 MiB) |
| bartowski Q4_0 | 4.78 | 14,899 | 262,144 max 262,144 max | 109,568 121,856 | no fit (−3,321 MiB) no fit (−2,797 MiB) | no fit (−7,397 MiB) no fit (−6,873 MiB) |
| bartowski Q4_K_S | 4.79 | 14,913 | 262,144 max 262,144 max | 109,568 121,856 | no fit (−3,335 MiB) no fit (−2,811 MiB) | no fit (−7,411 MiB) no fit (−6,887 MiB) |
| bartowski IQ4_NL | 5.10 | 15,942 | 262,144 max 262,144 max | 87,040 98,304 | no fit (−4,364 MiB) no fit (−3,840 MiB) | no fit (−8,440 MiB) no fit (−7,916 MiB) |
| bartowski Q4_K_M | 5.10 | 15,942 | 262,144 max 262,144 max | 87,040 98,304 | no fit (−4,364 MiB) no fit (−3,840 MiB) | no fit (−8,440 MiB) no fit (−7,916 MiB) |
| bartowski Q4_1 | 5.22 | 16,231 | 260,096 262,144 max | 79,872 92,160 | no fit (−4,653 MiB) no fit (−4,129 MiB) | no fit (−8,729 MiB) no fit (−8,205 MiB) |
| bartowski Q4_K_L | 5.51 | 17,256 | 236,544 248,832 | 57,344 68,608 | no fit (−5,678 MiB) no fit (−5,154 MiB) | no fit (−9,754 MiB) no fit (−9,230 MiB) |
| bartowski Q5_K_S | 5.73 | 17,816 | 224,256 235,520 | 45,056 56,320 | no fit (−6,238 MiB) no fit (−5,714 MiB) | no fit (−10,314 MiB) no fit (−9,790 MiB) |
| bartowski Q5_K_M | 6.12 | 19,110 | 195,584 206,848 | 16,384 27,648 | no fit (−7,532 MiB) no fit (−7,008 MiB) | no fit (−11,608 MiB) no fit (−11,084 MiB) |
| bartowski Q6_K_S | 6.69 | 20,793 | 157,696 168,960 | no fit (−942 MiB) no fit (−418 MiB) | no fit (−9,215 MiB) no fit (−8,691 MiB) | no fit (−13,291 MiB) no fit (−12,767 MiB) |
| bartowski Q6_K | 6.98 | 21,750 | 136,192 148,480 | no fit (−1,899 MiB) no fit (−1,375 MiB) | no fit (−10,172 MiB) no fit (−9,648 MiB) | no fit (−14,248 MiB) no fit (−13,724 MiB) |
| bartowski Q6_K_L | 7.30 | 22,798 | 112,640 124,928 | no fit (−2,946 MiB) no fit (−2,423 MiB) | no fit (−11,220 MiB) no fit (−10,696 MiB) | no fit (−15,296 MiB) no fit (−14,772 MiB) |
| bartowski Q8_0 | 8.52 | 26,469 | 30,720 43,008 | no fit (−6,618 MiB) no fit (−6,094 MiB) | no fit (−14,891 MiB) no fit (−14,367 MiB) | no fit (−18,967 MiB) no fit (−18,443 MiB) |
Vision, 2 slots (V2)
UNVERIFIED — DERIVED CONFIG. V2 (vision, 2 slots). Largest per-slot -c by card, fully resident: C = 3,677 MiB, 44.0 KiB per token, reserves 1,796 and 1,272 MiB as above. Flags: -ngl 99 -ctk q8_0 -ctv q8_0 --parallel 2 -c 2xN --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --mmproj mmproj-ukisai_Swift-1.5-Qwen3.8-27b-f16.gguf (p-min costs no memory, so the windows are unchanged from n4/p0.75 der).
| File | bpw | VRAM weights, MiB | 32 GB RTX 5090 32,607 MiB cited | 24 GB RTX 3090 24,576 MiB meas | 16 GB RTX 5080 16,303 MiB cited | 12 GB RTX 5070 12,227 MiB cited |
|---|---|---|---|---|---|---|
| bartowski IQ2_XXS | 2.60 | 7,938 | 223,232 † 229,376 | 129,024 135,168 | 32,768 38,912 | no fit (−1,184 MiB) no fit (−660 MiB) |
| bartowski IQ2_XS | 2.66 | 8,136 | 220,160 † 226,304 | 126,976 133,120 | 30,720 36,864 | no fit (−1,382 MiB) no fit (−858 MiB) |
| bartowski IQ2_S | 2.83 | 8,704 | 214,016 † 220,160 | 120,832 126,976 | 24,576 30,720 | no fit (−1,950 MiB) no fit (−1,426 MiB) |
| bartowski IQ2_M | 3.08 | 9,503 | 204,800 † 210,944 | 111,616 117,760 | 15,360 21,504 | no fit (−2,749 MiB) no fit (−2,225 MiB) |
| bartowski Q2_K | 3.17 | 9,789 | 201,728 † 207,872 | 107,520 113,664 | 11,264 17,408 | no fit (−3,035 MiB) no fit (−2,511 MiB) |
| bartowski IQ3_XXS | 3.60 | 11,057 | 186,368 † 192,512 | 93,184 * 99,328 | no fit (−227 MiB) no fit (3,072) | no fit (−4,303 MiB) no fit (−3,779 MiB) |
| bartowski Q3_K_S | 3.73 | 11,457 | 182,272 † 188,416 | 88,064 94,208 | no fit (−627 MiB) no fit (−103 MiB) | no fit (−4,703 MiB) no fit (−4,179 MiB) |
| bartowski IQ3_XS | 3.75 | 11,514 | 181,248 † 187,392 | 88,064 94,208 | no fit (−684 MiB) no fit (−160 MiB) | no fit (−4,760 MiB) no fit (−4,236 MiB) |
| bartowski Q3_K_M | 3.92 | 12,091 | 174,080 † 180,224 | 80,896 87,040 | no fit (−1,261 MiB) no fit (−737 MiB) | no fit (−5,337 MiB) no fit (−4,813 MiB) |
| bartowski Q3_K_L | 4.13 | 12,776 | 166,912 † 173,056 | 72,704 78,848 | no fit (−1,946 MiB) no fit (−1,422 MiB) | no fit (−6,022 MiB) no fit (−5,498 MiB) |
| bartowski IQ3_M | 4.35 | 13,480 | 158,720 † 164,864 | 64,512 70,656 | no fit (−2,650 MiB) no fit (−2,126 MiB) | no fit (−6,726 MiB) no fit (−6,202 MiB) |
| bartowski IQ4_XS | 4.53 | 14,104 | 151,552 † 157,696 | 57,344 63,488 | no fit (−3,274 MiB) no fit (−2,750 MiB) | no fit (−7,350 MiB) no fit (−6,826 MiB) |
| bartowski Q4_0 | 4.78 | 14,899 | 142,336 † 147,456 | 48,128 54,272 | no fit (−4,069 MiB) no fit (−3,545 MiB) | no fit (−8,145 MiB) no fit (−7,621 MiB) |
| bartowski Q4_K_S | 4.79 | 14,913 | 141,312 † 147,456 | 48,128 54,272 | no fit (−4,083 MiB) no fit (−3,559 MiB) | no fit (−8,159 MiB) no fit (−7,635 MiB) |
| bartowski IQ4_NL | 5.10 | 15,942 | 130,048 136,192 | 35,840 41,984 | no fit (−5,112 MiB) no fit (−4,588 MiB) | no fit (−9,188 MiB) no fit (−8,664 MiB) |
| bartowski Q4_K_M | 5.10 | 15,942 | 130,048 136,192 | 35,840 41,984 | no fit (−5,112 MiB) no fit (−4,588 MiB) | no fit (−9,188 MiB) no fit (−8,664 MiB) |
| bartowski Q4_1 | 5.22 | 16,231 | 125,952 132,096 | 32,768 38,912 | no fit (−5,402 MiB) no fit (−4,877 MiB) | no fit (−9,478 MiB) no fit (−8,953 MiB) |
| bartowski Q4_K_L | 5.51 | 17,256 | 114,688 120,832 | 20,480 26,624 | no fit (−6,426 MiB) no fit (−5,902 MiB) | no fit (−10,502 MiB) no fit (−9,978 MiB) |
| bartowski Q5_K_S | 5.73 | 17,816 | 107,520 113,664 | 14,336 20,480 | no fit (−6,986 MiB) no fit (−6,462 MiB) | no fit (−11,062 MiB) no fit (−10,538 MiB) |
| bartowski Q5_K_M | 6.12 | 19,110 | 93,184 99,328 | no fit (−8 MiB) no fit (5,120) | no fit (−8,281 MiB) no fit (−7,757 MiB) | no fit (−12,357 MiB) no fit (−11,833 MiB) |
| bartowski Q6_K_S | 6.69 | 20,793 | 73,728 79,872 | no fit (−1,690 MiB) no fit (−1,167 MiB) | no fit (−9,964 MiB) no fit (−9,440 MiB) | no fit (−14,040 MiB) no fit (−13,516 MiB) |
| bartowski Q6_K | 6.98 | 21,750 | 62,464 68,608 | no fit (−2,647 MiB) no fit (−2,123 MiB) | no fit (−10,920 MiB) no fit (−10,396 MiB) | no fit (−14,996 MiB) no fit (−14,472 MiB) |
| bartowski Q6_K_L | 7.30 | 22,798 | 50,176 56,320 | no fit (−3,695 MiB) no fit (−3,171 MiB) | no fit (−11,968 MiB) no fit (−11,444 MiB) | no fit (−16,044 MiB) no fit (−15,520 MiB) |
| bartowski Q8_0 | 8.52 | 26,469 | no fit (7,168) 13,312 | no fit (−7,366 MiB) no fit (−6,842 MiB) | no fit (−15,639 MiB) no fit (−15,115 MiB) | no fit (−19,715 MiB) no fit (−19,191 MiB) |
† beyond the deepest vision two-slot window measured (knee pick 8: 139,264 tokens per slot, 2026-10-03). * the configuration that anchors the vision constant (knee pick 8, run at n4/p0.75; p-min costs no memory); the knee figure shows its decode at that drafter.
Text, 1 slot, drafter off (T0)
UNVERIFIED — DERIVED CONFIG. T0 (text, 1 slot, drafter off). Largest per-slot -c by card, fully resident: C = 1,283 MiB, 40.867 KiB per token, reserves 1,796 and 1,272 MiB as above. Flags: -ngl 99 -ctk q8_0 -ctv q8_0 --parallel 1 -c N --spec-type none.
| File | bpw | VRAM weights, MiB | 32 GB RTX 5090 32,607 MiB cited | 24 GB RTX 3090 24,576 MiB meas | 16 GB RTX 5080 16,303 MiB cited | 12 GB RTX 5070 12,227 MiB cited |
|---|---|---|---|---|---|---|
| bartowski IQ2_XXS | 2.60 | 7,710 | 262,144 max 262,144 max | 262,144 max 262,144 max | 137,216 150,528 | 35,840 48,128 |
| bartowski IQ2_XS | 2.66 | 7,908 | 262,144 max 262,144 max | 262,144 max 262,144 max | 133,120 145,408 | 30,720 44,032 |
| bartowski IQ2_S | 2.83 | 8,476 | 262,144 max 262,144 max | 262,144 max 262,144 max | 118,784 131,072 | 16,384 29,696 |
| bartowski IQ2_M | 3.08 | 9,275 | 262,144 max 262,144 max | 262,144 max 262,144 max | 98,304 111,616 | no fit (−128 MiB) 9,216 |
| bartowski Q2_K | 3.17 | 9,561 | 262,144 max 262,144 max | 262,144 max 262,144 max | 91,136 104,448 | no fit (−413 MiB) no fit (2,048) |
| bartowski IQ3_XXS | 3.60 | 10,829 | 262,144 max 262,144 max | 262,144 max 262,144 max | 59,392 72,704 | no fit (−1,681 MiB) no fit (−1,157 MiB) |
| bartowski Q3_K_S | 3.73 | 11,229 | 262,144 max 262,144 max | 257,024 262,144 max | 49,152 62,464 | no fit (−2,082 MiB) no fit (−1,558 MiB) |
| bartowski IQ3_XS | 3.75 | 11,286 | 262,144 max 262,144 max | 254,976 262,144 max | 48,128 61,440 | no fit (−2,139 MiB) no fit (−1,615 MiB) |
| bartowski Q3_K_M | 3.92 | 11,863 | 262,144 max 262,144 max | 240,640 253,952 | 33,792 47,104 | no fit (−2,715 MiB) no fit (−2,191 MiB) |
| bartowski Q3_K_L | 4.13 | 12,548 | 262,144 max 262,144 max | 223,232 236,544 | 16,384 29,696 | no fit (−3,401 MiB) no fit (−2,877 MiB) |
| bartowski IQ3_M | 4.35 | 13,252 | 262,144 max 262,144 max | 205,824 219,136 | no fit (−28 MiB) 12,288 | no fit (−4,104 MiB) no fit (−3,580 MiB) |
| bartowski IQ4_XS | 4.53 | 13,876 | 262,144 max 262,144 max | 190,464 203,776 | no fit (−653 MiB) no fit (−129 MiB) | no fit (−4,729 MiB) no fit (−4,205 MiB) |
| bartowski Q4_0 | 4.78 | 14,671 | 262,144 max 262,144 max | 171,008 183,296 | no fit (−1,447 MiB) no fit (−923 MiB) | no fit (−5,523 MiB) no fit (−4,999 MiB) |
| bartowski Q4_K_S | 4.79 | 14,685 | 262,144 max 262,144 max | 169,984 183,296 | no fit (−1,462 MiB) no fit (−938 MiB) | no fit (−5,538 MiB) no fit (−5,014 MiB) |
| bartowski IQ4_NL | 5.10 | 15,714 | 262,144 max 262,144 max | 144,384 157,696 | no fit (−2,490 MiB) no fit (−1,966 MiB) | no fit (−6,566 MiB) no fit (−6,042 MiB) |
| bartowski Q4_K_M | 5.10 | 15,714 | 262,144 max 262,144 max | 144,384 157,696 | no fit (−2,490 MiB) no fit (−1,966 MiB) | no fit (−6,566 MiB) no fit (−6,042 MiB) |
| bartowski Q4_1 | 5.22 | 16,003 | 262,144 max 262,144 max | 137,216 150,528 | no fit (−2,780 MiB) no fit (−2,256 MiB) | no fit (−6,856 MiB) no fit (−6,332 MiB) |
| bartowski Q4_K_L | 5.51 | 17,028 | 262,144 max 262,144 max | 111,616 124,928 | no fit (−3,805 MiB) no fit (−3,281 MiB) | no fit (−7,881 MiB) no fit (−7,357 MiB) |
| bartowski Q5_K_S | 5.73 | 17,588 | 262,144 max 262,144 max | 97,280 110,592 | no fit (−4,365 MiB) no fit (−3,841 MiB) | no fit (−8,441 MiB) no fit (−7,917 MiB) |
| bartowski Q5_K_M | 6.12 | 18,883 | 262,144 max 262,144 max | 64,512 77,824 | no fit (−5,659 MiB) no fit (−5,135 MiB) | no fit (−9,735 MiB) no fit (−9,211 MiB) |
| bartowski Q6_K_S | 6.69 | 20,566 | 224,256 237,568 | 22,528 35,840 | no fit (−7,342 MiB) no fit (−6,818 MiB) | no fit (−11,418 MiB) no fit (−10,894 MiB) |
| bartowski Q6_K | 6.98 | 21,522 | 199,680 212,992 | no fit (−26 MiB) 12,288 | no fit (−8,299 MiB) no fit (−7,775 MiB) | no fit (−12,375 MiB) no fit (−11,851 MiB) |
| bartowski Q6_K_L | 7.30 | 22,570 | 174,080 187,392 | no fit (−1,073 MiB) no fit (−549 MiB) | no fit (−9,346 MiB) no fit (−8,822 MiB) | no fit (−13,422 MiB) no fit (−12,898 MiB) |
| bartowski Q8_0 | 8.52 | 26,038 | 87,040 100,352 | no fit (−4,542 MiB) no fit (−4,018 MiB) | no fit (−12,815 MiB) no fit (−12,291 MiB) | no fit (−16,891 MiB) no fit (−16,367 MiB) |
| GSQ-RCO IQ2_XS | 2.50 | 7,757 | 262,144 max 262,144 max | 262,144 max 262,144 max | 136,192 149,504 | 34,816 47,104 |
| GSQ-RCO IQ2_S | 2.75 | 8,432 | 262,144 max 262,144 max | 262,144 max 262,144 max | 119,808 133,120 | 17,408 30,720 |
| GSQ-RCO IQ3_XXS | 3.00 | 9,228 | 262,144 max 262,144 max | 262,144 max 262,144 max | 99,328 112,640 | no fit (−80 MiB) 10,240 |
| GSQ-RCO IQ3_S | 3.50 | 10,827 | 262,144 max 262,144 max | 262,144 max 262,144 max | 59,392 72,704 | no fit (−1,680 MiB) no fit (−1,156 MiB) |
In all four tables: every cell is derived from the per-token arithmetic below unless marked meas; the large figure assumes a 1,796 MiB desktop reserve, the small one 1,272 MiB; 'max' = the native 262,144; 'no fit (N)' = under 8,192 per slot, N what would be left; 'no fit (−X MiB)' = weights and constant alone exceed the budget. Highlighted rows are the two Swift files this page measured; in the 24 GB column their one-slot drafter-on commands carry n-max 3 / p-min 0 and Swift IQ3_XXS’s two-slot command n-max 4 / p-min 0, the pairs the 2026-10-08 sweep scored best (§06), as the §03 cards do (two slots were swept on Swift IQ3_XXS only). Their windows were computed at n-max 4, which holds about 150 MiB more than n-max 3, so they still fit. Hover a cell for its command. The other drafter-on commands run n-max 4 / p-min 0 der: those files, cards and slot counts were not swept, and p-min costs no memory, so each window computed at n-max 4 still fits. On the four files the sweep ran, p-min 0.75 at n-max 4 decoded shallow reasoning at 0.87–0.97 of p-min 0’s speed with one slot, reasoning and answers at 0.67–0.75 per slot with two (Swift IQ3_XXS: n4/p0 read 1.337 of it per slot on reasoning, 1.484 on answers), and 0.90–1.03 after a 57,540-token prefix; on UD-IQ4_XS deep answers it was slightly faster (n4/p0 read 0.973 of it, interval 0.969–0.996). So p-min 0 is the better default: measured on four files, derived for the rest.
The 12 and 16 GB bracket: which Swift files fit, and how tight the margin is
On a 12 GB card, only files at about 3 bits per weight or less fit, for text with one slot. The windows in this paragraph and the 16 GB one assume a 1,272 MiB desktop reserve (a 555 MiB desktop plus 0.7 GiB); at this page’s 1,796 MiB reserve each is smaller, as the tables above show. bartowski IQ2_XXS (2.60 bits per weight) holds a conservative window of 21,504 tokens with the drafter on, or 32,768 at the expected (measured-anchored) constant; drafter off, 48,128 conservative and 61,440 expected. bartowski IQ2_XS (2.66) holds 16,384 and 28,672 with the drafter on, 44,032 and 56,320 off. GSQ-RCO IQ2_XS (2.502 bits per weight standard file, 2.565 with the -mtp drafter head) holds 17,408 and 29,696 with the drafter on, or 47,104 and 60,416 drafter off. GSQ-RCO IQ3_XXS (2.999 bits per weight) fits only with the drafter off: 10,240 conservative and 23,552 expected. A standard Q3 at 3.60 bits per weight or larger does not fit in any configuration; GSQ-RCO IQ3_S (3.498 bits per weight) is 640 MiB short even drafter off at the expected constant. A window under 32,768 tokens leaves this model very little room to think (§09). Vision does not fit at the 1,272 or 1,796 MiB reserve; headless (717 MiB reserve), bartowski IQ2_XXS holds 14,336 tokens and bartowski IQ2_XS 9,216, both conservative. GSQ-RCO ships no vision projector, and its drafter needs the -mtp file.
bartowski IQ2_XXS, bartowski IQ2_XS and GSQ-RCO IQ2_XS are below 2.912 bits per weight, the floor below which this page recommends nothing for the base model (§08); that floor was measured on unsloth quantisations of Qwen3.8-27B, not on Swift-1.5 files. ukisai’s cards report mean KL divergence against Swift-1.5 BF16 of 0.2866 for ukisai’s own IQ2_XXS cited (bartowski’s file has a different importance matrix, so that figure does not transfer; below) and, on the same wikitext-2 protocol (100 × 512 tokens), 0.189979 for GSQ-RCO IQ2_XS and 0.097774 for GSQ-RCO IQ3_XXS, on the development set that informed their refinement, against 0.1523 for the standard IQ2_M and 0.1655 for Q2_K cited; GSQ-RCO IQ2_XS’s held-out C4-prose figure is 0.161594.
What -c 32768 needs against a 12,227 MiB card (RTX 5070), as an expected server footprint (dedicated + shared): bartowski IQ2_XXS with the drafter on needs 10,934 MiB der, leaving 1,293 for your desktop and about 127 MiB of load-to-load variation; drafter off, 9,785 MiB, leaving 2,442. GSQ-RCO IQ2_XS with the drafter on needs 11,085, leaving 1,142; drafter off, 9,832, leaving 2,395. GSQ-RCO IQ3_XXS drafter off needs 11,303, leaving 924; after the 127 MiB of variance, 797 MiB is left, more than the 258–539 MiB desktops of the October knee loads but less than the 841–855 MiB of the 2026-10-04 deep fills and every August reading (1,179–1,669 MiB, §02). Pass --parallel 1 explicitly. llama-server build 10502 defaults to 4 slots with a unified KV cache; with the drafter on that holds 4 × 748.125 = 2,992.5 MiB of recurrent state instead of 748.1, 2,244.4 MiB more, which costs 50,107 tokens der and removes every 12 GB drafter-on fit above.
On a 16 GB card (RTX 5080 at 16,303 MiB, the lowest cited 16 GB total; the same 1,272 MiB reserve), the bracket opens to files that do not fit at 12 GB at all. bartowski IQ2_XXS with the drafter on holds a conservative window of 112,640 tokens, expected 123,904; drafter off, 150,528 conservative and 163,840 expected. GSQ-RCO IQ3_XXS (2.999 bits per weight, above the 2.912 floor) holds 75,776 with the drafter on (conservative) and 87,040 expected, or 112,640 and 125,952 drafter off. Standard Q3 files (3.60–3.92 bits per weight) fit with the drafter on as well as off: bartowski IQ3_XXS holds 43,008 conservative and 54,272 expected with the drafter on, 72,704 and 86,016 off. No file in this bracket was measured for accuracy (§15), and the base model’s ladder found empty answers from the 2.481-bit rung (§08).
The per-token memory arithmetic: how each cell is computed
Every cell in the per-card tables follows one formula. The card’s total VRAM, minus the desktop reserve, minus the weights, minus a fixed constant that depends on the configuration (text or vision, 1 or 2 slots, drafter on or off), is divided by the per-token slope and the number of slots, floored to 1,024 tokens and capped at the native 262,144. If the result falls below 8,192 the configuration does not fit.
per_slot = floor_to_1024(
(card_MiB - reserve - W - C_cfg) * 1024 / kib_cfg / slots
), capped at 262,144; 'no fit' if under 8,192
W = file bytes - token_embd - GGUF header
basis = llama-server dedicated + shared VRAM (MiB)
Weights (W). The file’s size minus the token embedding layer (which llama.cpp places in host RAM, not on the card — llama-model.cpp:1605–1607) minus the GGUF header. The parameter count is 27,320,697,856 (MTP layer 424,699,392); bits per weight = (file bytes − header) × 8 ÷ the parameters in the file (the GSQ-RCO standard files carry no MTP layer). The measured file-to-file delta (Q4_K_M against UD-IQ4_XS, -c 32768, n4, August 2026) matched CPU placement exactly: 2,278 MiB measured against 2,278.34 predicted with the embedding on the CPU, where GPU placement predicts 2,439.38 meas.
Per-token slope. The KV floor by arithmetic is 38 KiB per token: 16 full-attention layers × (K+V) × 1,024 values × 1.0625 B (q8_0) = 34 KiB, plus 4 KiB for the MTP layer in f16 der. The measured slopes sit above that floor because buffers add to it. At 1 slot the slope is 45.87 KiB per token der, the maximum of three readings on this RTX 3090: an August 2026 load-to-load delta at 45.69 and an August vision fill series at 45.77 (unsloth files of the base model), and an August-to-October transplant at 45.87 der: an August UD-IQ4_XS load at 32,768 tokens carried to the October at-load reading of Swift IQ3_XXS at 253,952 tokens (knee pick 5), with the two files’ weights difference taken out. At 2 slots the slope is 44 KiB per token (total tokens) meas, measured on knee pick 8 (Swift IQ3_XXS) at 43.84 at load and 43.85 at depth, 16,384 to 114,688 per slot. With the drafter off the slope is 40.87 KiB der, 5,120 B per token less (§05).
Recurrent state. 48 recurrent (Gated DeltaNet) layers, each carrying 786,432 values of state and 30,720 values of convolution state, at 4 bytes: 149.625 MiB per recurrent row der. With --spec-type draft-mtp at --spec-draft-n-max 4, llama.cpp keeps 1 + n_max = 5 rows per sequence — 748.125 MiB per slot. Measured check (August 2026, UD-IQ4_XS, one slot, -c 32768, dedicated memory): n_max 4 to n_max 10 added 898 MiB meas n=1 against 6 rows at 149.625 = 897.75 predicted; n_max 2 to n_max 3 added 150 against 1 row at 149.6; n_max 4 to n_max 6 added 298 against 2 rows at 299.2.
The constant per configuration. Each configuration has a fixed constant, C, that absorbs everything the formula does not model per token: compute buffers, driver overhead, the vision projector. The anchor is measured on the RTX 3090; every other card assumes the same constant.
- C_V2 = 3,677.09 MiB der — anchored on a measured footprint of knee pick 8 (Swift IQ3_XXS, vision, two slots): 16,142 − W(Swift IQ3_XXS) 11,056.91 − 44 × 32,768 / 1,024 = 3,677.09 (15,930 dedicated + 212 shared at about 14,000 tokens per slot with a 1440p image).
- C_V1 = 2,928.96 MiB der — C_V2 minus one slot’s recurrent state: 3,677.09 − 748.125 = 2,928.96.
- C_T1 = 2,044.33 MiB der — C_V1 minus the F16 vision projector file (884.64 MiB): 2,928.96 − 884.64 = 2,044.33. Nothing else is subtracted, so C_T1 is conservative: the vision-encoder compute buffer (about 250 MiB) and the image-processing growth (about 386 MiB) stay inside it.
- C_T1x = 1,528.18 MiB (expected) der — the measured-anchored constant, the maximum of 3 readings: 1,407.18 (August 2026, UD-IQ4_XS n4
-c 32768at load plus 68 growth), 1,389.99 (August 2026, UD-IQ4_XS n4-c 131072at 90,862 depth), and 1,528.18 (Swift IQ3_XXS n4-c 253952at load plus 189 deep-fill growth). - C_T0 = 1,283.44 MiB der — C_T1 minus the drafter fixed cost without MTP weights (760.88 MiB, from paired on/off measurements at
-c 32768: UD-IQ4_XS 761.25, Q4_K_M 760.88). T0 weights also drop the MTP layer. - C_T0x = 767.30 MiB (expected) der.
Reserve. The tables print the large figure at 1,796 MiB, this page’s own desktop threshold (the August resting-desktop worst case of 1,669 plus 127 of variance), and the small figure at 1,272 MiB — a 555 MiB desktop measured 2026-10-03 on this rig, plus a 717 MiB buffer. A headless card (717) gains what the sensitivities below give.
Assumptions — each changes the cells if wrong:
- token_embd stays in host RAM. llama.cpp’s source says so (llama-model.cpp:1605–1607). The measured file-to-file delta confirmed CPU placement: 2,278 against 2,278.34 predicted, where GPU placement predicts 2,439.38. If it were on the GPU with the constant held fixed, the cells would lose 11,632 to 28,762 tokens, depending on the embedding type (sensitivity below).
- The footprint basis is dedicated + shared. Shared memory at rest (132–622 MiB at 1 slot) is host-pinned, so counting it as card memory is conservative.
- The 1-slot slope rests on the August 2026 deep fills (up to 218,233 tokens, unsloth files) plus one October at-load point: the arithmetic uses the at-load readings of two October text-only loads that a 32 GB job memory cap killed during the fill (their 38 GB re-runs, in the knee table, are not used by the formula).
- C_T1 is conservative. It keeps about 250 MiB of vision-encoder compute buffer and about 386 MiB of image growth that text-only use does not need. C_T1x removes them, anchored on 3 measurements.
- Drafter off does not load the MTP layer. The T0 total equals T1 minus the measured on/off pair (760.88 MiB fixed plus the MTP weights).
- Other cards behave like this one. Same driver overhead, same compute buffers at
-b 2048-ub 512. A larger-ub, another OS or another backend changes the constant. - Card totals are nvidia-smi readings. Only the RTX 3090 (24,576 MiB) is measured. The RTX 5090 (32,607), RTX 5080 (16,303) and RTX 5070 (12,227) are cited from user-posted nvidia-smi output in GitHub issues. The 16 GB column uses the lowest cited 16 GB total; the RTX 4080, 4070 Ti SUPER, 5060 Ti 16 GB and 4060 Ti 16 GB read 16,311–16,380 (cited or assumed), so it is conservative for them. An RTX 4090 is assumed to read 24,564, 12 MiB under the 3090.
- The drafter runs at n_max 4 and
--parallelis set explicitly. n_max 10 costs 897.75 MiB more per slot — 6 extra state rows; n_max 3, on the highlighted rows’ 24 GB one-slot cells, holds 149.625 MiB less (about 3,340 tokens at the 1-slot slope), so those cells are conservative der. The 4-slot default costs 2,244.4 MiB more at 1-slot use. - The reserve is the desktop you run. The large figure assumes 1,796 MiB and the small one 1,272; a heavier desktop than either shrinks every cell (sensitivities below).
Anchor check. The configuration that anchors the vision constant is Swift IQ3_XXS, vision, two slots, on the 24 GB RTX 3090 (knee pick 8) meas. With a reserve of 555 MiB (the measured desktop) and no buffer, the formula predicts 108,067 per slot, floored to 107,520. Dedicated plus shared memory at load grows at 43.84 KiB per token (total tokens) from 16,384 per slot, and dedicated memory stops at 23,670 MiB, its level at 114,688 with a 510 MiB desktop (it rises only 20 MiB more by 122,880). No step between 16,384 and 114,688 was loaded, so the spill start can only be bracketed der: about 111,500 per slot if the whole 274 MiB rise in shared memory at 114,688 were overflow, and about 113,700 if shared memory also grew at the 16 MiB per 8,192-token step this file shows below its spill at two slots (pick 6). The derived window lands 3.6–5.4% below that range and 12.5% below 122,880, the knee table’s spill onset (the first step judged against a near neighbour), so the formula is conservative. At the 1,272 MiB reserve the predicted window is 99,328. Decode, the knee’s test, held to 131,072 at 22.94 t/s and to 133,120 at 23.12 n=1, then collapsed to 11.59 at 135,168 n=1 and 5.55 at 139,264 (re-run 5.45); the decode-verified window is 123,904 per slot.
token_embd sensitivity. If the embedding layer sat on the GPU, the cells would move by an amount that depends on its type. With the constant held fixed, 1-slot cells lose 15,227 tokens at the Q4_K class (682.03 MiB embedding) and 14,381 at the IQ4_XS class (644.14 MiB). With the constant re-fitted, the Q3_K and IQ4_XS classes gain (+3,595 and +846 at 1 slot), the Q4_K class is unchanged, and the Q4_1 to Q8_0 classes lose 1,692 to 13,535. The file-to-file delta (2,278 against 2,278.34 predicted for CPU placement and 2,439.38 for GPU) rules GPU placement out, and the August-to-October transplant lands within 39 MiB.
Sensitivities. Moving the reserve from 1,272 to 1,796 MiB shifts the budget by −524 MiB: −11,699 tokens per slot at 1 slot, −6,097 at 2 slots. Moving to headless (717) adds 555 MiB: +12,391 tokens at 1 slot, +6,458 at 2. Setting --spec-draft-n-max to 10 instead of 4 adds 897.75 MiB per slot (898 measured): −20,043 tokens at 1 slot, −20,893 per slot at 2. Omitting --parallel 1 — the 4-slot default — holds 2,992.5 MiB of recurrent state instead of 748.1, 2,244.4 MiB more: −50,107 tokens off every 1-slot cell. ukisai’s files need the same weight memory (W) as bartowski’s for every quant except Q8_0, which differs by −66.09 MiB, adding 1,476 tokens at 1 slot.
Take bartowski’s files by default; take ukisai’s when you need published fidelity numbers for the exact file you run
Download from bartowski unless you specifically need ukisai’s per-file KLD figures. Both repos carry the same recipe in 21 of their 22 shared tiers meas — the same tensor names, shapes, types and byte sizes (866 tensors per LLM file) — so every window and memory figure on this page applies to both repos’ files at any of those tiers; decode speed should match with the drafter off, because the bytes read per token are identical (derived, not measured; between different files the size-based prediction missed: 0.917 and 1.175 of the anchor’s speed predicted, 1.005 and 1.034 measured, §06), while drafter-on speed also depends on draft acceptance, measured here only on bartowski’s files. A measured case of the same arithmetic, on two fine-tunes rather than two repos: bartowski’s IQ4_XS files of Swift-1.5 and of bytkim’s Qwen3.8-27B-pi have identical tensor names, shapes and types, and with the drafter off they decoded within 0.2% of each other (ratios 1.001 and 1.002, two passes each, 2026-10-07) der; each deep fill reached the same board VRAM peak for both (24,019 and 23,429 MiB) meas, and with the drafter on the ratio moved between 0.919 and 1.037 with content and token regime (Appendix B). bartowski’s importance matrix was computed on Swift-1.5, the model inside the file (a reasoned conclusion from five converging pieces of evidence; below); his metadata is correct (name, licence, base model); his files can be rebuilt from the published layouts, imatrix, calibration text and llama.cpp release; and this page’s own two Swift files are bartowski’s (their sha256, read 2026-10-04, matches his published files meas). Choose ukisai when you want published KLD against Swift-1.5 BF16 for every tier you run — ukisai’s card gives that for all 22 tiers at 512 and 32,768 context cited, and those numbers do not transfer to bartowski’s files because the imatrix differs — or when you want a Q8_0 that is 66.09 MiB smaller (ssm_alpha and ssm_beta stored at Q8_0 rather than F32; the quality effect is not measured). No measurement here ranks the two repos' answer quality.
Every shared tier differs by exactly 640 B, all of it in the header, except Q8_0 which differs by 69,305,280 B. The table below lists every tier in both repos. Sizes meas are from the Hugging Face file listing and HTTP Range reads on 2026-10-03; for every parsed file, tensor bytes plus header equal the file size exactly. GB is decimal (109), as on the download pages. GiB is binary (230), the unit a memory budget uses. The GiB column is for the bartowski file; ukisai is the same to two decimals except for Q8_0 (27.05 GiB).
| Quant | ukisai bytes | ukisai GB | bartowski bytes | bartowski GB | GiB | Difference |
|---|---|---|---|---|---|---|
| IQ2_XXS | 8,881,268,416 | 8.88 | 8,881,269,056 | 8.88 | 8.27 | +640 B |
| IQ2_XS | 9,088,526,016 | 9.09 | 9,088,526,656 | 9.09 | 8.46 | +640 B |
| IQ2_S | 9,684,043,456 | 9.68 | 9,684,044,096 | 9.68 | 9.02 | +640 B |
| IQ2_M | 10,522,248,896 | 10.52 | 10,522,249,536 | 10.52 | 9.80 | +640 B |
| Q2_K | 10,821,543,616 | 10.82 | 10,821,544,256 | 10.82 | 10.08 | +640 B |
| IQ3_XXS | 12,320,167,616 | 12.32 | 12,320,168,256 | 12.32 | 11.47 | +640 B |
| Q3_K_S | 12,739,884,736 | 12.74 | 12,739,885,376 | 12.74 | 11.86 | +640 B |
| IQ3_XS | 12,799,604,416 | 12.80 | 12,799,605,056 | 12.80 | 11.92 | +640 B |
| Q3_K_M | 13,404,051,136 | 13.40 | 13,404,051,776 | 13.40 | 12.48 | +640 B |
| Q3_K_L | 14,122,817,216 | 14.12 | 14,122,817,856 | 14.12 | 13.15 | +640 B |
| IQ3_M | 14,860,629,696 | 14.86 | 14,860,630,336 | 14.86 | 13.84 | +640 B |
| IQ4_XS | 15,475,951,296 | 15.48 | 15,475,951,936 | 15.48 | 14.41 | +640 B |
| Q4_0 | — | — | 16,348,768,576 | 16.35 | 15.23 | bartowski only |
| Q4_K_S | 16,363,513,536 | 16.36 | 16,363,514,176 | 16.36 | 15.24 | +640 B |
| IQ4_NL | 17,442,399,936 | 17.44 | 17,442,400,576 | 17.44 | 16.24 | +640 B |
| Q4_K_M | 17,442,399,936 | 17.44 | 17,442,400,576 | 17.44 | 16.24 | +640 B |
| Q4_1 | — | — | 17,825,458,496 | 17.83 | 16.60 | bartowski only |
| Q4_K_L | 18,820,703,936 | 18.82 | 18,820,704,576 | 18.82 | 17.53 | +640 B |
| Q5_K_S | 19,566,913,216 | 19.57 | 19,566,913,856 | 19.57 | 18.22 | +640 B |
| Q5_K_M | 20,923,877,056 | 20.92 | 20,923,877,696 | 20.92 | 19.49 | +640 B |
| Q6_K_S | 22,857,455,296 | 22.86 | 22,857,455,936 | 22.86 | 21.29 | +640 B |
| Q6_K | 23,860,565,696 | 23.86 | 23,860,566,336 | 23.86 | 22.22 | +640 B |
| Q6_K_L | 24,958,908,096 | 24.96 | 24,958,908,736 | 24.96 | 23.24 | +640 B |
| Q8_0 | 29,047,084,416 | 29.05 | 29,116,389,696 | 29.12 | 27.12 | +69,305,280 B |
| mmproj F16 | 927,606,912 | 0.93 | 927,607,552 | 0.93 | 0.86 | +640 B |
| mmproj BF16 | — | — | 931,146,496 | 0.93 | 0.87 | bartowski only |
| bf16 (2-part split) | — | — | 54,657,734,816 | 54.66 | 50.90 | bartowski only |
IQ4_NL and Q4_K_M are the same size on purpose: both formats store 4.5 bits per weight, and the two files hold the same 11,449,958,400 bytes of 4-bit tensors, typed IQ4_NL in one and Q4_K in the other; every other type total matches meas. This page’s bpw is (file − header) × 8 ÷ 27,320,697,856 parameters (the parameters in the file, MTP layer included). bartowski’s card prints "File bits/weight" about 1.5% lower (IQ3_XXS 3.55 against 3.60, IQ2_XXS 2.56 against 2.60), consistent with a larger parameter count in the divisor, but the card does not state its count.
Files only one repo has. bartowski adds Q4_0, Q4_1, bf16 (split into 2 parts: 39,960,210,144 + 14,697,524,672 B), mmproj BF16 (931,146,496 B), the imatrix (ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf, 13,642,688 B), the calibration text (ukisai_Swift-1.5-Qwen3.8-27b-calibration-v6.txt, 1,258,850 B), and a .layout.json and .tensor-types.txt for each of the 21 computed-layout files (none for Q8_0, Q4_0, Q4_1 or bf16). ukisai has no file without a bartowski counterpart; it ships one projector (mmproj F16) and no imatrix, calibration file or layout files.
In 21 of 22 shared tiers the tensor layout is the same recipe meas — the same tensor names, shapes, types and byte sizes (866 tensors per LLM file, 334 per projector), each pair differing by exactly 640 B, all of it in the header. In IQ3_XXS, 627 B of differing metadata, padded to 32-B alignment, moves data_start from 10,995,392 to 10,996,032; the tensor-data regions are the same length. The table below summarises what was compared and what differs.
| What was compared | Result (21 imatrix tiers; per-tensor samples: the IQ3_XXS pair) | Result (Q8_0) |
|---|---|---|
| Tensor names, shapes, types, byte sizes | Identical in all 866 tensors | 770 of 866 identical; 96 differ |
| sha256 of the first 1 MiB per tensor (F32 tensors: output_norm, blk.0.attn_norm, blk.0.ssm_a, ssm_alpha/ssm_beta, ssm_conv1d, ssm_dt.bias) | Bit-identical | Bit-identical for F32-in-both tensors |
| sha256 of the first 1 MiB per tensor (imatrix-quantized: output.weight, token_embd.weight, blk.40.ffn_down.weight) | Differ | Identical (Q8_0 does not use an imatrix) |
| Start of data region (1 MiB at offset 0) | Differs | Matches |
| mmproj F16 (3 samples at 0%, 37%, 81% of data region; 334 tensor types) | Identical | |
Q8_0 differs in 96 tensors. They are blk.N.ssm_alpha.weight and blk.N.ssm_beta.weight in the 48 recurrent blocks (0–63 minus every 4th). Each is [5120 × 48] = 245,760 values. ukisai stores each as Q8_0 (261,120 B) and bartowski stores each as F32 (983,040 B) — a data difference of +69,304,320 B (66.09 MiB) plus a header difference of +960 B because ukisai’s Q8_0 has no quantize.* keys, totalling exactly 69,305,280 B. The other 770 tensors match in name, shape, type and size; in every other file in both repos, ssm_alpha and ssm_beta are F32.
Tokenizer values (including the embedded chat template) and whole-file hashes were not compared (§15). What was compared: tensor names, shapes, types, byte sizes, per-tensor sha256 samples, metadata keys and the 10 tokenizer key names.
ukisai used bartowski’s Swift-1.0 imatrix, and bartowski used a Swift-1.5 imatrix — a reasoned conclusion from five converging observations, not a measurement. bartowski’s card says he used his own Swift-1.5 imatrix (calibration-v6, 583 chunks of 512 tokens) cited, and his metadata names ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf, the file in his 1.5 repo. ukisai’s imatrix path names ukisai_Swift-Qwen3.8-27b with no "1.5", its dataset path uses bartowski’s /models_out/ convention, and both file names exist in bartowski’s Swift-1.0 repo (created 2026-09-12). ukisai’s own card says the tiers "reuse the importance matrix and the per-tensor type layouts that bartowski computed for Swift-1.0" cited. ukisai’s initial-release commit (2026-09-24 16:20 UTC) is about 14 hours before bartowski’s Swift-1.5 upload began (2026-09-25 06:03 UTC), so bartowski’s published 1.5 imatrix did not exist yet. The measured data pattern fits: tensors quantized with an imatrix differ between the two repos, and those quantized without one are identical, consistent with both repos using a common BF16 source (ukisai: llama.cpp 6f41ac5 from BF16 cited; bartowski: release b11159, commit 224b8b72 cited).
An app that pairs a projector automatically may match it by general.basename, a reasoned inference from one ukisai commit about Unsloth Studio: ukisai’s is Miroslav and bartowski’s is Swift-1.5-Qwen3.8 meas. In such an app, take both files from the same repo. With llama-server --mmproj, either repo’s F16 projector carries the same sampled tensor data meas (3 samples at 0%, 37% and 81% of the data region, and all 334 tensor types identical).
ukisai’s separate GSQ-RCO repo ships four smaller tiers, 8.4–12.1 GB, each with and without a draft head, text only (no projector ships). These files are built on ISTA-DASLab’s per-tensor type allocation with a refinement re-run on Swift-1.5 cited: the allocation comes from ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO release, forced via llama-quantize --tensor-type-file (llama.cpp fee39dd), with an importance matrix from a mixed corpus (309 chunks × 4,096 tokens, 8 shards merged) and a refinement pass (refine_gguf_fast.py, 512 sequences of 2,048 tokens from the same corpus, 3 epochs). GSQ-RCO tier names are allocation profiles, not standard llama.cpp tiers: GSQ-RCO "IQ3_XXS" is 10.09 GB, against 12.32 GB for the standard IQ3_XXS. The -mtp file adds the MTP head (332.33 MiB); the card says the MTP files "require a runtime with support for this model's MTP implementation" and that "MTP decoding speed and quality have not been separately evaluated" cited.
| GSQ-RCO tier | Standard bytes | GB | GiB | bpw | -mtp bytes | GB | GiB | bpw | token_embd | Card KLD (dev) |
|---|---|---|---|---|---|---|---|---|---|---|
| IQ2_XS | 8,422,841,632 | 8.42 | 7.84 | 2.502 | 8,771,311,616 | 8.77 | 8.17 | 2.565 | IQ1_M 265.23 MiB | 0.189979 |
| IQ2_S | 9,259,511,072 | 9.26 | 8.62 | 2.751 | 9,607,981,056 | 9.61 | 8.95 | 2.81 | IQ2_S 388.38 MiB | 0.134751 |
| IQ3_XXS | 10,094,357,792 | 10.09 | 9.40 | 2.999 | 10,442,827,776 | 10.44 | 9.73 | 3.055 | IQ2_S 388.38 MiB | 0.097774 |
| IQ3_S | 11,771,546,912 | 11.77 | 10.96 | 3.498 | 12,120,016,896 | 12.12 | 11.29 | 3.546 | IQ2_S 388.38 MiB | 0.051265 |
The maker’s mean KL divergence against Swift-1.5 BF16 cited is on the development set (wiki.test.raw, 100 × 512 tokens) — that set "informed refinement and is not independent validation" — with a held-out C4-prose figure beside it:
| GSQ-RCO tier | Development-set KLD | Held-out C4-prose KLD |
|---|---|---|
| IQ2_XS | 0.189979 | 0.161594 |
| IQ2_S | 0.134751 | 0.107438 |
| IQ3_XXS | 0.097774 | 0.080238 |
| IQ3_S | 0.051265 | 0.041748 |
The same maker’s card for its standard files, on the same 100 × 512 wikitext-2 protocol, reports 0.0742 for the standard IQ3_XXS (12.32 GB) and 0.1523 for the standard IQ2_M (10.52 GB, about the size of the GSQ-RCO IQ3_XXS draft-head file at 10.44 GB) cited. All of these are the maker’s own measurements, not reproduced here, and the GSQ-RCO development-set figures are on the set that tuned them. The 12 GB windows for these files are above, in the 12 and 16 GB bracket.
bartowski’s IQ4_XS of bytkim’s Qwen3.8-27B-pi runs on the 24 GB RTX 3090 in the October launcher’s Swift IQ4_XS vision configuration (139,264 tokens, one slot), and the five auto-scored sets of the 175-prompt suite detect no quality difference from Swift IQ4_XS or the anchor — measured on the 24,576 MiB RTX 3090 on 2026-10-06/07 with llama.cpp build 10502 (commit 0adcc3bb5). Qwen3.8-27B-pi is a fine-tune of Qwen3.8-27B for the Pi coding agent, published under Apache-2.0 cited; bartowski’s IQ4_XS of it is 192 B larger than his Swift IQ4_XS and has the same tensor layout, tokenizer and chat template meas. On the 175-prompt suite (greedy, drafter off, single-turn) its 5-set composite is 80.6 at xhigh and 81.2 at medium, against 81.3 for the anchor and 82.7 for Swift IQ4_XS at xhigh meas; every paired interval includes zero, so there is no difference the suite can detect at 25 questions per set n=25; ALPACA and MT-Bench were not judged. Against Swift IQ4_XS the intervals only just include zero (upper ends +0.2) and both point estimates favour Swift, so the suite cannot rule out a gap of up to about 5 points. With the drafter off it decodes at Swift IQ4_XS’s speed (ratio 1.001 der); at the launcher’s 139,264-token vision window, filled to 126,838 tokens, it held decode at 0.974 of Swift IQ4_XS’s speed in the same sweep der n=2 (answer tokens, greedy, the n4/p0.75 drafter; the launch command below runs n4/p0, the same n-max and so the same memory). In the Pi agent it passed a three-probe smoke test n=1. No agent-task benefit from the fine-tune was measured. To run it on 24 GB, copy the launch command (IQ4_XS, medium, n-max 4 / p-min 0); 16 and 12 GB cards are after it.
The Pi file joins the comparison through Swift IQ4_XS, not through the anchor. Its speed and window steps ran Swift IQ4_XS (and, for one two-slot window, Swift IQ3_XXS) beside the Pi files in one sweep, so its speeds are published only as ratios to Swift, read pass by pass, and a ratio counts only when its two passes agree within 3%; every one did. Absolute decode on this rig moves about 13% between server starts (§03), so the absolute t/s below are not comparable with any other sweep on this page, the 2026-10-04 Swift figures included. Within the sweep, every file and drafter setting is its own server start in each pass, so the two-level effect can touch any ratio: a level change in one pass would break the 3% agreement and void the cell, and none was voided. For the drafter result the tokens kept per verify pass repeat exactly from start to start, and Swift IQ4_XS reproduces the result, so it is not a level artefact. Its suite cells are paired item for item with the anchor’s and Swift IQ4_XS’s cells of 2026-10-04, reused rather than re-run (§17). A difference against the anchor mixes fine-tune and quantizer, because the anchor is unsloth’s file with other tensor types; Pi IQ4_XS and Swift IQ4_XS share bartowski’s recipe and tensor types, so a difference between those two sits in the weights der. The study was pre-registered on 2026-10-06 before any measurement, except the sampled drafter follow-up (below), and its fixed 36 GB job memory cap was replaced by a recorded per-step cap after the first start was refused (30 or 31 GB, recorded with each step). The speed, window and follow-up steps ran on a quiet host; the server checks, the Pi smoke test and the start of the xhigh suite ran while a model download was in progress (below).
The model author’s agent results are the author’s, not this page’s. bytkim reports Terminal-Bench 2.1, GPQA Diamond and SciCode results, including that Pi’s medium setting matches the base model’s xhigh completion rate on Terminal-Bench 2.1 with about 41% fewer output tokens cited (model card, blog post). None of those was re-run here, and none is a GGUF result: the author ran them on FP8 weights cited. The author’s own GGUF agent smoke test, at medium with a 16,384-token output cap on subsets of the same benchmarks, ran 15 to 20 points below FP8 by the research notes’ reading, and the author calls it a smoke test, not a quant ranking cited. The research notes written before these measurements read the Terminal-Bench 2.1 scores as best observed across at most 3 attempts, on a task set also used to select the training checkpoint, with every Pi-against-base gap within about one standard error; the 41% compares medium with the base model’s xhigh, and at the same level the author’s own figures give about 20% (§15). This page measured fit, speed, single-turn suite quality and a smoke test, not agent-task success, turns, tool use or agent wall time (below).
The file has Swift IQ4_XS’s tensor layout, tokenizer and template, so Appendix A’s memory arithmetic applies to it
Take bartowski’s files: they embed the MTP draft head this page’s drafter flags use, and their tensor layout matches the Swift files this page already measured. bartowski’s IQ4_XS and IQ3_XXS of Qwen3.8-27B-pi are each 192 B larger than his Swift-1.5 file of the same quant, and in both pairs the tensor names and shapes (866 tensors, last block 64), the per-tensor quant types, the tokenizer arrays and the chat template are identical meas; each file embeds an MTP head of the same size and types as Swift’s (227.91 MiB, F32 + Q4_0) meas. Against lmstudio-community’s Q4_K_M of the base model, the tensor manifest, tokenizer arrays and template are identical and only the tensor types differ, plus one converter array (qwen35.attention.recurrent_layers) on which the Swift files differ in the same way meas. Each Pi file was checked by size and by its Hugging Face LFS sha256 before use. bytkim’s own GGUFs ship the MTP head as a separate file for --model-draft cited, so this page’s --spec-type draft-mtp flags, which use an embedded head, were run on bartowski’s files only.
| File (bartowski) | Bytes meas | GB | GiB der | Against bartowski’s Swift-1.5 file of the same quant meas |
|---|---|---|---|---|
| Qwen3.8-27B-pi IQ4_XS | 15,475,952,128 | 15.48 | 14.41 | +192 B (Swift IQ4_XS: 15,475,951,936); tensor names, shapes and types, tokenizer arrays, MTP head and chat template identical |
| Qwen3.8-27B-pi IQ3_XXS | 12,320,168,448 | 12.32 | 11.47 | +192 B (Swift IQ3_XXS: 12,320,168,256); the same checks identical |
| mmproj f16 (vision projector) | 927,607,712 | 0.93 | 0.86 | not compared |
Headers read offline (no GPU), the IQ4_XS on 2026-10-06 and the IQ3_XXS on 2026-10-07 after its download finished, on files verified by size and Hugging Face LFS sha256. GB is decimal (109), GiB binary (230). “Identical” means structure and types, not weight values. Only the IQ4_XS and IQ3_XXS pairs were compared; no other Pi quant was read. Whether bartowski’s embedded MTP head differs from the upstream one was not checked; bytkim’s notice describes its own separate MTP files as quantised upstream heads, “not a Pi-fine-tuned draft” cited.
Because the structure matches and the files differ by 192 B, the per-token memory and every window cell that Appendix A gives for bartowski’s IQ4_XS and IQ3_XXS apply to the Pi files unchanged der. The deep fills below check three configurations directly, and at each IQ4_XS window both files reached the same board VRAM peak.
The effort knob is the base template’s: medium adds no instruction, and Pi sends no effort, so the launch setting decides
Set the effort at launch, in --chat-template-kwargs. The launch command below uses medium: the October launcher’s default for this file, and the level at which the author reports Pi matching the base model’s xhigh on Terminal-Bench 2.1 and ran the GGUF agent smoke test cited. The author’s own llama.cpp quickstart sets no level, so the template’s default, xhigh, applies there. The 8,952-character template is byte-identical to Swift-1.5’s and to the base model’s meas, so the knob is not a Pi feature. Rendered with jinja2, medium adds no system line; low adds “Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.”; xhigh adds “Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.”, and so does a render with no effort set anywhere, because xhigh is the template’s default; high raises “Unexpected reasoning effort high. Supported types are xhigh (default), medium, and low.” meas. This is not the anchor’s template: unsloth’s file carries a 9,993-character template that accepts high (§14).
On the server (build 10502, /apply-template, 2026-10-06, launched with medium), a request carrying no effort rendered no line, and a request carrying low or xhigh in its own chat_template_kwargs rendered that level’s line meas n=1. That shows a per-request chat_template_kwargs reaching the template; no request generated text through /v1/chat/completions with it, and it is a different field from the request-body reasoning_effort that llama-server ignores (§09). high was not sent to the server. Pi’s provider block sets supportsReasoningEffort to false (§14), so Pi sends no effort and the level set at launch is the one the model gets; a level chosen in Pi’s own selector does not reach the template der. That reading rests on the provider setting: Pi’s request bodies were not captured.
§09’s advice, xhigh for hard code you need finished, was measured on the base model. medium here follows the October launcher and the author’s Terminal-Bench claim; on this suite the Pi file’s two levels show no difference the suite can detect (below), and completeness on agent tasks at medium against xhigh was not measured. If you follow §09, launch with xhigh.
At xhigh the Pi file reasons 18–27% slower than at medium: on short prompts each check takes the same time, but fewer guesses survive
On thinking-on requests the Pi file decoded 59.6 t/s at xhigh against 81.6 at medium with the launch command’s own drafter, n-max 4 / p-min 0 (0.731 of it), and 64.4 against 78.5 with n-max 3 / p-min 0 (0.820) meas. On short prompts the model itself is not slower at xhigh: each verify check took the same time at both levels (41.9 and 42.0 ms at n4/p0, 37.3 and 37.3 ms at n3/p0). What changes is the text the model writes, and so how often the built-in head’s guesses survive.
| Pi IQ4_XS, thinking-on requests | medium t/s | xhigh t/s | xhigh ÷ medium der | drafted tokens kept per check medium / xhigh | guesses kept (acceptance) medium / xhigh |
|---|---|---|---|---|---|
| n4/p0, the launch command, short prompt meas | 81.6 | 59.6 | 0.731 | 2.42 / 1.50 | 61% / 38% |
| n4/p0, after a 57,540-token prefix meas | 69.6 | 50.9 | 0.731 | 2.87 / 1.98 | 73% / 50% |
| n3/p0, short prompt meas | 78.5 | 64.4 | 0.820 | 1.93 / 1.40 | 65% / 47% |
| n3/p0, after the prefix meas | 64.6 | 53.2 | 0.825 | 2.26 / 1.77 | 76% / 60% |
medium: the n/p sweep’s Pi block (2026-10-07/08; n3/p0 its six reference loads, n4/p0 two loads). xhigh: two loads per drafter on 2026-10-08, 22:28–23:43, a follow-up after the sweep at the owner’s request, not pre-registered. Both used the October launcher’s arguments and temperature-1.0 sampler, -c 81920, text only, q8_0 KV, six thinking prompts of 768 tokens and, for the deep rows, the same 57,540-token prefix; job memory cap 26–28 GB at xhigh, 28–30 GB at medium. At xhigh every token of these requests was reasoning; at medium about a tenth of the characters were the answer that followed it (10.5% at n3/p0, 10.2% at n4/p0), which the next paragraph rules out as the cause. Absolute t/s here are not comparable with other sweeps on this page (the two-levels note). The levels ran in separate sessions, so the ratios cross sessions and are read descriptively, without an interval; every one of the six prompts was slower at xhigh (0.75–0.87 of medium at n3/p0, 0.63–0.84 at n4/p0). Two checks put the sessions on the same speed level: the time per check on short prompts above, and the short thinking-off requests and the greedy code probe, which the template renders identically at both levels and which produced byte-identical text (answers 83.9 and 83.8 t/s at n4/p0, 81.1 and 81.3 at n3/p0; greedy code 104.0 and 102.9, 98.1 and 96.9). The deep rows are less clean. At xhigh the effort line precedes the prefix, so each deep thinking request was 42 tokens longer and re-prefilled the prefix once per load, and after the prefix the xhigh checks ran 3% (n3/p0) and 5% (n4/p0) slower than at medium, mostly on the first deep request after that re-prefill; so a few points of both deep rows’ gaps are not the text (about 4 of the n4/p0 row’s 27, 2 to 3 of the n3/p0 row’s 18). The medium n4/p0 deep cell’s two loads also read 67.6 and 71.7 t/s, outside the sweep’s 3% band, so that row’s ratio runs 0.715–0.747 by load.
Why xhigh text is harder to guess. At xhigh the template adds a system line asking the model to “think carefully through the task, validate key assumptions, consider plausible alternatives”; medium adds none (above). The reasoning that follows is harder for the head to predict: at n4/p0 38% of its guesses survived against 61%, so each check kept 1.50 drafted tokens instead of 2.42, and fewer kept tokens per check means slower decoding (§06 explains the mechanism). It is not answer text mixed into medium’s output: on the four prompts whose medium output was all reasoning, the head kept 61–81% of its guesses at medium against 42–63% at xhigh (n3/p0) meas.
The drafter that suits xhigh is n3, not the command’s n4. At xhigh, n3/p0 decoded thinking tokens 1.080 as fast as n4/p0 (64.4 against 59.6) and 1.047 after the prefix, while answer tokens favoured n4/p0 (83.8 against 81.3); at medium, n4/p0 leads (1.038, §06) meas. The two xhigh drafters ran in separate sessions on the same level (time per check and thinking-off rows above), so the ratio is read descriptively. It agrees with the 2026-10-07 check at xhigh, where n3/p0 ran ×1.154 as fast as n4/p0.75 on this file (below). The launch command keeps n4/p0, the best measured at its own medium.
What it costs over the suite. Over the 175-prompt suite (greedy, drafter off, single-turn; below) the Pi file wrote 310,013 tokens at xhigh against 177,158 at medium meas, 1.75× as many der. Four long code items hold 115,655 of the 132,855-token excess: HumanEval items 14 and 5 ran to the 32,768-token cap (item 14 is a greedy repetition loop), item 6 stopped 111 tokens short of it, and MBPP item 8 reached 23,875; at medium the four took 891 to 2,116 tokens each meas. Without the two capped items the total is 1.40× der. The typical item was shorter at xhigh: 0.82 of its medium length at the median, with fewer tokens on 123 of 175 items (the inverse of the quality section’s 1.22 and 52 of 175) der. With the speeds above, a typical prompt takes about as long at either level with n3/p0 (0.82 ÷ 0.820) and about 1.1× as long at xhigh with the command’s n4/p0 (0.82 ÷ 0.731) der; the large extra cost of xhigh in this suite is a few very long answers. All of this is derived, not timed: dividing the suite’s token totals by the thinking-token speed ratios, as if every token were reasoning, gives about 2.1× (n3/p0) and 2.4× (n4/p0) the decode time over the suite, an upper estimate. No agent task or wall time was timed, and the long items come from greedy decoding, not tested at temperature 1.0. That adds a speed reason to the launch command’s medium (above), with no quality difference the suite can detect between the two levels (below); launch with xhigh when you want the template’s most careful level, ideally with n-max 3 (the paragraph above).
On the 175-prompt suite there is no difference the suite can detect, at xhigh or at medium
At xhigh the 5-set composite is 80.6, against 81.3 for the anchor and 82.7 for Swift IQ4_XS; at medium it is 81.2 meas. The differences are −0.66 against the anchor (paired 95% bootstrap interval −3.5 to +1.9) and −2.15 against Swift IQ4_XS (−5.1 to +0.2) at xhigh, and −0.02 (−3.2 to +3.1) and −1.51 (−4.0 to +0.2) at medium der; medium against xhigh on the Pi file is +0.64 (−2.6 to +4.0) der. Every interval includes zero: no difference the suite can detect at 25 questions per set n=25, a smoke-test scale (§10). That is not “the same quality”; it is as far as this suite can see. Against Swift IQ4_XS the intervals only just include zero (upper ends +0.2) and both point estimates favour Swift, so the suite cannot rule out a gap of up to about 5 points. At xhigh the file passed the gate set before the run: its worst set sits 4.0 points below the anchor’s (rule: no more than 8.0, two items in 25), and its composite is within 3 points of Swift IQ4_XS’s.
| Set (n = 25) or quantity | Pi IQ4_XS · xhigh | Pi IQ4_XS · medium | Anchor · xhigh2026-10-04 cells | Swift IQ4_XS · xhigh2026-10-04 cells |
|---|---|---|---|---|
| GSM8K exact match meas | 100 | 100 | 100 | 100 |
| MATH-500 exact match meas | 100 | 96 | 100 | 100 |
| HumanEval execution pass@1 meas | 88 | 100 | 92 | 100 |
| MBPP execution pass@1 meas | 92 | 88 | 92 | 92 |
| MeetingBank ROUGE-L F1 meas | 23.0 | 22.2 | 22.3 | 21.7 |
| 5-set composite meas | 80.6 | 81.2 | 81.3 | 82.7 |
| Against the anchor der difference (95% interval) | −0.66 (−3.5 to +1.9) | −0.02 (−3.2 to +3.1) | — | +1.49 (−0.2 to +3.9) |
| Against Swift IQ4_XS der difference (95% interval) | −2.15 (−5.1 to +0.2) | −1.51 (−4.0 to +0.2) | — | — |
| Total tokens, 175 items meas | 310,013 | 177,158 | 477,074 | 212,543 |
| Ratio of totals to the anchor der median per-item ratio; items with fewer tokens | 0.650 0.92; 123 of 175 | 0.371 1.10; 75 of 175 | — | 0.446 0.87; 139 of 175 |
| Ratio of totals to Swift IQ4_XS der median per-item ratio; items with fewer tokens | 1.459 1.03; 70 of 175 | 0.834 1.27; 47 of 175 | — | — |
| Items at the 32,768-token cap meas | 2 (HumanEval) | 0 | 4 | 0 |
Pi cells 2026-10-06/07; anchor and Swift IQ4_XS cells from the sweep of 2026-10-04 (§09), reused read-only and paired item for item. 175 single-turn prompts (7 sets × 25), greedy (temperature 0, top_k 1, seed 42), drafter off, q8_0 KV, 32,768-token cap, effort set through --chat-template-kwargs; job memory cap 30 GB. Composite over the five automatically scored sets; intervals by paired bootstrap, 10,000 resamples, seed 42. Token figures cover all seven sets: ALPACA and MT-Bench were not judged for the Pi file, so they contribute tokens only, and the anchor’s total includes its four capped items. Not measured at the cards’ temperature-1.0 sampler, in multi-turn use or in an agent.
At medium the suite total is lower, but the typical item is not. Pi at medium generated 0.571 of its own xhigh total, yet its median item ran 1.22 times as long and it used fewer tokens on only 52 of 175 items der; against every xhigh reference in the table its median per-item ratio is above 1. The saving is the long tail: two xhigh items ran to the cap and no medium item did, while on a character count medium wrote longer final answers on three sets (median ALPACA 1,258 against 257 characters, GSM8K 404 against 266, MATH-500 689 against 486), with reasoning length level outside ALPACA meas — a descriptive proxy, not a controlled test. So this suite does not show medium as cheaper per prompt, and it does not test the author’s 41% figure, which is a Terminal-Bench 2.1 agent measurement against the base model at xhigh.
At xhigh, the Pi file’s HumanEval is 88 against 92 and 100 — one item won and two lost against the anchor, three lost against Swift IQ4_XS meas. Item 3 is a short wrong answer (470 tokens), and items 5 and 14 ran to the 32,768-token cap. Item 14 is a greedy repetition loop (“Could test "I am. I am not. I am."” about 2,200 times) on an item the anchor passed at 22,914 tokens and Swift IQ4_XS at 1,390; on item 5 the anchor also capped and Swift IQ4_XS passed at 28,768. Whether this loop recurs at the cards’ temperature-1.0 sampler was not tested; the suite does not exercise it. At medium the same file scored 25 of 25 on HumanEval. These are single items at n=25; the automatic loop flags over the suite are not published as a loop count (§15).
With the drafter off it decodes at Swift IQ4_XS’s speed; with the drafter on, the gap moves with the content
With the drafter off, Pi IQ4_XS and Swift IQ4_XS decode at the same speed: ratios 1.001 at 1.5k of depth and 1.002 on the drafter prompt der, each from two passes that agree within 0.11%. That is what two files with the same architecture, quant types and bytes per token should do, and it is a measured case of Appendix A’s derivation that same-recipe files decode alike. With the drafter on, each model drafts its own text, so acceptance differs with the content and the ratio moved between 0.919 and 1.037 across these probes der: the Pi file was slower at 28k of depth, where it kept 3.11 drafted tokens per verify pass against Swift’s 3.412, and faster at 91k (3.428 against 3.24); on reasoning tokens at n4/p0.75 it kept 1.276 against 1.479. No probe here shows either file faster in general.
| Probe | Drafter | Swift IQ4_XS (t/s) meas pass 1 / pass 2 | Pi IQ4_XS (t/s) meas pass 1 / pass 2 | Pi ÷ Swift der | Pass ratios differ by | Kept per verify pass der Swift / Pi |
|---|---|---|---|---|---|---|
| 1.5k depth | off | 44.22 / 44.32 | 44.31 / 44.36 | 1.001 | 0.11% | — |
| Answer tokens, 1.5k depth | n4/p0.75 | 97.93 / 98.50 | 97.86 / 97.40 | 0.994 | 1.05% | 3.211 / 3.11 |
| Answer tokens, 28k depth | n4/p0.75 | 91.00 / 90.73 | 85.63 / 84.63 | 0.937 | 0.88% | 3.412 / 3.11 |
| Answer tokens, 91k depth | n4/p0.75 | 67.93 / 67.83 | 70.21 / 70.44 | 1.036 | 0.47% | 3.24 / 3.428 |
| Drafter prompt, reasoning off | off | 44.62 / 44.59 | 44.68 / 44.69 | 1.002 | 0.09% | — |
| Drafter prompt, reasoning off | n3/p0 | 102.59 / 102.58 | 101.98 / 101.87 | 0.994 | 0.10% | 2.784 / 2.763 |
| Drafter prompt, reasoning off | n4/p0.75 | 103.68 / 103.54 | 101.99 / 101.75 | 0.983 | 0.10% | 3.348 / 3.308 |
| Drafter prompt, reasoning off | n10/p0.5 | 108.08 / 107.11 | 105.31 / 104.51 | 0.975 | 0.14% | 5.667 / 5.336 |
| Drafter prompt, reasoning on | n3/p0 | 72.70 / 72.47 | 71.37 / 71.30 | 0.983 | 0.22% | 1.682 / 1.632 |
| Drafter prompt, reasoning on | n4/p0.75 | 66.93 / 66.66 | 61.46 / 61.27 | 0.919 | 0.09% | 1.479 / 1.276 |
| Drafter prompt, reasoning on | n10/p0.5 | 53.16 / 52.88 | 54.93 / 55.02 | 1.037 | 0.69% | 1.657 / 1.685 |
| Reasoning tokens, the cards’ sampler, 8 probes | n3/p0 | 66.46 / 66.47 | 64.14 / 64.08 | 0.965 | 0.11% | 1.481 / 1.389 |
| Reasoning tokens, the cards’ sampler, 8 probes | n4/p0.75 | 57.66 / 57.65 | 55.46 / 55.65 | 0.964 | 0.36% | 1.159 / 1.027 |
Rows 1–11: one sweep, 2026-10-07 07:49–08:58, job memory cap 31 GB; rows 12–13: a follow-up run, 10:12–10:34, cap 30 GB; both on a quiet host; not comparable with the 2026-10-04 Swift rows in §06 and §08, nor with the August tables. Text only (no projector), one slot, q8_0 KV, --load-mode mmap; two passes in alternating order, the first probe after each prefill discarded, 30 s settle; a ratio is void when its two passes differ by more than 3%, and none did. Rows 1–4: -c 131072, greedy, 400-token probes, rows 2–4 answer tokens at about 1.5k, 28k and 91k of prompt depth. Rows 5–11: one prompt (a JavaScript red-black tree), 700 tokens, -c 32768, greedy (temperature 0, top_k 1); reasoning on uses the template’s default, the xhigh line. Rows 12–13: reasoning on at the template’s default xhigh line, temperature 1.0, top_p 0.95, top_k 20, min_p 0, 4 prompts × 2 fixed seeds × 700 tokens, -c 32768, not pre-registered. Kept per verify pass is accepted ÷ (predicted − accepted), from the server’s own timings (§15).
n-max 3 / p-min 0 decodes reasoning tokens about 1.15× faster than n4/p0.75 on both IQ4_XS fine-tunes; on one answer-token prompt they tie
On bartowski’s IQ4_XS of either fine-tune, n-max 3 / p-min 0 decoded reasoning tokens faster than the n4/p0.75 drafter this page’s cards shipped until 2026-10-08: ×1.154 on the Pi file and ×1.153 on Swift IQ4_XS under the cards’ temperature-1.0 sampler der (4 prompts × 2 fixed seeds, 700 reasoning tokens each, -c 32768, text only, --load-mode mmap, at the template’s default xhigh line, two passes in reversed order, 2026-10-07; not measured at medium, the level the launch command uses), faster on 7 of 8 probes for the Pi file and 8 of 8 for Swift meas. The pre-registered test that chose the Pi configuration’s drafter — one prompt, greedy, two passes — gave ×1.162 on the Pi file and ×1.087 on Swift IQ4_XS der. On answer tokens the two settings were level: ×1.001 on the Pi file and ×0.990 on Swift (one prompt, greedy, reasoning off) der. n-max 3 also read 152 MiB less board VRAM at -c 32768 on both files meas, about the one recurrent-state row it does not keep (149.625 MiB by Appendix A’s arithmetic) der.
The pattern is consistent with the confidence gate der. In the reasoning stream p-min 0.75 stops drafts early: on the Pi file n4/p0.75 kept a larger share of what it drafted (acceptance 0.807 against 0.549) but fewer tokens per verify pass (1.276 against 1.632), and on Swift IQ4_XS 1.479 against 1.682 (greedy) der; under the sampler the order held (Pi 1.027 against 1.389, Swift 1.159 against 1.481). n-max and p-min changed together, so the gate is not isolated, and tokens kept per verify pass do not rank every pair here: on answer tokens they favour n4/p0.75 (3.308 against 2.763 on the Pi file) while the speeds tie, and with reasoning on the Pi file’s n10/p0.5 keeps the most (1.685) and decodes the slowest, as §06 records for the Swift files.
Whether the gain holds on the base model’s files was not measured in a matched sweep. The anchor did not run in this study. On the same prompt on 2026-10-04 its n4/p0.75 kept 2.684 drafted tokens per verify pass with reasoning on, against Swift IQ4_XS’s 1.479 (§08), so on that prompt these files kept far fewer than the anchor; n-max 3 was not run on the anchor then. The 2026-08-23 sweep on UD-IQ4_XS, on a different prompt with a fresh server per configuration, had n3/p0 at 82.0 against n4/p0.75 at 81.0 t/s on reasoning tokens (§06). The anchor’s draft head is also quantised differently (Q8_0, Q6_K and F32, 334.75 MiB, against Q4_0 and F32, 227.91 MiB; §17), so the gain may come from these fine-tunes, from the 4-bit head or from the prompt. 2026-10-08: the n/p sweep then ran the base model’s UD-IQ4_XS in its own sessions at the cards’ sampler and xhigh: n3/p0 decoded shallow reasoning tokens 1.111 of n4/p0.75 (paired 95% interval 1.075–1.156), so the gain is not specific to the fine-tunes or their 4-bit head; speeds across files are still not compared (§06).
| File | Tokens and sampler | n3/p0 (t/s) meas pass 1 / pass 2 | n4/p0.75 (t/s) meas pass 1 / pass 2 | n3 ÷ n4 der | Kept per verify pass der n3 / n4 | Acceptance meas n3 / n4 |
|---|---|---|---|---|---|---|
| Pi IQ4_XS | reasoning, temperature 1.0, 8 probes | 64.14 / 64.08 | 55.46 / 55.65 | 1.154 | 1.389 / 1.027 | 0.465 / 0.765 |
| Swift IQ4_XS | reasoning, temperature 1.0, 8 probes | 66.46 / 66.47 | 57.66 / 57.65 | 1.153 | 1.481 / 1.159 | 0.497 / 0.756 |
| Pi IQ4_XS | reasoning, greedy, 1 prompt | 71.37 / 71.30 | 61.46 / 61.27 | 1.162 | 1.632 / 1.276 | 0.549 / 0.807 |
| Swift IQ4_XS | reasoning, greedy, 1 prompt | 72.70 / 72.47 | 66.93 / 66.66 | 1.087 | 1.682 / 1.479 | 0.565 / 0.837 |
| Pi IQ4_XS | answer, greedy, 1 prompt | 101.98 / 101.87 | 101.99 / 101.75 | 1.001 | 2.763 / 3.308 | 0.928 / 0.919 |
| Swift IQ4_XS | answer, greedy, 1 prompt | 102.59 / 102.58 | 103.68 / 103.54 | 0.990 | 2.784 / 3.348 | 0.936 / 0.917 |
Rows 1–2 are the follow-up run of 2026-10-07 10:12–10:34 (cap 30 GB) and rows 3–6 the sweep of 07:49–08:58 (cap 31 GB), with the probes and conditions of the speed table above (text only, one slot, -c 32768, --load-mode mmap, 700 tokens, the reasoning rows at the template’s default xhigh line, two passes in alternating order). The sampled rows fix each probe’s seed, so the two passes repeat each probe’s stream and the comparison is across 8 probes; the Pi file’s one slower probe was the red-black-tree prompt at seed 22 (54.5 against 57.7 t/s). Board VRAM peaks at -c 32768, one slot, desktop inside: 17,878 MiB at n3/p0 against 18,030 at n4/p0.75 (greedy, reasoning off), and 17,112 against 17,264 under the sampler meas. Not measured at depth, with the projector loaded, or on Swift IQ3_XXS.
What it does not show. In this study (2026-10-07), every n-max 3 measurement is at -c 32768 on 700-token probes, text only, at the template’s default xhigh line: no file was measured at n-max 3 at depth, with the projector loaded or at medium, and the windows below were filled at n4/p0.75. A faster reasoning-token decode is not a measured agent wall-time saving, and the sampled run was not pre-registered. The study’s pre-registered rule chose n-max 3 for the Pi configuration on 2026-10-07, and the launch command ran it until 2026-10-08. The n/p sweep then measured both at the command’s own medium with the launcher’s arguments, at -c 81920, text only, shallow and after a 57,540-token prefix: n4/p0 decoded shallow reasoning tokens 1.038 of n3/p0 (paired 95% interval 1.000–1.092) and 1.073 after the prefix (range 1.041–1.116 over 3 prompts), a repeat block read 1.038 and 1.082, and n5/p0 and longer were slower. So the launch command below now uses n-max 4 / p-min 0, the n-max its window was filled at, and the Swift IQ4_XS cards run n3/p0, that sweep’s best score for their file (§06, §08).
The 139,264-token vision window held decode at depth, within 3% of Swift IQ4_XS; 159,744 text and the two-slot IQ3_XXS shape held too
Filled to 126,838 tokens with a 1440p image in front, on answer tokens at temperature 0 with the n4/p0.75 drafter, Pi IQ4_XS decoded 44.23 and 44.09 t/s against Swift IQ4_XS’s 45.14 and 45.55 at the same window, a ratio of 0.974 meas n=2, inside the pre-registered rule of 10%. Text-only at 159,744 it held at 1.054 of Swift, and Pi IQ3_XXS with vision at two slots of 123,904, both slots busy, at 0.963 per slot (0.969 with the desktops matched, the Pi file’s 500 MiB against Swift’s 572) der. At each IQ4_XS window both files reached the same board VRAM peak (24,019 and 23,429 MiB) meas, as the identical tensor layouts predict. No larger window was loaded, so these show that decode held, not where it collapses.
| Configuration | Per-slot window (tokens) | Fill (tokens) meas | Pi decode (t/s) meas load 1 / load 2 | Swift decode (t/s) meas load 1 / load 2 | Pi ÷ Swift der | Board VRAM peak (MiB) meas desktop inside | Shared at depth (MiB) meas | Desktop before (MiB) meas |
|---|---|---|---|---|---|---|---|---|
| IQ4_XS · vision · n4/p0.75 · 1 slot the launch command’s window | 139,264 | 126,838 | 44.23 / 44.09 | 45.14 / 45.55 | 0.974 | 24,019 | 444 | 1,217–1,247 |
| IQ4_XS · text · n4/p0.75 · 1 slot | 159,744 | 146,423 | 51.09 / 50.97 | 48.36 / 48.51 | 1.054 | 23,429 | 484 | 1,217 |
| IQ3_XXS · vision · n4/p0.75 · 2 slots the October launcher’s Swift IQ3_XXS two-slot vision configuration | 123,904 | 112,652 / 112,647 per slot | 23.43 / 24.04 per slot | 24.48 / 24.81 per slot | 0.963 0.969 desktop-matched | 24,197–24,231 | 1,620–1,748 | Pi 1,217 / 500 Swift 634 / 572 |
One sweep, 2026-10-07; not comparable with the knee sweep of 2026-10-03/04 or the deep fills of 2026-10-04. window-knee.py’s measure(), imported unedited: each window filled to 91–92% of itself (each slot, for the two-slot row), answer tokens with --reasoning off at temperature 0, two settled probes of 400 tokens, MTP n4/p0.75, --cache-ram 0, q8_0 KV; the vision rows load each file’s own f16 projector and hold a 1440p image. Order Pi, Swift, Swift, Pi per configuration, n=2 per file; job memory cap 31 GB; NVIDIA GeForce RTX 3090 (24,576 MiB). Board VRAM peak is memory.used with the desktop inside; the desktops were not matched across the two-slot loads. Each row’s title carries the Pi file’s server command.
The vision fill reproduced the 2026-10-04 footprint, which leaves 1,330 MiB for a desktop. It ran with a 1,217–1,247 MiB desktop, peaked at 24,019 MiB of board VRAM with the desktop inside, and kept 444 MiB in shared memory at depth, for both files meas. At a 1,217 MiB desktop that peak leaves 22,802 MiB dedicated to the server, the footprint of the 2026-10-04 Swift fill at an 841 MiB desktop, with the same 444 MiB shared der. The page reads that footprint as fully resident with a desktop of up to 1,330 MiB (the Swift IQ4_XS vision card, Appendix A), and these desktops were inside it. It is below this page’s 1,796 MiB desktop threshold, at which the vision window for this file shape is 128,000 der, and that applies to the Pi file through the identical tensor layout. Every fill here used n4/p0.75; the launch command runs n4/p0, the same n-max, and p-min costs no memory (18,872 MiB of server VRAM at p-min 0 against 18,871 at 0.75 on this file, -c 81920, 2026-10-08, §06), so the fill’s footprint applies; decode at this depth with p-min 0 was not filled. Decode was tested; quality at depth was not (no needle ran), and Pi IQ3_XXS has no quality measurement at all, so the two-slot row is a speed-and-fit result only. The Swift fills in this table also re-measure Swift IQ4_XS’s own windows; their absolute t/s are not comparable with the single fills of 2026-10-04 (§05).
In Pi it passed a three-probe smoke test, which shows the file works end to end and nothing more
Run headless in Pi against the Pi IQ4_XS server, it wrote and ran a FizzBuzz script and reported its last line correctly (PASS, 48.1 s), carried an attached screenshot to the model, which answered “171” (PASS, provisional, 9.4 s), and with the screenshot withheld did not produce the answer (5.9 s) meas n=1 per probe, 2026-10-06. The server ran the October launcher’s Swift IQ4_XS vision argv with the Pi file and projector (139,264 tokens, one slot, n4/p0.75, medium at launch). Pi was routed to llamacpp/qwen/qwen3.8-27b as pi --no-session -p, sent no effort (by its provider setting; the request bodies were not captured), and used its existing 123,904-token provider entry (§14), so a 139,264-token entry was not exercised. The page’s agent-probe harness ran unedited, and the server’s /slots showed temperature 1.0 on every task, the sampler the server was launched with (--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0). The server checks and the smoke test ran while a model download was in progress, so their t/s and wall times are single readings under that load and are not compared with anything else. These wall times are not comparable with the Swift IQ3_XXS run in §14 (584.4 s), which used a different file, window, effort and slot count. Three probes at n=1 say nothing about agent-task success, turns, tool use or agent wall time, and nothing here tests the author’s agent claims.
| Server check (one request each) | Result meas n=1 | What it does not show |
|---|---|---|
| Load at 139,264 with vision | /health in 10.6 s; board VRAM 1,313 MiB before load (desktop, with a download running), 23,571 after load and 23,971 at the end of the checks, leaving 605 MiB of the card free der | A load check, not a deep fill (above); n4/p0.75; the launch command runs n4/p0, the same n-max |
Coherence at medium | A correct is_prime function and a two-sentence explanation; finish stop, 264 completion tokens, 423 characters of reasoning; 61.3 t/s, drafter acceptance 0.841 (185 of 220) | One request in one server start, taken while a download ran; its t/s is not comparable with any other figure on this page |
Prefill, cache_prompt off | 9,340 tokens at 1,254 t/s; 10,003 at 1,275 t/s, against a pre-registered pass line of 500 | A synthetic prompt, not Pi’s first request in agent use; taken while a download ran |
| Screenshot attached | “171”, the number on the page’s detail target, at the full image budget (3,636 prompt tokens, 66 completion tokens) | One question; the seven-question instrument of §12 was not run |
| Screenshot withheld | “No screenshot was provided in your message, so I’m unable to identify the number. Please attach the image and I’ll be happy to help.” | One control question |
2026-10-06, the October launcher’s Swift IQ4_XS vision argv with the Pi file and its f16 projector: -c 139264, --parallel 1, --image-min-tokens 1024 --image-max-tokens 10580, -ngl 99, q8_0 KV, --load-mode none, n4/p0.75, medium, the cards’ sampler, --reasoning-preserve, --jinja; build 10502; job memory cap 30 GB. The effort-line check on the same server is in the effort section.
To run it on 24 GB: the October launcher’s Swift IQ4_XS vision configuration with the Pi file, medium effort and n-max 4 / p-min 0
Choose it for the Pi agent if you prefer this fine-tune; nothing measured here puts it ahead of Swift IQ4_XS on quality, on speed in general or on agent success, and the October launcher’s default stays its Swift IQ4_XS vision configuration. Reasons that need no measurement, such as the Apache-2.0 licence or the author’s tuning for Pi, are the reader’s to weigh; the author’s agent results are cited, not reproduced (above).
NVIDIA 24 GB · 3090 — Qwen3.8-27B-pi IQ4_XS (text + vision, 1 slot, 139,264-token window, desktop up to 1,330 MiB)
:: Qwen3.8-27B-pi IQ4_XS, text + vision, 1 slot, 139,264 :: tokens per slot. Measured 2026-10-06/07; not comparable :: with the August tables in section 06 or the 2026-10-04 :: Swift rows. The window was deep-filled at n4/p0.75; :: this command's n4/p0 has the same n-max (p-min costs no :: memory); n4/p0 at medium was measured at -c 81920, :: text only, not at this window (section 06). :: Choose this to run bytkim's Pi-agent fine-tune; for the :: Swift fine-tune choose the Swift IQ4_XS vision card, for :: the base model choose the default card. :: file: bartowski/bytkim_Qwen3.8-27B-pi-GGUF :: (IQ4_XS, 15,475,952,128 bytes) + f16 mmproj :: (927,607,712 bytes) llama-server.exe -m bytkim_Qwen3.8-27B-pi-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-bytkim_Qwen3.8-27B-pi-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 139264 :: measured: deep fill to 126,838 tokens :: with a 1440p image held decode at 0.974 :: of Swift IQ4_XS in the same sweep (answer :: tokens, n4/p0.75); footprint 22,802 + :: 444 MiB, leaving 1,330 MiB for a desktop, :: below this page's 1,796 MiB threshold. :: Desktop above 1,330 MiB: -c 128000 :: (DERIVED at the 1,796 MiB threshold, :: Appendix A). Text only: -c 159744 held :: too (drop the projector lines) -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --api-key dummy :: must match apiKey in Pi's provider file --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 :: the :: n4/p0 drafter since 2026-10-08 (n/p sweep, :: this file at medium, -c 81920, text only, :: main + repeat block): reasoning 1.038x :: n3/p0 (paired 95% 1.000-1.092), 1.073- :: 1.082x after a 57,540-token prefix; n5/p0 :: and longer slower. p-min written out as :: 0.0. Same n-max as the deep fill :: (section 06, Appendix B) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: medium: the level of the author's :: Terminal-Bench claim; it adds no system :: line. low and xhigh offered; high raises :: (jinja2 render). Pi sends no effort, so :: this line decides. Section 09 favours :: xhigh for hard code (base model); medium :: against xhigh in agent use: unmeasured :: suite (greedy, drafter off, 175 prompts, n=25 per set; :: ALPACA and MT-Bench not judged): 5-set composite :: 81.2 at medium and 80.6 at xhigh, against 81.3 :: (anchor) and 82.7 (Swift IQ4_XS) at xhigh: no difference :: the suite can detect (Appendix B) :: speed: ratios to Swift IQ4_XS only, inside one sweep :: (Appendix B); no absolute band is printed for this card :: FLAGS DELIBERATELY OMITTED: :: --parallel 2 — not measured on this file; for two users :: with vision, Pi IQ3_XXS held decode at 123,904 x 2 :: (speed and fit only; its quality is unmeasured) :: high — not an effort level of this template :: invalidation trigger: build 10502 / 0adcc3bb5, :: driver 596.36, or a new upload of the file :: fields that differ from the measured command: the drafter :: (the deep fill, the server checks and the Pi smoke test :: ran n4/p0.75, the same n-max; n4/p0 was measured at :: -c 81920, text only, shallow and after a 57.5k prefix); :: port (no effect on inference); the deep fill ran :: --reasoning off at temperature 0 with the host-RAM prompt :: cache off (--cache-ram 0); this card keeps it on; :: load mode (the deep fill and the drafter runs memory- :: mapped the file, the auto default; this card and the :: server checks use --load-mode none). :: Quality at depth: not verified (no needle ran)
Copy-paste version — the same command with the comments removed
llama-server.exe -m bytkim_Qwen3.8-27B-pi-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-bytkim_Qwen3.8-27B-pi-f16.gguf --image-min-tokens 1024 --image-max-tokens 10580 -c 139264 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --api-key dummy --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --reasoning-preserve --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}"The file: bartowski/bytkim_Qwen3.8-27B-pi-GGUF, IQ4_XS — 15.48 GB on the download page, 14.41 GiB (15,475,952,128 bytes), plus the f16 mmproj (927,607,712 bytes). The window: -c 139264 meas — deep-filled to 126,838 tokens at a 1,217–1,247 MiB desktop, with the 2026-10-04 footprint (22,802 MiB dedicated + 444 MiB shared), which leaves 1,330 MiB for a desktop der. That is below this page’s 1,796 MiB threshold: this card departs from the rule the Swift cards follow because 139,264 is the one vision window deep-filled with this file and the window the October launcher runs it at. With a desktop above 1,330 MiB take 128,000 der, as the Swift IQ4_XS vision card does; for text only 159,744 held at depth too meas. The flags: those of the October launcher’s Swift IQ4_XS vision configuration with the Pi file and projector — the Swift IQ4_XS vision card’s flags except the window (139,264 here, 128,000 there) and the effort (medium here, xhigh there) — plus --api-key dummy, which must match the apiKey in Pi’s provider file (§14), and the n-max 4 / p-min 0 drafter (§06, above). The effort levels it offers: medium at launch, and low and xhigh; high raises under a jinja2 render (above). Pi sends no effort, so the launch setting is what the model gets; in Pi’s provider file set contextWindow to the server’s -c (§14). A thinking block with a defaultLevel in that file does nothing: Pi 1.0.2 reads no such field (no defaultLevel or requiresEffort in its installed code, 2026-10-08), and with supportsReasoningEffort false it sends no effort, so --chat-template-kwargs at launch is the only effort setting.C45 The smoke test used the existing 123,904-token entry. Speed: published only as ratios to Swift IQ4_XS inside one sweep (above): with the drafter off the two decode alike, and on reasoning tokens this card’s n4/p0 ran ×1.038 n3/p0 at medium (2026-10-08, -c 81920, text only, shallow, §06), and on 2026-10-07 n3/p0 had run ×1.154 this file’s n4/p0.75 at -c 32768, text only, at the template’s xhigh line der. No absolute speed band is printed for this card, because none was measured in a sweep with the anchor. Room left for your desktop: 1,330 MiB at 139,264, from the fill’s footprint der; this page’s 1,796 MiB threshold gives 128,000 der.
On 16 and 12 GB cards nothing was measured: take the page’s own cards
On a 16 GB card use the page’s measured 16 GB card, the base model’s UD-Q2_K_XL at -c 65536 (§03). If you want the Pi fine-tune anyway, the derived, unmeasured option is Pi IQ3_XXS, text only, drafter off, --parallel 1, -c 59392 at this page’s 1,796 MiB desktop reserve der. Pi IQ4_XS does not fit a 16 GB card in any configuration, and nothing here ran on one. Pi IQ3_XXS (12,320,168,448 B) has Swift IQ3_XXS’s tensor layout and is 192 B larger, so Appendix A’s bartowski IQ3_XXS row applies to it. On an RTX 5080 (16,303 MiB cited), text with one slot and the drafter on gives 30,720 tokens at this page’s 1,796 MiB desktop reserve and 43,008 at 1,272 MiB (54,272 at Appendix A’s expected, measured-anchored constant); with the drafter off, 59,392 and 72,704 (86,016 expected); with vision, 11,264 and 22,528, too small to be practical. Pass --parallel 1 explicitly (Appendix A). Three cautions: Pi IQ3_XXS was never run on a quality instrument; a drafter-on window under 32,768 tokens leaves this model little room to think (§09), and Pi’s own prompt takes part of it; and Appendix A’s drafter-on constants assume n-max 4 (measured at n4/p0.75; p-min costs no memory, so they hold at n4/p0 der). Nothing here shows it beats that card.
On a 12 GB card nothing for this fine-tune was measured, and this page does not recommend it. By Appendix A’s arithmetic, for text and one slot, only bartowski’s IQ2_XXS and IQ2_XS (2.60 and 2.66 bits per weight) fit with the drafter on; with it off IQ2_S (2.83) fits too, and IQ2_M (3.08) only at the 1,272 MiB reserve, at 9,216 tokens der. That holds only if bartowski’s Pi files at those tiers follow the same recipe as his Swift ones, which was checked here for IQ4_XS and IQ3_XXS only. None was downloaded, read or measured; all but IQ2_M sit below the 2.912 bits per weight this page sets as a floor for the base model (§08), and the author’s own results give the 2-bit files the lowest agent scores, from a smoke test the author says is not a quant ranking cited. Use the page’s 12 GB guidance (§03).
What this appendix cannot say
- That the fine-tune helps an agent. Only a three-probe smoke test ran in Pi; no pass rate, turn count, tool use or wall time was compared with Swift IQ4_XS or the base model. The study’s pre-registered agent benchmark would close it (§15).
- “The same quality” as the anchor or Swift IQ4_XS. Every interval includes zero at 25 questions per set, on single-turn greedy prompts with the drafter off; ALPACA and MT-Bench were not judged for the Pi file; and against Swift IQ4_XS a gap of up to about 5 points is not ruled out. A larger suite would narrow it (§10).
- The author’s Terminal-Bench 2.1, GPQA Diamond and SciCode results, or the “about 41% fewer output tokens”. Cited, not reproduced; this suite measures none of them.
- Token savings at the cards’ temperature-1.0 sampler, in multi-turn use or in an agent. Every token figure here is greedy and single-turn, and
medium’s typical item ran longer thanxhigh’s. - That the Pi file is faster or slower than Swift IQ4_XS in general. With the drafter off they decode alike; with it on, the ratio moved between 0.919 and 1.037 with the content.
- That the launch command reasons faster than the Swift IQ4_XS cards because of the model. On 2026-10-07 the ×1.15 came from the drafter setting, and Swift IQ4_XS gained the same with it. Since 2026-10-08 each runs its own best score, n4/p0 on the Pi file and n3/p0 on Swift IQ4_XS; the two were not compared in one session, because the sweep’s ratios are within one file.
- How large the drafter gain is across files. Whether the n-max 3 gain belongs to the fine-tunes is answered: on 2026-10-08 the anchor gained too, n3/p0 reading 1.111 of its n4/p0.75 on shallow reasoning tokens against 1.189 on Swift IQ4_XS, each in its own sessions, so the two sizes are not a same-session comparison (§06).
- The launch command as shipped (n4/p0) at its 139,264 window with an image. n4/p0 at
mediumwas measured at-c 81920, text only, shallow and after a 57,540-token prefix (§06); the fills, the server checks and the smoke test ran n4/p0.75, the same n-max. - The 139,264 window at this page’s 1,796 MiB desktop threshold. Its fill reproduced the 2026-10-04 footprint, which leaves 1,330 MiB for a desktop; at 1,796 MiB the derived window is 128,000.
- The quality of Pi IQ3_XXS or any other Pi quant, and header identity beyond IQ4_XS and IQ3_XXS. Only the IQ4_XS ran the suite, and only those two files were read.
- Server behaviour with
high, and generation with a per-request effort.highraised only under jinja2; the per-requestchat_template_kwargswas checked at/apply-template, not in generation. - Whether the embedded MTP head is the upstream one or tuned for Pi. The heads were not compared.