Field guide · Local inference
Qwen3.8-27B on a 24 GB RTX 3090: measured settings, speeds and energy
Copy-paste server configurations, the speed and power each one delivers, and the measurement behind every recommendation. Settings for other cards — RTX 30/40/50, DGX Spark (GB10), Intel Arc Pro B50/B70, Intel Arc B390-class iGPU — are calculated from the same arithmetic and marked as calculated.
Evidence tier. Everything here was measured on one machine — a 24 GB RTX 3090, driver 596.36, Windows 11 Pro 26200, i5-13600KF, llama.cpp build 10502 (commit 0adcc3bb5) — on 2026-08-21 through 2026-08-25. 2026-08-23 alone spent about 12.9 hours of GPU time across five completed rounds: a two-hour independent re-measurement of this page's own figures, run without sight of them; a 1.7-hour follow-up round; an 8.48-hour benchmark sweep of 525 generations over seven benchmark sets; a 43.5-minute power matrix of 19 configurations; and two zero-GPU energy analyses. A sixth round, a quantization ladder, ran from 14:13 on 08-23 to 12:35 on 08-24 for about six hours more, and a seventh pass used no GPU at all — the blind judge panel of §09. The ladder ranked nine files of this model on 294,912 scored token positions each, from 4.40 down to 1.83 bits per weight — how many bits a file spends on each of the model's numbers, averaged over the whole file — with automated checks beside every file for the failures a score cannot see: an empty answer, output that repeats until it hits the cap, a broken chat template. Eight of those files were then put through a second instrument on 2026-08-24 — a frozen three-benchmark suite scored on the identical 75 items per file, so that every claim comparing one file with another is paired, meaning only the questions the two answered differently count — and on 2026-08-25 three further rounds measured what the ladder could not: the drafter on and off on both candidate files; all three candidate configurations deep-filled to 218,233 real tokens at the model's full native window; and a requirement sweep of three files × three windows × the drafter on and off, each window deep-filled to about 90% of itself, which is the source of §08's memory-and-speed table and of the finding that the fastest file changes with the window. That whole chapter is now in §08 rather than summarized in a sentence, and §15's run log dates each round. The 08-21 and 08-22 rounds that produced the perplexity, quantization and drafter tables were not hour-logged. Solid, meaning measured more than once or on more than one file: the VRAM budget model (17 server loads with the whole model and its cache on the card, all reproduced within 127 MiB), the drafter's speed and its energy, the depth curve, decode energy per token, and the perplexity ranking. Smoke-tier, meaning one run or a small sample decided it: every quality judgement (n=1 or n=2 per effort level; the 08-22 pages were graded blind by independent reviewers, the 08-23 single runs were not), the 20-question and 25-question benchmark cells, and the whole vision section. And that rule binds this page's own newest table: the quantization ladder's accuracy column — the ladder being nine versions of this one model, largest file to smallest — is n=25 per benchmark set, which is exactly the sample size this page tells readers not to rank files with — so it is used to locate where the model breaks and never to order one file against another, and every "tie" or "worse" in §08 comes from a paired test on identical items instead. Not measured at all, and named rather than estimated: power capping, system RAM under either load mode, any card that is not this 3090 — so every statement about how fast a 16 or 12 GB card runs this model is derived, and always will be, while how much memory a given file and window need is a property of the model and the flags rather than of the card and is measured here (§08) — board power for any file below 4 bits, and any external drafter beyond DFlash2 (which was measured and lost). The concurrency question that sat here in earlier editions was closed on 2026-08-25: two slots with the drafter on gain 22.0% of aggregate throughput and cost 35.4% of each user's own speed. This page corrects its own earlier editions in twenty-nine places; the numbered list is in §15. One instrument is not a measurement of this machine at all: the two open-ended benchmark sets are scored by a blind panel of three Claude Opus 5 judges reading the kept transcripts (§09) — a judgement, labelled as one wherever it appears.
-c 32768, 86.3 at a 1,458-token fill measured after the card had settled rather than on the first probe after a long prompt (§11)xhigh reasoning stream at that same depth01
What this page is, and what it measured
This page tells you how to run one model — Qwen3.8-27B — on one graphics card, a 24 GB RTX 3090, and what to expect when you do. Every setting it recommends was measured on that machine, and the measurement that decides each recommendation is printed beside it. It also carries settings for other cards, marked clearly as calculated rather than measured. It is not a review, not a comparison between models, and not a set of numbers you should expect to match exactly on different hardware. If all you want is a working configuration, go to §03 and copy the block for your card.
The limits, stated plainly: one machine, one model family, one operating system. Every measured row on this page is the same RTX 3090 on Windows 11 with driver 596.36. Nothing here has been checked on Linux, on AMD, on Apple silicon, or on a second copy of the same card. Where a number was calculated instead of measured, it says so in the cell.
Three labels, and nothing on this page is a fourth thing. measured means it came off this machine on a stated date, with its conditions attached. derived means it was calculated from measured numbers by arithmetic printed nearby. cited means it came from somebody else's document, with a link. A single-sample figure carries n=1 beside it. Where a published block or cell differs in any field from the command actually measured — a different port, a different -c, a different effort level — the difference is named in that cell together with the reason it does not change the answer.
Conditions travel with numbers. A speed on this page is meaningless without four things: the file, the drafter flags, the token regime (see §02), and the prompt depth. A watt figure is meaningless without its instrumentation tier. Both are printed with every table.
Speed. One configuration was probed four times, each probe fired immediately after its prefill, and read 18.27, 18.82, 19.21 and 26.60 t/s (2026-08-23). The highest is 45.6% above the lowest; half of that spread is ±22.8%, which this page rounds to a ±25% band for any single probe taken right after a long prefill. The "45% swing" in §11 and the "±25%" band are therefore the same measurement stated two ways, and there is only one speed noise floor on this page. Under the cooled protocol (§02, documented in §11) — throw away the probe that comes back with the prefill, wait 30–45 s, take the median of several short repeats — repeat probes on this machine agree to 0.4%. The drafter-off decode floor holds inside 3.6% across five contents and both token regimes, which is why 3.6% is the pass band on the reproduction check in §15.
Energy. One configuration measured three times in the 2026-08-23 power matrix reproduces joules per decode token to 2.9% on the slow arms and 5.6% on energy-delay product, tightening to 0.03% on the fast speculating arms. Separately, board power drifted ±6% between arms measured hours apart at constant throughput, so a 6% energy gap between two arms run at different times is the instrument, not the setting.
What survives those bands: shapes, ratios and rankings. Individual levels do not. Do not believe a slow-arm energy gap under about 3%, and do not believe a speed difference between two single post-prefill probes under about 25%. Printed precision on this page follows the band, with one deliberate exception. Single probes are rounded; cooled and multi-probe speeds are given to one decimal. The energy tables print four significant figures (8.104, 3.743, 3.210 and the rest) because that is what the integrator reports and because the ratios between them are the finding — but the energy noise floor is 2.9%, so the third digit of any of those is not meaningful and only their ratios and ranking are. Two figures are printed to four and five places on purpose, and neither is a level claim: 81.71 t/s and perplexity 6.5956 are reproduction matches between two independent runs, quoted at that precision because matching to the decimal is the point.
02
Plain words: every term used here
These are the technical words this page uses without stopping to explain them again. Each is defined once, here, with one of this campaign's own numbers attached wherever a number makes it concrete. Every later section links back to this box instead of defining anything again.
- token
- A piece of a word. The model reads and writes tokens, not letters. English text runs roughly 3–4 characters per token, so a 700-token answer is about half a page. Speeds on this page are tokens per second, written t/s.
- context window
- The total number of tokens the server will hold at once — your prompt, the model's thinking, and its answer, all in one pool. You set it with the
-cflag. This model can go up to 262,144; what actually limits you is memory. The recipes here ship 122,880 as the everyday value. Depth on this page means how much of that window is already filled when a measurement is taken: "64.8 t/s at 91k of depth" means the prompt already held 91,000 tokens. - slot
- One conversation the server will work on at a time, set with
--parallel. One slot means one request is answered at a time and it gets the whole window and the whole card; two slots serve two people at once, splitting both. Measured here, two slots deliver 22.0% more tokens in total and 35.4% fewer to each person. - VRAM
- The memory on the graphics card itself. The reference card has 24,576 MiB of it. Your card's own memory is reported as dedicated; memory the card borrows from the system is reported as shared. When the model plus its context window need more than the card has, the extra spills into ordinary system memory and everything slows down (see spill below).
- resident
- Living in the card's own memory rather than in system memory. A file's resident size is what it occupies once loaded, and it is not the size printed on its download page: UD-IQ4_XS is 14.25 GB there and 13.3 GiB resident. A window is fully resident when every byte of its cache is on the card too. Past that point the server still starts and still answers a short prompt — which is exactly what makes it dangerous (§05).
- the three memory counters
- Three different numbers describe the same card, they do not agree, and this page always says which one it means. The server's own report at load counts only the server process — budget with this one; it reproduced to 127 MiB across 17 loads. Board VRAM, from
nvidia-smi, is the whole card, so it also contains your desktop and your browsers. Dedicated GPU memory in Task Manager is what the card holds in its own memory, and it cannot see a spill; the shared figure beside it is the one that can. This page writes board VRAM and board power in full, and never lets the word "board" stand alone as a number. (Board here is the graphics card itself — the circuit board with the chip, the memory and the fans on it.) - headless
- No graphical session is using this card, so the model gets all of its memory. Unplugging a monitor does not achieve this on Windows: the Desktop Window Manager keeps running and keeps its share, measured here at 1,179–1,669 MiB in direct no-server readings. What does achieve it: close everything that draws, drive the display from a second graphics card, or connect from another machine with nobody logged in locally. This page uses the word for that condition and for nothing else — a browser or a coding agent running with no window on screen is a different thing entirely, and is written out in plain words wherever it appears.
- desktop VRAM share
- How much of the card's memory your own screen is using: the window manager, the browser, anything else drawing. Measured here at 1,179–1,669 MiB in direct no-server readings on one machine. It is a quantity that depends on what is on screen, not a switch you can throw, and the three ways to shrink it are the three listed under headless above.
- slack
- The VRAM left on the card once the server has taken its share. It is what your desktop, your browser and load-to-load variation have to fit into. This page judges slack against one threshold, 1,796 MiB — earlier editions used 1,308 SUPERSEDED, corrected 2026-08-25 — and it is important to be exact about what that number is. The old figure was built on a desktop maximum of 1,181 MiB that an audit showed was not the maximum: the campaign's own power-matrix round carries a direct reading with no server loaded at all of 1,669 MiB, and an earlier bare-idle reading the same day of 1,179 — so this desktop moved 490 MiB between two idle measurements. The old 133 MiB floor was worse than stale: it came from a board at 24,296 of 24,576 MiB that was already spilling, so it recorded a desktop being evicted rather than a desktop at rest. Two measurements go into it: the desktop's own share of board VRAM, which measured 1,179–1,669 MiB in direct no-server readings depending on what was on screen, and 127 MiB of load-to-load variation in the server's own report. The 1,796 itself is derived — the worst case of the first plus the second — and it is this page's own threshold for calling a configuration usable with a screen attached. The desktop does not "need" 1,796 MiB. It needed anywhere from 1,179 to 1,669 on this machine, and 1,796 is the ceiling this page plans against so that a bad day still fits. Subtract your own desktop instead if you know it: the number that matters is what your screen holds, not this one's worst case.
- bandwidth
- How fast a card can read its own memory, in gigabytes per second. It is the single number that decides generation speed here, because the card must read the whole model once per token. The reference card is 936 GB/s; a laptop's integrated graphics is nearer 150.
- projector
- The extra file (
mmproj) that lets the model see images. It is optional at start-up. Measured here it occupies 1,138 MiB of VRAM and costs nothing in speed. - quantization
- Storing the model's numbers with fewer bits so the file is smaller and faster to read. A 4-bit file of this model is about 13–15 GiB instead of 54. It costs a little quality: measured here, the gap between the two 4-bit files this page recommends is 0.9% of perplexity. Bits per weight is the unit: how many bits the file spends on each of the model's numbers, averaged over the whole file. It is measured from the file rather than read off its name — a file called
Q2_K_XLmeasures 2.912 (§08). Two families appear on this page and the filename tells you which is which: a name containing IQ (UD-IQ4_XS, UD-IQ3_XXS, UD-IQ2_S) is an IQ-quant; a name containing K without IQ (Q4_K_M, UD-Q3_K_XL, UD-Q2_K_XL, Q6_K) is a K-quant. The two decode at different efficiencies, which is the one place the distinction changes a number you work out for yourself (§04). - KV cache
- What the model remembers about the tokens already in the window, kept in VRAM so it does not have to re-read them. It grows with the window. Measured on this model, a server allocates 39,936 bytes per window token with no drafter and 45,056 with one.
- prefill and decode
- Two different phases with two different speeds. Prefill is reading your prompt: measured here at about 1,000 t/s on short prompts, slowing to 816 t/s on a 92,679-token one, which took 1.9 minutes. Decode is writing the answer, one token at a time: 40–90 t/s here. Every speed on this page is decode unless it says prefill.
- prose, and novel code
- Prose is ordinary written English — sentences and paragraphs, the kind of text in documentation, an article or an email. Code is programming text. This page keeps saying which of the two a speed was measured on, and that is not padding: the drafter guesses far better on code than on prose, because code is predictable. After
function foo(almost certainly comes); indentation repeats; brackets close in known ways. The next word of an English sentence could be any of hundreds. So the same drafter setting that wins by 12 % on code loses by 10 % on prose (§06), and a speed figure without its content type is not a number you can plan with. “Novel code” means a program the model has not effectively memorised — textbook algorithms like a red-black tree inflate the drafter’s hit rate, so this page’s drafter measurements deliberately ask for something less familiar. - wikitext
- A standard, freely available body of English text (Wikipedia articles) used as the fixed reading material for perplexity scoring, so that every file on this page is judged on identical words. It is prose, which is why the perplexity ranking and the speed rankings can disagree: they are measured on different kinds of text.
- acceptance, and mean draft length
- Two numbers the server reports about the drafter, and the second is the one that predicts speed. The drafter proposes several tokens at once; the full model then checks them and keeps the ones it agrees with. Acceptance is the share it kept. Mean draft length is how many tokens it proposed per attempt. High acceptance is not the same as fast — if the drafter only dares propose two tokens at a time, it can be right almost always and still save you little. Measured on this page: a setting with acceptance 0.897 ran 33 % slower than one with acceptance 0.611, because the second proposed much longer runs. Read draft length first.
- answer tokens vs reasoning tokens
- With thinking switched on, this model writes two things: a private working-out pass, then the reply you actually read. Reasoning tokens are the working-out; answer tokens are the deliverable. They run at different speeds — answers about 1.7× faster than reasoning on code — so a “tokens per second” figure means two different things depending on which it counted. Every speed band on this page says which. If you are waiting at a keyboard for a reply, the answer-token rate is your experience; the reasoning rate is what decides how long you wait first.
- drafter (speculative decoding)
- A small helper that guesses several tokens ahead so the full model only has to check them in one pass. The output is exactly what the model would have written anyway — only the speed changes. This model has one built in, called an MTP head, switched on with
--spec-type draft-mtp. Measured here it roughly doubles decode speed and cuts energy per token by 2.52×. Two numbers describe how well it is doing: draft acceptance is the share of guesses that survive checking, and mean draft length is how many tokens it dares to guess per pass. Acceptance tells you whether the guessing was right; mean draft length tells you how fast you go. Two flags tune it. n-max (--spec-draft-n-max) is how many tokens the helper may guess before the full model checks them — this page ships 4 or 10. p-min (--spec-draft-p-min) is a confidence floor: the helper stops guessing the moment it is less sure than that — 0.75 or 0.5 here. The page writes the pair as n-max 4 / p-min 0.75, shortened to n4/p0.75. - token regime
- Which kind of token is being counted. Reasoning tokens are the model thinking; answer tokens are what you keep. This model thinks by default, so most speed numbers published anywhere are reasoning tokens. On code they run about 1.7× slower than answer tokens on the same machine with the same flags. A speed without its regime cannot be checked.
- effort
- A dial on this model —
low,mediumorxhigh— that sets how long it thinks before answering. Measured across a 175-prompt benchmark suite, it moved wall-clock from 1.0 to 2.7 hours and the composite score by 1.6 points, which is a tie at that sample size. - perplexity
- A quality measure for a model file, not a search engine. Run the model over real text and score how well it predicted each next token. A perplexity of 6.5 means it was on average as uncertain as choosing between about 6.5 equally likely words. Lower is better, and it can only be compared between files of the same model family scored on the same text with the same tokenizer.
- spill
- When the card runs out of VRAM, the driver quietly moves part of the model or its cache into system memory instead of refusing. Nothing errors. The signature is shared GPU memory rising during a run while dedicated sits at its ceiling (§11). Measured here: a configuration that should decode at 50 t/s delivered 8.0 t/s on a 91k-token document with 2,364 MiB spilled.
- percentage point vs percent
- A point is an absolute unit of a score: 94% to 95% is one point. "Percent better" is relative and means something else. Points and questions are chained: in a scored run of n=20 each question is worth 5 points, at n=25 each is worth 4 points, and at n=200 each is worth 0.5. That is why a 25-question benchmark cell cannot separate two healthy files (§10).
- board power, in-band, J/token, tokens/kWh, EDP
- Board power here means what the graphics card itself draws. It is read in-band — from the card's own sensor (NVML), not from a meter at the wall. So it is not wall power: the power supply, the processor and the rest of the machine are excluded and were never measured. J/token is joules of board energy per decode token — measured here from 3.21 (best drafter settings) to 9.29 (worst quant, no drafter). tokens/kWh is the same fact upside down: one kilowatt-hour buys 387,000 to 1,121,000 tokens on this card. EDP is energy-delay product, energy multiplied by how long it took, in joule-seconds — it is the number that punishes a slow setting twice.
- arm
- One configuration inside a comparison run. Two arms differ in exactly one setting, so whatever separates them can be blamed on that setting. A 19-arm power matrix is nineteen such configurations measured one after another.
- rung, and the ladder
- The ladder is this page's series of one model squeezed to different sizes, largest file to smallest; each file in it is a rung. §08 measures eight of them.
- GGUF
- The single-file format llama.cpp loads. One
.gguffile holds the whole model, including the draft head on the files this page recommends. - prefix cache
- The server keeps the processed form of a prompt it has already read, so a later request whose text starts identically skips re-reading that part. Measured here, reading a 92,679-token prompt takes 1.9 minutes and is 90.7% of that request's energy (§05) — a reused prefix pays none of it.
- cooled protocol
- How this page takes a speed measurement that repeats. Fire the prompt, throw away the decode timing that comes back with it, wait 30–45 s, then send several short repeat requests and report the median of those. The first probe after a long prompt reads up to 25% low, because the card is still raising its clock (§11); under this protocol repeat probes agree to 0.4%.
- blind
- Used on this page for judging only: a judge who cannot see which setting wrote the answer being rated (§09). A second kind of blinding also happened here, and it is deliberately called something else — the 2026-08-23 round that re-measured this page's own figures without sight of them is called the independent re-measurement throughout, so the two can never be confused.
- functional detector
- An automated check for the failures a score cannot see: an answer that comes back empty, output that repeats until it hits the cap, a broken chat template. In §08 they are the empty-answer and truncation columns printed beside perplexity and accuracy.
- empty answer · at cap · silent
- Three words this page's failure columns use, and they name different failures. An empty answer is a reply with no characters in it: you asked, and nothing came back. At cap means that reply stopped because it ran out of its token budget and was cut off — the server records that as a truncation, so it shows up in a log. Silent means the opposite: the model ended the reply by itself, in the ordinary way, and handed back nothing. A silent empty is the dangerous one, because no truncation counter can see it and nothing in your logs reports a problem at all. Measured here, empty answers are exactly zero down to 2.912 bits per weight and then rise; five of the five at 1.994 bits are silent (§08).
- paired test, discordant, and
p - Paired means the same questions were put to both files and compared one by one, so only the questions they answer differently count; those are the discordant ones.
pis how often a difference this large would turn up by chance if the two were really equal —p=1.00 means constantly,p=0.0001 means almost never. - greedy decoding, pass@1, ROUGE-L F1
- Three scoring words that appear in this page's benchmark tables. Greedy: always take the most likely next token, so the same prompt gives the same answer every time. pass@1: the share of problems whose first generated program actually runs and passes its tests. ROUGE-L F1: how much of a reference summary's wording a generated summary reproduces, scored out of 100.
- Ampere, Blackwell, Battlemage
- Names for generations of graphics chips, used here because the vendors and the other documentation you will read use them. Ampere is NVIDIA's RTX 30 series — this page's own 3090 is an Ampere card. Blackwell is the RTX 50 series, GB10 and B200. Battlemage is Intel's Arc B series, including the B580 and the Arc Pro B50 and B70.
xhigh run whose thinking wants 61,000–76,000 tokens fails by arithmetic inside a 49,152-token window, and returns nothing (2026-08-22 and 2026-08-23, RTX 3090).03
The recipes — copy one for your card
One recipe per class of card. Each is the single best configuration this page can defend from its own measurements, not a menu of options. The 24–32 GB class is the one exception: two files won there, split by what you value, so both are printed and the rule for choosing between them is stated first. The reasoning behind every flag is in §04 to §09; here you copy a block.
The choosing rule for a 24 GB card, in one sentence: take the default recipe (UD-IQ4_XS) unless you want the highest measured quality at a 122k window, in which case take the Q4_K_M recipe (Q4_K_M) and accept that it leaves only about 0.3 GiB of the card for your desktop — less than a browser typically holds. The two files are statistically tied on quality — perplexity 6.535 against 6.596, a 0.9% gap against a combined error bar of ±0.063, so the difference is smaller than the uncertainty on it — while B is 2.1 GiB smaller and measured faster in every comparison run: 42.97 against 39.99 t/s with no drafter, 93.9 against 81.7 at the best drafter settings, both pairs from the same matched sweep on the same prompt. The smaller file is what buys you both vision and a big window at the same time.
the Q4_K_M recipe's quality claim rests on two metrics leaning the same way: perplexity (measured, 2026-08-22) and a 200-question GSM8K comparison that scored Q4_K_M 94.0% against UD-IQ4_XS 93.0%. That second number is now unaudited. The 2026-08-23 benchmark sweep found a grader bug of exactly the shape that would produce it — the grader compared the whole answer line instead of the number, which only bites the arm that writes units — and moved one score by 8 points when fixed. The GSM8K runs behind the Q4_K_M recipe used a different grader and have not been re-scored, so treat the 94.0 / 93.0 ordering as unresolved and the Q4_K_M recipe's edge as resting on perplexity alone (§08). It also costs about 8% more energy per token than the default recipe (8.853 against 8.198 J per decode token, measured with the drafter off, 2026-08-23).
NVIDIA 24–32 GB · 3090 / 4090 / 5090 — THE DEFAULT · UD-IQ4_XS (vision, speed, big window)
n-max 4 when the text-only picks get n-max 10 — measured 2026-08-25
A reader asked the right question: if the wider drafter is faster, why does the vision recipe ship the conservative one — and would trading window for it be worth it? Earlier editions answered by assertion, and priced the trade against a reasoning stream at 5.6 % SUPERSEDED when the reader in question was writing code. It has now been run.
First: there is no window to trade. At this same 122,880 window, deep-filled to 112,735 real tokens, the wide drafter delivers 62.47 t/s against 53.33 — +17.1 % — and still leaves 1,776 MiB. Shrinking -c to buy it back is pure loss: at 98,304 it measured 87.04 against 87.49 t/s, inside noise, for 24,576 fewer tokens of context. So on text-only work the wide drafter simply wins, which is why picks 2 and 3 carry it.
Then the image goes in, and that is what decides it. With a real 1440p screenshot in flight on top of 105,397 tokens, board VRAM was sampled every 0.5 s across the whole request — because a peak between a load reading and a depth reading is exactly what an end-point pair cannot see. It peaks at 23,450 MiB, leaving 1,126 MiB against the 1,796 MiB this page reserves for a desktop.
It is short by 182 MiB, and the same configuration varied 397 MiB between two loads on the same day. So the honest verdict is not “fails” but “marginal” — and marginal against a safety reserve is a no, because the failure mode is not an error. It is a silent spill into system memory that costs half your speed and logs nothing (§11). The conservative drafter stays on this pick. If you serve text-only, take pick 2 and the wider drafter with it.
And if you want the wide drafter WITH vision anyway, here is the one that works meas — because “no” at 122,880 is not “no” everywhere. Four routes were measured on the same image-at-depth test:
| route to n10/p0.5 with vision | peak | slack | verdict |
|---|---|---|---|
as shipped, -c 122880 | 23,450 | 1,126 | no — 182 short |
plus -ctkd q8_0 -ctvd q8_0 (quantise the draft cache, which -ctk never reaches) | 23,321 | 1,255 | no — 53 short |
-c 98304 keep vision, keep the wide drafter, spend 24,576 tokens of window | 22,258 | 2,318 | YES — clears by 1,010 |
| headless (no desktop, so no reserve to clear) | 23,450 | 1,126 | YES |
So the answer is -c 98304. It was verified the same way as the refusal above — a real 1440p screenshot on top of 77,194 tokens, board VRAM sampled every 0.5 s — and the model still read the 15 px table cell correctly at that depth. Quantising the draft cache is worth having in the ledger and is not enough on its own: it saved 129 MiB where the per-token arithmetic predicted roughly 300, which is a reminder that the drafter’s memory is not all cache.
Before you take it, check you want it. The wide drafter is the code peak, not the speed peak in general. On prose the conservative setting wins — 48.4 against 43.8 t/s — and speculation is worth only 1.16× there at all, so writing work is better served by --spec-type none and the memory back. And the advantage shrinks with depth: +33.2 % on an empty window, +17.1 % at 112k of real fill. A long agent session gets about half the figure that sells the setting.
Two things worth keeping from that run. The model answered a fine-detail question about the screenshot correctly at 105,397 tokens of depth — reading a 15 px table cell — so vision is not degraded by context pressure, only memory is. And the wide drafter's advantage measured +33.2 % on an empty window against +17.1 % at real depth: a shallow probe overstates speculation by about double, which is worth knowing before trusting any drafter comparison taken at an empty context.
The file: unsloth/Qwen3.8-27B-GGUF, UD-IQ4_XS — 14.25 GB on the download page, 13.3 GiB resident (download pages count in decimal GB, memory budgets in binary GiB; the ~7% gap is real and matters whenever memory is tight). Plus the BF16 mmproj projector from either repository if you want vision. The window: -c 122880. The flags: all layers on the GPU, one slot, an 8-bit KV cache, and the conservative drafter. The effort levels it offers: all three — low, medium and xhigh. Nothing is excluded here: xhigh's measured thinking appetite is 61,500–75,800 tokens and a 122,880-token window holds its upper tail with room for a prompt and an answer. Expected speed, measured with these exact flags: 83–86 t/s of answer tokens on short-context code — 83.5 on a novel rate-limiter prompt at -c 32768 and 86.3 at a 1,458-token fill under the cooled protocol, the two endpoints differing by content rather than by configuration. Then 64.8 at 91k of depth, and 37–39 t/s on the xhigh reasoning stream at 91k — which is where a long agent run actually spends its wall clock. VRAM at the top of this window: 20,658 MiB of server process by the measured budget model, leaving 2.7–3.7 GiB of board VRAM depending on what your desktop is holding. Room left for your desktop: 2.7–3.7 GiB — this is the row chosen so that a browser and an agent's web interface do not push it over.
:: TAKE THIS ONE for everyday use: quality statistically TIED with Q4_K_M (6.596 vs :: 6.535, +0.9%, smaller than the +/-0.063 combined error bar), measured FASTER :: everywhere (decode 42.97 vs 39.99 with no drafter and 93.9 vs 81.7 at n10/p0.5, :: both from one matched sweep; prefill 1365 vs 1303), 2.1 GiB smaller — and :: this block ships VISION on. The Q4_K_M block below is kept only because you have :: 122k, and only if you can free the VRAM your desktop is using - A leaves it about :: 0.3 GiB. Caveat: IQ-format decode (files whose name contains IQ) was measured on :: CUDA only - verify on Vulkan or Metal before trusting these speeds there. :: file: unsloth/Qwen3.8-27B-GGUF (UD-IQ4_XS, 14.25 GB = 13.3 GiB) + BF16 mmproj llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 :: min 1024 is a FLOOR only — it lifts small images and never :: shrinks a big one (measured: 720p 922 -> 1,077 tokens, 1440p :: unchanged at 3,602). max 10580 = 4K-class detail. Per-image :: token costs: section 12. The projector costs 1,138 MiB of :: VRAM and 0% of decode (0.04-0.09% at a 91k fill, section 06) -c 122880 :: 20,658 MiB of server VRAM by the measured budget model :: (section 05), leaving 2.7-3.7 GiB for a desktop that itself :: measured 1,179-1,669 MiB idle, no server. The largest window :: whose every byte still lives on the card in THIS configuration :: is 163,840; 122,880 is the daily value, chosen to leave a :: desktop room -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: -ngl 99 is not a guess: llama.cpp counts the output layer as :: layer 65, so -ngl 64 leaves the output layer on the CPU and :: costs 29.5% of decode with NO memory signature (section 11) --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: the CONSERVATIVE :: drafter, kept here on purpose. n-max 10 / p-min 0.5 is the :: measured PEAK on code — 93.9 vs 83.5 t/s on this file, 81.7 :: vs 69.8 on Q4_K_M, same ranking on both (section 06) — but it :: costs 898 MiB, and THIS row already spends 1,138 MiB on the :: projector. It also buys least where this default actually :: runs: on xhigh REASONING tokens at 91k it is worth 5.6% :: (38.7 vs 36.6 t/s). Put n10/p0.5 on the text-only rows. :: On PROSE, drop speculation entirely: 1.16x for ~1.8 GiB --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: all three levels fit this window ONLY IF you leave :: room for them. The window is one POT: prompt + thinking + :: answer share it. xhigh wants 61,500-75,800 tokens of :: THINKING, so your prompt and history must stay under about :: 47,000 of these 122,880 - 38%% of the window. Past that, :: xhigh truncates and returns NOTHING (section 09). medium :: when you are waiting at the keyboard, and whenever the :: context is deep :: FLAGS DELIBERATELY OMITTED HERE: :: --parallel 2 — SETTLED 2026-08-25 by the matched drafter-ON pair: +22.0% aggregate :: throughput but -35.4% per slot (82.98 -> 101.25 t/s aggregate, 85.79 -> 55.41 :: per slot). The old drafter-OFF +60% was a measurement error, not a real gain - :: with the drafter on for one user, because the drafter is already saving most of :: the repeated reading of the model that a second slot would otherwise save. :: Serving ONE person, one slot is faster; raise it only to serve two users at :: once (section 03's axis table). :: -ctk q4_0 — measured at +0.693% perplexity against fp16 (q8_0 costs +0.309%). :: A knowing trade for a bigger window, not a default (section 08). :: nvidia-smi -pl — MEASURED 2026-08-25. Capping the board is a real :: efficiency win and the only setting here that lowers wattage: 300 W costs :: 5.0% of throughput for 11.2% less power, 250 W costs 12.8% for 24.8%. :: Left out of the recipe because it is a persistent hardware setting and :: needs an elevated shell, not because it does not work (section 11). :: text-only variants — drop the --mmproj/--image-* line and the window grows: :: -c 180224 + n-max 10 / p-min 0.5: the measured text-only ceiling WITH that :: drafter - 23,729 MiB of board VRAM, 847 MiB of slack left for a desktop. :: (196,608 leaves 415 MiB and 212,992 leaves 280, with decode already :: sagging 56.6 -> 51.9 -> 49.2 on a :: SHORT probe. Earlier editions shipped 196,608 and claimed ~3 GiB of slack; :: that was wrong by about 2.6 GiB.) :: -c 262144 THE FULL NATIVE CONTEXT - but ONLY with --spec-type none AND with no :: graphical session using this card ("headless", section 02: closing everything :: that draws, a second card driving the screen, or a remote login with nobody :: logged in locally - NOT just unplugging the monitor). :: Drafter off it derives to 23,216 MiB (22.7 GiB), leaving 1,360 MiB of board :: VRAM - which is 179 MiB once a desktop takes its measured worst case. WITH the :: drafter on it does not fit at all: 2,364 MiB measured living in system RAM, :: and a 91k-token document then decodes at 8.0 t/s (section 05). :: The rule: a ceiling belongs to the whole configuration — file + drafter flags + :: projector + the desktop you run. Quote all four or the ceiling is not portable.
NVIDIA 24–32 GB — THE BIG-WINDOW PICK · UD-Q2_K_XL (vision and the drafter and 196,608 tokens)
The file: unsloth UD-Q2_K_XL — 9.83 GB, 9.154 GiB resident. The window: -c 196608. The flags: the universal set, the conservative drafter, and the full image budget. VRAM at depth: 22,014 MiB meas — peak, with a 1440p screenshot in flight on 163,124 tokens, sampled every 0.5 s. Room left for your desktop: 2,562 MiB, clearing this page’s 1,796 MiB reserve by 766. Effort: medium is what this pick ships, and the reason is the window rather than the clock. The window is one pot — prompt, thinking and answer share it. An ordinary xhigh answer only wants about 2,217 tokens of thinking, which fits here easily; but this pick exists to fill 196,608, and at the 163,124 tokens it was measured at only 33,484 remain. A complete-build task — the one thing measured wanting 61,500–75,800, four samples of one task (§09) — would truncate there and return nothing. So: xhigh for questions of a full window, medium when you are asking for a whole program. Earlier editions blamed the 278-second prefill SUPERSEDED: effort changes how much the model writes, not how long it takes to read your prompt.
Why this pick exists, added 2026-08-25. A reader wanted the largest window they could get while keeping both vision and the drafter, and this page had no answer for them — the 4-bit file cannot do it, and the full 262,144 fails in every drafter state once an image is in flight. The 2-bit file can, and three separate measurements say it costs less than its name suggests: it ties the 4-bit reference on accuracy (one of 75 paired items different, p=1.00), it reads images identically (7/7 against 7/7 on a generated detail target), and its window genuinely retrieves — 5 of 5 needles found at every depth out to 241,655 tokens, with a clean control (§08). What you actually trade is decode speed at short windows, where the 4-bit file drafts better.
196,608 is the ceiling, measured. -c 229376 peaks at 23,529 MiB and keeps only 1,047 — it is under the reserve at load, before an image is even sent. So this window is not a round number chosen for looks; it is the largest that survived the test. One honest note on that failing run: its answer to the image question came back as prose rather than the bare number, which the scoring here cannot grade either way. It is reported as a memory failure only, because that is what was measured.
The cost that matters more than VRAM: prefill. Filling this window took 278 seconds meas, against 210 s at -c 163840 and roughly 8 minutes at the full 262,144. Prefill is compute-bound and does not get cheaper. If your agent appends to its context you pay this once and the prefix cache carries you; if it rebuilds context every turn you pay it every turn, and -c 163840 saves you a minute for 32,768 fewer tokens. Choose on that, not on the bigger number.
:: THE BIG-WINDOW PICK: vision + drafter + 196,608 tokens, which no 4-bit :: configuration on this page can hold. MEASURED 2026-08-25 with a 1440p image :: in flight at 163,124 tokens: peak 22,014 MiB, 2,562 MiB left for a desktop. :: 196,608 IS THE CEILING, and it is measured rather than assumed: :: -c 196608 peak 22,014 MiB 2,562 MiB slack CLEARS by 766 :: -c 229376 peak 23,529 MiB 1,047 MiB slack FAILS by 749 :: -c 262144 peak 22,877 MiB 1,699 MiB slack FAILS by 97 (drafter OFF!) :: 229,376 is already under the reserve AT LOAD (1,516) before an image lands. :: file: unsloth/Qwen3.8-27B-GGUF (UD-Q2_K_XL, 9.83 GB = 9.154 GiB) llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 :: full image budget - the 1024 cap makes the model :: confidently MISREAD fine detail (section 05) -c 196608 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: 22,014 MiB peak MEASURED at depth with an image, not :: at load. Drop to -c 163840 for 3,852 MiB of slack and :: a minute less prefill --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: the wider :: n10/p0.5 is the CODE peak but costs ~900 MiB and is :: not measured at this window - do not assume it fits --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: the window is one POT :: an ordinary xhigh answer averages 2,217 tokens (175 benchmark :: prompts), which the 33,484 left at the 163,124-token depth above :: holds 15 times over. medium is shipped because a COMPLETE BUILD :: is the one thing measured wanting 61,500-75,800 (four samples of :: one task, section 09) and that would truncate here. Ask questions :: of a full window at xhigh; ask for a whole program at medium. :: This has nothing to do with the 278 s prefill: effort changes how :: much the model WRITES, not how long it takes to READ your prompt.
NVIDIA 24–32 GB — THE ONE YOU HAVE HEARD OF · Q4_K_M (no measurable quality advantage, 2.1 GiB larger, ~0.3 GiB left for your desktop)
Q4_K_M is the best-known 4-bit GGUF and that is the main reason it is still here. It is kept so that an option you have already heard of is answered rather than absent. But it is not a close second, and earlier editions of this page overstated it: they labelled it “maximum measured quality” SUPERSEDED, and it does not have a measurable quality advantage.
The perplexity gap is 0.061 (6.535 against 6.596). This page's own yardstick for comparing one file with another is a combined error of ±0.063 (§08) — and that passage names this very pair when it says so. The gap is inside the bar. It is not resolved. The second leg of the old argument, a 200-question GSM8K comparison, was measured with a grader later found to carry three bugs and has never been re-scored. So the case rests on one unresolvable difference and one withdrawn instrument.
What it costs to keep believing it: 2.1 GiB, which is either the vision projector or about 50,000 more tokens of window; roughly 0.3 GiB left for your desktop, against the 1,796 MiB this page reserves for one; and no room for the wider drafter. Take the default above unless you have a specific reason not to, and if you already downloaded this file, you have lost nothing worth measuring.
The file: lmstudio-community/Qwen3.8-27B-GGUF, Q4_K_M — 16.5 GB on the download page, 15.4 GiB resident. The window: -c 122880. The flags: the same universal set as the default recipe, with the conservative drafter — there is no room here for the wider one. The effort levels it offers: all three, for the same reason as the default recipe. An ordinary xhigh answer averages 2,217 tokens of thinking across 175 benchmark prompts; the 61,500–75,800 figure quoted elsewhere on this page is a ceiling from four samples of one task — a complete build — and a 122,880-token window holds even that, leaving about 47,000 for your prompt. Expected speed, measured with these exact flags: 69.8 t/s of answer tokens on short-context code and 57.9 t/s on the reasoning stream; the 81.7 t/s figure this file is famous for on this page belongs to the wider drafter, which this row cannot afford — it is a cross-reference, not this recipe's speed. VRAM at the top of this window: about 23.7 GiB of board VRAM with a desktop up, measured, leaving roughly 0.3 GiB. Room left for your desktop: about 0.3 GiB — less than a browser typically holds. §11 measured this very configuration spilling about 1.9 GB with two browsers open and falling to 20–35 t/s.
:: TAKE THIS ONE only if you believe an unresolvable 0.9% perplexity edge is :: worth 2.1 GiB AND you can free the VRAM your desktop uses - this row leaves :: about 0.3 GiB, less than a browser typically holds. The edge is 0.061 of :: perplexity against this page's own +/-0.063 comparison error: INSIDE THE BAR. :: The GSM8K half of the old argument is unaudited. THE DEFAULT ABOVE is better :: on context, speed, vision and desktop headroom, and no worse on any :: measurement this page can resolve. :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M, 16.5 GB = 15.4 GiB) :: before first run: set "Prefer No Sysmem Fallback" (NVIDIA Control Panel) and :: never write -ngl 64 — both explained in section 11 llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 122880 :: 24 GB: the largest window whose every byte still lives on the :: card is ~131k, so this leaves slack (section 05's two :: ceilings). 5090 32 GB: -c 262144 :: fits — budget ~11 GiB of KV at the MEASURED slope, not the :: 8.5 GiB the q8 arithmetic suggests -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: NOT a per-file :: setting — the 2026-08-23 matched sweep ranks n10/p0.5 first on :: BOTH files (81.7 vs 69.8 here). It is kept at n4 only because :: n10 costs 898 MiB and this ~0.3 GiB-slack row cannot spare it. :: BUDGET the drafter: 1,008 MiB fixed + 11.4% of your window :: (section 05). --spec-type none frees all of it; on prose it is :: worth only ~1.16x, so that is the easiest 1.8 GiB to reclaim :: vision: add --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 :: --image-max-tokens 10580 when you need it. The projector costs 1,138 MiB :: once loaded - not the 0.867 GiB its file weighs - and with this file at -c 122880 :: that is the measured ~23.7 GiB reference configuration: it fits, but leaves only :: ~0.3 GiB of the card for your desktop. With a desktop running, drop to -c 65536, :: or -c 49152 if you keep browsers open. :: Budget from the MEASURED slope: each 32k of window costs ~1.375 GiB :: with the drafter on and ~1.22 GiB with it off; the 1.06 GiB q8 arithmetic is a :: FLOOR. Vision + a desktop + a BIG window = the default recipe above. :: On the 1 GiB-larger UD-Q4_K_XL, projector + full context needs more than 24 GB - :: use -c 98304 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: quality-first default; :: medium when the wait matters (section 09's price list) :: FLAGS DELIBERATELY OMITTED: the same three as the default recipe — --parallel 2 (+22.0% :: aggregate but -35.4% per slot with the drafter on, measured 2026-08-25 — slower :: for one user), -ctk q4_0 (+0.693% perplexity, a window trade), and nvidia-smi -pl :: (unmeasured, needs an elevated shell).
Every number above is tied to llama.cpp build 10502, commit 0adcc3bb5, and NVIDIA driver 596.36. Re-measure the speed band and the VRAM ceiling when any of these changes: a new llama.cpp build (the MTP implementation moved twice in the two days this page covers), a driver update, a new quantization of this model, or a change to the model's own draft head. The cheapest re-check is the one command in §15; it takes about two minutes and tells you whether anything moved.
The full menu for a 24 GB card, losers included
The reference machine's launcher offers eight configurations, and this is all of them — including the ones that lost, each with the number that beat them, so that an option you have already heard of is answered rather than absent. Hover or tap a row to see the flags that produce it.
| # | Configuration | Window | Expect (regime stated) | When to pick it — or why not |
|---|---|---|---|---|
| 1 | UD-IQ4_XS + hi-res vision (max-tokens 10580 · n-max 4 / p-min 0.75) | 122,880 | 83–86 t/s answer tokens, short context · 64.8 at 91k · 37–39 on xhigh reasoning at 91k · 2.7–3.7 GiB slack | the daily default — screenshot loops to 4K detail (§12); a full xhigh cycle plus about ten 1440p shots fit the window. Keeps the conservative drafter, and that is now measured rather than argued — see the box below |
| 2 | UD-IQ4_XS text-only (n-max 10 / p-min 0.5 — the code peak) | 180,224 | 93.9 t/s answer tokens at -c 32768 (the matched sweep's peak; a different probe on novel JavaScript read 95.4) · 847 MiB of slack, measured · little on screen | the largest text-only window that stays resident with this drafter, and the row where the longer guesses pay for themselves (+12% on code on this file, +17% on Q4_K_M). Earlier editions printed 196,608 here and claimed ~3 GiB of slack; it measures 415 MiB |
| 3 | UD-IQ4_XS text-only + --spec-type none | 262,144 | ~43 t/s (no drafter) at load · 15.96 t/s and 755 MiB left when actually filled to 218,233 tokens · drafter left on = 2,364 MiB spilled and 8.0 t/s at a 91k fill | full native context on this file — needs the card free of any graphical session (§02), drafter off, and now measured at depth rather than at load: it keeps only 755 MiB, which is 553 MiB short of the 1,796 MiB this page reserves for a desktop. If you actually need this window, use UD-Q2_K_XL instead — same -c 262144, drafter on at n-max 4 / p-min 0.75, 21.33 t/s at the same depth with 1,717 MiB of slack, and it ties this file on accuracy (§08). The price is a seven-minute prefill |
| 4 | Q4_K_M + vision | 122,880 | 69.8 t/s answer tokens · 57.9 reasoning · ~0.3 GiB slack | maximum measured quality — leaves about 0.3 GiB for your desktop, less than a browser typically holds (browsers spill it to 20–35 t/s, §11); with a desktop up, drop -c or go text-only |
| 5 | UD-Q4_K_XL text-only | 122,880 | ~55–65 t/s reasoning | loser: perplexity 6.682 against Q4_K_M's 6.535 (+2.3%) while being 1 GiB bigger. On 24 GB the extra gigabyte gains you nothing |
| 6 | UD-Q4_K_M + vision | 122,880 | ~55–65 t/s reasoning | loser: measured worse than plain Q4_K_M at the same size — perplexity +1.8%. Skip |
| 7 | DFlash2 external drafter | — | ~46 t/s on real code (best at n-max 2) | loser: beaten by the built-in MTP head's 57.9 on the same content, and it costs a 1.14 GB download plus a source build. Historical interest (§06) |
| 8 | NVFP4-HIGH | — | ~48–54 t/s reasoning | loser on this card's chip generation (Ampere — the RTX 30 series, which the reference 3090 belongs to): software dequantization fallback, perplexity +4.4%, and 13.4% more energy per token than UD-IQ4_XS. It runs natively only on Blackwell — RTX 50, GB10, B200 (§08) |
Every "Expect" band carries its token regime, because answer tokens run about 1.7× the speed of reasoning tokens on code and a band without its regime is not a number you can plan with. Speed also falls with depth, not only with spill: on the shipped flags, answer tokens measure 86.3 → 80.2 → 64.8 t/s at 1.5k / 28k / 91k of fill under the cooled protocol, while the reasoning stream on the same file and flags runs 51.2 → 47.1 → 36.6. An agent session holding 10–30k of context therefore lives near 80 t/s of deliverable; the same session left on the xhigh default over a 91k document spends most of its wall clock at 37–39 t/s, whichever drafter flags it carries. Slack figures are the board VRAM measured left over with little on screen; the desktop's own share swung 1,179–1,669 MiB in direct no-server readings on one machine (§05), so a row keeping under about 1 GiB has room for a near-idle desktop and nothing beyond it.
The launcher behind this table — the reference machine's own serve-qwen.bat, with paths genericized (tap to expand)
@echo off
rem ============================================================================
rem serve-qwen.bat - llama.cpp server for Qwen3.8-27B on RTX 3090 24GB
rem ============================================================================
rem Measured 2026-08-22, corrected 2026-08-23. Every pick carries the
rem measurement that justifies it, because this file travels to other
rem machines without the guide.
rem - -ngl 99 always (llama.cpp counts the output layer as layer 65; -ngl 64
rem leaves the output layer on CPU: 25.7 vs 39.7 t/s on Q4_K_M,
rem 29.8 vs 42.3 on UD-IQ4_XS, with NO memory signature)
rem - MTP flags are PER PICK, and the split is NOT per file. A matched sweep
rem (2026-08-23, both quants, novel code, thinking OFF) peaks at n-max 10 /
rem p-min 0.5 on BOTH files: IQ4_XS 93.9 t/s (2.18x over its 43.0 no-drafter
rem floor), Q4_K_M 81.7 (2.04x over 40.0). The shipped n-max 4 / p-min 0.75
rem reads 83.5 and 69.8 on the same probe - 11%% / 15%% off the peak, and
rem 898 MiB cheaper. So: n10/p0.5 on the text-only picks where that VRAM is
rem spare, n4/p0.75 wherever the projector is loaded or the slack is thin.
rem Acceptance is a property of the MTP head, not of the quant - the two
rem files land within 1.6 points at six of seven configs (3.7 at the seventh).
rem - the drafter is also the cheapest ENERGY setting here: 3.21 J per decode
rem token at n10/p0.5 against 8.10 with --spec-type none (2.52x), and
rem the board draws the SAME power while decoding either way (344.6 / 341.7 /
rem 341.0 W): the entire win is throughput, not wattage.
rem - TOKEN REGIME decides the number. Answer tokens (thinking off) run ~1.7x
rem reasoning tokens, and mean DRAFT LENGTH is what predicts speed, not
rem acceptance: at 91k depth the reasoning stream drafts 2.99 tokens per
rem pass (36.6 t/s) and the answer stream 4.31 (62.0) - at the same ~0.90
rem acceptance. Highest-acceptance config (n2/p0.75, 96.5%%) is the slowest.
rem - Depth, answer tokens, n4/p0.75, cooled probes: 86.3 t/s at 1.5k, 80.2 at
rem 28k, 64.8 at 91k. On the xhigh REASONING stream at 91k expect 37-39 t/s
rem (n10/p0.5 buys only 5.6%% there - not worth 898 MiB).
rem - UD-IQ4_XS (13.3 GiB): quality tied with Q4_K_M (perplexity 6.596 vs 6.535;
rem the GSM8K half of that comparison is unaudited - see the guide),
rem FASTER everywhere measured, and 8%% cheaper per token in energy.
rem Measured ceilings for a window whose every byte still lives on the card:
rem -c 180224 text-only WITH n10/p0.5; the full native -c 262144 ONLY with
rem --spec-type none (1,360 MiB of board VRAM left, and that window needs no
rem graphical session on the card at all); -c 163840 with vision at n4/p0.75.
rem 122880 + vision is the daily default, chosen to leave a desktop room.
rem Filling a window is the only test that counts - a short
rem probe at 262144 with the drafter on looks fine and delivers 8 t/s on a
rem 91k document.
rem - Q4_K_M (15.4 GiB): the quality leader, by one metric and narrowly; +mmproj
rem at -c 122880 (~0.3 GiB of slack - room for a near-idle desktop, no more).
rem - the BF16 mmproj costs 1,138 MiB of VRAM and 0%% of decode: measured at
rem 90,862 tokens of fill, projector on vs off, 0.04-0.09%% apart.
rem - KV cache: q8_0 costs +0.309%% perplexity, q4_0 +0.693%% (measured
rem 2026-08-23, super-linear in bits). q8_0 is what these recipes ship.
rem - reasoning_effort is THE wall-clock knob: ~4x medium on one long
rem authoring task, 1.8x across a 175-prompt benchmark suite where the
rem three levels scored within 1.6 points of each other. llama-server
rem ignores per-request effort, so it must be set here.
rem ----------------------------------------------------------------------------
rem USAGE: serve-qwen.bat [low^|medium^|xhigh] [context] [1-8]
rem arg1 = reasoning effort, default xhigh (quality-first)
rem arg2 = context override (otherwise each choice's safe default)
rem arg3 = model choice 1-8, skips the menu (for scripts)
rem No args: menu below, auto-picks [1] after 8 seconds, so an
rem unattended restart never blocks on a human.
rem ----------------------------------------------------------------------------
set EFFORT=%1
if "%EFFORT%"=="" set EFFORT=xhigh
if /i "%EFFORT%"=="low" goto effort_ok
if /i "%EFFORT%"=="medium" goto effort_ok
if /i "%EFFORT%"=="xhigh" goto effort_ok
echo Unknown reasoning effort "%EFFORT%". Usage: serve-qwen.bat [low^|medium^|xhigh] [context] [1-8]
pause
exit /b 1
:effort_ok
set CTXARG=%2
set PICK=%3
set MODELS=C:\path\to\your\models
set MMPROJ=%MODELS%\mmproj-Qwen3.8-27B-BF16.gguf
rem the drafter is PER PICK (it is a load-time flag): n10/p0.5 is the code peak
rem on both files but costs 898 MiB, so it rides the text-only picks; every pick
rem carrying the projector keeps n4/p0.75.
set SPEC_FAST=--spec-type draft-mtp --spec-draft-n-max 10 --spec-draft-p-min 0.5
set SPEC_SAFE=--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75
if not "%PICK%"=="" goto pick_%PICK%
echo.
echo Qwen3.8-27B - pick a model (effort: %EFFORT%):
echo [1] UD-IQ4_XS + HI-RES vision -c 122880 n4/p0.75 ~3 GiB slack (DEFAULT)
echo ~83-86 t/s answer tokens on code,
echo ~65 at a 91k fill, 37-39 on xhigh
echo reasoning at 91k;
echo screenshots to ~4K detail (1080p=2.0k /
echo 1440p=3.6k / 4K=8.2k ctx tokens each);
echo full xhigh cycle + ~10 shots fit
echo [2] UD-IQ4_XS text-only -c 180224 n10/p0.5 code-regime peak 93.9 t/s
echo the measured text-only ceiling WITH this
echo drafter: 847 MiB slack, little on screen
echo [3] UD-IQ4_XS text-only -c 262144 DRAFTER OFF - required, not optional
echo full native - needs the card free of any
echo graphical session;
echo no drafter = ~43 t/s short-context floor
echo [4] Q4_K_M + vision -c 122880 n4/p0.75 highest measured quality,
echo reference config
echo ~0.3 GiB slack - a near-idle desktop only
echo [5] UD-Q4_K_XL text-only -c 122880 n4/p0.75 quality LOSES to Q4_K_M
echo (perplexity +2.3%%) and is 1 GB bigger:
echo gains nothing
echo [6] UD-Q4_K_M + vision -c 122880 n4/p0.75 measured worse than plain Q4_K_M
echo (perplexity +1.8%%) - skip
echo [7] DFlash2 build ~46 t/s on real code (loses to built-in MTP)
echo [8] NVFP4 HIGH ~48-54 t/s dequant fallback, lower quality,
echo +13%% J/token vs IQ4_XS
echo.
echo t/s above are ANSWER tokens (thinking off) at short context unless said
echo otherwise. Reasoning tokens run ~1.7x slower on the same server, and
echo decode falls with DEPTH: 86 shallow / 80 at 28k / 65 at 91k (answer),
echo 51 / 47 / 37 (reasoning). Acceptance actually RISES with depth
echo (0.80 -^> 0.92): the cost is KV reads, not the drafter failing.
echo.
choice /C 12345678 /T 8 /D 1 /M "Choice"
goto pick_%ERRORLEVEL%
:pick_1
set MODEL=%MODELS%\Qwen3.8-27B-UD-IQ4_XS.gguf
set CTX=122880
set MM=--mmproj "%MMPROJ%" --image-min-tokens 1024 --image-max-tokens 10580
rem n4/p0.75 on purpose: this pick already spends 1,138 MiB on the projector,
rem and the n10 flags would take another 898 MiB out of the slack that
rem absorbs the desktop (measured 1,179-1,669 MiB). It also buys least here -
rem 5.6%% on the xhigh reasoning stream this default actually decodes at depth.
set SPEC=%SPEC_SAFE%
goto launch
:pick_2
set MODEL=%MODELS%\Qwen3.8-27B-UD-IQ4_XS.gguf
set CTX=180224
set MM=
rem text-only: the 898 MiB is spare here, so take the code peak (+12%%).
set SPEC=%SPEC_FAST%
goto launch
:pick_3
set MODEL=%MODELS%\Qwen3.8-27B-UD-IQ4_XS.gguf
set CTX=262144
set MM=
rem drafter OFF is required, not a preference: with it on this window needs
rem ~25.5 GiB and 2,364 MiB was measured living in system RAM.
set SPEC=--spec-type none
echo.
echo WARNING: -c 262144 fits ONLY with the drafter off, and then only just:
echo 23,216 MiB derived, 1,360 MiB of board VRAM left - 179 MiB once a worst-case
echo desktop takes its measured share. WITH the drafter on it does not fit at
echo all, and a 91k-token document then decodes at 8.0 t/s. A browser or agent
echo web UI will spill it either way. Run this with NO graphical session using
echo the card: close everything that draws, drive the screen from a second card,
echo or log in from another machine. Unplugging the monitor does not do it.
echo.
goto launch
:pick_4
set MODEL=%MODELS%\Qwen3.8-27B-Q4_K_M.gguf
set CTX=122880
set MM=--mmproj "%MMPROJ%" --image-min-tokens 1024
rem ~0.3 GiB slack with the projector loaded: no room for the n10 flags here.
set SPEC=%SPEC_SAFE%
goto launch
:pick_5
set MODEL=%MODELS%\Qwen3.8-27B-UD-Q4_K_XL.gguf
set CTX=122880
set MM=
rem kept in the menu as an answered loser: perplexity 6.682 vs Q4_K_M's 6.535.
set SPEC=%SPEC_SAFE%
goto launch
:pick_6
set MODEL=%MODELS%\Qwen3.8-27B-UD-Q4_K_M.gguf
set CTX=122880
set MM=--mmproj "%MMPROJ%" --image-min-tokens 1024
rem answered loser: perplexity 6.654, +1.8%% over plain Q4_K_M at the same size.
set SPEC=%SPEC_SAFE%
goto launch
:pick_7
rem DFlash2 needs its own build (llama.cpp PR 27342) and a separate 1.14 GB
rem drafter; best on real code measured 46.5 t/s at n-max 2, below MTP's 57.9.
call serve-qwen-dflash2.bat %EFFORT%
exit /b
:pick_8
rem NVFP4 on Ampere runs through a software fallback: slower, +4.4%% perplexity,
rem +13%% J/token. Kept for Blackwell users and for the record.
call serve-qwen-nvfp4.bat HIGH %EFFORT%
exit /b
:launch
if not "%CTXARG%"=="" set CTX=%CTXARG%
if not exist "%MODEL%" (
echo Model not found: %MODEL%
pause
exit /b 1
)
echo Serving %MODEL%
echo effort=%EFFORT% ctx=%CTX%
echo spec=%SPEC%
llama-server.exe ^
-m "%MODEL%" ^
%MM% ^
--alias qwen/qwen3.8-27b ^
-c %CTX% ^
-ngl 99 ^
--parallel 1 ^
--load-mode none ^
--api-key dummy ^
-ctk q8_0 -ctv q8_0 ^
%SPEC% ^
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 ^
--reasoning-preserve ^
--chat-template-kwargs "{\"reasoning_effort\":\"%EFFORT%\"}" ^
--jinja ^
--host 127.0.0.1 ^
--port 1234
Three structural differences from the file that actually runs on the reference machine, named so that nobody assumes this is a byte-for-byte copy. The model paths are genericized to %MODELS%; picks 7 and 8 hand off to sibling scripts whose paths are local; llama-server.exe is unqualified here where the real file gives an absolute path, so the real file's matching "binary not found" guard is dropped with it. The eight picks, their windows, their drafter split, the timed default and every flag are identical. The launcher on disk was synced to this listing on 2026-08-23 (comments and echo text only; behavior verified identical across all eight picks), so the remaining differences are the three structural ones above.
NVIDIA 16 GB · 5080 / 4080 / 4070 Ti S / 5060 Ti
The file: unsloth UD-Q2_K_XL — 9.83 GB, 9.154 GiB resident. The window: -c 65536. The flags: the universal set plus the conservative drafter. VRAM at the top of this window: 13,982 MiB meas. Room left for your desktop: 606 MiB against the 14,588 MiB this page budgets for a 16 GB card — enough for a light desktop, not for a browser full of tabs; drop to -c 49152 if you need more. The effort levels it offers: low and medium. xhigh is not offered: its measured thinking appetite is 61,500–75,800 tokens against a 65,536 window, so it would fit on its shortest runs and truncate on its longest — which is worse than a clean no, because you would not know which you got. Expected speed: 25–50 t/s drafter-off, derived from the bandwidth formula in §04 der; the drafter adds to that. No speed on a 16 GB card is measured on this page and none can be, from a machine that owns one 24 GB card. The memory figure is measured and transfers to any card.
This recipe changed on 2026-08-25, and the reason is worth more than the change. It used to ship UD-Q3_K_XL at -c 49152, priced at “about 15.3 GiB” SUPERSEDED. That figure was derived from a bandwidth estimate, and when the requirement was finally measured file by file (§08) it came out wrong in the direction that breaks a machine: UD-Q3_K_XL with this drafter needs 15,530 MiB at -c 32768 and 16,906 MiB at -c 65536, so -c 49152 lands near 16,218 MiB — 1,142 MiB over the budget on this very page. It does not fit. The smaller file gives you twice the window inside less memory, and on the paired accuracy test it ties the 4-bit reference (§08), so almost nothing is being traded away. The old recipe is left below rather than deleted, because a derived number that was never checked against a measurement is exactly the failure this page exists to document.
:: best daily experience on 16 GB, MEASURED 2026-08-25: the 2-bit file, not the :: 3-bit one. UD-Q2_K_XL at -c 65536 with this drafter needs 13,982 MiB; UD-Q3_K_XL :: at the same window needs 16,906 and does not fit. See section 08's requirement table. :: Q4-with-CPU-offload preserves more fidelity but crawls at ~8-12 t/s (section 07). :: file: unsloth/Qwen3.8-27B-GGUF (UD-Q2_K_XL, 9.83 GB = 9.154 GiB) llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: 13,982 MiB MEASURED on the reference 3090 at this exact :: window with this exact drafter - not arithmetic. Against :: the 14,588 MiB this page budgets for a 16 GB card that :: leaves 606 MiB for your desktop. Drop to -c 49152 if :: you run a browser, or --spec-type none to buy back ~1.3 GiB. :: A >131k window cannot fit resident on 16 GB - that needs :: the Q4 offload path (section 07) --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: start at 4 and :: sweep on your own card (section 06): the optimum moved with :: content and hardware in every run measured here --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: xhigh NOT OFFERED: :: measured appetite 61.5-75.8k tokens against a 65,536 window, :: so it would fit on short runs and truncate on long ones. You :: would not know which you got - see section 09's ceiling table :: WHAT THIS RECIPE USED TO BE, kept because the error is instructive: UD-Q3_K_XL :: at -c 49152, priced at "~15.3 GiB" from a DERIVED bandwidth estimate. Measured, :: it needs ~16,218 MiB there - 1,142 MiB OVER this page's own 16 GB budget. :: FLAGS DELIBERATELY OMITTED: --parallel 2 (settled 2026-08-25: +22.0% aggregate, :: -35.4% per slot with the drafter on - slower for one user), :: -ctk q4_0 (+0.693% perplexity - with the 2-bit file the window no longer needs :: buying, so the one card class that could defend that trade no longer has to).
12 GB class · RTX 3060 / RTX 5070 / Arc B580
The honest framing first: this class runs a 27B model poorly in every configuration measured or derived here, and if the exact model is not a requirement, a 14B-class Qwen at 4-bit fits resident and is several times faster. The file, if it must be this model: Q4_K_M with most layers on the processor. The window: -c 112640. The effort levels it offers: all three fit the window — but the recipe ships medium, and here the binding constraint is wall clock, not the window. How much wall clock depends entirely on what you ask for. At 6–8 t/s an ordinary xhigh answer — 2,217 tokens of thinking, averaged over 175 benchmark prompts — takes about 5 minutes. A complete build, the one task measured wanting 61,500–75,800, takes about three hours against one for medium. Earlier editions quoted the three hours as the cost of “an xhigh answer” SUPERSEDED, which overstates ordinary work by roughly thirty times. Ask questions at xhigh freely here; save the overnight run for whole deliverables. Expected speed: 6–8 t/s, derived, and set by your system memory bandwidth rather than by the graphics card. Room left for your desktop: nearly all of the card — almost none of the model is on it.
:: honest framing: this class runs the 27B poorly in every configuration. If the :: exact model is not required, a 14B-class Qwen at Q4 fits resident and flies. :: If it must be this model: Q4 with most layers on CPU (~6-8 t/s; speed is set by :: your system RAM. The ~6-8 t/s assumes TWO sticks of DDR5 in dual channel :: (~90 GB/s); one stick = single channel = half the bandwidth = half the t/s). :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M — on a bandwidth-bound :: offload path the smaller Q4 file means fewer GB re-read per token, and the :: measured perplexity in section 08 favours it too) llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 112640 -ngl 28 :: ~28 of the 65 offloadable layers (64 + the output :: layer — section 11); raise until <500 MB VRAM is free --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: CPU layers need ~10 GB of real RAM :: no MTP flags here on purpose: speculation is UNMEASURED on the CPU-offload path. :: The measurement that would justify adding them is one paired probe with and :: without --spec-type draft-mtp at this -ngl, on your machine. It costs ten minutes. --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: unlike the 16 GB card :: this is NOT a window limit — the 112k window (mostly in system RAM) holds xhigh's :: 61.5-75.8k appetite fine. It is a PATIENCE limit: at ~6-8 t/s an xhigh answer runs :: ~3 HOURS against ~1 h for medium (section 09). Switch to xhigh for unattended runs. :: Arc B580: use the Vulkan build (same SYCL caveat as the Arc Pro cards - they are :: all Battlemage, Intel's Arc B series)
Intel Arc Pro B70 · 32 GB
The file: Q6_K — 22.4 GB, 20.9 GiB resident, near-8-bit quality, which the extra VRAM affords. The window: -c 229376 with the drafter, or the full 262144 with --spec-type none. The effort levels it offers: all three; the window holds xhigh's appetite several times over. Expected speed: 18–26 t/s, derived from bandwidth and unmeasured on this card's chip generation (Battlemage — Intel's Arc B series). VRAM at the top of the window: about 30.5 GiB derived with the drafter off. Room left for your desktop: not measured — every figure for this card is derived, so treat the last gigabyte as unverified.
:: VULKAN llama.cpp build - currently beats SYCL on Battlemage, Intel's Arc B series, :: which this card belongs to (SYCL has known perf :: bugs there; IPEX-LLM is archived — do not use it). Re-benchmark after driver updates. :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q6_K, 22.4 GB = 20.9 GiB) llama-server.exe -m Qwen3.8-27B-Q6_K.gguf --alias qwen/qwen3.8-27b -c 229376 -ngl 99 :: NOT the full 262144 with a drafter loaded: at the :: MEASURED 45,056 B/token the native window costs ~11 GiB of :: KV, not the 8.5 the q8 arithmetic suggests, and Q6_K + MTP + :: that needs more than 32 GB. For the FULL 262144 here, add --spec-type :: none and drop the MTP line below (~30.5 GiB). Both figures :: are DERIVED from the 3090's constants — unmeasured on Battlemage --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: verify KV-quant support in your build --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: start at 4, sweep (section 06) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: all three fit, with room to spare at THIS window (229,376). :: An ordinary xhigh answer averages 2,217 tokens of thinking :: (175 benchmark prompts). Only a COMPLETE BUILD has been :: measured wanting 61,500-75,800 - four samples of one task, :: section 09 - and even that leaves ~153,000 here for prompt :: and history. Earlier editions of this comment quoted 122,880 :: and ~47,000: copied from the 24 GB recipe, never re-done :: against this window. SUPERSEDED 2026-08-25 :: FLAGS DELIBERATELY OMITTED: --parallel 2 (+22.0% aggregate, -35.4% per slot with :: the drafter on, measured 2026-08-25 — see section 03's axis table).
Intel Arc Pro B50 · 16 GB
The file: UD-Q2_K_XL — 9.83 GB, 9.154 GiB resident. A 4-bit file does not fit, and at 224 GB/s this card is very slow as soon as any layer has to sit on the processor, so every byte must be resident. The window: -c 65536. VRAM: 13,982 MiB meas on the reference 3090 at this window with this drafter — a memory requirement is a property of the model and the flags, so it transfers to this card even though no speed here does. Room left for your desktop: 606 MiB of the 16 GB. The effort levels it offers: low and medium. xhigh is not offered, on the same window constraint as the NVIDIA 16 GB recipe: 61,500–75,800 tokens of appetite against a 65,536-token window, so it fits on short runs and truncates on long ones. Expected speed: 8–12 t/s der from this card’s 224 GB/s and the bandwidth formula in §04 — no Intel card was ever measured for this page.
Changed 2026-08-25, and this recipe was the one that had no warning on it. It shipped UD-Q3_K_XL at -c 49152 budgeted at “~15.3 GiB” SUPERSEDED, copied from the NVIDIA 16 GB recipe. When that budget was measured rather than derived it came out at ~16,218 MiB, 1,142 MiB over what a 16 GB card has to give (§08). The NVIDIA recipe got a correction note when that was found; this one was missed for a day, which is exactly how a copied recommendation outlives the measurement it was copied from.
:: Q4 does not fit, and at 224 GB/s any layer left on the CPU is very slow - so run :: the 2-bit file with every byte on the card, Vulkan build. MEASURED 2026-08-25: :: UD-Q2_K_XL at -c 65536 needs 13,982 MiB; UD-Q3_K_XL at that window needs 16,906 :: and does not fit at all. Section 08 has the full requirement table. :: file: unsloth/Qwen3.8-27B-GGUF (UD-Q2_K_XL, 9.83 GB = 9.154 GiB) llama-server.exe -m Qwen3.8-27B-UD-Q2_K_XL.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: 13,982 MiB measured, leaving 606 MiB of a 16 GB card for :: your desktop. --spec-type none buys back ~1.3 GiB if you :: would rather have headroom than the drafter. A >131k window :: cannot fit resident on this card --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: start at 4, sweep (section 06) --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: xhigh NOT OFFERED: 61.5-75.8k :: appetite against a 65k window - a WINDOW limit (section 09)
DGX Spark · GB10, 128 GB unified memory
The file: Q4_K_M — 16.5 GB, 15.4 GiB resident. With 128 GB of unified memory, capacity is not the constraint; bandwidth is, at 273 GB/s, so the smallest strong file wins. The window: the full -c 262144. The effort levels it offers: all three — but note the wall clock: the reference 100k-token run derives to about 2–2.8 hours here, so xhigh is an unattended-run setting. Expected speed: 10–15 t/s, derived from bandwidth. Room left for your desktop: not applicable — this machine has 128 GB of memory shared between the chip and the system.
:: llama-server (ARM64 Linux build — no .exe here), same stack as every other card. :: Bandwidth (273 GB/s) is the ceiling, so the smallest high-quality file wins: :: Q4_K_M (16.5 GB) over UD-Q4_K_XL (17.6 GB) is ~6% more t/s on a bandwidth-bound :: box — and the measured perplexity favours it too (section 08). :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M, 16.5 GB = 15.4 GiB) llama-server -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 262144 -ngl 99 :: 128 GB unified: full native context is trivial (~11 GiB of :: KV at the MEASURED slope with the drafter on, not the 8.5 :: the arithmetic suggests) --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: start at 4, sweep :: (section 06) — UNMEASURED on GB10; the 3090 gained +45% on :: real code from the same flags --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}" :: window holds it; wall clock :: is the cost — ~2-2.8 h for the 100k reference run (section 07) :: faster vendor alternative: Blackwell-native NVFP4 on vLLM (~1.5x BF16 speed, FP8 KV, :: 92-97% accuracy) — vllm serve unsloth/Qwen3.8-27B-NVFP4 --max-model-len 131072. :: llama.cpp cannot load that release (FP8 lm_head); switching stacks buys the FP4 path
Intel Core Ultra iGPU · Arc B390-class
The file: Q4_K_M — 16.5 GB, 15.4 GiB resident; every gigabyte re-read per token hurts at 153.6 GB/s. The window: -c 65536 as an interactive default. This is not a memory limit — an integrated graphics chip borrows system memory — but a practical one. The effort levels it offers: low and medium at this window; xhigh is not offered here for two independent reasons, and both need naming because they have different fixes: the 65,536-token window cannot hold a 61,500–75,800-token thought (raise -c to fix that), and at 5–7 t/s an xhigh answer takes three to four hours (only patience fixes that). Expected speed: 5–7 t/s, derived, and it scales with your memory bandwidth. Room left for your desktop: ample — an integrated graphics chip borrows system memory — but leave the processor and operating system some working memory.
:: llama-server.exe, VULKAN build (SYCL has known perf bugs on current Intel GPUs; :: IPEX-LLM is archived). RAM is the spec that matters: the B390 tier (Core Ultra X9 :: 388H) ships dual-channel LPDDR5X-9600 = 153.6 GB/s, up to 96 GB shared. :: Intel's formal Arc-branding floor for B390/B370 is 7,467 MT/s — slower RAM relabels :: the iGPU as generic "Intel Graphics" — and other Core Ultra configs may ship slower :: RAM or ONE module (single channel = half the bandwidth = half the t/s). The ~5-7 t/s :: scales with YOUR bandwidth via section 05's formula; the model fits easily. :: file: lmstudio-community/Qwen3.8-27B-GGUF (Q4_K_M, 16.5 GB = 15.4 GiB) llama-server.exe -m Qwen3.8-27B-Q4_K_M.gguf --alias qwen/qwen3.8-27b -c 65536 -ngl 99 :: interactive default, NOT a memory limit: an iGPU borrows :: system RAM, ~57% by default. To raise it: Intel Graphics :: Software >=25.26.1602.2 with driver >=32.0.101.6974 -> the :: GPU's General tab -> "Shared GPU Memory Override" slider :: (~87%; up to ~93% on B390/B370 with current Arc Pro :: drivers, RAM-dependent) -> reboot. Needs >=10 GB RAM, select :: Core Ultra systems. With 96 GB RAM even -c 262144 fits --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 :: verify KV-quant support in the Vulkan build --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: start at 4, sweep :: (section 06); UNMEASURED on iGPUs — drop the spec flags if :: a paired probe shows no gain on your machine --jinja --host 127.0.0.1 --port 1234 --chat-template-kwargs "{\"reasoning_effort\":\"medium\"}" :: xhigh NOT OFFERED, for TWO :: separate reasons with two separate fixes: 61.5-75.8k of appetite against this 65k :: WINDOW (fix: raise -c, you have the RAM), and 3-4 HOURS per answer at 5-7 t/s :: (fix: nothing but patience). Raise -c and effort together for overnight runs only. :: possibly-faster vendor alternative: Intel's OpenVINO Model Server on the iGPU with :: OpenVINO/Qwen3.8-27B-int4-ov (OVMS LLM quickstart; OpenAI-compatible /v3 endpoint) — :: Intel tunes that stack for its own silicon; benchmark both if the t/s matters
Standardized industry metrics
What each recipe costs in electricity, in the units the industry compares on. Read the two plain-language conclusions first: turning the drafter on is the single cheapest energy decision on this page — it cuts the energy per token by 2.52× — and a slower setting is a more expensive setting in exactly the same proportion, because while the model is generating, this card draws essentially the same power whatever it is doing, so the joules per token are just those watts divided by the tokens per second. The whole energy win is throughput. Nothing here saves power; it saves time at the same power, which is the same thing arriving by a less interesting route.
Instrumentation tier for every figure below: in-band GPU board power — read from the card's own sensor rather than from a meter at the wall (NVML, nvidia-smi --query-gpu=power.draw). Inside the number: the graphics chip, its memory, the board's voltage regulators and fans. Excluded and never measured on this machine: power-supply conversion loss, the processor, system memory, drives, chassis fans, the display, and any datacentre overhead. Do not call any of these figures system power or wall power, and do not divide an electricity bill by them.
The idle pair, dated. Settled, with little on screen, first 60 seconds of each idle window discarded (2026-08-23 power matrix): 29.9 W with no server running, 34.1 W with the model resident and answering nothing. A resident model costs +4.2 W — small, and in the physically sensible direction. Earlier editions of this page printed 33.2 W with no server against 30.7–31.1 W loaded, which is backwards; those readings were taken on a board that was still cooling from the previous job. The 34.1 W figure is what the idle-subtracted columns below remove.
Sustained board power is a range, not a constant. Across 2.5 hours of drafter-off benchmark work the board drifted 305.5 → 341.1 W (+11.7%) at constant throughput, constant 77–79 °C and constant memory clock, tracking its own SM clock from 1,453 to 1,606 MHz. Decode at one request at a time is limited by memory bandwidth, so the extra clock bought no tokens and cost about 6% more energy per token. Plan with the band — 306–341 W drafter-off, 338–345 W on the long drafter-on authoring runs, 302–339 W across the short power-matrix arms — or with your own run's mean. Earlier editions printed a flat "~344 W"; that is the top of a range.
| Recipe | Tier | Mean W whole window / during decode | J / decode token, gross | J / decode token, net of 34.1 W idle | J / prompt token (prefill) | tokens / kWh | EDP (J·s) | Wh per 700-token answer, gross | Wh, net |
|---|---|---|---|---|---|---|---|---|---|
| [1] UD-IQ4_XS + vision · n4/p0.75 energy arm B2 | in-band NVML | 308.4 / 341.7 meas | 3.743 meas | 3.369 der | 0.144 meas | 961,846 der | 20,091 meas | 0.786 meas | 0.699 der |
| [2] UD-IQ4_XS text-only · n10/p0.5 energy arm B3 | in-band NVML | 302.4 / 341.0 meas | 3.210 meas | 2.889 der | 0.146 meas | 1,121,380 der | 14,811 meas | 0.683 meas | 0.606 der |
| [3] UD-IQ4_XS text-only · no drafter energy arm B1 | in-band NVML | 325.2 / 344.6 meas | 8.104 meas | 7.302 der | 0.099 meas | 444,218 der | 93,380 meas | 1.616 meas | 1.447 der |
| [4] Q4_K_M + vision · n4/p0.75 energy arm C2 whole row measured drafter-OFF — not this recipe's flags | in-band NVML | 326.9 / 345.8 meas | 8.853 meas | 7.980 der | 0.102 meas | 406,644 der | 111,069 meas | 1.763 meas | 1.579 der |
| Every recipe on a card that is not this 3090 | Not measured. No power instrumentation exists on any other machine in this campaign, and board power does not scale with bandwidth the way decode speed does. Do not derive it | ||||||||
| E_comm (interconnect energy) | N/A — single GPU, no interconnect. Every recipe here serves from one card, so no energy is spent moving activations between accelerators. On a multi-GPU machine this row would have to be measured or split out | ||||||||
Conditions, and the ways each row differs from the recipe it prices. All four energy arms ran on 2026-08-23 at -c 32768 with a 1,458-token prompt and 700 generated tokens, thinking off, no projector, three requests per arm, 500 ms power sampling, 100% coverage. Row [1] and [2]: the recipes serve at 122,880 and 180,224 tokens of window with the projector loaded. Window allocation does not change joules per decode token; prompt depth does, and the depth series is its own set of arms with its own shallow reference point — 3.919 J/token at a 1.5k fill, 4.453 at 28k, 5.598 at 91k. Read that series against its own 3.919, not against this row's 3.743: the two shallow points sit 4.7% apart in energy and 6.4% apart in throughput, both outside the 2.9% floor, because they are different arms on different prompts. The projector is decode-free (0.04–0.09% at a 91k fill), so it changes nothing here. Rows [1] and [2] both decoded faster than the matched sweep did at the same flags — 91.3 against 83.5 for row [1] (9.3% apart) and 106.2 against 93.9 for row [2] (13% apart), unexplained and an open item in the campaign log. Plan speed from the sweep's 83.5 and 93.9; use these rows for energy. The energy ratios survive because every arm in the matrix shares whatever drift this is. Row [4]: this arm ran with the drafter off while the recipe ships n4/p0.75. On UD-IQ4_XS the same drafter cut joules per token by 2.17× (8.104 → 3.743), so recipe [4]'s own figure is expected near 4.1 and is unmeasured — the printed 8.853 is this file's no-drafter cost and is the right number for comparing files, not for pricing this recipe. "Wh per 700-token answer" means exactly that: every arm generated 700 tokens and stopped, so these are not complete-answer figures. Complete-answer energy, for one real authoring task, is in §09. Net columns subtract the dated 34.1 W loaded idle over the covered seconds and are derived; the campaign's own runner subtracted 31.0 W instead, a uniform shift of about 1% that changes no ranking.
Energy per setting: the J/token table
Every knob this page recommends on, priced in energy, so that a setting chosen for speed can be checked for electricity too. The pattern underneath all of it is simple and was verified with no residual left over: joules per decode token equals the watts drawn while decoding divided by the decode tokens per second. That qualifier matters — the mean-watt column in the table above spans each request's whole window, including a short low-power prefill, so it runs 6–11% below the decode figure and dividing it gives the wrong answer. Anything that raises throughput at constant power lowers energy per token in the same proportion. That is why the drafter is an energy feature and why depth is an energy problem.
| Axis | Arms compared (all 2026-08-23, in-band NVML) | J / decode token | What it says |
|---|---|---|---|
| Drafter | B1 --spec-type none · B2 n4/p0.75 · B3 n10/p0.5, same server, same prompt | 8.104 → 3.743 → 3.210 meas | 2.52× less energy per token at n10/p0.5, 2.50× the throughput, 6.3× better EDP. Power while decoding is unchanged — 344.6 → 341.7 → 341.0 W — so the energy saving is the throughput gain, exactly and with nothing left over. The whole-window means fall further (325.2 → 308.4 → 302.4 W), but that is the low-power prefill segment taking a larger share of a shorter run, not the card using less power |
| Quant | C1 UD-IQ4_XS · C2 Q4_K_M · C3 NVFP4-HIGH, drafter off | 8.198 · 8.853 · 9.293 meas | Q4_K_M costs +8.0% and NVFP4-HIGH +13.4% against UD-IQ4_XS — both outside the 2.9% noise floor, so both are real. Choosing a file is choosing an energy bill. (The campaign's summary line prints +7.8% and +13.1%. Those are not reproducible from either the gross or the idle-subtracted columns of the arm table, both of which give +8.0% and +13.4%; the ratios above are re-derived from the gross cells printed here) |
| KV cache precision | D1 -ctk/-ctv f16 · D2 q8_0 | 8.284 · 8.341 meas | 0.7% apart — a clean null, well inside the 2.9% floor. Half the KV memory for no measurable energy, which independently supports the q8_0 default that every recipe ships |
| Token regime | E1 thinking on · E2 thinking off, one server, one prompt | 6.066 · 3.744 meas | 1.62×, against the 1.69× throughput ratio the mean-draft-length measurement produced on a different instrument entirely (§06). Two independent measurements of the same mechanism |
| Depth | F1 1.5k · F2 28k · F3 91k of prompt fill, same 700-token answer | 3.919 · 4.453 · 5.598 meas | Decode energy rises 43% across the span — but the real story is prefill: at a 91k fill 90.7% of the arm's joules are prefill, and the same 700-token answer costs 0.83 Wh at 1.5k against 11.71 Wh at 91k, a factor of 14. Caveat: F3 decoded 10.4% slower than the cooled reference ladder at the same depth, so its level is suspect while its shape is not |
| Effort level | One authoring task, drafter on, temp 1.0, n=1 per level — low · medium · xhigh truncated · xhigh completed | 4.26 · 5.18 · 6.60 · 6.13 meas | Rising energy per token with effort — but this is a throughput effect, not an effort effect. On 25-question arms with the drafter off, where all three levels decode at the same speed, energy per token is flat to 0.3% (7.921 at xhigh against 7.947 at medium). Effort changes joules per answer, not joules per token (§09) |
--parallel 1 vs 2 | G1 · G2, drafter off — and the matched drafter-on pair that replaced them, 2026-08-25 | 8.59 → 5.19 superseded | Settled 2026-08-25. This page had refused to use either figure until a decisive test could separate them — it called that refusal a quarantine — and the test has now run: the drafter-off figure was wrong for the shipped recipe. The arm in this matrix measured +60.3% aggregate throughput, −39.6% energy per token and −62% EDP with the drafter off, contradicting an earlier ~+11% drafter-on measurement. The matched pair has now run on UD-IQ4_XS with the drafter on at n10/p0.5: 82.98 → 101.25 t/s aggregate (+22.0%) and 85.79 → 55.41 t/s per slot (−35.4%), acceptance flat at 0.618 → 0.620. So the truth is much nearer the older figure, and the reason is the one the quarantine named: the drafter is already saving most of the repeated reading of the model that batching would otherwise save, leaving batching far less to win. The energy column is not re-derived here — the matched pair measured throughput, not board power — so the 5.19 figure is retired rather than corrected. Every recipe still ships --parallel 1, now because one user gets their tokens about 55% faster that way (85.79 against 55.41 t/s per slot) (§06) |
Power cap (nvidia-smi -pl) | 350 · 300 · 250 W, one server load each | 4.033 · 3.771 · 3.479 meas | The one setting on this page that lowers the board’s draw at all. Measured 2026-08-25 once an elevated shell was available — this row read “unmeasured, requires an elevated shell” for the whole campaign SUPERSEDED. Capping to 300 W costs 5.0 % of throughput and saves 11.2 % of power; 250 W costs 12.8 % and saves 24.8 %. The throughput cost is about half the power saving at both levels, so energy per token improves 6.5 % and 13.7 %. Every other lever here saves energy only by finishing sooner at the same wattage; this one lowers the wattage. The mechanism is visible in the clock: throughput falls less than the SM clock does (−5.0 against −5.8 %, then −12.8 against −22.1) because decode is partly memory-bandwidth-bound and memory clock is untouched. And the stock arm never reaches its own cap — 305.4 W mean against a 350 W limit — which is why the first 50 W costs so little: it removes headroom the workload was not using. Decode only; prefill is compute-bound and would plausibly lose more (§11) |
04
Five model facts that decide every setting
Five things about how this model is built explain almost every recommendation in §03. It stores far less memory per token of context than its size suggests, so long contexts are cheap. Its context limit is 262,144 tokens, and what you will actually run out of is card memory, not model capability. It thinks before answering by default, and that thinking comes out of the same pool as your prompt. Vision is optional at start-up and costs memory rather than speed. And it wants a specific sampling setup that some clients quietly override, which causes the model to repeat itself. Everything below is those five facts with their numbers.
One more fact governs speed rather than settings, and it is worth having in plain words before the arithmetic: with speculation off, a bigger file is a slower file. To write each token, the card has to read the whole model out of its memory once. A file that is 2 GiB larger takes proportionally longer to read, every single token, forever.
The drafter breaks that rule, and it is why this page no longer simply recommends the smallest file that holds its quality. meas Turn speculation on and the ordering can invert: UD-IQ4_XS goes 42.34 → 86.91 t/s (2.05×) while the smaller UD-Q2_K_XL goes 45.66 → 77.01 (1.69×) — so the file that is faster drafter-off is slower drafter-on, by 12.9%. The draft head is quantized along with the model, so a more damaged file drafts worse and gets less back from speculation. The window moves the ordering too (§08). Since every recipe here ships with the drafter on, “smallest that holds its quality” is the right rule for memory and the wrong one for speed.
- Hybrid attention. Most of its 64 layers use Gated DeltaNet, a linear-attention design whose memory use stays constant no matter how long the context gets; only 16 of the 64 are standard attention layers (and those use grouped-query attention with 4 KV heads). The result: the KV cache costs just 64 KiB per token at fp16 — 4.0 GiB at 122k context with an 8-bit cache, 7.5 GiB unquantized — where a dense 27B storing KV in every layer would need exactly 4× as much. Those are the arithmetic, and arithmetic is a floor: what a real server allocates measured 15–29% steeper (§05). Huge contexts are cheap in memory here, and speed falls off gently as they fill rather than collapsing.
- Native 262,144 context. No rope or YaRN tricks are needed below that; the limit you will actually hit is VRAM. The model card lists extension to 1M tokens via YaRN, which is out of scope here: the KV cache alone would run about 34 GiB even at
q8_0, by the same arithmetic. - Thinking is always on, and defaults to
xhigh. Reasoning tokens share the context window with your prompt and your output, so a 100k-token thinking run needs a window bigger than 100k (§09). This is also why every speed number on this page states its token regime. - Vision is optional at serve time — and it costs context, not speed. The BF16
mmprojprojector occupies 1,138 MiB of VRAM once loaded — not the 0.867 GiB its file weighs, which is what earlier editions of this page quoted. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) measured the resident cost two independent ways and they agreed to the megabyte, and a follow-up pair at 91k of depth reproduced the same 1,138 twice more. At the measured window slope that is about 26,500 tokens of 8-bit window with the coding drafter loaded, 29,900 without it (§05). What it does not cost is speed: at 90,862 tokens of fill, decode with and without the projector differed by 0.04–0.09% (§06) — the only question the projector asks is about window. Serving text-only converts vision you are not using directly into context, but not all the way to the native maximum: with the drafter on, UD-IQ4_XS's measured text-only ceiling is 180,224, and 262,144 fits only with the drafter off, and then only just (§03 prints both variants). - Official thinking-mode sampling: temperature 1.0 · top-p 0.95 · top-k 20 · min-p 0. Clients that silently send temperature 0 — aider does — cause repetition loops. Override them; §14 shows how, per agent.
config.json; what a server really allocates is measured in §05 and runs 15–29% higher.The arithmetic that predicts decode speed
Generation speed comes down to one ratio: how fast your card can read its memory, divided by how much it must read per token. Every token requires reading the entire model — about 16.5 GB for a 4-bit K-quant, 14.25 GB for UD-IQ4_XS. So take your card's bandwidth in GB/s, divide by your file's GB, and multiply by an efficiency constant, and that is your speed ceiling in tokens per second. That constant is format-specific, which earlier editions of this page flattened into a single ~0.7: on CUDA it measures about 0.70 for K-quants and about 0.65 for IQ-quants. You can tell which family a file belongs to from its name: a name containing IQ — UD-IQ4_XS, UD-IQ3_XXS, UD-IQ2_S — is an IQ-quant, so use 0.65; a name containing K without IQ — Q4_K_M, UD-Q3_K_XL, UD-Q2_K_XL, Q6_K — is a K-quant, so use 0.70. The independent re-measurement measured 0.649 on UD-IQ4_XS — 936 GB/s ÷ 14.25 GB = 65.7 t/s theoretical against 42.6 measured — and Q4_K_M lands at 0.700: 936 ÷ 16.5 = 56.7 against 39.7 measured. This one calculation predicts every row of the matrix in §07, and it doubles as a fault detector: land far below your format's constant at short context and something is wrong (§11).
05
What fits in 24 GB, and where speed collapses
A 24 GB card has two ceilings, not one, and the difference between them is the most expensive misunderstanding on this page. The first ceiling is how large a window you can set and still have every byte of it live on the card — fully resident, in the words of §02: about 131,000 tokens for the reference configuration. The second is how large a window you can set before the server refuses to start at all: about 213,000. Between those two numbers the server starts happily, answers a short prompt at full speed, and then collapses to a fifth of that speed the first time somebody actually fills the window. Set your window under the first ceiling, not the second, and never trust a ceiling that was checked with a short prompt. The rest of this section is the memory arithmetic that tells you where your own first ceiling is.
The second law of this section: speed against context is a cliff, not a curve. It is flat while the weights and the cache fit in VRAM, and it collapses once anything spills across the PCIe bus — with the magnitude set by how much spilled: about 10× when the weights file itself lands in system memory; 1.1–2× from a 1.9 GB spill of weights answering a short prompt (§11); and 2.5–3.8× once the spilled pages are cache that a deep prompt actually reads (the table below). What decides the size is not how much spilled but how often the spilled pages are touched. Measured on the reference 3090:
q8_0 KV, drafter on, projector loaded); the dashed tail to 256k is an older run with no drafter where the 4-bit weights file itself spilled, so it is a different configuration and is drawn dashed for that reason. The 256k point is not a verdict on 256k — a smaller resident file gets closer to full native context on the same card. The dashed vertical at 131k marks where dedicated VRAM fills: everything to the right of it is fast only while the deep context stays untouched, which is exactly what a short probe never tests.Allocated is not the same as used: a big context window with a small prompt in it costs VRAM, not speed. And as long as you stay below the cliff, shrinking the window gains you nothing — there is no hidden optimal point to search for.
That corollary now has a measured face. Probed 2026-08-22 with Q4_K_M (q8_0 KV, drafter on, projector loaded — the 1 GiB-larger UD-Q4_K_XL shifts every ceiling about 30k tokens lower, while dropping --mmproj frees 1,138 MiB and raises them about 26,500) by stepping -c upward with short temperature-0 decode probes: speed holds at 50–55 t/s all the way to -c 212992, turns borderline at 217,088 (41.2 t/s), and collapses at 221,184 (19.5 t/s). But nvidia-smi tells the other half of the story: dedicated VRAM hits its ~24 GB ceiling already at -c 131072 (measured 24,052 MiB). Past that point the KV cache is overcommitted, and the run stays fast only while a short prompt leaves the deep pages untouched. So: about 131k is the practical fully-resident ceiling — fast even when a long run really fills it, though it sits right at the limit, which is why the recipe ships 122,880 — while up to about 213k will load and answer a short prompt. What earlier editions of this page got wrong is what happens next.
Earlier editions said the overcommitted zone "degrades progressively as real tokens land in the tail". It does not degrade — it collapses, and only once real tokens land there, which is exactly what a short probe never does. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS, projector on, drafter at n-max 4 / p-min 0.75, identical prompts, only -c differing) filled both configurations instead of probing them shallowly:
| Prompt actually filled | -c 131072 · fully resident | -c 262144 · projector loaded, ~3.5 GiB in system RAMa different configuration from the text-only 262,144 row in the budget table, which spills 2,364 MiB | Penalty |
|---|---|---|---|
| 1,531 tok | 53.5 t/s | 42.9 t/s | 1.25× |
| 29,837 tok | 47.3 t/s | 18.6 t/s | 2.54× |
| 90,886 tok | 30.6 t/s | 8.0 t/s | 3.82× |
Prefill collapses with it — 819 → 183 t/s at 90.9k, 1,085 → 408 at 29.8k — while draft acceptance is unchanged (0.907 against 0.849 at 90.9k), which rules out any drafting explanation: this is purely the KV cache being read across the PCIe bus. The practical reading: a reader who trusted the shallow ceiling gets 26% of the promised speed the first time they open a 90k-token document — 8 t/s on a card that does 50 — and nothing in the server log says why. Never label a window resident or safe on the strength of a short probe; fill it once, deeply, before you believe it.
One caveat on the table's own numbers, added 2026-08-23 by the follow-up round: every cell in it is a single decode probe fired immediately after its prefill, and that is the one moment this card is not at its boost clock — the ±25% band declared in §01 and measured in §11. Re-measured under the cooled protocol, the resident 90.9k row reads 36.6 t/s rather than 30.6, about 20% higher. The spilled column was taken the same way and carries the same uncertainty, so read the penalty ratios as approximate. Nothing about the verdict moves: a 2.5–3.8× collapse is an order of magnitude outside a ±25% band.
-cThe same re-measurement round measured the same -c 163840 configuration at 22,418 MiB dedicated with the n-max 4 drafter and 23,189 MiB with n-max 10 — a 771 MiB swing from a flag, on top of a 474 MiB swing in board VRAM from nothing but what was on screen. (That 771 and the 898 MiB in the budget table are the same quantity read two ways: 898 is the constant fitted across all 17 resident loads, 771 is what this one pair of loads happened to show. They differ by 127 MiB, which is exactly the budget model's worst residual — so the disagreement is the model's stated error, not a second finding. Budget with 898.) A window is resident or not as a property of the whole configuration: file + drafter flags + projector + desktop. Quote all four whenever you quote a ceiling — reports of "same command, different ceiling" are almost always one of those four moving.
The VRAM budget table
The cliff chart says where speed dies; this table says why — and the independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) replaced most of its estimates with measurements. Earlier editions budgeted by arithmetic: 64 KiB per token at fp16, 34,816 bytes per token at q8_0 with block scales, plus "1–1.5 GiB of compute buffers" — the scratch memory a server needs while it is working — and an assumption of "22–23 GiB usable". Every one of those is a floor, not a figure. The campaign instead read llama-server's own dedicated-VRAM counter across 17 fully-resident server loads and fitted two constants per configuration — a fixed base plus a per-window-token slope. The model reproduces all 17 loads to within 127 MiB, so it is what this page now budgets with. (Those 17 are the fully-resident subset of 26 loads in total; the other nine were overcommitted configurations, which the model does not describe. Later measurement rounds loaded further configurations that have never been checked against it.)
| Line item | Measured | What it means |
|---|---|---|
| Per window token, drafter off | 39,936 B meas | +15% over the 34,816 arithmetic — alignment and padding that the config.json maths cannot see |
| Per window token, drafter on | 45,056 B meas | +29% over arithmetic. The 5,120 B/token difference is the draft path: speculation costs 11.4% of your window |
| Weights + buffers (UD-IQ4_XS, text-only, no drafter) | 13,232 MiB meas | the fixed base. Q4_K_M runs about 2.1 GiB above it, UD-Q4_K_XL about 3.1 |
Vision projector (BF16 mmproj) | 1,138 MiB meas | ≈26,500 tokens of window with the drafter on, 29,900 without. The file is 0.867 GiB — quoting the file size, as earlier editions did, under-books it by about 0.25 GiB |
--image-max-tokens 1024 — shrinking a screenshot | saves ~2,600 tokens of window meas | Measured 2026-08-25, and it is not the free win it looks like. On a 1440p dashboard it cuts the image from 3,635 tokens to 1,043 — 3.49×. The model still reads the layout, the headline and both 4-digit values. It loses the smallest type and does not know it: a latency of 207 came back as 287, and a serial QUB-85731-3D3B came back as QLN-15731-3053. It does not refuse and it does not hedge — a confident misread, which is worse than a blank, because nothing downstream can see it. A control run with the image withheld scored 0/7, so the readings above are perception and not guesswork. Use it for layout questions, never for reading numbers off a screenshot. n=1 per question, 7 questions |
--spec-type draft-mtp — turning the drafter on | 1,008 MiB fixed meas | earlier editions said the draft head costs no VRAM. It costs this, plus the per-token line above. The head's own weights are 334.7 MiB; the rest is draft-path buffers. --spec-type none loads none of it — the server logs unused tensor blk.64.* |
--spec-draft-n-max 10 instead of 4 | +898 MiB meas | a flag moves your ceiling by about 0.9 GiB — roughly 20,900 tokens of window — without -c changing at all |
| The desktop's own share of board VRAM | 1,179–1,669 MiB meas | measured in direct no-server readings on one machine, and it depends on what is on screen. Not yours to spend (§02, §11) |
Conditions: RTX 3090, 24,576 MiB, UD-IQ4_XS, q8_0 KV, -ngl 99, --parallel 1, llama.cpp build 10502. Worked example derived — §03's the default recipe with vision at -c 122880, if you put the n-max 10 coding drafter on it: 13,232 + 1,138 + 1,008 + 898 + (122,880 × 45,056 ÷ 1,048,576 = 5,280) = 21,556 MiB of server-process VRAM, leaving 3,020 MiB (2.95 GiB) of board VRAM for the desktop and slack. Drop the 898 and the same row keeps 3,918 MiB (3.83 GiB) — which is why §03 ships the vision recipe at n-max 4: the wider drafter setting fits here, it just spends the slack that absorbs a desktop measured anywhere from 133 to 1,181 MiB.
Two rules follow, and both change how you read every ceiling in this guide. First, budget on the server's own report, not on the board VRAM total — the three counters and what each can see are in §02. The nvidia-smi board VRAM figure includes your window manager and your browsers; the campaign watched one identical configuration read 474 MiB apart on the board VRAM — about 11,000 tokens of window — with nothing changing but what was on screen, while llama-server's own number moved by at most 127 MiB. That spread is the whole explanation for "same command, different ceiling". Second, the drafter is a line item, not a freebie. At -c 163840 with vision, its fixed and per-token costs together book roughly 1.8 GiB — more than the vision projector. On a pure prose workload, where speculation was measured worth only about 1.16× (§06), that is the easiest 1.8 GiB on the card to reclaim: --spec-type none.
With those constants in hand, what fits on a 24 GiB card by file and window — the KV column stated as the q8_0 arithmetic floor, so add 15% (drafter off) or 29% (drafter on) for what a server really allocates:
| Context | KV type | KV size (arithmetic floor) | Largest unsloth file that fits |
|---|---|---|---|
| 262,144 | fp16 | 16.0 GiB der | nothing — even the smallest file (UD-IQ1_S, 5.8 GiB) misses once buffers are counted. That file is now measured, and it is not a fallback: at 1.83 bits per weight it is +35% perplexity against this page's default and it is the worst of four rungs that fail a functional detector (§02) — 2, 3, 5 and 28 empty replies out of 75 at 2.481, 2.153, 1.994 and 1.835 bits per weight (§08) — and the only one that also stops terminating, at a median 932 tokens against 424. Earlier editions called it “the one file that fails” SUPERSEDED: that was written from the pre-ladder screen, before the empty-answer audit found three more. Its code does not run either, and neither does that of the two rungs above it (§08) |
| 262,144 | q8_0 | 8.5 GiB der | UD-IQ4_XS (13.3 GiB) — text-only, no projector, and only with --spec-type none: at the measured drafter-off slope the window costs 9,984 MiB (9.75 GiB), not 8.5, for 23,216 MiB total — leaving 1,360 MiB of board VRAM, which is 179 MiB once a desktop takes its measured worst case of 1,181 MiB. Marginal, and only with no graphical session using the card (§02). With the coding drafter on it does not fit at all: 25,504 MiB (24.9 GiB) required against 24,576 available, and 2,364 MiB measured living in system RAM. Earlier editions called this row "verified by shallow probes" — see the probe-artefact correction above. Applying the same measured slope, UD-Q3_K_XL (12.2 GiB) lands near 22.2 GiB — about 1 GiB more comfortable, though unmeasured — and UD-IQ3_XXS (10.15 GiB) fits with real room to spare |
| 262,144 | q4_0 | 4.5 GiB der | UD-Q4_K_M would fit — and the q4_0 cache now has a measured price rather than a warning: +0.693% perplexity against fp16 (§08). Small, but nothing here argues for spending it on a 24 GB card: q8_0 already reaches every ceiling this page publishes. On a 16 GB card it is the one place the trade earns its keep |
| 131,072 | q8_0 | 4.25 GiB der | Q4_K_M (15.4 GiB) — 19.7 GiB by this table's own two columns, and that is the floor: apply the measured constants and the same file at -c 122880 with the projector and the n-max 4 drafter books about 22.3 GiB of server VRAM, which is why §03's the Q4_K_M recipe measures ~23.7 GiB of board VRAM with a desktop up. This is the reference recipe's row; budget it from the constants, not from this column. The alternative UD-Q4_K_XL (17.6 GB = 16.39 GiB) also fits when VRAM is spare |
| 65,536 | q8_0 | 2.125 GiB der | UD-Q5_K_S (~17.4 GiB) — the quality play when you do not need long thinking runs |
Two footnotes the table earns. First, the q4_0 KV flag, which earlier editions ruled out on mechanism alone and which has now been measured (2026-08-23, same corpus and flags as the q8_0 run): a q4_0 K/V cache costs +0.693% perplexity against fp16 — 6.6413 against 6.5956 — where q8_0 costs +0.309%. The degradation is super-linear in bits: the second halving of the cache costs slightly more than the first (+0.383% on top of q8_0), which is the shape the mechanism predicts and the reason the cut stops paying. That mechanism is unchanged and still unmeasured directly: the problem is not that quantization error accumulates — each cache entry is quantized once — it is that at long context every new token attends over hundreds of thousands of perturbed entries, and retrieval gets unreliable at exactly the 200k-plus lengths a q4_0 cache would enable. What the number changes is the tone. 0.69% is small — smaller than the perplexity gap between the two 4-bit files this page recommends (0.93%) — so "never q4_0" was too strong. q8_0 remains what every recipe ships because it already buys the context ceilings here; q4_0 is a window-extension trade to take knowingly (§08 carries the row and its caveat). Second, prefill: these budgets make a 200k-plus prompt fit, and earlier editions estimated that filling one costs "several minutes at 100k, potentially tens of minutes at 200k+". The independent re-measurement measured it instead, across a 60× span of prompt lengths on a fully-resident -c 131072 server:
| Prompt filled | Prefill | Wall to first token | Prefill energy at this depth |
|---|---|---|---|
| 1,533 tok | 1,082 t/s meas | 1.4 s | 0.164 J/prompt-token meas |
| 10,658 tok | 1,198 t/s meas | 8.9 s | — |
| 28,423 tok | 1,097 t/s meas | 25.9 s | 0.306 J/prompt-token meas |
| 56,840 tok | 951 t/s meas | 59.8 s | — |
| 92,679 tok | 816 t/s meas | 1.9 min | 0.421 J/prompt-token meas |
Conditions: UD-IQ4_XS, -c 131072 fully resident, q8_0 KV, drafter at n-max 4 / p-min 0.75, no projector, prompts prefixed with a unique identifier so the prefix cache (§02) cannot be reused (2026-08-23). The energy column comes from a separate 2026-08-23 arm at 1.5k / 28k / 91k of fill, not from these five prompts, and its depths are close but not identical. Use about 1,000 t/s as the planning figure: prefill decays only 25% over a 60× longer prompt, because the quadratic term lives on just 16 of the 64 layers (§04) — the hybrid dilutes it. A genuinely full 100k prompt is about 2 minutes, not "several"; extrapolating the same slope, 200k lands nearer 5 minutes than tens of them. Overcommit the window and this number collapses with everything else: 183 t/s at 91k in the correction above.
With a short prompt and a long answer, reading the prompt costs almost nothing: measured at 0.12–0.18 J per prompt token against 4.3–6.6 J per generated token, which is 0.06–0.38% of an answer's energy. That flips completely at depth. At a 91,000-token fill, 90.7% of the whole request's joules are prefill, and the same 700-token answer costs 0.83 Wh at 1.5k of depth against 11.71 Wh at 91k — fourteen times more for identical output. The energy per prompt token itself climbs 0.164 → 0.306 → 0.421 J across those three depths, which is attention's quadratic term showing up as electricity. Practical consequence for anyone running agents: reusing a cached prefix — starting a follow-up request with exactly the same opening text, so the server does not have to read it again (§02) — is not a speed optimization, it is an energy one, and it is the largest single saving available at depth. (All figures 2026-08-23, in-band board power; the 91k arm decoded 10.4% slower than the cooled reference ladder at the same depth, so treat its level as soft and its shape as solid.)
06
Speed: which flags to set, and what the drafter buys
Six flags belong on every card, and each one is here because a measurement put it here. Put every layer on the graphics card with -ngl 99: writing -ngl 64 looks complete and is not, and it costs about 30% of your speed with no warning anywhere. Serve one request at a time with --parallel 1. Store the memory of the conversation in 8-bit with -ctk q8_0 -ctv q8_0: it halves that memory and costs 0.3% of perplexity. Turn the built-in drafter on with --spec-type draft-mtp: it roughly doubles your speed, changes nothing about what the model writes, and cuts the electricity per token by two and a half times. Choose --load-mode none if system memory is tight and --load-mode mmap if you restart the server often. And whatever number you end up quoting, say which kind of token you counted, because on this model that alone is worth a factor of 1.7.
- All layers on the GPU, or accept the cliff —
-ngl 99. llama.cpp counts the output layer as one more layer than the model has transformer layers, so on this 64-layer model-ngl 64reads as complete and leaves that output layer — a 5120 × ~151k-vocabulary matrix multiplication — on the processor, running for every generated token. Measured cost: 25.7 against 39.7 t/s on Q4_K_M, and independently 29.84 against 42.31 on UD-IQ4_XS — −29.5% — with no memory signature at all (§11). Partial offload is set by system memory bandwidth, not by the graphics card, so a 5080 and a 4060 Ti are predicted to offload at nearly the same 8–12 t/s der — that band is §04's formula run at system-memory bandwidth, not a measurement: neither of those cards has ever been in this machine, and the only offload figure measured here is the 25.7 t/s above, on this 3090. And this is a dense model — every part of it runs on every token — so the trick of parking a model's rarely-used parts on the processor has nothing to target here. - One slot —
--parallel 1, and this is now settled rather than quarantined. Some interfaces default to 2, which halves the window each request gets and doubles the KV cache. Two earlier measurements of running two requests at once disagreed badly — one found +11% aggregate throughput, another found +60.3% — and the second of those ran with the drafter off, so this page refused to use either figure until a decisive test could settle them — it called that refusal a quarantine — and named the matched pair that would run it. That pair ran on 2026-08-25. On UD-IQ4_XS with the drafter on at n-max 10 / p-min 0.5 — the faster shipped setting — three reps each, first post-prefill probe discarded:--parallel 1gave 82.98 t/s aggregate and 85.79 per slot;--parallel 2gave 101.25 t/s aggregate and 55.41 per slot — +22.0% aggregate, −35.4% per slot. The +60.3% figure was an artefact of measuring with the drafter off: once the drafter is already saving most of the repeated reading of the model, batching has far less left to save, which was the stated hypothesis for the quarantine and is now measured rather than assumed. Draft acceptance is unchanged across the pair (0.618 → 0.620), so the second slot does not disturb the drafter at all; the whole cost is that each user waits about 55% longer for their own tokens — 12.6 s against 8.2 s for a 700-token answer. Every recipe here still ships--parallel 1, and now for a reason a single user can check: serving one person, one slot is the faster setting (§03). - 8-bit KV cache —
-ctk q8_0 -ctv q8_0, verified twice on two files. Measured wikitext perplexity with the 8-bit cache: 6.5498 against 6.5348 at fp16 on Q4_K_M — +0.23%, inside the error bars — and the re-measurement replicated the effect independently on UD-IQ4_XS at 6.6160 against 6.5956, +0.31%, likewise about one standard error. Half the KV memory for statistically nothing, on both files. It also costs nothing measurable in energy: 8.341 against 8.284 J per decode token, 0.7% apart and well inside the 2.9% floor. The one price that is not nothing is speed: fp16 KV measured 43.19 againstq8_0's 42.31 t/s at-c 32768, so 8-bit costs about 2% of decode — when VRAM is genuinely spare, fp16 is marginally the faster cache. The step below is measured too:q4_0K/V at 6.6413, +0.693% over fp16 (§08). - The drafter on —
--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75. Lossless by construction: every drafted token is verified by the full model before it counts, so quality is never on the table. Measured here it doubles decode on code and cuts energy per token by 2.52× at its best settings. It is not free in memory — 1,008 MiB fixed plus 11.4% of your window (§05) — and on prose it is worth only about 1.16×, which is the one workload where--spec-type noneand the memory back is the better answer. The rest of this section is how far it goes and why. - Load mode —
nonefor tight RAM,mmapfor restart-heavy work. Without--load-mode none(which replaces the deprecated--no-mmap) the whole GGUF stays cached in system memory even after the weights are in VRAM — about 15 GB for the 4-bit file. The cost is a slower model load: measured 10–30 s to first health check across dozens of restarts in the campaign proper, and it is storage-dependent. The independent re-measurement ran about 30 restarts on--load-mode mmapinstead and measured warm reloads at 6–10 s with decode unaffected (repeat probes agreed to 0.7%) — because at-ngl 99every weight lands in VRAM either way, so load mode cannot touch decode speed. So: if you are sweeping flags or restarting dozens of times a day,mmapis the better choice. Every recipe in §03 nonetheless shipsnone, and that is a deliberate default rather than a contradiction:noneis the safe assumption about a machine this page cannot see, because the failure it prevents (system memory exhausted by a cached model file) is worse than the one it causes (a load 20 seconds slower). Change it if you know your machine has the memory to spare. The one thing that campaign did not re-verify is the 15 GB saving itself; it remains untested rather than recommended. - Say which tokens you counted. This model thinks by default, so a probe labelled "code generation" can spend every timed token reasoning about code and return an empty answer field. On code, answer tokens run about 1.7× faster than reasoning tokens on the same server with the same flags. Every speed row below declares its regime.
The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) saved the generated text from its speed probes and found that several returned an empty content field: 700 tokens of thinking, no deliverable. The probe labels had been describing the task, not the token stream. That does not invalidate any decode measurement — the tokens are real and the server timed them correctly — but it changes what they mean, and this page's earlier tables inherited the same ambiguity. Every row below now declares its regime: reasoning tokens (thinking on, the model's default) or answer tokens (enable_thinking:false, what a reader keeps). They are not interchangeable, and on code the answer tokens are the faster ones.
Flash Attention: you are already using it, and the recipes depend on it
None of the recipes on this page passes -fa, and none of them needs to. This build documents -fa, --flash-attn [on|off|auto] with a default of auto — so a recipe that omits the flag is not running with Flash Attention off, it is letting the server decide. On this machine it decides on meas, and that is not an inference: with -fa unset and -ctk q8_0 -ctv q8_0 the server loads fine, while the same configuration with -fa off refuses to start. Acceptance also reads identically across auto and on — 0.611 shallow and 0.571 deep, to three decimals — and deep decode differs by 0.1 %.
What Flash Attention actually is. It is a different way of computing attention, not a quality setting. Ordinary attention builds the whole attention matrix in memory and then reads it back, so its memory traffic grows with the square of how much context it is attending over. Flash Attention fuses those steps so the large intermediate is never written out at all. The arithmetic is the same; the memory traffic is far smaller, and the saving grows with depth.
Every recipe here ships -ctk q8_0 -ctv q8_0, which halves the cache and costs 0.23 % of perplexity. That setting does not work without Flash Attention. Asked for both, the server exits during model load:
llama_init_from_model: quantized V cache requires flash_attn to be enabled
srv load_model: failed to create_context with model ...
srv llama_server: exiting due to model loading error
This is the good kind of failure — loud, immediate, and it names its own cause. You cannot silently end up with a degraded cache. But it does mean the recipes have a dependency they never state: the KV setting they ship only works because Flash Attention is available and auto turns it on.
Where that dependency could actually fire. Not on this 3090. The risk is a backend where auto resolves differently — and this page carries recipes for Intel Arc Pro B50 and B70 under Vulkan that also ship q8_0 KV, measured on none of them. If Flash Attention is unavailable there, those recipes will stop at the line above rather than run slowly. If that happens to you, drop to -ctk f16 -ctv f16 and re-check the window against the budget table, because an f16 cache is roughly twice the size per token.
-ctk q8_0 does not quantise the drafter’s cache
Chasing the flag above turned up something no recipe on this page states. -ctk and -ctv apply to the target model only. The draft model has its own pair, -ctkd and -ctvd, and they default to f16. No recipe here passes them — so on every pick that runs a drafter, the draft cache is unquantised even though the main one is not.
This is visible in the server’s own log meas. With -fa unset, the two contexts resolve Flash Attention by different routes: the draft context reports flash_attn = auto and then runs a probe that turns it on, while the target context reports enabling flash_attn since it is required for quantized V cache — forced, because its cache is quantised. Both end enabled, which is why auto and on measure the same; but they get there differently, and only the target context is forced.
Whether -ctkd q8_0 -ctvd q8_0 is worth setting is UNMEASURED here and is not a recommendation. The arithmetic says it would roughly halve the drafter’s per-token window cost, which the budget table books at 5,120 B per token — on the order of 300 MiB at a 131k window. Whether it costs acceptance, and therefore speed, is exactly the kind of thing this page does not guess at. It is written down here so the next person measuring has somewhere to start.
So should you set it? No — leave it alone. Adding -fa on buys nothing over the default on this card (0.1 % at depth, inside noise), and pinning a flag you have not measured on your own backend removes the one mechanism that would have chosen correctly for you. The reason to know about it is diagnostic, not performance: if a recipe here refuses to load and mentions flash_attn, that is this dependency, and the fix is the KV width, not the flag.
One measurement from this run that is not trustworthy, stated so nobody uses it: the prefill figures. Probes reused a cached prefix, so the server reported a full prompt_n against a near-zero prompt_ms and the resulting “prefill t/s” is meaningless. Only the decode, acceptance, VRAM and load-success results from this arm are reported anywhere on this page.
Every speed this page measured, in one table
Read the band, not a number. The floor is set by how fast the card reads its own memory, and everything above that floor is the drafter, so the right row for you is the one whose content and token regime match your work.
| Workload | Regime | Acceptance | Decode | What it means |
|---|---|---|---|---|
| Ceiling — copying supplied text verbatim (IQ4_XS · n-max 10 / p-min 0.5 · thinking off) | answer | 0.99 | 148.7 t/s | the real ceiling. Earlier editions printed 119.8 at 0.93 for "the same" task with thinking left on — which timed the model reasoning about copying, not copying |
| Novel JavaScript class (IQ4_XS · n-max 10 / p-min 0.5) | answer | 0.67 | 95.4 t/s | answer tokens on code beat reasoning tokens on the identical task (88.8) — written code is more predictable than free-form reasoning about it |
| Novel Python module + tests (same flags) | answer | 0.59 | 79.1 t/s | against 63.1 for the reasoning stream on the same task |
| Short maths, greedy (GSM8K — §08's scored table's build and workload, not this table's) | mixed | 0.90 | 63–70 t/s | short, predictable reasoning |
| Realistic code, temperature 0 (Q4_K_M · n-max 4 / p-min 0.75) | reasoning | 0.81 | 57.9 t/s | the tuned-flags benchmark number |
| Production coding, temperature 1.0 | mixed | — | 55–61 t/s | what a normal session feels like |
Long xhigh run, sustained across a 61–76k-token thought | reasoning | 0.84–0.88 | 48.8–55.8 t/s | KV deepens, thinking drafts worse |
| English prose explainer (IQ4_XS · thinking off) — the drafter's worst content | answer | 0.44 at n10/p0.5 0.80 at n4/p0.75 | 43.8 t/s at n10/p0.5 48.4 t/s at n4/p0.75 | speculation is nearly worthless here. The best configuration on this content is n4/p0.75 at 48.35 t/s, which is 1.16× a 41.55 floor; the wide n10/p0.5 tree manages only 43.80, or 1.05×. The drafter's ~1.8 GiB (§05) is hard to justify on a pure writing workload; --spec-type none buys that window back. Correction: earlier editions, and the campaign log's own summary line, attached the 1.16× to the 43.80 row. 43.80 ÷ 41.55 = 1.05; the 1.16× belongs to 48.35 |
| Floor — speculation off | either | — | 39.7–43.0 t/s | pure bandwidth (Q4_K_M 39.7–40.0; UD-IQ4_XS 41.46–42.97 across five contents and both token regimes — 41.46 copying, 42.17 novel JavaScript, 41.84 Python, 41.55 prose, 42.97 the rate-limiter prompt, plus 42.47–42.81 on the reasoning stream). Content moves it by ≤3.6%: without a drafter, what you are writing barely matters at all |
All rows measured 2026-08-21 to 2026-08-23 on the reference RTX 3090 at short context, from the server's own timings. The floor row is also the diagnostic threshold, and it is published as a band rather than a point for that reason: it was established across five contents and both token regimes and holds inside 3.6%.
How speed falls as the prompt grows
Speed drops as the conversation gets longer, gently and predictably, and the drop is about 25–30% between an empty window and a 91,000-token one. It is not a cliff — that only happens when the window overflows the card (§05). What matters far more than the depth is which kind of token you are counting: at the same 91k depth, the answers a reader keeps arrive at 64.8 t/s while the thinking that precedes them runs at 36.6. Plan an agent session on the first number and a long xhigh run on the second.
All three series below are UD-IQ4_XS at -c 131072, fully resident, with prompts prefixed by a unique identifier so nothing is served from the prefix cache. The first column is the one to plan against: it was measured 2026-08-23 under the cooled protocol (§02, documented in §11) — the probe that comes back with the prefill discarded, three probes per depth off the cached prefix — while the other two columns are single probes taken right after their prefill and therefore carry the ±25% band from §01.
| Prompt depth | Answer tokens · cooled ladder no projector · n-max 4 / p-min 0.75 · thinking off | Answer tokens · coding flags + vision vision on · n-max 10 / p-min 0.5 · single probes, ±25% | Reasoning tokens no projector · n-max 4 / p-min 0.75 · thinking on | Naive tokens ÷ wall-clock the WRONG number, shown so you recognise it |
|---|---|---|---|---|
| ~1.5k | 86.3 t/s · acc 0.89 | 94.5 t/s · acc 0.60 | 51.2 t/s · acc 0.80 | — |
| ~10.7k | — | — | 50.3 t/s · acc 0.89 | — |
| ~28–30k | 80.2 t/s · acc 0.93 | 81.0 t/s · acc 0.61 | 47.1 t/s · acc 0.92 | 9.2 t/s |
| ~57k | — | — | 41.1 t/s · acc 0.91 | — |
| ~91k | 64.8 t/s · acc 0.92 | 70.1 t/s · acc 0.65 | 35.8 t/s · acc 0.86 (36.6 cooled) | 2.4 t/s |
The last column is not a measurement of the model: it is the same reasoning-series runs divided the wrong way, tokens ÷ total wall-clock, which buries a 26-second or 105-second prefill in the denominator. It is printed here so that a reader who computes 9.2 t/s on a card that decodes at 47.1 recognises their own arithmetic instead of blaming their hardware. Quote the server's own timings, and quote prefill separately (§05). Why the acceptance columns differ so much: the cooled ladder re-used the campaign's red-black-tree prompt, a textbook algorithm the draft head predicts extremely well, while the middle column generated novel code. Acceptance ranks the content, not the protocol — and, as the next part measures, it does not rank speed at all.
Three things to take from it. First, the old "34–44 t/s at agent depth" band understated a real coding session by about 2× — it was a reasoning-token number, and the tokens an agent actually receives run 65–70 t/s at 91k of depth. Second, "acceptance never moves" was wrong in the safe direction: acceptance rises with depth, 0.80 → 0.92 in the reasoning series and 0.60 → 0.65 in the answer series — two independent confirmations. Decode still falls 25–30% across the same span (86.3 → 64.8 on the cooled ladder), which settles the mechanism: the cost is per-token KV reads growing with depth, not the drafter failing. Depth and acceptance are independent axes. Third, the decline is gentle and it is not a cliff — as long as the window is genuinely resident. Overcommit it and the same depths give 18.6 and 8.0 t/s instead (§05).
Earlier editions carried this as a loose end: one cross-check had put the projector's runtime cost at 91k near 15% (30.6 t/s with the projector loaded against 35.8 without), which would have been a real reason to serve text-only. It was a measurement artefact, and the paired probe now exists. Byte-identical prompts filling 90,862 tokens, loads ordered A-B-B-A, cooled probes, the only difference being --mmproj: with the drafter off, with-projector 26.965 t/s against no-projector 26.975 (n=10 each) — a 0.04% difference against a 0.4% within-arm spread. Repeated with the drafter on: 62.71 against 62.655 t/s (n=8 each), 0.09% apart, with draft acceptance matching probe for probe. The projector costs exactly what §05 books for it and nothing else: 1,138 MiB of VRAM, 0% of decode — 19,376 − 18,238 in one pair and 21,034 − 19,896 in the other, the same figure to the megabyte a third and fourth time.
The mechanism under the two regimes: mean draft length
Line up the cooled ladder against the reasoning column — same file, same flags, different protocols — and the reasoning side sits a near-constant 1.7× below at every depth: 86.3/51.2 = 1.69, 80.2/47.1 = 1.70, 64.8/35.8 = 1.81 — so call it 1.69–1.81×. A ratio that stable across a 60× span of depth is not clock noise, so it was isolated properly (2026-08-23): one server, one 90,894-token prompt, the same n-max 4 / p-min 0.75 flags, cooled probes, the only change being the chat template:
| Token stream at 91k | Decode | Draft acceptance | Mean draft length | Energy per decode token |
|---|---|---|---|---|
reasoning (enable_thinking on — the default) | 36.62 t/s | 0.8951 | 2.99 | 6.066 J (separate arm) |
answer (enable_thinking:false) | 62.02 t/s | 0.9073 | 4.31 | 3.744 J (separate arm) |
either, --spec-type none | 27.0 t/s | — | — | 8.104 J (separate arm) |
A 1.69× speed difference produced by nothing but the token stream — and acceptance says nothing about it (0.895 against 0.907, essentially tied). The mean-draft-length column is where the difference actually lives. On uncertain reasoning tokens the p-min 0.75 confidence gate keeps cutting the draft tree short — a mean of 2.99 tokens per pass against 4.31 on the answer stream — so acceptance stays high precisely because the gate is discarding the part that would have missed. That makes acceptance blind to the loss. The refinement this page now carries everywhere: acceptance tells you whether a draft was right; mean draft length tells you how fast you will go. It is the same fact that makes the highest-acceptance configuration the slowest one in the sweeps below.
The energy column comes from a different pair of arms on a different day's prompt — 2026-08-23's power matrix, thinking on against thinking off at short context — and is printed beside these rows because it measures the same mechanism on a completely independent instrument: 1.62× in joules per token against 1.69× in throughput. Two instruments, one effect. Note also the floor row: 27.01 t/s thinking-on against 26.95 thinking-off — identical, which confirms once more that content dependence on this page is almost entirely a speculation effect: without a drafter, what you are writing moves decode by at most 3.6%.
The flags cannot buy the difference back. At the same depth with thinking on, --spec-type none gives 27.01, the shipped n4/p0.75 gives 36.62 (1.36×) and the code-optimal n10/p0.5 gives 38.67 (1.43×) — a 5.6% gain where the same flag change is worth 12% on code. So the practical number for the configuration this page ships by default: anyone running xhigh over a 91,000-token document should plan on 37–39 t/s, not the 62–70 t/s the code probes advertise. The deliverable arrives at code speed; the thinking that precedes it does not.
What the built-in draft head is
Qwen3.8-27B ships with a small helper network inside the same file — a multi-token prediction (MTP) draft head — that guesses several tokens ahead so the full model only has to check them, which is much faster than generating each one. The standard K-quant GGUF files (§08's table) already include it, and the head itself always stays at 8-bit whatever the main quantization is. llama.cpp added support in July 2026, and because every guess is verified by the full model before it is accepted, the output is exactly what the model would have written anyway — you only gain speed. It works on any GPU llama.cpp supports.
--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: a safe starting point, not the peak — the optimum moves with the CONTENT, and NOT :: with the file: n-max 10 / p-min 0.5 wins every code row on BOTH quants (sweep below). :: And it is not free: 1,008 MiB + 11.4% of your window, +898 MiB at n-max 10 (section 05).
How much speed it adds depends heavily on hardware and workload — published reports range from +33% to +145%, and one RTX 5090 tester saw 148 average / 203 peak t/s on the NVFP4 tiers. On the reference 3090, generation went from 40.3 (the 08-21 build's cold-start baseline; the canonical 08-22 baseline is 39.7–39.9) to 58–60 t/s — +45% at the default-ish settings — and on the original benchmark prompt, tuning pushed it to 81.7 t/s, +105%, a figure a 2026-08-22 re-sweep appeared to overturn and a 2026-08-23 matched sweep reproduced to the decimal once the token regime was pinned. The warning all these numbers carry: benchmark on your own machine. That 5090 tester found n-max 4 fastest with 6–8 slower; this 3090 agreed on real code yet kept climbing to n-max 10 on its near-ideal prompt. The optimum moves with hardware and content, so there is no universal best value.
The cleanest demonstration measured here (2026-08-22): same card, same flags (n-max 10, p-min 0.5, temperature 0), only the content differs. Generating novel code: 51.5 t/s at acceptance 0.47, mean accepted run 4.0 of 10 drafted. Copying a provided snippet verbatim: 119.8 t/s at acceptance 0.93, mean run 9.9 of 10 — on the reasoning stream, so this timed the model thinking about copying rather than copying, which is the caveat the band table at the top of this section attaches to the same figure. A 2.3× speed difference from identical settings. The independent re-measurement ran the same demonstration on answer tokens and got a wider spread still: a 41.5–42.2 t/s floor against 148.7 t/s copying — a 3.59× swing on one card with one file and one set of flags, produced by nothing but what was being written. Which is why a speculative benchmark number quoted without its acceptance rate cannot be checked. With one refinement the sweep section below carries: acceptance ranks your content, but it does not rank your flags.
DFlash2, and why it lost
There is a second way to draft, and on this machine it was slower than the one already inside the file. DFlash2 predicts a whole seven-token block in a single pass, then picks the most plausible sequence out of the candidates it considered for each position. Like the built-in head, every draft is verified before it is accepted, so the output is identical to running the model alone. The costs are real: a separate 1.14 GB drafter to download (incoai/Qwen3.8-27B-DFlash2-GGUF, about 1.2 GB of VRAM) and — until llama.cpp PR 27342 is merged — building llama.cpp yourself from that branch. On its best benchmark prompt it reached 69–77 t/s; on realistic code its best was 46.5 t/s at n-max 2, below the built-in head's 57.9. That is the whole verdict.
| Configuration (hover or tap a row for its exact flags) | Decode | vs baseline | Draft acceptance |
|---|---|---|---|
| Q4_K_M, no speculation (08-21 cold-start build — the canonical 08-22 baseline is 39.7–39.9) | 40.3 t/s | — | — |
| Q4_K_M + built-in MTP (n-max 4, p-min 0.75 — the everyday pick; the 58–60 / 89% here is the 08-21 run, and the 08-22 re-sweep measured the same flags at 57.9 t/s, 0.81 acceptance, on realistic code) | 58–60 t/s | +45% | 89% |
| Q4_K_M + MTP at n-max 10, p-min 0.5 (the peak, on answer tokens) | 81.7 t/s† | +105% | ≈80% (longer drafts) |
| Q4_K_M + DFlash2 (n-max 7; benchmark prompt — the 2026-08-22 re-sweep on real code measured its best at 46.5 t/s, n-max 2, below tuned MTP's 57.9) | 69–71 t/s | +75% | 73% (7-token blocks) |
| NVFP4-HIGH + built-in MTP (2026-08-21 build) | 46.8 t/s | +16% | 90% |
| NVFP4-VERY-LOW + built-in MTP (2026-08-21 build — the 2026-08-22 re-measure at these same flags flipped the tier order: VERY-LOW 54.6 against HIGH 48.2 t/s, both accepting 0.83 — §08) | 43.0 t/s | +7% | 86% |
The exact llama-server command for each row (tap to expand)
:: shared flags in every run below: llama-server -c 32768 -ngl 99 --parallel 1 -ctk q8_0 -ctv q8_0 --jinja --port 1235 :: row 1 — baseline: just the model, no speculative flags -m Qwen3.8-27B-Q4_K_M.gguf :: row 2 — built-in MTP, n-max 4 / p-min 0.75 — the everyday pick (the head ships inside the main GGUF): -m Qwen3.8-27B-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: row 3 — MTP at n-max 10 / p-min 0.5 — the peak on answer tokens († = measured warm-cache): -m Qwen3.8-27B-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 10 --spec-draft-p-min 0.5 :: row 4 — DFlash2 (separate 1.14 GB drafter via -md; needs the PR 27342 build): -m Qwen3.8-27B-Q4_K_M.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 7 :: rows 5-6 — NVFP4 files with their embedded MTP head (same MTP flags): -m Qwen3.8-27B-NVFP4-MTP-HIGH.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -m Qwen3.8-27B-NVFP4-MTP-VERY-LOW.gguf --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 :: what each speculative flag means: :: --spec-type which drafting method: draft-mtp (built-in head) or draft-dflash (external drafter) :: -md / --model-draft path to the external drafter GGUF (DFlash2 only) :: --spec-draft-n-max how many tokens to draft per step - the main tuning knob, sweep 2-12 on your machine :: --spec-draft-p-min drafter confidence floor 0-1; below it, drafting pauses for that step
That table says two things. First, speculation is real free speed: the re-swept everyday flags measured 57.9 t/s on realistic code, and the temperature-1.0 xhigh sweep sustained 48.8 t/s across a 73,000-token thinking run — putting the 100k reference run at roughly 34 minutes, down from 45. Second, the NVFP4 rows show that the quantization and the drafter interact: speculative decoding makes the GPU check several drafted tokens at once, and that batch-checking is exactly the work NVFP4's software fallback does slowest. So on cards older than Blackwell, pair speculation with normal K-quants, not NVFP4.
Tuning the drafter: three sweeps, two reversals, and one acquittal
The settings advice is short and it is at the end of this part, in the "So what should you set?" box: use n-max 10 / p-min 0.5 where the memory is spare and the work is code, and n-max 4 / p-min 0.75 everywhere else. What takes the space is the story of how this page got there, because it went wrong twice on the way and both mistakes are ones any careful person could repeat. An 18-configuration sweep measured one answer; a second sweep the next day appeared to demolish it; a third, run 2026-08-23 with the token regime finally pinned, reinstated the first and retired the demolition. All three stay on the page, in order.
Two knobs control drafting behaviour, and both mattered more than the defaults suggest.
The comparison table above used the settings each method's authors recommend. Then the campaign swept --spec-draft-n-max across both methods, same prompt, temperature 0, warm cache — a different run state from the cold-start table, so absolute numbers shift a few t/s in either direction: this sweep's baseline reads 39.8 against the table's 40.3, and comparisons belong inside one chart, not across them. The result contradicted both recommendations, and then a second sweep on realistic content contradicted the first:
Three findings, all measured 2026-08-21:
- On this prompt — and only on prompts that draft like it — the tuned built-in head won outright: n-max 10 with p-min 0.5 gave 81.7 t/s, +105% over baseline, beating DFlash2's best (77.6 at n-max 4–5). No drafter download and no source build — but not "no VRAM cost", which is what this bullet used to say. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) isolated the built-in head with an on/off pair and measured 1,008 MiB fixed, plus 5,120 B per token of window — 11.4% of whatever
-cyou set — plus a further 898 MiB for n-max 10 over n-max 4. At-c 163840that is roughly 1.8 GiB, more than the vision projector costs. The head's own weights are only 334.7 MiB; the rest is draft-path allocation.--spec-type nonefrees every byte of it, and the server then logsmodel has unused tensor blk.64.* -- ignoringto prove the head was never loaded. --spec-draft-p-minis the hidden second knob. At n-max 10, lowering it from 0.75 to 0.5 was worth +8 t/s (73.5 → 81.7); at n-max 4 the difference was small. A high confidence floor keeps pausing exactly the long drafts that a big n-max exists for.- DFlash2's recommended n-max 7 was not its own best setting here — 4 and 5 both beat it (77.6 and 77.5 against 74.0). Every recommendation in this space, including this page's, is one machine's measurement.
Those findings are one prompt's truth, and that prompt drafted near-ideally. Re-swept 2026-08-22 — 10 configurations, temperature-0 code-generation probe, thinking left on, which nobody recorded at the time — the ranking inverted: n-max 4 / p-min 0.75 won at 57.9 t/s (acceptance 0.81), while the benchmark prompt's champion n-max 10 / p-min 0.5 fell to 48.3. Read the rest of this box as measured history: the numbers stand, the verdict did not survive 2026-08-23's matched sweep, and the case study below is the post-mortem. DFlash2 flipped with it: its best on real code was 46.5 t/s at n-max 2 — below the tuned built-in head, the reverse of its benchmark-prompt 69–77, and hard to justify against the extra download and source build it costs. The explanation offered at the time — "at ~80% acceptance long drafts compound into free tokens; at ~50% they compound into wasted verification" — is the part that has since been retired: the matched sweep below runs at 60% acceptance and is the fastest configuration measured on either file.
Both sweeps above ranked configurations by decode speed and explained them with acceptance. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) swept eight configurations and found the two metrics pulling in opposite directions — the highest-acceptance configuration is not the fastest one:
| Config | Acceptance | Drafted per pass | Accepted per target pass | Decode |
|---|---|---|---|---|
| n3 / p0 | 87.1% | 2.97 | 2.59 | 82.0 t/s |
| n4 / p0.75 | 89.9% | 2.94 | 2.65 | 81.0 t/s |
| n6 / p0.5 | 69.0% | 5.22 | 3.61 | 80.2 t/s |
| n10 / p0 | 46.8% | 9.83 | 4.60 | 87.3 t/s |
| n10 / p0.5 | 59.2% | 7.69 | 4.56 | 87.4 t/s |
| n10 / p0.75 | 72.3% | 4.59 | 3.32 | 80.3 t/s |
| n16 / p0.5 | 47.1% | 10.36 | 4.88 | 80.3 t/s |
Conditions: UD-IQ4_XS, -c 32768, -ngl 99, q8_0 KV, --parallel 1, temperature 0 / top-k 1, an 83-token novel-JavaScript prompt, 700 predicted tokens, a fresh server per configuration, reasoning tokens. accepted per target pass = draft_n_accepted ÷ (predicted_n − draft_n_accepted), both fields from the server's own timings.
Rank configurations by accepted tokens per target pass, not by acceptance rate. n-max 4 / p-min 0.75 accepts a near-perfect 89.9% and still loses, because a shallow draft caps how much a single verification pass can deliver. But the metric is a ridge, not a straight line: n16 accepts the most per pass (4.88) and is 8% slower than n10, because drafting deeper costs real time. The same reversal shows up at the extreme: on the verbatim-copy row, n-max 4 accepted a flawless 100.0% (424 of 424) and ran 97.0 t/s while n-max 10 accepted 99.0% and ran 148.7 — 3.96 accepted per pass against 10.54.
Earlier editions split the flags by file: n-max 10 / p-min 0.5 for UD-IQ4_XS, n-max 4 / p-min 0.75 for Q4_K_M, with the deciding experiment named and unrun — one matched sweep across both files on the same code prompt. It has now been run. Seven configurations × two files × two probes each, same server flags, same 149-token prompt asking for a novel sliding-window rate limiter (not a textbook algorithm — those inflate acceptance), thinking off, so all 700 timed tokens are the deliverable:
| Config | UD-IQ4_XS | acceptance | Q4_K_M | acceptance |
|---|---|---|---|---|
--spec-type none | 42.97 t/s | — | 39.99 t/s | — |
| n-max 2 / p-min 0.75 | 73.41 | 96.5% | 65.76 | 96.7% |
| n-max 3 / p-min 0.75 | 79.57 | 93.3% | 66.74 | 94.2% |
| n-max 4 / p-min 0.75 (shipped default) | 83.50 | 89.7% | 69.82 | 88.6% |
| n-max 6 / p-min 0.5 | 87.78 | 75.8% | 66.38 | 72.1% |
| n-max 10 / p-min 0.5 | 93.86 · 2.18× | 61.4% | 81.71 · 2.04× | 60.0% |
| n-max 10 / p-min 0.75 | 90.84 | 76.8% | 78.15 | 75.4% |
Conditions: -c 32768 -ngl 99 --parallel 1 -ctk q8_0 -ctv q8_0 --jinja, enable_thinking:false, temperature 0 / top-k 1, 700 predicted tokens, a fresh server per configuration (the speculative flags are load-time) plus a discarded warm-up, two probes per server, agreeing to about 1%. Answer lengths varied by about 1% between configurations despite temperature 0 — the usual CUDA batch-shape reduction-order non-determinism, with every stream in a file sharing its opening 120 characters. "Identical output regardless of speculative flags" is true in principle and not quite literally true on this build.
Four readings, and the first one retires a claim this page shipped. Same winner, same loser, same shape on both files — n10/p0.5 is fastest on each, n2/p0.75 is the slowest speculating configuration on each, and the ranking between them matches. There is no per-file optimum to split; the split earlier editions published was a token-regime difference wearing a file's name. Acceptance is a property of the draft head, not of the quantization: at six of the seven configurations the two files land within 1.6 points of each other (96.5/96.7, 93.3/94.2, 89.7/88.6, 61.4/60.0, 76.8/75.4), and the seventh — n6/p0.5 at 75.8/72.1 — is 3.7 points apart. The campaign log's own summary says "within 1.6 points at every configuration"; the table above shows that is true of six of them, and this page reports the exception rather than inheriting the rounder claim. The quantization does not change how well the head predicts — it changes how fast the target model verifies. UD-IQ4_XS is faster at every configuration, and gains more from a wider drafter (+7.5% with no drafter, +19.6% at n4, +14.9% at n10). And the shipped default leaves 11% on the table on UD-IQ4_XS and 15% on Q4_K_M for this workload. The one genuine difference is of degree, not kind: UD-IQ4_XS climbs steadily to n6 while Q4_K_M flattens and dips there (69.8 → 66.4), because its verification step is expensive enough that a wide, low-acceptance tree stops paying sooner.
This page has printed 81.7 t/s in half a dozen places since the first sweep, always with an apology attached. The 2026-08-21 sweep measured it on Q4_K_M at n-max 10 / p-min 0.5 and called it the tuned peak; the 2026-08-22 re-sweep on "realistic code" measured the same flags on the same file at 48.3 and concluded the peak was an artefact of a near-ideal benchmark prompt; the number was demoted to "best case", footnoted as unreproducible on real work, and this page's default moved to n-max 4. The re-sweep's verdict was wrong, and the table above says so: the same file at the same flags on a deliberately novel code prompt measures 81.71 t/s. The two decimal places are a reproduction match between two independent runs, not a claim that the level is stable to 0.01 t/s.
The original mistake was not the prompt. It was an unlabelled token regime: the 2026-08-22 re-sweep left thinking on, so it timed the model reasoning about code while its label said code generation, and reasoning tokens draft short (the mean-draft-length mechanism above). Switch the regime off and the "irreproducible" headline reproduces to two decimal places. Three lessons, and the third is the expensive one. A speed number without its token regime cannot be checked — §10's checklist line was added for exactly this. A disproof is a measurement and needs its conditions stated as carefully as the claim it overturns; this one was published as a correction, propagated into a default, and stood for a day. And a result that refuses to reproduce is a lead, not a verdict — the discrepancy was pointing at a real mechanism the whole time.
n-max 10 / p-min 0.5 is the speed pick wherever the VRAM is there to spend — it wins every code row on both files, by 12% over the shipped default, it costs 898 MiB (about 20,900 tokens of window, §05), and it is also the cheapest setting in energy: 3.210 J per decode token against 3.743 at n4/p0.75 and 8.104 with no drafter. Text-only serving is where that trade is easy, which is why §03's text-only rows carry it. n-max 4 / p-min 0.75 is still right in two places. The first is a tight memory budget: the shipped vision default at -c 122880 is already spending 1,138 MiB on the projector and needs the rest of its slack for a desktop that measured anywhere from 133 to 1,181 MiB — 898 MiB is not free there. The second is reasoning-heavy work at depth: on the thinking stream at 91k the wider drafter setting recovers only 5.6% (38.7 against 36.6 t/s), which does not buy a gigabyte. And on prose, neither flag matters much — n-max 4 wins by 10% (48.4 against 43.8) over a drafter that is only worth 1.16× at all, so --spec-type none and the 1.8 GiB back is the better answer there.
Every drafting number this page has measured still sits on a single line. What the 2026-08-23 round changes is what that line runs along. Draft acceptance was the obvious candidate, and within one configuration it works: same flags, different content, 0.47 → 51.5 t/s, 0.93 → 119.8, and with thinking off 0.99 → 148.7. Across configurations it inverts. The highest-acceptance configuration in the matched sweep — n2/p0.75 at 96.5% on both files — is the slowest speculating one on both; and at 91k depth two streams with statistically identical acceptance (0.895 and 0.907) run 1.69× apart. The quantity that survives both tests is the product: how many tokens the drafter dares to propose × how many of them survive = accepted tokens per verification pass. Stated as a rule: acceptance tells you whether the drafting was right; mean draft length tells you how fast you go; only their product ranks a configuration. The old sweep was never wrong as a measurement — it was wrong as a promise, because it reported one point on this curve as if it were the curve.
Published speculative-decoding tables — like the DFlash2 blog's "mean acceptance length" — are speed numbers, not quality scores. The verification step guarantees every configuration writes exactly what the base model would have written; a higher acceptance number just means more tokens got through per check. Quality only changes when the weights change (a different quantization); speed only changes when the drafter changes.
07
What the same run costs on other cards
Only one row of this section was measured: the RTX 3090. Every other card's number is the same arithmetic from §04 — bandwidth divided by file size, times the efficiency constant for that file's format — applied to the card's published bandwidth. That arithmetic predicted the measured 3090 row correctly, which is the only reason to trust it elsewhere, and it is still a prediction. The pattern that matters more than any individual number: how much memory a card has decides which strategy you can use; how fast that memory is decides how long you wait. A card with enormous bandwidth and too little memory loses to a slower card that fits the model, because the moment one layer lands on the processor the whole calculation changes to the processor's memory speed.
The figures below are for the reference run: about 2,000 tokens of prompt and about 100,000 tokens of thinking and output, with each card using the settings that keep quality highest. All of them exclude speculative decoding, which varies per system but helps a lot: on the reference 3090 it moved this run to roughly 34 minutes from 45. Treat the drafter as moving a card up this chart — furthest on predictable content — rather than as a new baseline.
| Card | VRAM | Bandwidth | Strategy | Decode | 100k run |
|---|---|---|---|---|---|
| RTX 5090 | 32 GB | 1792 GB/s | UD-Q4_K_XL, all on GPU, 262k window — or NVFP4 via vLLM | ~65–80 t/s der | ~25 min |
| RTX 4090 | 24 GB | 1008 GB/s | Q4_K_M, all on GPU, ~122k window | ~40–45 t/s der | ~40 min |
| RTX 3090 | 24 GB | 936 GB/s | Q4_K_M, all on GPU, 122,880 window | 40 t/s meas | 45 min meas 100k ÷ 40 t/s is 41.7 min of decode; the rest is prefill and turnaround |
| RTX 5080 / 4080 / 4070 Ti S / 5060 Ti | 16 GB | 448–960 GB/s | 4-bit partial offload (fidelity) or UD-Q2_K_XL resident at -c 65536 (speed) — 13,982 MiB meas. UD-Q3_K_XL resident SUPERSEDED: measured, it does not fit a 16 GB card at this window | ~8–12 / 25–50 t/s der | ~3 h / ~45 min† |
| RTX 3060 / 5070 | 12 GB | 360–672 GB/s | heavy offload — better served by a smaller model | ~6–8 t/s der | ~4–5 h |
| Intel Arc Pro B70 | 32 GB | 608 GB/s | Q5/Q6 GGUF (Vulkan) or OpenVINO int4; 229k with the drafter, full 262k with --spec-type none (§03) | ~18–26 t/s der | ~1.1–1.5 h |
| Intel Arc Pro B50 | 16 GB | 224 GB/s | UD-Q2_K_XL resident at -c 65536 — 13,982 MiB meas, 606 MiB spare; or OpenVINO int4. UD-Q3_K_XL no longer recommended here SUPERSEDED — measured, it needs ~16,218 MiB at -c 49152 and does not fit. 4-bit does not fit either | ~8–12 t/s der | ~2.5–3.5 h† |
| Intel Arc B580 | 12 GB | 456 GB/s | even 3-bit does not fit — heavy offload (Vulkan) or 2-bit; better served by a 14B model. The 2-bit option now has numbers: the ladder measured every file down to 1.8 bits, and quality turns sharply below 2.91 bits per weight — a 2-bit file costs about 21% perplexity against this page's 4-bit default, five times the price per gigabyte of any step above it | ~6–8 t/s der | ~4–4.5 h |
| DGX Spark (GB10) | 128 GB unified | 273 GB/s | Q4_K_M on llama-server (§03's recipe) · NVFP4 via vLLM for maximum speed · Q8_0 for reference quality | ~10–15 t/s der | ~2–2.8 h |
| Intel Arc B390-class iGPU (Core Ultra) | shared, ≤96 GB | 153.6 GB/s (LPDDR5X-9600, 2ch) | Vulkan Q4_K_M on llama-server (§03's recipe) or OpenVINO int4; effort ≤ medium | ~5–7 t/s der | ~4.5–5.5 h |
† the 3-bit rows trade quality for that speed — a different model for benchmark purposes (§08). They also trade window: on 16 GiB the resident 3-bit file caps context near 49,152 tokens by §05's budget arithmetic, so a single run longer than 100k tokens must either checkpoint across windows or take the offload path.
Every row here is bandwidth ÷ file size × the format constant, so the bandwidth figure carries the whole prediction. The independent re-measurement verified each one against primary vendor documents and found three ways the public numbers mislead. Precision that has no source. The RTX 3090 is 936 GB/s — NVIDIA's own GA102 whitepaper figure, 384-bit × 19.5 Gbps. The widely-copied "936.2" is a third-party derivation from an unrounded memory clock: not wrong, but false precision with nothing behind it. Marketing bandwidth that is not bandwidth. The RTX 4060 Ti's much-quoted "554 GB/s" is NVIDIA's effective-bandwidth argument about the card's 32 MB L2 cache; its peak is 288 GB/s, and only peak belongs in a bandwidth column — feed 554 into the formula and you will promise a reader roughly double the speed the card delivers. Silicon ceilings quoted as machine figures. Apple's M4 Max is 410 or 546 GB/s depending on the bin, so "up to 546" describes the part number, not the laptop in front of you. And a sourcing note for anyone re-checking this table: TechPowerUp and AnandTech both sit behind bot challenges, and a naive scrape of TechPowerUp returns a perfectly valid-looking page with zero specification content in it — which is exactly how a wrong number gets a citation.
System-memory assumption for every offload and shared-memory row: those t/s figures are derived at dual-channel bandwidth — two sticks: desktop DDR5-5600 ≈ 90 GB/s, the integrated-GPU row's LPDDR5X-9600 ≈ 153.6 GB/s. A single stick runs single-channel at half that bandwidth — halve the t/s and double the hours, and slower DDR5 grades scale down proportionally. If your machine reads slower than this table, check that assumption before blaming the table: Task Manager → Performance → Memory shows your speed and slots used, and §04's formula predicts your number — your GB/s ÷ your file's GB × 0.70 for a K-quant or 0.65 for an IQ-quant. Read the family off the filename: a name containing IQ (UD-IQ4_XS, UD-IQ3_XXS, UD-IQ2_S) is an IQ-quant and takes 0.65; a name containing K without IQ (Q4_K_M, UD-Q3_K_XL, UD-Q2_K_XL, Q6_K) is a K-quant and takes 0.70. Rows where the model is resident on a discrete card are unaffected; their bandwidth is the card's own.
08
Which file to download
On a 24 GB card, download UD-IQ4_XS from unsloth for everyday use, or Q4_K_M from lmstudio-community if you want the highest measured quality and can accept that it leaves only about 0.3 GiB of the card for your desktop. Two files of the same size from different quantizers really can differ in quality, so the choice is not arbitrary — but between these two the measured difference is 0.9% of perplexity, which is smaller than the uncertainty of the measurement itself, while the size difference is 2.1 GiB and buys you either the vision projector or about 50,000 more tokens of window. If you have a 16 GB card, take UD-Q2_K_XL — the memory each file and window needs is now measured rather than estimated, and it is the 2.9-bit file, not the 3-bit one, that fits a 65,536-token window with its drafter still switched on (the requirement table below); if you have 32 GB, Q6_K; and if you have a Blackwell card and can run vLLM, the NVFP4 build is the one that runs natively there. If you have 12 GB, the answer is no — not a smaller quantization of this model, but a smaller model. And the 24 GB answer is the interesting result of this revision: the right file now depends on how big a window you want, and there are three answers rather than one. Up to about 98,000 tokens, stay on UD-IQ4_XS. At about 131,000, UD-Q3_K_XL decodes 20% faster on this page's fill and ties the reference file on accuracy. At the full native 262,144, UD-Q2_K_XL is the faster one and the only one that can still speculate. All three are measured; the ladder that settles the quality half starts at the eight-rung table below, the memory and speed each configuration needs is in the requirement table, and the card-by-card verdicts are in the last part of this section.
The advice has not changed: stop at about 9 GiB. The reason it was given has, and the old reason was wrong. Earlier editions of this page said that quality falls about 1% of perplexity per gigabyte saved down to roughly 2.9 bits per weight and about five times that below it, and inferred "stop at 9 GiB" from the shape of that curve. The per-gigabyte figures were right. The inference was not. 2.9 bits per weight is not the point after which quality starts to degrade — it is the last size that still performs like the full-quality file.
Four instruments now sit beside every rung, and they have different sensitivities. Perplexity falls with every step from the very top, so it never had a flat stretch that could mark a beginning. Task accuracy is flat from 4.22 bits per weight down through 2.91 — the same 75 questions score 97.30, 96.00, 93.30 and 96.00 — and one rung further down, at 2.48, it is still a statistical tie with the 4-bit file. Empty answers, the failure a reader actually notices, number exactly zero at every rung down to and including 2.91, then rise 2, 3, 5, 28 below it (the full table). Take the second and third together and 2.9 bits is where "still behaves like the big file" ends, not where quality begins to fall.
Where it stops working is lower still, and it fails in a way no quality curve predicted. The smallest file measured, at 1.835 bits per weight, does not merely answer badly: it loses the ability to stop. Twenty of its 75 answers ran into the token cap, its median answer is 932 tokens against 424 at the top of the ladder, and one arm took two and a half hours where the next longest took thirty-six minutes.
The fourth instrument settles where “stops working” actually begins, and it is two rungs higher than that. meas On 2026-08-25 the JavaScript each file had written was rebuilt and run. Everything down to 2.481 bits per weight executes and prints its answer; at 2.153 and below none of it does — one throws a TypeError, two do not even parse — while the prose around the code still reads perfectly well (the execute probe). Earlier editions put the working floor at 1.835 SUPERSEDED, on detectors that only ever read the output. Note what this does to the four answers: the two instruments that examine text disagree with each other, and the two that ask whether the model works — a paired accuracy test and a parser, sharing no machinery at all — break at the same rung. Four instruments, and the agreement is as informative as the disagreement.
This is the summary. The eight-rung table and the paired tests behind the word "tie" are in the part below, the empty-answer audit in full is in the one after it, what each file and window needs in memory is in the requirement table, which file is fastest at which window is in the part after that, and what to download for a 24, 16 or 12 GB card is in the last part of this section. §15 still lists what is measured and what is not — including the two things this ladder does not establish: speed and fit on any card that is not this 24 GB 3090.
| Size class | Best file | Weights (GB · GiB) | Notes |
|---|---|---|---|
| 4-bit GGUF (the default class — the comparison at the end of this section splits Q4_K_M against UD-IQ4_XS) | Q4_K_M · lmstudio-community (or UD-Q4_K_XL) | 16.5 GB 15.4 GiB | settled by measurement 2026-08-22: wikitext-2 perplexity (the table below) came out 2.3% in Q4_K_M's favour against UD-Q4_K_XL, and a 200-question GSM8K comparison agreed within noise — though that GSM8K comparison is now unaudited, see the warning below. Unsloth's claim that the XL file stays about 10% closer to the full-precision model's own next-token probabilities is measured on their own calibration text and is relative, worth 1–3 true accuracy points at most. With quality tied, the 1 GiB decides: Q4_K_M's extra gigabyte of slack holds about 23,800 more resident context tokens at the measured slope, or the vision projector (§03) |
| 3-bit GGUF — and, on 24 GB, the pick at about a 131,072-token window | UD-Q3_K_XL · unsloth | 13.1 GB 12.2 GiB | Measured here as of 2026-08-24: perplexity 6.7691 ± 0.047, which is +2.63% against the 4-bit UD-IQ4_XS file this page ships — the cheapest quality step down on the whole ladder, at 1.03 GiB saved — and it ties that file on 75 paired benchmark items, p=1.00. Every functional detector passes (§02). New on 2026-08-25: at -c 131072, deep-filled with prose, it decodes 41.74 t/s against UD-IQ4_XS's 34.68 and needs 19,724 MiB against 20,848, which makes it the 24 GB pick at that window (§08). Qwen's own AD-IQ3_S reportedly scores slightly better (next-token probability data) but you have to request access to download it, and it is not measured here. On a 16 GB card, what it needs is now measured and what it would do is not: at -c 65536 with the drafter it asks for 16,906 MiB, more than such a card has in total (§08) — the memory figure transfers to any card, the speed figure never can |
| 2-bit GGUF — the 16 GB pick, and on 24 GB the full-window pick | UD-Q2_K_XL · unsloth | 9.83 GB 9.154 GiB | Measured here 2026-08-24 and 2026-08-25, and the surprise of this revision: at 2.912 real bits per weight it ties the 4-bit reference file on 75 paired benchmark items — they answered exactly one of the 75 differently, p=1.00 (§02) — and returns zero empty answers, for +6.07% of perplexity. It is the smallest file on the ladder that still performs like the full-quality one. On a 24 GB card it is not the daily pick — the drafter makes UD-IQ4_XS 12.9% faster at a short window — but at the full 262,144-token window it is 34% faster and the only one of the two that can still speculate (§08). On a 16 GB card its requirement is measured: 13,982 MiB at -c 65536 with the drafter on, which fits inside such a card with 606 MiB left after this page's desktop reserve (§08). How fast it would run there is not measured and cannot be measured on this machine. And one limit that matters more than speed: everything on this page was measured on single-turn prompts of at most 16,384 tokens. Nothing here tested a long agentic loop where a model calls tools dozens of times in a row. An independent tester (§08) reports this file completing a web app in one shot but locking into an infinite tool-calling loop on a game-engine task. For agentic coding, treat this file as untested rather than recommended. |
| 6-bit GGUF (quality, 32 GB cards) | Q6_K · lmstudio-community | 22.4 GB 20.9 GiB | wins its size class; near-8-bit quality. Unmeasured here |
| 8-bit GGUF (reference) | Q8_0 · lmstudio-community | 29 GB 27.0 GiB | the reference-quality choice, for machines that can hold it (DGX Spark; 32 GB cards at short context). Unmeasured here |
| NVFP4 (native on Blackwell) | Qwen3.8-27B-NVFP4 · unsloth, via vLLM | ~15 GB ~14.0 GiB | for Blackwell cards (RTX 50 / GB10 / B200): about 1.5× the speed of the unquantized model while keeping 92–97% of its accuracy, per the vendor. Runs only in vLLM — its FP8 output layer is not supported by llama.cpp or SGLang. cited, unmeasured here |
| NVFP4 GGUF (any GPU — measured) | NVFP4-MTP GGUF tiers · community, via llama.cpp | 14.8–24.9 GB 13.8–23.2 GiB | The model card claims these need a Blackwell GPU. That claim was tested on an RTX 3090 and it is wrong — llama.cpp falls back to software decoding and the files run. The catch: on older cards you keep the small file and lose the speed. Prompt processing is about 20% slower than Q4_K_M (1058 against 1288 t/s on the 08-21 build), speculative decoding is slower too (54.6/48.2 t/s for VERY-LOW/HIGH against Q4_K_M's 57.9 with the same head), and the energy is worse: 9.293 J per decode token against UD-IQ4_XS's 8.198, +13.4%. The one reason to pick NVFP4 on an older card: the 13.8 GiB VERY-LOW file is about 1.6 GiB smaller than Q4_K_M, which buys roughly 38,000 extra tokens of window at the measured slope. One quirk: llama-bench reports these files as "Q8_0" — a metadata bug, so tell them apart by file size |
| OpenVINO int4 (Intel Arc / iGPU) | Qwen3.8-27B-int4-ov · Intel official | ~16 GB ~14.9 GiB | INT4_ASYM g128; first-party Intel runtime, OpenAI-compatible server; experimental — validate long runs. cited, unmeasured here |
| OpenVINO int8 (32 GB Arc quality play) | Qwen3.8-27B-int8-ov · Intel official | ~28 GB ~26.1 GiB | Intel's equivalent of the 6-bit/8-bit quality tier; needs a 32 GB card. cited, unmeasured here |
Units, because they bite: download pages list decimal GB while VRAM budgets are binary GiB, and the gap is about 7%. The 17.6 GB UD-Q4_K_XL is 16.39 GiB resident, which is why "17.6 GB of weights + 4.25 GiB of cache" still fits inside 24 GiB (§05). Every GiB figure in the table above is the GB figure divided by 1.073741824.
For a benchmark, use the same quantization on every machine — a different quantization is a different model and its scores are not comparable. For daily use, pick the best quantization that fits entirely in your VRAM, and remember that the file you pick is also an energy decision: the three files measured here span 8.198 to 9.293 J per decode token, a 13% range, because a bigger or slower-to-decode file spends more board-seconds per token.
Does FP4 beat INT4? A scored smoke test
A common belief says floating-point number formats always beat integer formats at the same bit width. It did not survive the test here: all three 4-bit files scored the same on one benchmark, and on the other the integer-grid file answered two more questions than either floating-point file. Twenty questions cannot rank healthy files — that is what the rest of this section and §10 are about — but it is enough to say the belief is not a rule you can apply blind.
First, what the two formats actually are. A 4-bit quantization stores each weight as one of only 16 allowed values; the formats differ in where those 16 values sit:
The three quantizations above were graded for correctness on GSM8K and MATH-500, 20 problems each, greedy decoding — always take the most likely next token, so the same prompt gives the same answer every time (§02) — identical prompts, a 4,096-token budget, with answers that ran out of budget mid-thinking counted wrong and reported as truncations. One framing rule before the numbers: at 20 problems this is a smoke test, not a ranking — each question is worth 5 points (§02). Its job is to catch a broken file: a bad conversion or a mangled chat template collapses a score by 20–40 points or more, which even 20 questions detects reliably. It cannot rank healthy files, whose true gaps are a point or three (the arithmetic is in §10). Measured on the reference 3090:
| Quant | Format | GSM8K | MATH-500 | Decode (drafter on) | Draft accept len | J / decode token |
|---|---|---|---|---|---|---|
| Q4_K_M | K-quant — integer grid; a filename with K and no IQ | 85% (1 trunc) | 75% (5 trunc) | 63–70 t/s | 4.7–4.9 | 8.853 meas |
| NVFP4-HIGH | FP4, most extras upgraded | 85% (1 trunc) | 65% (5 trunc) | ~55 t/s | 4.7–4.8 | 9.293 meas |
| NVFP4-VERY-LOW | FP4, smallest tier | 85% (0 trunc) | 65% (5 trunc) | ~49 t/s | 3.33 | — not measured |
| UD-IQ4_XS (not in the scored run) | IQ-quant — integer grid too; a filename with IQ | — | — | — | — | 8.198 meas |
The MATH-500 column is cap-taxed and its grader is unaudited. Five of 20 answers per file — 25% — ran out of the 4,096-token budget and were graded wrong. This page's own rule, stated in §10, is to raise the budget rather than shrink the test, and that rule was applied to the effort comparison's GSM8K run and never to this table. Worse, the 2026-08-23 benchmark sweep found a MATH-500 grader bug of exactly the shape that would bite here — it compared how an answer was written rather than what it was worth, marking 145 wrong against a reference of 145^\circ — and fixing it moved one arm by 32 points at n=25. Read this column as "no file collapsed", nothing finer. The decode column comes from a different llama.cpp build and workload than §06's configuration table (both are listed in §15) — compare speeds within this table only. The energy column comes from a third run, the 2026-08-23 power matrix, measured with the drafter off on a different prompt, so it ranks the files against each other and does not price the configurations in the decode column.
What the numbers say:
- "Floating point always wins" did not survive the test. All three tied on GSM8K; on MATH-500 the integer-grid K-quant answered two more problems than either FP4 file. At 20 samples that edge is weak evidence — but it is the opposite direction from the belief. What actually decides 4-bit quality is how cleverly a format spends its bits — block sizes, scale precision, which layers stay at higher precision — and mature K-quants are very good at that. Both families use fine-grained scaling; the float-versus-integer grid underneath matters far less than people assume.
- The two NVFP4 tiers were indistinguishable at this sample size — which is what noise looks like, not evidence of equivalence. The file structure does predict a tie: the family shares one byte-identical NVFP4 backbone, with tiers differing only in the surrounding layers. The surprise is where VERY-LOW's savings actually landed: on these maths prompts at n-max 10 its draft head drafted noticeably worse (accepted length 3.33 against about 4.8) — a speed cost, not an accuracy one. The penalty is configuration-dependent: re-measured 2026-08-22 at n-max 4 / p-min 0.75 on code generation, both tiers accepted identically (0.83) and VERY-LOW was the faster file (54.6 against 48.2 t/s) — smaller weights, less memory traffic per token.
- Sample-size honesty: each problem is worth 5 points. Every gap in this table is one or two questions wide. Treat it as "no meaningful quality difference, mild edge to the K-quant" — not a ranking. The same prompts were also run through a different serving stack and grader version, and individual cells moved by up to 10 points, which means measurement variance is the same size as the differences being measured. A one-run quantization ranking published without error bars is not evidence, whoever published it.
How the files rank on perplexity
Accuracy at 20 questions cannot separate healthy files, but a per-token measure can, because it feels quantization damage at every scored token instead of collecting one right-or-wrong bit per question. Perplexity is defined in §02; the run below scores 294,912 token positions — 36 chunks × 8,192, forced by the corpus, the tokenizer and -c 8192 — not the "~330k" earlier editions of this page printed, which was an unsourced estimate about 12% high. Measured 2026-08-22 on the reference 3090 with llama-perplexity over the full wikitext-2-raw test set:
| File | Perplexity (lower is better) | vs best |
|---|---|---|
| Q4_K_M | 6.535 ± 0.044 meas | — |
| UD-IQ4_XS (13.3 GiB) | 6.596 ± 0.045 meas | +0.9% — a gap of 0.061 against a ±0.063 combined error bar, so unresolved |
| UD-Q4_K_M | 6.654 ± 0.045 meas | +1.8% |
| UD-Q4_K_XL | 6.682 ± 0.046 meas | +2.3% |
| NVFP4-HIGH | 6.826 ± 0.047 meas | +4.4% |
| NVFP4-VERY-LOW | 6.877 ± 0.047 meas | +5.2% |
Read it with three limits attached. One corpus. Wikitext rewards plain-text prediction while unsloth's Dynamic calibration targets chat-style data, so the ordering among the top files may not carry over perfectly to instruction-following work. The scored-position count is tokenizer-bound. A later run on the same corpus files recorded 297,193 tokens under Qwen's tokenizer and 295,216 under Gemma's, both at 36 chunks — so 294,912 is a count of scored windows (36 × 8,192), not of corpus tokens, and it moves with the tokenizer. The practical consequence is a hard rule: perplexity is never comparable across model families. And the reference file reproduces. A separate 2026-08-23 run re-measured the UD-IQ4_XS figure as a gate before further work and returned 6.5956 against an expected 6.5956 — bit-identical, which is what pins the error bars and the chunk count.
Two robust findings survive all three limits: both NVFP4 tiers trail the K-quants — the same direction the scored table's MATH-500 column hinted, reached by an independent measurement — and none of the unsloth Dynamic variants beat the plain lmstudio Q4_K_M on this corpus, including at identical size. Two cheap metrics agreeing beats one expensive one.
Alongside perplexity, a 200-question GSM8K run compared the same files: Q4_K_M 94.0%, UD-Q4_K_XL 94.5%, UD-IQ4_XS 93.0%, NVFP4-HIGH 92.5%, NVFP4-VERY-LOW 90.5%, all ±3–4 points and therefore a tie by §10's arithmetic. Those numbers now need re-grading before they are quoted again. The 2026-08-23 sweep found that its GSM8K grader had been comparing the whole answer line instead of the number, so #### 156 kg failed against a reference of 156 — and that single bug was worth 8 points, all of it on the arm that reasons in units. The runs above used a different grader (chinkeong/benchmark), so this is a re-grade order, not a proven error. The blast radius matters: §03's the Q4_K_M recipe was argued on "two independent metrics agreeing", and one of the two is this comparison. Until the saved transcripts are re-graded with a number-extracting grader, the Q4_K_M recipe rests on perplexity alone, and that caveat is printed at the Q4_K_M recipe as well as here.
The same corpus, the same 36 chunks and the same error bars also price the cache instead of the weights — the one setting this page used to argue about from mechanism alone. Measured 2026-08-23 on UD-IQ4_XS, every other flag identical to the runs above (-c 8192 -fa on -ngl 99), on a corpus file byte-verified identical to the one the fp16 and q8_0 baselines used:
| K/V cache | Perplexity | vs fp16 | What it buys |
|---|---|---|---|
-ctk f16 -ctv f16 | 6.5956 ± 0.045 meas | — | the reference — and about 1.6× the window cost of q8_0 on the one measured pair (16,428 against 15,661 MiB at -c 32768, which puts the fp16 slope near 64,500 B per token against q8_0's measured 39,936). The arithmetic floors would say 1.88×; budget from a load, not from the floors (§05) |
-ctk q8_0 -ctv q8_0 (shipped) | 6.6160 ± 0.045 meas | +0.309% | halves the cache. Every context ceiling on this page is built on it |
-ctk q4_0 -ctv q4_0 | 6.6413 ± 0.045 meas | +0.693% | halves it again, for a further +0.383% on top of q8_0 |
Three readings, and the third is the one that keeps this honest. Degradation is super-linear in bits — the second halving costs slightly more than the first (0.383 against 0.309), so the cut stops paying rather than scaling. Both are small next to the model quantization: the Q4_K_M-to-UD-IQ4_XS gap on this same corpus is 0.93%, larger than q4_0's entire cache penalty, so anyone who chose the smaller weights file on quality grounds has already accepted a bigger hit than moving to a 4-bit cache would add. And the error bars overlap: ±0.045 on each of two independent estimates gives ±0.063 on their difference, so the fp16-to-q4_0 gap of 0.0457 is about 0.7 of one standard error and a single 36-chunk run does not resolve it on its own. (That combined figure, not the ±0.045 on a single estimate, is the right yardstick for every file-against-file comparison on this page, including the 0.061 gap between the two 4-bit files above.) the direction is consistent with the q8_0 point and with theory, and that is all it is. Two limits on how far to carry this. It was measured at -c 8192, so it says nothing about the long-context retrieval failure that is the actual argument against a 4-bit cache at 200k-plus (§05's mechanism footnote — still unmeasured). And q4_0 was not faster in any way that matters: 266 s against 298 s of perplexity wall-clock is a bandwidth effect on a prefill-shaped workload, not a decode result. Nothing here argues for changing the default; it argues for stating the price instead of the prohibition.
The quantization ladder: eight files of one model, four instruments
This is the table the rest of this section was missing. One model, squeezed to eight sizes from 13.3 GiB down to 5.8 GiB, every file measured the same way on the same day. Three instruments read the table below — perplexity, task accuracy and the empty-answer count — and a fourth one, added 2026-08-25, runs the code each file wrote (further down). It answers the question people actually ask — how small can I go before it stops being the same model? — and it answers it in more than one voice, because the instruments disagree about where the damage starts and each is right about something the others cannot see. Where two of them agree turns out to matter as much as where they differ.
Read the columns like this. Perplexity is a per-token measure of how surprised the model is by ordinary text: it feels a little damage at every rung and is the earliest warning, but a file can lose perplexity and still do your work. Accuracy Mean is the composite of three graded benchmarks — GSM8K, HumanEval and MBPP, 25 questions each, 75 items in total — and it is the only column in the reader's own units. The empty-answer column counts the times the model returned nothing at all, and it is kept separate from truncations because the two are different failures: a truncation hit the token cap, an empty often terminated normally and simply emitted zero characters, which no truncation counter can see. Those last two columns are the functional detectors that §01 counts as the third of this campaign's instruments — the automated checks for what a score cannot see (§02). The two vocabularies name one thing.
| File (all unsloth/Qwen3.8-27B-GGUF) | Weights GiB | Bits per weight | Perplexity vs the 4-bit reference file | Accuracy Mean n=25 × 3 sets | Empty (silent) | Trunc | Plain-language verdict |
|---|---|---|---|---|---|---|---|
| UD-IQ4_XS the reference file · this page's daily file | 13.274 | 4.223 | 6.5956 the reference | 97.30 | 0 (0) | 0 | Full quality. Everything below is measured against this file, on the same 75 questions |
| UD-Q3_K_XL | 12.244 | 3.895 | 6.7691 +2.63% | 96.00 | 0 (0) | 0 | The cheapest step down on the whole ladder. Ties the reference file on the paired test, no empties, no truncations |
| UD-IQ3_XXS | 10.184 | 3.240 | 6.9187 +4.90% | 93.30 | 0 (0) | 0 | Still a tie with the reference file, still clean. Saves 3.1 GiB against it |
| UD-Q2_K_XL where the perplexity curve turns | 9.154 | 2.912 | 6.9957 +6.07% | 96.00 | 0 (0) | 0 | The last rung that still performs like the reference file — ties it on the paired test, and it is the last file with zero empty answers. Below here the perplexity cost per gigabyte jumps about fivefold and the empty count leaves zero for good |
| UD-IQ2_S | 7.797 | 2.481 | 7.5481 +14.44% | 90.70 | 2 (1) | 1 | Still a tie on accuracy, but the weakest one on the ladder and right at the edge — and the first rung that ever returns nothing at all. The empty count sees damage a full rung before the benchmark can |
| UD-IQ2_XXS | 6.767 | 2.153 | 8.0079 +21.41% | 78.70 | 3 (2) | 1 | Measurably worse than the reference file, and the first rung where the paired test says so rather than shrugging. Not recommended |
| UD-IQ1_M | 6.267 | 1.994 | 8.1418 +23.44% | 85.30 | 5 (5) | 0 | Also measurably worse than the reference file. Its higher Mean than the rung above is not a reversal — see below — and it has the worst silent-empty record short of the bottom: five questions answered with nothing, while its truncation counter read zero |
| UD-IQ1_S | 5.767 | 1.835 | 8.9265 +35.34% | 34.70 | 28 (10) | 20 | Broken, and in an unusual way: it stops being able to stop. Median answer 932 tokens against 424 at the reference file, 20 of 75 answers running into the cap, one arm taking two and a half hours where the next longest took thirty-six minutes. This is the ladder's "clearly no" rung |
Conditions, and they differ between the two scored columns. Perplexity meas: frozen wikitext-2-raw test set, 36 × 8,192 = 294,912 scored positions per file — the same count and corpus as the table above it — -ngl 99 -c 8192 -fa on --load-mode mmap, fp16 KV, llama.cpp build 10502. Accuracy and the two failure columns meas: a frozen suite — fixed once and never edited again, so every file is asked the identical questions — 1cdf54f8eb9d3f8f, GSM8K + HumanEval + MBPP at n=25 each, greedy, seed 42, --max-tokens 16384, -c 32768, q8_0 KV, reasoning_effort=low, and no drafter — which turns out to matter, and is dealt with in the next part. "Silent" empties are the ones that did not hit the cap: they terminated normally and returned zero characters, so no truncation counter ever saw them. Means are printed at the harness's own precision; what they can be read to is the subject of the next two paragraphs.
Twenty-five questions per set detects a collapse and nothing finer. That is §10's own first row — about 25–30 samples resolves a ~20-point gap — and the accuracy column above is built from exactly that sample size. So it is in the class this page warns readers about, and the warning applies to it in full: the Mean column may say where this model breaks; it may not rank one rung against another. Every "tie" and every "worse" below comes from a paired test on the identical 75 items, not from the difference between two Means, and no ordering is claimed anywhere that the paired test does not support.
So what can be said, arm against arm, is whatever a paired McNemar test says. It compares the two files question by question on the same 75 items and counts only the ones they disagree about — b is the number the left file got right and the right file got wrong, c the reverse, and p is how often a gap this large would turn up by chance if the two files were really equal (§02). The exact test, counting a difference in either direction:
| Pair | b | c | p | What it licenses |
|---|---|---|---|---|
| UD-IQ4_XS vs UD-Q3_K_XL | 1 | 0 | 1.0000 | Tie. One discordant item of 75 |
| UD-IQ4_XS vs UD-IQ3_XXS | 3 | 0 | 0.2500 | Tie |
| UD-IQ4_XS vs UD-Q2_K_XL | 1 | 0 | 1.0000 | Tie. One discordant item of 75 — the strongest tie on the ladder, and the basis of every 2.9-bit recommendation below |
| UD-IQ4_XS vs UD-IQ2_S | 5 | 0 | 0.0625 | Tie, and the weakest one measured — right at the edge of resolving |
| UD-IQ4_XS vs UD-IQ2_XXS | 14 | 0 | 0.0001 | Different. The 4-bit file is measurably better |
| UD-IQ4_XS vs UD-IQ1_M | 9 | 0 | 0.0039 | Different. The 4-bit file is measurably better |
| UD-IQ3_XXS vs UD-Q2_K_XL | 0 | 2 | 0.5000 | Tie — and perplexity ranks these two the other way, so no ordering claim between them exists on this page at all |
| UD-IQ2_S vs UD-IQ2_XXS | 10 | 1 | 0.0117 | Different |
| UD-IQ2_XXS vs UD-IQ1_M | 5 | 10 | 0.3018 | Tie. The Mean column appears to rise again here; that is noise, and it is never to be written as a reversal |
The one sentence that reconciles a falling perplexity curve with a flat accuracy curve. Look at the c column: it is zero in six of the nine rows. The smaller file almost never wins a question the larger one lost, and in every comparison against the 4-bit reference file the losses run in one direction only. So quality is falling at every step down the ladder; it is simply below what 25 questions per set can resolve until the cliff arrives. The flat accuracy column is not evidence that the files are equal. It is evidence that this instrument cannot see a gap this small, which is what §10 said before the ladder ran.
Where the boundary is, stated as an interval because that is all the data supports. UD-IQ2_S at 2.481 bits per weight ties the reference file. UD-IQ2_XXS at 2.153 does not. The boundary therefore lies between 2.48 and 2.15 bits per weight, and 25 questions per set cannot place it more finely than that. Anyone quoting a single number for where this model breaks — including a future edition of this page — is quoting something the measurement does not contain.
A file called Q2_K_XL measures 2.912 bits per weight, not the 2.5625 that stock Q2_K spends, and Q3_K_XL measures 3.895 against stock Q3_K's 3.4375. That is not an error in either place: k-quants (llama.cpp PR #1684, Kawrakow, June 2023) fix a bit width for each kind of the model's internal number tables, and a vendor's "UD"/"XL" mix then spends about 0.35–0.46 bits per weight more than the label by keeping the tables that matter most at higher precision. cited for the PR and the stock widths; measured for the file sizes this column is computed from. The curve above is a function of the real bits, not of the name on the file — which is exactly why the column is printed.
Two limits to carry out of this table. Every step is resolved on perplexity, and the smallest one is not close: the 4.223-to-3.895 step moves perplexity by 0.1735, about 2.8× the ±0.063 combined error bar this section uses for any file-against-file difference derived, and every step below it is larger. And a cross-model reference point, because "how small before I should just use a smaller model?" is the real question underneath. On the identical frozen suite, gemma-4-12B-it-QAT-Q4_0 — 6.497 GiB at 4.65 bits per weight, within 4% of UD-IQ2_XXS's download size — scored 73.30 with 19 truncations, against the 2.153-bit 27B's 78.70 with one. At an equal weight budget the crushed big model won, which is worth knowing before assuming that a smaller model at higher precision is automatically the safer trade. Conditions differ between those two arms (the Qwen ran at reasoning_effort=low with a q8_0 cache; gemma ran at its own defaults, having no effort knob), and that is a disclosed asymmetry, not a fair fight.
The empty-answer table, in full
The ladder above prints the empty count in one narrow column. It has earned a table of its own, because it is the cheapest instrument in this whole campaign and the only one of the three that draws a usable line. Cheapest, literally: it costs no GPU time at all — every number below was counted from transcripts that were already sitting on disk when the accuracy run finished. And it draws a line where the other two cannot: perplexity falls at every step from the very top, so it never has a flat stretch that could mark a beginning, while the paired accuracy test still says "tie" a full rung below the point where empty answers first appear.
Three words in plain language before the numbers, because they name three different things. An empty answer is a reply with no characters in it — you asked, and nothing came back. At cap means that particular empty reply had run out of its token budget and was cut off; the server records that as a truncation, so it appears in a log where somebody might see it. Silent means the opposite, and it is the one to worry about: the model ended the reply by itself, in the ordinary way, and handed back nothing — an answer that finished normally and returned zero characters, so no truncation counter can see it and nothing in your logs reports a problem at all. That is why the two are separate columns rather than one. Every row is n=75: three benchmark sets — GSM8K, HumanEval and MBPP — at 25 questions each, the identical 75 questions put to every file.
| File (all unsloth/Qwen3.8-27B-GGUF) | Bits per weight | Empty of n=75 | At cap cut off | Silent ended normally | Median answer tokens |
|---|---|---|---|---|---|
| UD-IQ4_XS the reference file | 4.223 | 0 | 0 | 0 | 424 |
| UD-Q3_K_XL | 3.895 | 0 | 0 | 0 | 488 |
| UD-IQ3_XXS | 3.240 | 0 | 0 | 0 | 454 |
| UD-Q2_K_XL the last clean rung | 2.912 | 0 | 0 | 0 | 416 |
| UD-IQ2_S first detection | 2.481 | 2 | 1 | 1 | 475 |
| UD-IQ2_XXS | 2.153 | 3 | 1 | 2 | 472 |
| UD-IQ1_M | 1.994 | 5 | 0 | 5 | 502 |
| UD-IQ1_S | 1.835 | 28 | 18 | 10 | 932 |
Conditions meas: the frozen suite 1cdf54f8eb9d3f8f, greedy at temperature 0 / seed 42, --max-tokens 16384, -c 32768, q8_0 KV, reasoning_effort=low, no drafter — the same eight arms that produced the accuracy column above, read back from their saved artifacts rather than from a counter. How this table relates to the ladder's truncation column. "At cap" counts only the empties that were cut off; the ladder's Trunc column counts every answer that was cut off, empty or not. At the bottom rung the two read 18 and 20 — so 18 of that file's 20 truncated answers came back with nothing in them.
The finding, and it is a clean one. The empty-answer rate is exactly zero at every rung down to and including 2.912 bits per weight — the same rung where the perplexity curve turns and the last one that ties the reference file on the paired test — and then it rises without going back down: 2, 3, 5, 28. It notices damage at 2.481 bits, a full rung above where the paired accuracy test can resolve anything at all. At 2.481 the paired test returns p=0.0625 and says "tie"; the first verdict of "different" does not arrive until 2.153. The empty column had already stopped reading zero one rung earlier, for no GPU time.
How to read it, and how not to. The useful reading is zero against not-zero, and the shape of the rise. The difference between 2 empties and 3 out of 75 is not a ranking and must not be used as one — the same sample-size limit that governs the accuracy column governs this one. What the column does carry is a threshold, and it is the reason this page recommends nothing below 2.912 bits per weight even though the paired test would still license 2.481.
The median-answer column is a second detector hiding inside the first, and it finds a different failure. Median answer length holds inside a band of 416–502 tokens all the way down to 1.994 bits per weight — seven of the eight rungs barely move how long an answer is — and then jumps to 932 at 1.835. That is the rung that stops terminating: it is not answering worse at the same length, it is failing to stop, which is why 20 of its 75 answers ran into the cap and why that one arm took two and a half hours where the next longest took thirty-six minutes. A quality curve would never have predicted it, and neither would the empty count on its own.
A fourth instrument: we ran the code instead of reading it
Every test above scores text. Perplexity scores it per token, the accuracy suite scores whether the answer is right, and the repetition screens check whether it has degenerated into a loop. None of them ever executed anything. So we did meas: each file had been asked to finish a JavaScript Dijkstra implementation, and every one of those programs was rebuilt with the prompt's own opening lines and run under node v24.15.0.
Six of the ten programs run. Four do not. The line falls in a place worth knowing:
- Down to 2.481 bits per weight the code runs —
UD-IQ4_XS,UD-Q3_K_XL,UD-IQ3_XXS,UD-Q2_K_XLandUD-IQ2_Seach build the graph, run the search and print a shortest path with its cost. - Below that, none of it runs.
UD-IQ2_XXSthrowsTypeError: Cannot read properties of undefined.UD-IQ1_MandUD-IQ1_Sdo not even parse.
This does not change what this page recommends, and that is the point of reporting it. The recommendation floor here is 2.912 bits per weight, set by the empty-answer column. The executed floor is 2.481 — one rung below the recommendation. A fourth instrument, aimed at something none of the others measure, landed inside the margin the other three had already left. UD-Q2_K_XL writes code that works, which is the most direct answer this page can give to whether a 2-bit file is a real daily driver.
The failures arrive together, which is the useful part. The prompt said to continue the file and to emit no markdown fences. Every file that ran followed both instructions. Every file that failed broke them — restarting the program from scratch, or wrapping it in the fences it had been told not to use. At the bottom of the ladder, following instructions and producing working output stop being separate skills. That is a better early warning than any single score, and you can watch for it yourself without a benchmark: a model that starts ignoring the shape you asked for is telling you something about the code it is about to hand you. n=1 per file — one task, one language — so read this as a threshold, not a rate.
And a caution about a rule of thumb you will see everywhere. The usual advice is that a smaller model at more bits beats a bigger model at fewer bits. On this one task it went the other way: gemma-4-12B-QAT-Q4_0 at 4.651 bits per weight produced code that does not run — it opened a helper function inside a method, never closed the method, and declared the next one as though it had — while Qwen3.8-27B at 2.912 bits produced a working program. One sample is not a refutation and this page does not offer it as one. It is a reason to run the substitution on your own workload rather than accept either rule of thumb: the comparison is cheap, and it does not always land where the slogan says.
UD-IQ4_XS reference, so that instruments in different units can share one axis: perplexity added = (PPL − 6.5956) ÷ 6.5956, accuracy lost = (97.30 − Mean) ÷ 97.30, empty rate = empties ÷ 75. The measured values behind all three are printed in the tables above. The run / does-not-run strip along the bottom is meas directly. Perplexity moves at the very first step down and never stops — the most sensitive line here, and the one that says least about whether the model still works. The empty-answer rate is exactly zero down to 2.912 bits per weight and first flinches at 2.481: it is the early warning, and it costs no GPU time because it is counted from answers already on disk. Then the two instruments that ask whether the model actually works — the paired accuracy test and a JavaScript parser, which share no machinery at all — break at the same rung. Everything down to UD-IQ2_S at 2.481 runs; nothing below it does. The functional boundary is between 2.481 and 2.153 bits per weight, and it has two independent witnesses. Execution is n=1 per file.What each file and window actually needs
Every recommendation on this page ends in the same small piece of arithmetic, so here is the table that lets you do it yourself instead of trusting a verdict. Find your file and the window you want, read what it needs, and subtract your own desktop from your card's total. What is left over is your slack (§02).
The memory column transfers to any card. The speed column is this 3090's and nobody else's. How much memory a configuration needs is a property of the model, the window and the flags — the same file at the same -c asks for the same bytes on a 5080 as it does here — so a 16 GB owner can plan against these figures directly, and that is the one thing about a smaller card this page can state as measured. How fast it then decodes is a property of this card's memory bandwidth. No speed on this page was ever measured on any other card, and none can be.
| File · weights | Drafter | -c 32768board VRAM · decode | -c 65536 | -c 131072 |
|---|---|---|---|---|
| UD-IQ4_XS 13.274 GiB | on | 16,586 MiB 41.35 t/s | 17,962 MiB 42.69 t/s | 20,848 MiB 34.68 t/s |
| UD-IQ4_XS | off | 15,376 MiB 35.92 t/s | 16,678 MiB 30.73 t/s | 19,188 MiB 24.77 t/s |
| UD-Q3_K_XL 12.244 GiB | on | 15,530 MiB 39.59 t/s | 16,906 MiB 42.98 t/s | 19,724 MiB 41.74 t/s |
| UD-Q3_K_XL | off | 14,322 MiB 36.44 t/s | 15,568 MiB 31.28 t/s | 18,064 MiB 23.91 t/s |
| UD-Q2_K_XL 9.154 GiB | on | 12,606 MiB 43.19 t/s | 13,982 MiB 40.30 t/s | 16,800 MiB 29.44 t/s |
| UD-Q2_K_XL | off | 11,396 MiB 33.48 t/s | 12,644 MiB 32.18 t/s | 15,140 MiB 24.54 t/s |
Conditions, all cells meas (2026-08-25, the reference RTX 3090, board total 24,576 MiB): -ngl 99 -fa on --parallel 1 -ctk q8_0 -ctv q8_0 --jinja, reasoning off, and text-only — there is no --mmproj among the recorded flags, which matters because the projector is the fourth determinant of any ceiling on this page (§05); where the drafter is on it is --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75, the flags §03's vision recipe ships. Every window was deep-filled to about 90% of itself with ordinary English prose from the wikitext-2-raw test set — not probed shallowly, which is the mistake §05 documents — the first probe after the prefill was discarded because the card is still raising its clock (§11), and the figure printed is from two settled probes. Memory is board VRAM, so it contains this machine's desktop as well as the server. The three arms at the full -c 262144 are a separate deep fill and are in the table two parts below.
Your card's total, minus what the table says, is your slack. On this 24,576 MiB card, UD-Q3_K_XL at -c 131072 with the drafter leaves 4,852 MiB and UD-IQ4_XS at the same window leaves 3,728 derived, both comfortably clear of the 1,796 MiB this page reserves. On a 16 GB card the total is 16,384 MiB, so keeping that reserve free leaves 14,588 MiB to spend derived; on a 12 GB card it is 12,288 and 10,980 derived. Hold those two numbers against the column you want and the card verdicts below follow by arithmetic rather than by assertion. And be exact about the reserve itself (§02): the desktop's own share measured 1,179–1,669 MiB in direct no-server readings depending on what was on screen, load-to-load variation measured 127 MiB, and 1,796 is this page's own derived threshold built from the worst case of the first plus the second. The desktop does not need 1,796 MiB. If you know what your own screen holds, subtract that instead and you will get a better answer than this page's worst case gives you.
What the drafter costs, and a cross-check worth having. Turning speculation on costs a consistent 1,208–1,660 MiB across these arms, and the cost grows with the window rather than with the file — which is exactly what §05's two-constant model predicts, since the drafter books 1,008 MiB fixed plus 5,120 B per window token. Subtracting the pairs above gives 1,208–1,210 MiB at -c 32768, 1,284–1,338 at 65,536 and 1,660 at 131,072 derived, against 1,168, 1,328 and 1,648 predicted derived — agreement within 44 MiB, well inside that model's own 127 MiB worst residual. That model is fitted at ordinary windows and it is not reliable at the top of the native window: at -c 262144 it under-predicts the measured requirement by about 1,213 MiB, because it does not carry the compute buffers (below).
Which file is fastest changes with the window — and with what you are writing
The speed column of the table above holds a result that is worth separating out, because it contradicts the way almost everyone reasons about quantization. With the drafter on, a different file reads fastest at different windows. Not a different file per card, or per workload — per window, on the same card, on the same content, with the same flags. Here is that column on its own, with a fourth window added, all arms drafter-on and deep-filled with prose:
| File · drafter on (n-max 4 / p-min 0.75) | -c 32768 | -c 65536 | -c 98304 | -c 131072 |
|---|---|---|---|---|
| UD-IQ4_XS 4.223 bpw | 41.35 t/s | 42.69 t/s | 36.09 t/s | 34.68 t/s |
| UD-Q3_K_XL 3.895 bpw | 39.59 t/s | 42.98 t/s | 32.90 t/s | 41.74 t/s |
| UD-Q2_K_XL 2.912 bpw | 43.19 t/s | 40.30 t/s | 34.08 t/s | 29.44 t/s |
Read down the columns and the highest reading keeps moving: UD-Q2_K_XL at 32,768, UD-Q3_K_XL at 65,536, UD-IQ4_XS at 98,304, UD-Q3_K_XL again at 131,072 — three different files across four windows. Before that becomes a claim about which file is fastest, here is how much of it the measurement actually supports. At -c 32768, -c 65536 and -c 98304 the three files sit inside a narrow band — 3.6, 2.7 and 3.2 t/s from top to bottom der — and at those windows only UD-Q3_K_XL was loaded twice. Those three columns are bands, not rankings, and this page does not order the files inside them. One column is different, and it is the one the recommendation rests on.
At -c 131072 the separation is wide and it reproduced across independent server loads. Each file was loaded twice, from scratch, with five settled probes per load: UD-Q3_K_XL read 41.44 then 42.04 (standard deviation 0.36), UD-IQ4_XS 34.66 then 34.70 (0.16), UD-Q2_K_XL 29.46 then 29.41 (1.07). Every file landed within 1.4% of its own first load. The gap between the top two is 20% — 41.74 against 34.68 der — which survives a 1.4% reproduction spread comfortably. UD-Q3_K_XL was loaded twice at -c 98304 as well — 32.86 then 32.95 — which matters for the mechanism below. On the two decimal places: they are the harness's own precision, not a claim that a cell is stable to 0.01 t/s. The band any cell can be read to is the reproduction spread printed here — 0.16 to 1.07 t/s — and at the three windows where no second load was taken, the band is unknown and the column is not ranked.
Why a file can decode faster at a deeper window. For a plain decoder this is impossible, and the impossibility is the reason the result was doubted: more context means more cache to read per token, so speed falls. A speculating decoder is not a plain decoder. What you actually measure is the model's own decode rate divided by how many tokens it commits per verification pass, and that second quantity belongs to the drafter — specifically to mean draft length, which §06 already establishes is the thing that ranks a configuration while draft acceptance is not. Watch UD-Q3_K_XL move between two neighbouring windows:
| UD-Q3_K_XL, drafter on, prose fill | Draft acceptance | Mean draft length | Decode |
|---|---|---|---|
-c 98304 | 1.00000 | 2.89 | 32.90 t/s |
-c 131072 | 0.885 | 3.30 | 41.74 t/s |
Acceptance falls, draft length rises, and throughput follows the draft length. An acceptance of exactly 1.00000 sitting beside a short draft is not a triumph — it is the confidence floor cutting the draft off early. --spec-draft-p-min 0.75 tells the drafter to stop guessing the moment it is less than 75% sure (§02), so on content it finds hard it proposes almost nothing, and everything it does propose survives checking. You get a perfect score on a nearly empty exam, and you go slowly. This is the same fact §06 measured on the reasoning stream, arriving from a completely different direction.
One more disagreement, and it is the sharpest evidence in this section rather than an embarrassment. The drafter pair below also measures UD-IQ4_XS against UD-Q2_K_XL at -c 32768, and it ranks them the other way round: 86.91 t/s against 77.01, the 4-bit file 12.9% ahead. The table above, at the same -c, reads 41.35 against 43.19 with the 2.9-bit file ahead. Neither number is wrong and neither is the general answer, because two conditions differ between them: that pair generated novel code into a nearly empty window, while every row here is prose filling about 90% of the window. Change the content and change the depth and a speculating decoder reorders itself. If you take one thing from this part, take that — the ordering belongs to the workload, not to the file — and then go and measure your own.
The confidence floor gates on the drafter's confidence, and confidence is a property of the content, not of the file and not of the window. Every number in this part was measured with the windows filled with ordinary English prose from the wikitext-2 test set. Your own work is not that. Fill a window with your own code, or your own codebase, and the drafter is confident in different places, the gate cuts in different places, and the ordering can move with it.
So this table is published as "measured on prose fill" and never as a fixed property of these three files. The proof that it is content-sensitive is inside the table itself: UD-Q3_K_XL reverses its own reading between two neighbouring windows of the same prose. If a change of 32,768 tokens of the same text can do that, a change of subject matter certainly can.
The check costs about ten minutes and this page recommends you spend them. Load your two candidate files at the window you actually use, fill each one with your own content, throw away the first probe after the prefill, and take two settled ones. That is exactly the protocol behind this table, and running it on your own material is worth more than any ordering published here.
The 131,072 result above did not arrive cleanly, and both wrong turns are worth keeping. It was published; then withdrawn when a neighbouring window appeared to contradict it; then reinstated when both windows reproduced across independent server loads.
The retraction's own reasoning was wrong. It argued that decode cannot rise with depth. That is true of a plain decoder and false of a speculating one, for the reason set out above — the observed rate is decode divided by tokens committed per verification pass, and the draft length that sets the divisor is gated by content confidence, not by depth. A retraction is a claim, and it needs its conditions stated as carefully as the claim it retires. This one was published on a mechanism that did not apply to the measurement it overturned. This page has made that mistake before and recorded it before (§06's 81.7 t/s case study), which is the reason it is recorded again rather than quietly fixed.
And the replication that mattered was not the one that had already been done. The suspect number had passed every internal check, twice: five settled probes, a tight spread, a repeat load. None of that could help, because replicating within a condition cannot test whether the condition itself is anomalous. Only a neighbouring window could interrogate it, and only re-running both windows across independent loads settled it.
Which file for which card — and, on 24 GB, for which window
The three parts above cover quality, memory and speed at ordinary windows. Two further measurements finish the picture, and both of them changed an answer this page had already published. The first is that the whole ladder was scored with the drafter off — and the drafter is on in every recipe this page ships. The second is what happens when you actually fill the model's full 262,144-token window instead of arguing about it from arithmetic.
“Shipped” is two settings, not one. §03's launcher carries both: n-max 4 / p-min 0.75, the conservative pair, on four of its picks, and n-max 10 / p-min 0.5, the faster pair, on one. The table below was measured at the faster pair. Earlier editions called each of them “the shipped” setting in different places SUPERSEDED, which is ambiguous and is now stated once, here. The ranking inverts when you turn the drafter on. With no drafter the smaller file is faster, which is the obvious result and, for anyone following this page's recipes, the wrong one. Measured on 700-token novel-code generations at -c 32768, q8_0 KV, thinking off, temperature 0, three settled probes with the first post-prefill probe discarded:
| File | GiB | Drafter off | Drafter on (n-max 10 / p-min 0.5 — the faster of the two shipped settings) | Speculation is worth | Acceptance | Mean draft length |
|---|---|---|---|---|---|---|
| UD-IQ4_XS | 13.274 | 42.34 t/s | 86.91 t/s | 2.05× | 0.611 | 5.70 |
| UD-Q2_K_XL | 9.154 | 45.66 t/s | 77.01 t/s | 1.69× | 0.551 | 5.08 |
Drafter off, the 2.9-bit file is 7.8% faster, exactly as fewer gigabytes per token predicts. Drafter on — which is how every recipe in §03 runs — the 4-bit file is 12.9% faster despite being 45% larger. The mechanism is the draft head: it is quantized along with everything else, so it degrades with bit width. Acceptance falls 0.611 → 0.551 and mean draft length falls 5.70 → 5.08, a 10.9% drop against an 11.4% drop in throughput — draft length predicting throughput again, to within half a point, exactly as §06's measurement says it does. What the bigger file keeps in speculation — 2.05× against 1.69× — is worth more than everything the smaller file wins by being smaller.
Replicated 2026-08-25, after a reader challenged it. meas The argument against this table is a good one: UD-Q2_K_XL reads 31 % fewer bytes per token, so bandwidth alone says it must win. It was therefore re-measured from scratch by a second implementation in a different language, against the original script’s conditions taken verbatim — same prompt, same flags, same probe protocol. All four arms landed inside 1.3 % (86.91 → 86.58 and 77.01 → 76.97), and both acceptance figures reproduced exactly to three decimals (0.611 and 0.551). The bandwidth argument is right about bytes and still loses: what speculation returns to the less-damaged file is worth more than what the smaller file saves by being small.
And a condition this page was under-stating. The first two replication attempts read 73.71 t/s for the 4-bit file instead of 86.9. An earlier version of this paragraph blamed -fa on together with --reasoning off SUPERSEDED — that was wrong, and measuring it (§05) showed why: -fa defaults to auto, auto resolves to on, and those runs already had Flash Attention. The whole gap is --reasoning off, which changes what the model emits and therefore what the drafter has to guess — acceptance moved 0.523 → 0.611 with it. Every speed figure on this page is measured with reasoning off; benchmark your own card with reasoning on and you will read roughly 15 % lower, and the gap will look like hardware when it is a regime difference.
Both files were probed at -c 32768, so this pair ranks them at shallow context — which is the range the recommendation table's first row covers, and no further. That row's boundary is now about 98,000 tokens, and it comes from the depth series above rather than from this pair: at -c 131072 a third file overtakes both of these. Do not confuse it with UD-IQ4_XS's 180,224-token text-only ceiling (below), which is a statement about how much window that file can hold with a drafter aboard, not about which file is fastest inside it.
The general lesson is worth more than this pair: sweep at the recipe you ship, not at a clean-room default. The ladder was run drafter-off for clean determinism and, for the configuration anyone actually runs, it produced the wrong ordering. Carrying the shipped flags from the first arm would have cost nothing and would have had the right answer in hand from the start.
Then the second measurement, at the other end of the window. All three arms below were deep-filled to 218,233 real tokens — 83% of the full 262,144-token window, not a shallow probe — on the card's 24,576 MiB of board VRAM, with the first post-prefill probe discarded. Each is judged against the desktop reserve defined in §02, and it is worth restating which half of it is measured and which is a choice: the 1,181 MiB worst case of the desktop's own share is measured, the 127 MiB of load-to-load variation is measured, and the 1,796 MiB reserve built by adding them is this page's own derived threshold — the amount of slack it wants a configuration to keep before calling it usable with a screen attached, not an amount the operating system demands.
Arm at -c 262144, filled to 218,233 tokens | Board VRAM at depth | Slack | Decode at depth | Verdict against the 1,796 MiB desktop reserve |
|---|---|---|---|---|
| UD-Q2_K_XL, drafter on (n-max 4 / p-min 0.75) | 22,859 MiB | 1,717 MiB | 21.33 t/s | This verdict flipped on 2026-08-25. Keeps 1,717 MiB, published as 409 MiB more than the reserve SUPERSEDED back when that reserve was 1,308. Against the corrected 1,796 MiB it is 79 MiB SHORT: this configuration no longer leaves room for a desktop at its measured worst. Run it headless, or turn the drafter off — that drops the requirement to 20,567 MiB and leaves 4,009, which clears comfortably. Tight, and published as tight — but the whole native window is resident and speculating |
| UD-Q2_K_XL, drafter off | 20,567 MiB | 4,009 MiB | 17.96 t/s | Keeps 4,009 MiB — comfortably more than the reserve. The drafter is still worth +18.8% here even though mean draft length has collapsed from 5.70 to 2.43 at this depth |
| UD-IQ4_XS, drafter off | 23,821 MiB | 755 MiB | 15.96 t/s | Keeps only 755 MiB — 553 MiB short of the reserve. It will run with no graphical session using the card (§02), and there is no room left to turn speculation on at all |
At the full window the 2.9-bit file beats the 4-bit file on every axis at once: 34% faster at depth, 962 MiB more slack, and it can run speculation at all, which the 4-bit file cannot. That last point is the one that decides it — UD-IQ4_XS reaches 23,821 MiB with the drafter already off, so there is nothing to trade. This measures what earlier editions of this page asserted from arithmetic alone, and the number now says it: full native context on the 4-bit file needs the card free of any graphical session.
Two honest costs ship with that recommendation. Filling 218,233 tokens took 424 seconds at 514.7 t/s of prefill — seven minutes before the first token of the answer. This is a long-document configuration, not an interactive one. And the two-constant VRAM model in §05 under-predicts this window by about 1,213 MiB — it predicts 19,353 and 21,647 MiB against measured 20,567 and 22,859 — because it does not count the compute buffers — the scratch memory the server needs while it is actually working, on top of the weights and the cache. §05 already says its arithmetic is a floor rather than a budget; this measures how big the gap between floor and reality gets at the top of the window, and it is big enough to matter. A reader budgeting from the formula alone would think the drafter-on arm kept 2.9 GiB when it keeps 1.7.
Putting all of it together — the ladder, the empty-answer audit, the requirement table, the depth series, the drafter pair and the full-window trial — gives this. Every "fits" below means fits and still leaves this page's 1,796 MiB desktop reserve free, which is a stricter test than fitting on the card:
| Your card, and what you want from it | Take | Because | Tier |
|---|---|---|---|
| 24 GB · everyday windows, up to about 98k | UD-IQ4_XS — stay where you are | On novel code into a nearly empty window it is 12.9% faster than the 2.9-bit file with the drafter on, and that beats the size advantage outright. On prose filling about 90% of the window the picture is not an ordering at all: across -c 32768, 65536 and 98304 the three candidate files sit inside a band of 3.6 t/s or less, and this page does not rank files inside a band that narrow (§08). Either way there is no speed argument for moving, and of the three the 4-bit file has the lowest measured perplexity. Nothing about this recommendation changed | measured |
| 24 GB · a window of about 131k | UD-Q3_K_XL — switch | 20% faster at that window — 41.74 t/s against UD-IQ4_XS's 34.68 — and this is the one window where the separation is wide and reproduced: each file was loaded twice from scratch and every one landed within 1.4% of its own first load (§08). It needs 19,724 MiB with the drafter on, leaving 4,852 MiB free where the 4-bit file leaves 3,728 der, and it ties the reference file on accuracy — the two answered exactly one of 75 questions differently, p=1.00. The condition travels with the recommendation: the ordering was measured on prose fill, and the drafter's confidence gate is a property of your content. Spend ten minutes checking it on your own material before you commit | measured on prose fill |
| 24 GB · you need close to the full 262,144 window | UD-Q2_K_XL — switch | 34% faster at depth, keeps speculation, keeps 409 MiB more slack than the 1,796 MiB this page reserves for a desktop, and ties the 4-bit file on accuracy (the two answered exactly one of 75 questions differently, p=1.00). The 4-bit file cannot hold this window with a desktop running at any speed. Price: a seven-minute prefill | measured |
| 16 GB | UD-Q2_K_XL at -c 65536, drafter on | This answer is now settled by arithmetic on measured requirements rather than by judgement. A 16 GB card holds 16,384 MiB; keep the 1,796 MiB reserve free and 14,588 MiB is what you have to spend der. UD-Q2_K_XL at -c 65536 with the drafter needs 13,982 MiB — it fits, with 606 MiB spare der. UD-Q3_K_XL does not fit at that window at all: 16,906 MiB with the drafter, more than the whole card. The largest configuration of it that fits inside 14,588 is -c 32768 with the drafter off, at 14,322 MiB. So the smaller file gives you twice the window, a drafter you can keep, and 2,924 MiB less memory at the same window and drafter setting (13,982 against 16,906). It also ties the reference file on accuracy, p=1.00, for +6.07% of perplexity. On this 3090 those two configurations decode at 40.30 and 36.44 t/s — but that is this card's bandwidth, not a 16 GB card's, and no measurement here can tell you what yours would do. §03's 16 GB recipe still ships UD-Q3_K_XL at a smaller window and should be read beside this row | VRAM requirement measured · the subtraction derived · speed on a 16 GB card not measured |
| 12 GB | No. Use a smaller model | Not marginal — over the line, and the arithmetic is short. A 12 GB card holds 12,288 MiB; keep the reserve and 10,980 MiB is what you have der. The smallest configuration measured anywhere on this page — UD-Q2_K_XL at -c 32768 with the drafter off — needs 11,396 MiB. It is 416 MiB over der. There is no smaller file to fall back to either: everything below 2.912 bits per weight is on the wrong side of the empty-answer boundary (§08). Nothing below -c 32768 was measured, and a window smaller than that leaves this model very little room to think (§09). The answer on this card is a smaller model, not a smaller quantization of this one. §03's existing 12 GB recipe — Q4_K_M with most layers on the processor, 6–8 t/s — remains the only way this page can defend running this model there, and it is slow on purpose | VRAM requirement measured · the subtraction derived |
| 8 GB | No | The smallest file that still performs like the full-quality one is 9.154 GiB. It does not fit, and the files that do fit are on the wrong side of the boundary | derived |
| Anything below 2.48 bits per weight | Not recommended, on any card | Not because the score looks worse — because the model starts failing to answer. The empty count is exactly zero at every rung down to 2.912 and then rises 2, 3, 5, 28 (the empty-answer table). Empty replies first appear at the 2.481-bit rung itself, which is why this page recommends nothing below 2.912 either. Many are silent: of those four counts, 1, 2, 5 and 10 ended normally and returned zero characters, so nothing in your logs reports a problem at all. And at 1.835 bits the model loses the ability to stop — 20 of 75 answers ran into the cap, and the median answer jumps from a 416–502-token band held all the way down the ladder to 932 tokens | measured |
The boundary of what this machine can say, and it is sharper than earlier editions drew it. This machine owns one 24 GB card and has never loaded any of these files on anything else, so the 16 and 12 GB rows have to be split into three kinds of statement rather than dismissed as one. Quality is measured and it transfers — a file's perplexity and its benchmark answers do not depend on which card decodes them, so the accuracy tie between UD-Q2_K_XL and the 4-bit reference file is a measurement a 16 GB owner can rely on directly. The VRAM requirement is also measured, and it also transfers, because how many bytes a configuration asks for is a property of the model, the window and the flags rather than of the card: the 13,982 and 16,906 MiB figures above were read off this 3090 and are what those same configurations will ask for on any card (§08). The arithmetic that turns a requirement into a verdict is derived and its steps are printed in the cells: card total, minus the reserve, against the requirement. The speed is not measured, and on a card this machine does not own it never can be — every t/s in this section is this 3090's memory bandwidth. How fast UD-Q2_K_XL runs on a 5080 or a 4080 is not known here at all, and §15 lists it as a gap rather than filling it with an estimate.
UD-IQ4_XS against Q4_K_M — the two files this page recommends
These two files are the whole decision on a 24 GB card, and the honest summary is that they are tied on quality and not tied on anything else. At 13.3 GiB, UD-IQ4_XS is 2.1 GiB smaller than Q4_K_M, matches it on perplexity to within the measurement's own error bar, and — because decode is limited by memory bandwidth — is faster everywhere it was measured: prefill 1365 against 1303 t/s, plain decode 42.97 against 39.99 with no drafter, the same maths benchmark with speculation 73 against 63, the matched code sweep 93.9 against 81.7 — the first and last pairs measured on the same prompt in the same sweep. It also costs about 8% less energy per token. The saved memory becomes context or vision.
The measured ceilings for the smaller file, each row naming all four determinants (file · drafter flags · projector · desktop), and each one lower than earlier editions of this page published, because the re-measurement filled the windows instead of probing them shallowly: with vision and the n-max 4 drafter, 163,840 is fully resident (996 MiB slack); with vision and the n-max 10 drafter the same window keeps only 537 MiB, which is less than a busy desktop holds, so 131,072 is the largest window that still leaves a desktop comfortable (2,426 MiB slack) and 122,880 is what §03 ships. Text-only with the n-max 10 drafter the ceiling is 180,224 (847 MiB slack) — not the ~196k the projector arithmetic implied: 196,608 leaves 415 MiB and 212,992 leaves 280, with decode already sagging 56.6 → 51.9 → 49.2 across those three on a short probe. Do not try to read the budget model out of those three slack figures: the board VRAM total is saturating, so the growth goes somewhere the board VRAM figure cannot show — the server's shared memory use across the same three windows is 478, 638 and 744 MiB, which is the spill beginning before the board VRAM slack runs out. The full native 262,144 fits only text-only and drafter-off, at 23,216 MiB with 1,360 MiB of board VRAM left — and that figure was read at load rather than at depth: filled to 218,233 real tokens on 2026-08-25 the same configuration reaches 23,821 MiB and keeps only 755, which is 1,041 MiB short of this page's 1,796 MiB desktop reserve, so the window now needs the card free of any graphical session by measurement rather than only by arithmetic (§08). Earlier editions said shallow probes "stayed at 65–70 t/s all the way to the full native 262,144 with no collapse anywhere" — that is precisely the probe artefact §05 now documents: fill that same configuration to 91k of real tokens and it delivers 8.0 t/s.
So why does Q4_K_M keep being named a default anywhere? Because there are two defaults with two scopes. Q4_K_M is this page's measurement base and its maximum-quality pick — perplexity leans slightly its way, it is the file most of the 08-21 and 08-22 numbers were taken on, and IQ-format decode support varies more across backends (everything here was measured on CUDA). UD-IQ4_XS with vision at -c 122880 is the reference machine's shipped daily default (§03, menu row 1). Short version: context-hungry or speed-hungry, take UD-IQ4_XS; maximum measured quality at 122k with little on screen, stay on Q4_K_M. Either is defensible, and the caveat on the Q4_K_M recipe's quality argument above applies to this paragraph too.
The limit of everything above: none of it tested a long agentic loop
Every number on this page comes from single-turn prompts of at most 16,384 tokens. You ask once, the model answers once, and the answer is scored. That is not how a coding agent works. An agent calls a tool, reads the result, calls another, and keeps going for dozens of turns — and a model that goes wrong once in fifty turns will fail an agent while scoring perfectly here. So the smaller files on this page are recommended for what was measured, and are untested for agentic work.
A video published as "Qwen 3.8 27B Quantizations Q1 - Q8 compared" (channel reported as Luke's Dev Lab; the title was confirmed but the video was not watched from this machine, so everything below is cited from a viewer's written breakdown rather than measured here) ran the same unsloth files at high reasoning effort against three multi-turn agentic tasks: a Kanban web app, a Blender 3D asset job over MCP, and a Godot 4 game.
Where it agrees with this page. It found UD-Q2_K_XL one-shotting a full web app with no intervention and recommends it for 16 GB cards — which is this page's 16 GB pick, reached independently from perplexity, paired accuracy and a memory requirement. It found the 4-bit class the overall sweet spot, which is what this page ships. And it found IQ1_M unusable on every task, failing by locking into infinite loops — the same failure this page detected from short prompts at no GPU cost, as five empty answers out of 75, every one of them terminating normally and returning nothing (§08). That agreement matters in a specific way: this page's accuracy score for that file was 85.30, which looks survivable. The empty-answer count was the number that saw the problem.
Where it finds something this page cannot see. It reports UD-Q2_K_XL and UD-Q3_K_XL both failing the Godot task — one locking into a tool-calling loop immediately, the other generating crashing scripts until the project was unusable. Both of those files return zero empty answers here and tie the 4-bit reference on paired accuracy. There is no contradiction: those are long multi-turn runs, and nothing on this page tested one. Error compounding across turns is precisely the failure mode a single-turn benchmark is blind to.
What to take from it. If you are running chat, single-file coding, or web work, the recommendations above stand and now have independent support. If you are running an agent, treat every file below 4 bits on this page as untested — try it on your own workflow before trusting it, and watch specifically for a loop that never terminates rather than for wrong answers. That tester also went above 4 bits, to Q5, Q6 and Q8, and reports those needed for the hardest visual and game-logic work; this page measured nothing above 4.4 bits per weight and has no opinion there at all.
09
Reasoning effort: what each level costs and buys
This model has a dial that decides how long it thinks before it answers, and the honest finding is that the dial moves hours, not scores. On a 175-prompt benchmark suite run on this machine — all seven sets now scored, the last two by the blind judge panel added on 2026-08-24 (§09) — the three levels scored 80.2, 80.3 and 79.7 out of 100, a spread of 0.6 points at a sample size where a single benchmark cell is worth about ±16 points, which is a tie — while the wall clock went from 1.0 to 1.5 to 2.7 hours and the average answer grew from 830 to 2,217 tokens. So: use medium when you are waiting at the keyboard, and xhigh when you are not. This page ships xhigh as the default in every recipe whose window can hold it and whose speed makes the wait sane — the 12 GB recipe's window holds it and still ships medium, because at 6–8 t/s an answer takes three hours — for one reason that the benchmark suite cannot measure and the smaller judged tests could: on complex single-file tasks, xhigh was the only level that delivered 100% of the specification with no fatal defect, twice out of two. Buy it as insurance against an incomplete answer, not as a higher score. And not for prose: the two open-ended sets, once judged, are the only place in this whole campaign where the levels separate at all, and they separate against xhigh — it lost to medium on 14 of 25 MT-Bench prompts and won on 3 (§09). For writing, medium.
The strongest evidence: a 175-prompt suite, three levels, a tie
Measured 2026-08-23 on the reference 3090: five scored benchmark sets at 25 questions each, per effort level, greedy decoding, seed 42, no drafter (to keep the scores clean), 525 generations in 8.48 hours of GPU time. Two further sets — ALPACA and MT-Bench — needed a judge, and for a long time this page had none, so it published a five-of-seven index and said so. They are now scored (2026-08-24), by a blind three-seat panel described in §09 below. Both indices are kept: the five-set one because every earlier comparison on this page is against it, and the seven-set one because that is what the benchmark protocol actually specifies. They are comparable only with another run using the same sets.
| Benchmark (n=25 each) | Scorer | low | medium | xhigh |
|---|---|---|---|---|
| GSM8K | exact match | 100 | 100 | 100 |
| MATH-500 | exact match | 96 | 100 | 100 |
| HumanEval | execution pass@1 | 100 | 96 | 92 (2 trunc) |
| MBPP | execution pass@1 | 92 | 84 | 92 (1 trunc) |
| MeetingBank | ROUGE-L F1 | 22.6 | 22.4 | 22.3 |
| ALPACA | judge 1–10 (3-seat panel) | 70.2 | 74.1 | 72.1 (1 empty) |
| MT-Bench (turn 1) | judge 1–10 (3-seat panel) | 80.7 | 85.3 | 79.7 |
| Composite index over the 5 mechanically scored sets | mean | 82.1 | 80.5 | 81.3 |
| Composite index over all 7 sets (what rule 21 actually specifies) | mean | 80.2 | 80.3 | 79.7 |
Conditions: UD-IQ4_XS, llama.cpp build 10502, q8_0 KV, -ngl 99, --parallel 1, temperature 0 / top-k 1 / seed 42, no speculative decoding. This table is a merge of two caps, and it matters which cell came from which. Six cells were re-run at --max-tokens 32768 with -c 65536 — those are the ones that had truncated: MATH-500 at low and at xhigh, and HumanEval and MBPP at medium and at xhigh. The other nine scored cells are the original run at --max-tokens 16384 with -c 32768, carried across unchanged. What licenses carrying them is not an argument but the byte-comparison in §10: 139 of 139 untruncated prompts reproduced identically when the cap and the window were both raised. Decode speed, and which instrument read it: the benchmark harness reports 42.2 / 42.0 / 41.9 t/s as an unweighted mean over requests, while the server's own token-weighted timings read 38.4–40.2 t/s for the same arms. The harness figure is pulled up by short, shallow answers. The two are never averaged together on this page (§15). The three scorer names in that table are defined in §02: execution pass@1 is the share of problems whose first generated program runs and passes its tests, and ROUGE-L F1 measures how much of a reference summary's wording a generated summary reproduces. MeetingBank's ~22 is a property of the metric pairing rather than a failure: the model writes the 60–120-word summary the prompt asks for — measured across all 75 MeetingBank generations at 79–129 words, mean 105 — while the reference is a ~40-word resolution title, so ROUGE-L F1 is structurally capped. It is flat across the three arms and contributes no signal. Note also that mixing a ROUGE score into a mean of percentages is what makes this an index rather than an accuracy.
Every individual cell above is 25 questions, so one flipped question moves that cell by 4 points and a single cell carries roughly ±16 points of confidence interval. The composite pools about 125 scored samples per arm, which is why it is the interpretable number — and it still cannot resolve a 1.6-point difference. Two of the five sets are also saturated: GSM8K returned 100 for all three levels, which means it discriminates nothing at any sample size, and MATH-500 returned 100 for two of them. The correct reading is that this suite could not detect a quality difference between the three effort levels, not that there is none. §10 puts numbers on how many questions detecting one would take.
The two sets a machine cannot score, and what a judge found in them
Five of the seven benchmark sets can be marked by a computer: there is a right answer, or the code runs or it does not. The other two — ALPACA and MT-Bench — ask for open-ended writing, and the only way to score writing is to have something read it. For most of this campaign nothing suitable was available, so this page did the honest thing and published a five-of-seven index while keeping the transcripts. On 2026-08-24 those transcripts were finally read.
The reader was a panel of three independent Claude Opus 5 judge seats, working blind: every answer carried an opaque identifier, the mapping from identifier to effort level was sealed in a file no seat could open, and each seat received the answers in its own shuffled order so that ordering effects could not line up across seats. All three seats rated all 150 answers, on the standard 1-to-10 single-answer rubric, giving 450 ratings with none missing. Ratings convert to the 0–100 scale as (rating − 1) ÷ 9 × 100.
The answers were written by Qwen and read by Claude, so nothing here is a model grading its own work — that was the rule this page was waiting on, and it is satisfied. But this is not independence in the strongest sense: the judge and the author of this page are both Claude models, which makes them a correlated instrument. That is why the panel is three seats rather than one, why the spread between seats is published beside every mean, and why the conclusions below lean on the paired test rather than on the raw scores. Seat-to-seat spread averaged 0.28–0.92 rating points out of ten, so the seats largely agreed; agreement is not the same as being right.
| Set (n=25 each) | Measure | low | medium | xhigh |
|---|---|---|---|---|
| ALPACA | mean rating, 1–10 | 7.32 | 7.67 | 7.49 |
| score, 0–100 | 70.2 | 74.1 | 72.1 | |
| MT-Bench (turn 1) | mean rating, 1–10 | 8.27 | 8.68 | 8.17 |
| score, 0–100 | 80.7 | 85.3 | 79.7 | |
| Mean spread between the three seats (rating points) | 0.28–0.92 | 0.56–0.64 | 0.44–0.56 | |
Conditions. MT-Bench's generations are the 16,384-cap ones described above; it never truncated. ALPACA's xhigh answers come from the raised-cap re-run — one of its 25 answers spent its whole budget inside a reasoning block and returned nothing, so the budget rule required a re-run at 32,768. The re-run reproduced it exactly: the same answer consumed all 32,768 tokens and again returned nothing, while the other 24 came back token-for-token identical. That is not a budget shortfall, it is a generation that does not terminate, so the remedy is exhausted and the score is final, carrying one disclosed non-terminating answer that all three seats rated 1. Because the re-run replaced one arm's answers, all three ALPACA arms were re-judged together, so the three numbers come from one session rather than two. Judging protocol, packet builder and scorer: scripts/bench/judge-panel.py; every rating and the sealed key are kept with the run data.
The finding, and it is the only thing in this campaign that separates the effort levels at all. Because the same 25 prompts were put to every level, the arms can be compared prompt-by-prompt rather than only mean-against-mean, which is a far more sensitive test. Doing that (a bootstrap: the same set of per-prompt differences is re-scored 20,000 times with the prompts drawn at random and with replacement, and the range the answer falls in 95% of the time is reported): on MT-Bench, medium beats xhigh by 0.51 rating points, interval +0.21 to +0.80, and the count is lopsided — xhigh wrote the better answer on 3 prompts, the worse on 14, and tied on 8. All five other pairings are ties, ALPACA included.
The first judging pass put medium ahead of low on ALPACA by 0.44 rating points, interval +0.013 to +0.907. It cleared zero by thirteen thousandths, and this page called it marginal rather than a finding. The re-judge settled it: the same comparison is now +0.35, interval −0.04 to +0.73 — a tie. The marginal result was noise, and saying so in advance is the only reason it did not become a claim.
The re-judge also measured something this page had no number for: how repeatable the judge itself is. Seventy-four of the seventy-five answers were byte-identical between the two passes, so any change in their ratings is the instrument, not the model. Across those 74: mean absolute change 0.33 rating points, largest change 1.33, and 25 answers scored exactly the same both times. Averaged over 25 items an arm's score moves by at most 1.0 point on the 0–100 scale — which is why a 0.51-point paired gap on MT-Bench is a result and a 0.44-point one on ALPACA was not. Measured on ALPACA; MT-Bench was not re-judged, so applying this band to it is an assumption, stated as one.
xhighThis does not overturn the recommendation in §09; it sharpens it. The case for xhigh was never a higher score — it was completeness on complex code tasks, where it was the only level to deliver 100% of the specification with no fatal defect, twice out of two. That finding stands and is categorical. What the judged pair adds is the other half: on MT-Bench, more thinking made the answers measurably worse, not better — while on ALPACA the same pairing is a tie. Longer deliberation produced more places to go wrong: invented specifics, padded lists, and in the worst cases output that stopped being an answer at all. One of six comparisons survived a 95% test, so this is a set-specific result, and overthinking is a candidate explanation for it rather than an established one. So: xhigh for hard code you need finished, medium for writing. Five of the nine recipes ship medium and four ship xhigh (§03), the 24 GB default among them — earlier editions said “the recipes ship medium” SUPERSEDED. And you cannot switch level per request: llama-server ignores a reasoning_effort field in the request body, which is why every recipe sets it at launch. Changing level means relaunching with a different --chat-template-kwargs, and raising max_tokens to about 120,000 for xhigh, which fits one turn per 131,072-token window. An earlier version of this sentence told you to override per request SUPERSEDED: a reader who did got medium while believing they had xhigh, with nothing to show it.
What the seats caught that no mechanical scorer could. The value of a judge is not only the score; it is the failure modes a right-answer test is blind to. Three kinds turned up, and all three matter more than the point differences above.
Confident invention. The travel-writing prompts drew fluent, well-structured answers containing things that are simply not true — a “Hawaiian $2 bill”, a claim that the first surfers ever to ride a wave did so at Sunset Beach, a misdated battle, a hula centre that does not exist, and mistranslated Hawaiian words. A Chicago-transit prompt produced invented Metra line names, expressways matched to the wrong interstate numbers, and a count of “14 L lines”. A finance prompt turned a 1-for-4 bonus issue on 1,000 shares into 1,500 shares. None of this is detectable by perplexity, by an exact-match scorer, or by running code.
Degeneration well short of any limit. One low MT-Bench answer, asked for a short story, began spelling out numbers and never stopped — “one hundred and one, one hundred and two…” — and ended without a story. It used 1,682 tokens, nowhere near a cap, so nothing in the harness flagged it; all three seats rated it 1. One xhigh ALPACA answer expanded into 100 near-duplicate list entries that collapsed into nonsense, and another repeated the same two words three times each to pad a list to the requested length.
A failure that belongs to the model, not the setting. On one ALPACA prompt demanding a full sentence, low and medium independently returned the same bare noun phrase, and every seat rated both a 3. When two settings fail a prompt the same way, the effort dial is not the variable.
| Per arm (full 7-set run at the 16,384 cap) | low | medium | xhigh |
|---|---|---|---|
| Wall clock | 1.00 h | 1.47 h | 2.70 h |
| Mean output tokens (all 7 sets) | 830 | 1,228 | 2,217 |
| Decode, mean (harness, unweighted over requests — the server's token-weighted figure for the same arms is 38.4–40.2) | 42.2 t/s | 42.0 t/s | 41.9 t/s |
| Truncated at the 16,384 cap (scored sets) | 1 | 2 | 8 |
xhigh costs 1.84× medium's wall clock on this suite and produces 1.81× the tokens — the same ratio, because decode speed is flat across the three levels. That is a different multiplier from the one earlier editions of this page printed: on a single long authoring task, xhigh cost about 4× medium's wall clock. Both are measured and neither is wrong; the multiplier is task-shaped, not a constant, and it grows with how much room the task gives the model to keep thinking.
At the 16,384-token cap the same three arms scored 81.3 / 80.5 / 77.3, which reads as quality falling as effort rises — an attractive, quotable, wrong headline. The entire xhigh penalty was the cap: it truncated 8 times in the scored sets against 1 and 2 for the other levels, and by the grading rule a truncated answer scores zero. Raise the cap to 32,768 and re-run only the affected sets and the ranking becomes the 82.1 / 80.5 / 81.3 above. One prompt needed 18,273 tokens and was correct; another needed 17,025 and was correct; both had scored zero at 16,384. Three xhigh prompts still exceed 32,768 and are genuine runaway loops rather than long solutions, and they are counted in the table.
The other evidence: judged deliverables, at n=2
The suite above measures whether the model gets short, self-contained questions right. It does not measure whether it produces a complete working deliverable, and that is where the effort dial visibly earns its keep. Measured 2026-08-22: the same ~1,700-token coding task — build an animated aquarium page — run twice per level at temperature 1.0, with the six resulting pages graded blind against the specification by independent reviewers.
| Effort | Tokens (2 runs) | Wall clock | Blind quality score /100 n=2 | GSM8K n=200 |
|---|---|---|---|---|
low | 16.7k–29.8k | 4:34–8:34 | 40, 47 — both runs produced pages whose JavaScript crashes (one never draws a frame, one freezes on its first) | 96.0% (0 truncated) unaudited grader |
medium | 20.7k–23.4k | 5:54–7:04 | 86, 93 | 97.5% (0 truncated) unaudited grader |
xhigh | 73.1k–75.8k | 24:58–26:48 | 92, 95 · 100% of the specification, both runs | 95.0% (94.0 under a 4,096 cap†) unaudited grader |
† the first run's 4,096-token cap cut 5 of xhigh's 200 answers mid-thinking; re-run at a 16,384 cap it scored 95.0% with zero truncations, so the budget had cost about one point. Only the xhigh arm needed re-running, because greedy decoding is deterministic and the never-truncating arms are byte-identical under any larger cap — a licence that is no longer an argument but a measurement, 139 of 139 prompts reproduced byte-for-byte (§10). One scope warning that earlier editions of this page did not attach loudly enough: "16,384 proved sufficient" was true on GSM8K. On the seven-benchmark suite above, the same 16,384 cap truncated xhigh eight times in the scored sets, and three prompts exceeded even 32,768. Size your cap against your hardest set, not against GSM8K. The GSM8K percentages in this table also carry the grader caveat from §08 and should be re-graded before they are quoted again.
The findings from the judged runs, stated plainly. The dial changes the answer sharply at one end and hardly at all at the other, and neither is where you would guess. The sharp change is between low and medium, and it moves with task difficulty: low matched the others on short maths, then shipped fatally broken JavaScript on the complex page twice out of two, with one broken run burning more tokens than medium. That claim needs its evidence base attached, and earlier editions of this page did not attach it — it is two runs at temperature 1.0, on one task. The independent re-measurement (2026-08-23, same machine, UD-IQ4_XS) ran the same task once per level and got the opposite defect direction: its low page rendered correctly while its medium page shipped a real initialisation bug — the resize routine ran after every creature array was built, so every mobile creature spawned in the top-left corner, violating the specification's "distribute them across the tank". Three runs across two campaigns, both directions observed. What is solid is that a fatal defect is a live risk at every effort level at n=1 — not that low specifically manufactures them. Above that point the change is small: medium and xhigh overlap on score, and medium's own run-to-run spread beats the between-level gap. Only xhigh delivered 100% of the specification with zero fatal defects on both runs, and it won both blind head-to-head rankings. Its edge is completeness insurance, not average score — and the 175-prompt suite above says the same thing from the other side: on short scored questions, there is no score to win.
Both bodies of evidence above are built from greenfield tasks: write this page from scratch, solve this self-contained problem, implement this function. Exactly one read-then-edit task against existing code was ever run on this machine, and it failed. An agentic pipeline was validated end to end on one real sandboxed repair task in 2026-08-23: 22 model calls, 9 minutes 36 seconds of wall clock, and 0 of 96 target tests turned green, though all 561 previously-passing tests kept passing — so the harness worked and the fix did not. That is one sample of plumbing, not a quality measurement, and no scored read-then-edit suite was run at any effort level. Open a file, understand it, change one thing without breaking the rest is the shape of work an agent actually does, and it is the use this page recommends the model for in §13 and §14. So the quality verdict here is scoped to greenfield single-file work, and the measurement that would close the gap is named in §15's register of what was not measured.
The independent re-measurement read the model's chat template and then proved the knob reaches it. Four things worth knowing before you set it: only low, medium and xhigh are accepted (anything else raises an error) — and high is silently aliased to xhigh, so a client sending the familiar "high" gets the most expensive setting, not a middle one. Only low and xhigh inject a reasoning-instruction system line; medium injects nothing at all — it is the template's unmodified default path, not a third instruction. The proof that the setting is not being swallowed is the prompt-token differential: identical task, prompt_n of 1,689 / 1,659 / 1,701 for low / medium / xhigh — medium's is the short one, exactly as the template predicts. And separately from effort, --chat-template-kwargs "{\"enable_thinking\":false}", which works per request and not only at load, is the switch between the two token regimes §06 keeps separating.
Your window sets an effort ceiling
Effort has a second price besides time: thinking tokens live inside the context window (§02), so a window that cannot hold a level's thinking cannot offer that level at all. Measured appetite on one hard task: xhigh's appetite is a distribution, not a number. Earlier editions quoted "73–76k" from this page's two completed runs; pooling those with the re-measurement's two samples of the same task gives 61,500, 73,100, 75,800 and one run that hit a 65,536-token cap and returned nothing usable. Read it as 61,500–75,800 straddling 65,536 — which is the operationally important part: the re-measurement's re-run at a 120,000-token cap then wanted only 61,476 tokens, fewer than the cap the previous sample had blown through. A cap sitting near the middle of that spread looks generous and is a truncation machine. Size caps and windows against the distribution's upper tail, never its middle.
Your -c window | Levels offered | Level not offered, and why |
|---|---|---|
| ≥ 100k (24 GB+, DGX Spark, Arc Pro B70, 12 GB offload) | low · medium · xhigh | none — the 61,500–75,800 appetite distribution fits with room for a prompt and an answer, including its upper tail |
| < 76k (16 GB cards and Arc Pro B50, now at 65,536) | low · medium | xhigh not offered — a WINDOW limit. Its measured appetite alone can overflow the window mid-thought, and a truncated xhigh run scores worse than a completed medium one and often returns nothing at all. The 2026-08-25 switch to UD-Q2_K_XL raised these cards from 49,152 to 65,536, which reaches into the 61,500–75,800 appetite band rather than clearing it — so xhigh would now fit on its shortest runs and truncate on its longest, and you would not know which you got. Still not offered, for a slightly better reason than before |
| 12 GB offload at ~112k | low · medium shipped; xhigh for unattended runs | a WALL-CLOCK limit, not a window one. The window holds the appetite fine; at 6–8 t/s an xhigh answer takes about three hours against about one for medium. The fix is patience, not memory |
| integrated GPUs at 65,536 | low · medium | both limits at once, and they have different fixes. 61,500–75,800 of appetite against a 65,536-token window (fix: raise -c — an integrated GPU borrows system memory, so it can), and 3–4 hours per answer at 5–7 t/s (fix: nothing but patience) |
Set it server-side in llama.cpp with --chat-template-kwargs "{\"reasoning_effort\":\"xhigh\"}"; graphical front-ends usually expose it as a reasoning-effort setting. llama-server ignores a per-request reasoning_effort field, which is why every recipe in §03 sets it at launch and why the multi-agent setup in §14 needs two servers to route between two levels.
What each level costs in electricity
A local card's tokens are free but its electricity is not quite. Measured with a power log integrated over each generation window (2026-08-23, one run per level, with xhigh measured twice — once at a 64k cap that truncated and once at 120k that completed), an answer to the same task costs 20.6 Wh at low, 36.0 at medium and 104.8 at xhigh. Those are board joules only, so they are not a household electricity figure and must not be treated as one (§15). The number worth noticing is not the total but the rate: energy per token rises with effort, because the thinking stream decodes more slowly than the answer stream and the board draws the same watts either way.
| Effort level | Tier | Mean load W | J / decode token, gross | J / decode token, net of 34.1 W idle | J / prompt token | tokens / kWh | EDP (J·s) | Wh per complete answer, gross | Wh, net |
|---|---|---|---|---|---|---|---|---|---|
low | in-band NVML | 344.6 | 4.26 meas | 3.84 der | 0.166 meas | 844,778 der | 1.57e7 meas | 20.58 meas | 18.54 der |
medium | in-band NVML | 345.0 | 5.18 meas | 4.67 der | 0.120 meas | 694,661 der | 4.84e7 meas | 36.01 meas | 32.45 der |
xhigh · 64k cap — truncated, returned no file | in-band NVML | 338.7 | 6.60 meas | 5.94 der | 0.162 meas | 545,326 der | 5.52e8 meas | 120.25 meas | 108.15 der |
xhigh · 120k cap — completed, delivered the file | in-band NVML | 341.8 | 6.13 meas | 5.52 der | 0.178 meas | 586,830 der | 4.16e8 meas | 104.84 meas | 94.38 der |
| E_comm (interconnect energy) | N/A — single GPU, no interconnect | ||||||||
Conditions: RTX 3090, UD-IQ4_XS, -c 131072, no projector, drafter at n-max 4 / p-min 0.75 on, temperature 1.0 / top-p 0.95 / top-k 20, n=1 per level, one HTML-authoring task, mixed regime (thinking and answer tokens both timed and both counted), 1 Hz power sampling with no clock or power-state columns, idle not subtracted in the gross columns. Instrumentation tier: in-band GPU board power (NVML); the power supply, processor and the rest of the machine are excluded and unmeasured. All four gross Wh figures were independently re-integrated from the raw logs on 2026-08-23 and reproduced twice: the original rectangle-rule method returned the originally published 20.55 / 35.96 / 120.21 / 104.9 Wh to within 0.05%, and a trapezoid integration with interpolated edges — which is what the four values printed above are — agrees to within 0.15%. Either way they may be cited without hedging; the net columns subtract the dated 34.1 W loaded idle and are derived. The first-minute clock ramp makes each of these runs look 0.01–0.63% better than its settled rate, which is the right sign and a negligible size, so no ramp correction is applied — but it is a 17–20% artefact on 10-second probes, which is why short probes on this page are read differently (§01).
Three things hide inside that table. First, energy per token rises with effort — 4.26 → 5.18 → 6.13 J — but effort is not the cause. The mechanism is arithmetic: joules per token equals watts divided by tokens per second, and it predicts these four numbers as 4.24 / 5.16 / 6.60 / 6.13 against the measured 4.26 / 5.18 / 6.60 / 6.13, with no residual at all. The thinking stream simply decodes more slowly. The control that proves it: on the 25-question arms, where all three levels decoded at the same drafter-off speed, energy per token was flat to 0.3% (7.921 at xhigh against 7.947 at medium) while answers differed 3.1× in length. Effort changes joules per answer, not joules per token. Second, energy-delay product spans 1.57e7 to 5.52e8 — a factor of 35 across three levels, where energy alone spans 5.8×, because a slow setting is punished twice by that metric. Third, the 120 Wh row returned nothing. That xhigh run hit its 65,536-token cap mid-document: 5.8× low's energy for zero deliverable. The completed run beneath it cost less — 104.84 Wh — because a cap that fits lets the model stop when it is finished. Wall clock is still the bill that matters; truncation is the one that wastes it entirely.
Turning the drafter on cuts energy per token by roughly half, and it is free in quality: measured 7.884 ± 0.307 J per decode token with the drafter off across 115 requests, against 4.26–6.60 with it on. On the controlled comparison — one server, one prompt, only the flags changing — it is 8.104 → 3.210 J, a factor of 2.52 (§03). The drafter-off figure and the drafter-on figures come from different tasks with different sampling, so treat that pair as a reference comparison rather than a controlled experiment; the controlled version is the one in §03, and both point the same way. Anyone optimizing electricity here should reach for the drafter and for cached prefixes at depth — starting follow-up requests with exactly the same opening text, so the server does not read it twice (§02, §05) — long before touching the effort dial.
Which level to use for which kind of work
Within your window's ceiling, the trade is time. Pay it knowingly, and drop one level when you are sitting there waiting.
| Effort | Use for | Cost |
|---|---|---|
xhigh | the quality-first default this page ships — 100% specification coverage and zero fatal defects in the two judged runs that completed; a third xhigh run on the same task hit a 65,536-token cap and returned no file at all, which is the failure this level is most exposed to; worth it whenever a shipped bug or a missed requirement costs more than about twenty minutes of GPU time. Buy it as completeness insurance: on the 125 scored prompts it won nothing | 1.8× medium's wall clock across a benchmark suite, ~4× on one long authoring task. Raise max_tokens or thinking eats the budget |
medium | the value setting — most of the judged quality in a fraction of the wall clock, and statistically tied with the other two on the scored suite; the right choice for interactive sessions where the wait matters | minutes, not hours |
low / none | short, self-contained tasks. On the scored suite it posted the highest composite (82.1) and the fastest wall clock — though the suite cannot resolve a 1.6-point spread, so read that as "tied, and cheapest", not as a win. Either way it breaks any simple effort-quality ladder | not reliably cheaper on hard tasks, and the deliverable may not run: one campaign saw fatal defects at low in 2 of 2 runs while another saw the fatal defect at medium instead. At n=1 a fatal defect is a live risk at every level |
10
How many benchmark questions are enough?
Enough for what? That is the whole question, and it has an arithmetic answer. A 25-question benchmark can tell you a model file is broken and nothing finer. A 25-question benchmark can tell you a setting has not collapsed. Telling two healthy 4-bit files apart, where the real difference is one to three points, needs thousands of questions and days of machine time — which is why nobody ranks healthy files by accuracy scores, and why this page ranks them by perplexity instead. Before you believe any benchmark number, including the ones on this page, ask three things: how many questions, what was the token budget, and how many answers were cut off. The rest of this section is those three questions with numbers on them.
A worked example from this page's own data. A GSM8K score of 85% at 20 questions carries a 95% confidence interval of roughly 64–95% — a thirty-point window around a table where every score sits within ten points of every other. One flipped question moves that score 5 points. Healthy 4-bit files differ by perhaps 1–3 true points; inside that window, such a gap is invisible.
The arithmetic is unforgiving because sampling noise shrinks with the square root of the sample count: halve the gap you want to detect and you need four times the samples. What that costs on this 3090 — paired evaluation on the same questions, temperature 0, about 80% base accuracy, 95% confidence at 80% power, about 1,200 tokens per sample at about 55 t/s with the drafter on:
| True gap to detect | Samples needed | Wall clock per model | Board energy per model |
|---|---|---|---|
| ~20 pts — a broken file | ~25–30 der | ~9–11 min der | ~51–61 Wh der |
| ~10 pts | ~100–150 der | ~35–55 min der | ~0.20–0.31 kWh der |
| ~5 pts | ~300–500 der | ~2–3 h der | ~0.61–1.0 kWh der |
| 1–3 pts — typical 4-bit against 4-bit | ~2,000–10,000 der | ~12–60 h der | ~4.1–20.4 kWh der |
The energy column is derived at the measured 6.13 J per decode token for a drafter-on run at 55.8 t/s (§09), in-band GPU board power only. Two measured reference points for calibration, both from the 2026-08-23 sweep with the drafter off and real answer lengths rather than a 1,200-token assumption: 25 MATH-500 questions at low cost 126.2 Wh, and 50 questions at medium cost 165.9 Wh. Real benchmark work costs more than the table's assumption because real answers are longer than 1,200 tokens and because a drafter-off run spends about 7.9 J per token instead of 6.1.
The first row is the one job a local accuracy run does well, and it is the only claim §08's scored smoke test makes. The last row is why nobody ranks healthy files by accuracy: the run that could would take days per file and about twenty kilowatt-hours. And one failure mode the table cannot show: a saturated benchmark discriminates nothing at any sample size. On the 2026-08-23 suite, GSM8K returned 100% for all three effort levels. No amount of extra questions from that set would have separated them, because the set had stopped measuring anything about this model. Check for a ceiling before you buy more samples.
The other way a benchmark lies: the token budget
Sample size is not the only silent corrupter. A thinking model spends tokens reasoning before it answers, and a fixed max_tokens cap silently converts its hardest questions into wrong answers. The run does not error; the score just drops. This page has the cleanest possible demonstration of it, because the same measurement was run twice and the two runs disagree about which setting is best.
| Token cap | low | medium | xhigh | Truncations (scored sets) | What the table appears to say |
|---|---|---|---|---|---|
| 16,384 | 81.3 | 80.5 | 77.3 | 1 / 2 / 8 | "quality falls as effort rises" — a clean, quotable, wrong headline |
| 32,768 | 82.1 | 80.5 | 81.3 | 0 / 0 / 3 | all three levels within 1.6 points — indistinguishable at n=25 |
Same model, same prompts, same grader, same machine, same day. The only thing that changed is how many tokens each answer was allowed. The entire xhigh penalty was the cap, because by the grading rule a truncated answer scores zero, and xhigh truncated eight times against one and two for the other levels. Two individual prompts make it concrete: one MATH-500 problem needed 18,273 tokens and was correct; one HumanEval problem needed 17,025 and was correct. Both scored zero at 16,384. And the failure is worse than a cut-off answer: 11 of the 12 truncated generations came back with an empty answer field (eleven of the twelve are the scored-set truncations counted above; the twelfth is an ALPACA prompt, which was unscored at the time and is now scored — the judge panel rated that empty answer 1 out of 10 on all three seats, and the raised-cap re-run reproduced it exactly — the same answer consumed all 32,768 tokens and again returned nothing, §09), because the runaway happens inside the reasoning block, so the model never emits its end-of-thinking marker and never reaches an answer at all. Three xhigh prompts still exceed 32,768 and are genuine non-terminating loops rather than long solutions; they are reported as truncations rather than hidden.
The tempting fix is the wrong one: "re-run using only the questions that did not truncate." That selects the question set based on one arm's behaviour — dropping precisely the hard items where a quality difference could live — and quietly changes the question from "which setting scores better?" to "which setting scores better on easy questions?". It also breaks comparability with every published number on the same benchmark. The correct fix is to raise the budget, not to shrink the test, and to re-run only the arms and datasets that actually truncated.
The licence for a partial re-run is that greedy decoding is deterministic, so an answer that had room to finish under the old cap is byte-identical under a larger one. That used to be an argument on this page. It is now a measurement: every prompt that did not hit the old cap was compared byte-for-byte between the 16,384 run and the 32,768 re-run — low 24/24, medium 48/48, xhigh 67/67, a total of 139 of 139 with zero drift — and the re-run doubled the serving window as well as the cap. Raising --max-tokens and -c changed nothing about any generation that had room to finish. That is the evidence that re-running only the truncating datasets loses no information relative to a full re-run, and it is the cheapest way to fix a capped benchmark honestly.
Before trusting any thinking-model benchmark score — yours or a leaderboard's — ask two questions: what was max_tokens, and how many answers truncated? A score published without its cap cannot be checked, for the same reason a speculative-decoding speedup published without its acceptance rate cannot (§06). And one more question, added after a 2026-08-23 campaign caught this page's own probes on the wrong side of it: which token regime was being timed? On a model that thinks by default, a "code generation" probe with a 700-token cap can spend every timed token reasoning about the task and return an empty answer field — the label describes the task, the number describes the thinking. A speed number without its token regime is exactly as uncheckable as one without its acceptance rate. Pin enable_thinking or report the regime beside the number, and name probes after the token stream rather than the task.
Tokens ÷ wall-clock is not decode speed at depth — it silently averages the prefill in. The same measured run reads 47.1 t/s of decode and 9.2 t/s naive at a 28,000-token-deep prompt, and 35.8 against 2.4 at 92,000. Both are arithmetically correct; only one describes the model, because the other is mostly a statement about how long the prompt was. Quote the server's own timings, and quote prefill as its own number (§05, §06).
Perplexity and how close a file stays to the full-precision model's own next-token probabilities rank model files; a small-sample accuracy pass smoke-tests everything those metrics cannot see. Per-token metrics compare probabilities against the full-precision model over 294,912 scored positions on this page's corpus (36 chunks × 8,192 at -c 8192 — and note the tokenizer caveat in §08: that is a count of scored windows, not of corpus tokens, and it moves with the tokenizer), in an estimated 15–30 minutes per file on a 3090-class card. That estimate is not a benchmark from this page: it covers the per-file pass against saved full-precision baseline probabilities, and generating that baseline needs the unquantized model, which exceeds a 3090's memory and is a slower one-time cost. Those metrics resolve differences no feasible accuracy run can, which is why quantization authors compare files this way. But they run below the sampling and template layer, so a broken chat template or a mangled tool-call format sails through them unnoticed while collapsing scored answers by 20 points or more. Run both: per-token metrics to pick between files, one cheap accuracy pass to confirm the file you picked actually works end to end. §08's perplexity table is this advice in practice.
11
Why it sometimes runs at half speed
Four things on this machine produce the same complaint — the command that did 40 tokens per second yesterday does 25 today — and the server starts cleanly and reports nothing wrong in all four cases. Two are real slowdowns you can fix: the card quietly ran out of memory and spilled part of the model into system memory, or the command left one layer on the processor. One is not a slowdown at all but a client that never delivered your request. And the fourth does not slow anything down at all: it makes your own measurement read up to 25% low, so the server is fine and the number is wrong. Before any of them, though, comes a class of problem whose fix is never in the server command at all, so it is stated first.
Failure class 0 · the client contract
These are the failures a reader most often blames on the model, and every one of them is fixed in a different file from the server command. They belong together because they share one cause: the model has no awareness of its own limits and cannot budget its own tokens. Your configuration is the only thing enforcing them.
- One token pool.
prompt + thinking + output ≤ context. A 100,000-token thinking run inside a 32,768-token window fails by arithmetic, not by bug (§02). - Cut-offs are client-side. The server generates until it is told to stop. A truncation comes from your client's
max_tokens, its request timeout, or Ctrl+C. Sizemax_tokensso the whole run fits one response, and check that your client's context setting matches the server's-c— those two numbers drifting apart is the classic silent failure (§14 shows the setting per agent). - A cut generation cannot resume. With OpenAI-compatible interfaces there is no seamless continuation: the thinking is gone and the next request starts over. This is why a cap placed near the middle of the thinking-appetite distribution is a truncation machine rather than a saving (§09).
- Timeouts outlive defaults. A long
xhighrun outlasts most default client timeouts; the server log shows one asClient disconnected. Stopping generation. - A signature that separates client from server in one look. Check the llama-server log for the task. If the server logged no task at all, the request never left your machine and nothing about the model is implicated. If the server logged the task and returned something empty, that is the token-budget trap above, not a transport problem.
Failure 1 · the silent VRAM spill
The NVIDIA driver does not refuse an allocation that exceeds VRAM — by default it quietly overflows into shared system memory across the PCIe bus. On the reference machine, two open browsers held 2–3 GB of VRAM while the 122,880-token configuration needed about 22 GiB (§05), and about 1.9 GB spilled. The server loaded without a single warning and decoded at 20–35 t/s — and stayed there no matter how far the context was shrunk, because the weights stay spilled either way. This is §05's cliff in its most treacherous form: the cliff is invisible to the server. Nothing errors, nothing is logged, everything just runs at half its speed or worse.
The signature lives in Task Manager: dedicated GPU memory pinned at 23.5–24.0 / 24.0, shared GPU memory growing, processor use around 70% from driver paging. The fix is freeing the VRAM. The prevention is making the failure loud: NVIDIA Control Panel → Manage 3D Settings → Program Settings → llama-server.exe → CUDA — Sysmem Fallback Policy → Prefer No Sysmem Fallback. With that set, an over-budget load fails with an out-of-memory error at start-up instead of running slow all day.
Failure 2 · the -ngl off-by-one
llama.cpp counts the output layer as one more "layer" than the model has transformer layers. This model has 64, so -ngl 64 reads as complete — and leaves the output layer, a 5120 × ~151k-vocabulary matrix multiplication that runs for every generated token, on the processor. Measured cost: decode 25.7 against 39.7 t/s, prefill 106 against 290 t/s (server-log prefill on the short test prompt — a different probe from llama-bench's, which reads about 1300 on this card), GPU utilisation 53–67% instead of about 90%, and a stack of processor threads pinned doing the vocabulary multiplication. Independently replicated on a second file: the independent re-measurement (2026-08-23, same machine) measured UD-IQ4_XS at 29.84 t/s with -ngl 64 against 42.31 with -ngl 99 — −29.5% — on a different quantization. The Q4_K_M pair above is a steeper −35.3%; call the penalty 30–35% of your decode speed. The detail that makes it dangerous: the load succeeded and board VRAM read 15,256 against 15,661 MiB, near enough identical. There is no memory signature. The only signature is the speed — which is why §04's bandwidth figure for your file is the number to check against. Every recipe in §03 says -ngl 99 for exactly this reason: any number past the real count means "everything", with no off-by-one to get wrong.
High processor use alone proves nothing. llama.cpp worker threads spin while waiting on the GPU, so 50–70% processor use during perfectly healthy all-GPU decode is normal on a 20-thread machine. The trouble signal is the combination — high processor use and low t/s. Both failures above show both. And one log line that looks like an error and is not: failed to fit params to free device memory: n_gpu_layers already set by user only means that an explicit -ngl overrode the automatic fit. Any value triggers it, -ngl 99 included. The signals that actually separate the two failures are in the table below.
| Signal | VRAM spill | -ngl off-by-one |
|---|---|---|
| Shared GPU memory (Task Manager) | grows during the run | stays flat |
| Decode speed | 20–35 t/s, flat at every context size | 25.7 t/s, fixed (against 39.7) |
| Dedicated VRAM | pinned at 23.5–24.0 of 24.0 | normal — 15,256 against 15,661 MiB, no signature at all |
| Persistence | vanishes when other applications release VRAM | survives a reboot — it lives in your command line |
| Fix | free the VRAM · set Prefer No Sysmem Fallback | -ngl 99 |
The diagnostic commands that actually work on Windows
One tooling trap cost this page's own debugging session real time: on Windows, nvidia-smi's per-process listing shows [N/A] or "Insufficient Permissions" for memory, and its memory.used counts only dedicated VRAM — it cannot see the spill. The tools that exposed both failure modes:
# totals + who is on the GPU (dedicated only - the spill is NOT visible here): nvidia-smi --query-gpu=memory.used,memory.total,utilization.gpu --format=csv # the spill itself - total shared GPU memory in use (PowerShell): (Get-Counter '\GPU Process Memory(*)\Shared Usage').CounterSamples | Measure-Object CookedValue -Sum # divide Sum by 1MB; growing = spilling # per-process dedicated commitment - this is what caught the compositor (dwm) # holding 3.6 GiB during a browser-UI session: (Get-Counter '\GPU Process Memory(*)\Dedicated Usage').CounterSamples | Where-Object { $_.CookedValue -gt 200MB } # or simply Task Manager - Performance - GPU: the dedicated/shared split and, # under Performance - Memory, your RAM speed and "Slots used" (the section 07 channel check)
Failure 3 · PowerShell 5.1 silently drops your screenshot
One more Windows trap, and it is the kind that looks like a model problem. The independent re-measurement (2026-08-23, same machine) posted a 1440p screenshot to the vision endpoint from Windows PowerShell 5.1 and got nothing back — no answer, no error worth reading. Invoke-RestMethod cannot POST the roughly 261 KB body that a 1440p PNG data-URI produces. The identical request from Python succeeded in 3.8 s.
The diagnostic signature is the one from Failure class 0, and it separates client bugs from server bugs in one look: check the llama-server log for the task. If the server logged no task at all, the request never left your machine — nothing about the model, the projector, or --mmproj is implicated, and re-tuning any of them is wasted time. The fix is the transport: use Python's requests, curl, or PowerShell 7 or later for anything carrying a base64 image. Keep PowerShell 5.1 for the counters above, where it is excellent.
Failure 4 · the first probe after a long prefill — a 45% measurement swing
This one does not slow your server down; it makes your stopwatch lie, and it has already cost this page one published finding (§06's retired 15% projector cost). Measured 2026-08-23 while trying to settle that finding: four runs of one identical configuration read 18.27, 18.82, 19.21 and 26.60 t/s — the highest 45.6% above the lowest, with nothing changing but what the card had been doing beforehand. Every one of those was the probe fired immediately after a roughly 105-second prefill. This is the measurement that produces the ±25% band declared once in §01, and the two figures are the same measurement stated two ways: half of a 45.6% end-to-end spread is ±22.8%, rounded to ±25%.
The nvidia-smi samples taken before and after each prefill say why, and it is not heat. The first load of a session starts on a genuinely idle card — 56 °C, 225 MHz, 34 W — and its prefill drives the SM clock to a full 1,455 MHz boost. Every later load starts while the card is still winding down from the previous server's teardown and model load — 63–66 °C, 315–435 MHz, 62–136 W drawn at 1% utilisation — and its prefill only reaches 900–990 MHz. The decode probe that follows inherits whatever clock state that was. With the drafter on the same effect is milder but present: post-prefill probes read 54.3–60.0 against a settled 62.7.
The obvious suspect is thermal throttling, and it is wrong. Within a single cooled load, five consecutive probes ran the card from 57 °C to 82 °C and decode moved from 27.18 to 26.90 t/s — 1.0%. Once settled, this card's decode is nearly temperature-insensitive. All of the damage lives in the boost-clock ramp during and just after a long prefill, which is exactly where a benchmark script naturally puts its first measurement.
The fix is a protocol — this page's cooled protocol (§02) — and it costs about a minute per data point. Fire the prompt, discard the decode timing that comes back with it, wait 30–45 s, then send several short repeat requests that re-use the cached prefix and report the median of those. Under that protocol repeat probes on this machine agree to 0.4% — which is the only reason a 0.04% difference between two configurations was measurable at all. The same ramp is a power caveat: a prefill fired at a cold board draws unrepresentative watts over an unrepresentative duration, which is why 10-second power probes read 277–287 W against a sustained 306–341 (§03).
Two traps that only affect people taking measurements
Both were found while producing this page, and both are the kind of mistake that yields a plausible number rather than an obvious error.
- "The server is down" is not "the GPU is idle." A settled idle window measured after the last job read a mean of 37.8 W against a median of 31.1, because five short excursions to 121–124 W punched through it — each one the desktop rendering plot images, at 0–14% GPU utilisation, minutes after llama-server had exited. If you measure idle power, measure it with the desktop quiet and discard the first 60 seconds while the board is still cooling. The settled figures on this page (29.9 W with no server, 34.1 W with the model resident) were taken that way; the earlier 33.2 W and 34.6 W readings were not, which is how they came out above the loaded figure.
- Never charge a whole token count against a partly covered power log. Joining request timings to a power log that started mid-run, and crediting each request its full token count, produced a physically impossible 5.32 J per token and 59.8 t/s on a card that cannot do either with the drafter off. The fix is to count only requests whose entire window lies inside the log — 115 of 150 in that case. Any per-arm energy join across a log boundary needs the same filter.
12
Vision: sending screenshots to the model
This model can look at images if you start the server with the optional projector file, and on a 24 GB card that is a real option rather than a compromise: it costs 1,138 MiB of memory and, measured carefully, nothing at all in speed. What it does cost is context. A 1440p screenshot occupies about 3,600 tokens of your window and a 4K one about 8,244, so a feedback loop that keeps ten screenshots around has spent 36,000 tokens before the model writes anything. Use 1440p, set --image-min-tokens 1024 and --image-max-tokens 10580, send each picture as base64 data inside an ordinary request (that is the standard way to put an image into a text request), budget max_tokens generously, and clear old screenshots between iterations. One honest limit up front: everything in this section is a single demonstration each, and the control that would prove the model is actually using the image was never run.
The end-to-end run below is real and is reported exactly as it happened, but it is not evidence of a critique loop, because the experiment that would make it evidence was never done: ask the same question with the image withheld and see whether the answer is as good. A model that has been told "this is an animated aquarium page" can produce plausible criticism of an aquarium page without seeing one. Until that withheld-image control is run, treat this section as proof that images reach the model and cost what the arithmetic says, and treat the quality of the critique as unverified. The control is listed in §15's register of what was not measured.
What was demonstrated end to end on the reference 3090, 2026-08-22: Chrome captured a 4K screenshot (3840×2160, a 2.3 MB PNG) of a test page — an animated canvas aquarium — and the model, served with the configuration below, named the visible creatures individually, identified the weakest visual element (kelp drawn as wire-thin lines), and wrote working replacement drawing code with tapered, phase-offset swaying blades. Round trip: about 39 s, of which the 4K image cost 8,244 prompt tokens. n=1
Image cost is resolution, not file size — and it obeys an exact law. The vision tower uses patch 16 with a spatial merge of 2, so one visual token covers a 32 × 32-pixel block: tokens = pixels ÷ 1024 + 2, the +2 being the start and end markers. The independent re-measurement (2026-08-23, same machine) confirmed it to within 2 tokens at two resolutions — 720p measured 922 against 920 predicted, 1440p measured 3,602 against 3,600. The 4K row needs the rounding stated to match: 3840×2160 divided flat gives 8,102, while rounding each dimension up to whole 32-pixel blocks (120 × 68) gives 8,162 against the measured 8,244 — 1.0% high on the block form, 1.8% on the flat one, so quote the block form and expect about a percent of slack at 4K. Earlier editions interpolated the 1440p and 1080p rows and ran 28–30% high; only the 4K row was ever measured. Corrected:
| Screenshot | Context cost | Source | Feedback rounds alongside a full xhigh cycle at -c 122880 |
|---|---|---|---|
| 1080p (1920×1080) | ~2,040 tokens | der arithmetic — unmeasured; the height rounds up to 34 blocks | ~16–21 der |
| 1440p — recommended (2560×1440) | 3,602 | meas (earlier editions printed "~4.7k") | ~9–12 der |
| 4K (3840×2160) | 8,244 | meas | ~4–5 der |
1440p under --image-max-tokens 1024 | 1,010 | meas — 3.6× cheaper, but its effect on the answer is not measured (box below) | ~33–43 der |
Rounds are recomputed from the corrected costs rather than carried over: a full xhigh cycle wants 80,000–90,000 of the 122,880-token window (§09), leaving 33,000–43,000 for images, and the column is that remaining window divided by the row's cost. The old rows inherited their error into this column — "~7–9" for 1440p was arithmetic on a token count 30% too high.
--image-max-tokens 1024 takes a 1440p screenshot from 3,602 tokens to 1,010, measured — a genuine 3.6× saving and the knob to reach for when a feedback loop is eating the window. What it does to the answer was never measured. The experiment is small and obvious — the same image at both budgets, asked a question whose answer depends on a detail only the full-resolution version can show — and it was not run here. "Minimum for grounding, maximum for detail" is a mechanism, not a measurement, so treat the 1024 setting as an unverified default rather than a validated trade.
The two budget flags behave differently, and it is worth knowing which one actually saves you anything. --image-max-tokens is a real cap that bites, as above. --image-min-tokens is only a floor: it lifted a 720p shot from 922 to 1,077 tokens and left the 1440p shot at 3,602, untouched. It raises small images and does nothing to anything above the floor — so --image-min-tokens 1024 costs you nothing on the screenshots you care about, and is not a lever for shrinking them. The default cap is 4,096 tokens, which by the law above is about 4.2 megapixels: that is exactly why a 1440p shot (3,602) passes through untouched at default flags while a 4K one does not, and why this page's vision configuration raises the cap to 10,580 for 4K-class detail. (4K handling upstream still has open quirks, so 1440p remains the recommended resolution — and it is now measured rather than interpolated.)
The serving configuration, and the capture command that feeds it:
llama-server -m Qwen3.8-27B-UD-IQ4_XS.gguf --alias qwen/qwen3.8-27b -c 122880 -ngl 99 --parallel 1 --load-mode none -ctk q8_0 -ctv q8_0 --mmproj mmproj-Qwen3.8-27B-BF16.gguf --image-min-tokens 1024 --image-max-tokens 10580 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 # n4/p0.75, matching section 03's recipe [1]: with the projector loaded, the wider # n10 drafter's 898 MiB comes out of the slack a desktop needs. Measured live at # these flags: ~3 GiB of board VRAM left over with a desktop up. An older live check # of # this configuration read 71.3 t/s short-context, but its token regime was never # recorded, so this page does not use it as a speed band - plan on the 83-86 t/s # measured on answer tokens at these same flags (sections 03 and 06); # the projector itself books 1,138 MiB of that (section 05) --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --jinja --host 127.0.0.1 --port 1234 # capture the page your code just rendered, using Chrome's no-window mode. The flag # below is spelled --headless, but it has NOTHING to do with the graphics card or # with "headless" as section 02 defines it: it only means the browser draws no # window on screen. It does not free any VRAM. chrome --headless --disable-gpu --window-size=2560,1440 ^ --screenshot=shot.png --virtual-time-budget=6000 file:///path/to/page.html
Multi-image works, and it can compare versions. Measured: three screenshots of three different aquarium implementations in one request — 4K, 1440p, 1080p — cost 13,913 prompt tokens with a 74-second round trip, and the model ranked them by visual richness correctly and flagged the third as "broken / nearly empty… looks like the animation failed to spawn its fish". That page is the same low-effort output whose code was blind-scored 40/100 in §09 for freezing on its first frame — two independent reviews reaching the same verdict, though again without a withheld-image control. That 13,913 is also an unplanned third confirmation of the corrected token law: 8,244 + 3,602 + ~2,040 = about 13,890, some twenty tokens from the measurement once the text prompt is counted. Under the old interpolated figures the same three shots should have cost about 15,500 — this page's own multi-image measurement had been contradicting its own table all along.
Which coding agents can carry a screenshot to the model
All five were driven from a script, with no interactive terminal session, against this server with the same 2.3 MB 4K PNG and a question only answerable by seeing it (2026-08-22, one attempt each). The encouraging result: no agent hallucinated sight. Every failure was an honest "I see no image", which is the failure mode you want, because the alternative — an agent that describes a picture it never received — is undetectable from the transcript. Two of the five needed configuration before they passed, and in both cases the default behaviour was to silently send the image as text.
| Agent | Verdict | How — and the trap |
|---|---|---|
| OpenCode | PASS with configuration · FAIL-honest without it | a bare model entry silently reads the PNG as bytes — tested, and the model honestly reported seeing nothing. Declare "attachment": true and "modalities": {"input": ["text","image"]} on the model (the configuration in §14 includes them), and, run from a script, opencode run "@shot.png …" passes — verified on v1.18.21 after tracing issue #15728 |
| aider | PASS | image path as a plain file argument (aider … --message "…" shot.png); requires supports_vision: true in the model metadata file |
| Qwen Code | PASS with configuration · FAIL-honest without it | an environment-variables-only setup silently drops images. Declare the model in settings.json under modelProviders with capabilities: {vision: true}, then qwen -p "@shot.png …" works |
| Pi | PASS | image path as a positional argument after -p; the provider's models.json must list "image" in input |
| DeepSeek Harness | PASS | no attach flag exists or is needed — reference the path in the task text; settings.yaml must declare input: [text, image]. Gave the richest correct answer of the five |
One attempt per agent, one image, one question — so this table says "the plumbing works", not "this agent is good at vision". The image-attachment matrix at scale, across resolutions and multi-image requests, was never run.
For a scripted screenshot loop, the OpenAI-compatible interface itself is the zero-dependency path: send the image as a data:image/png;base64,… entry in an image_url content part alongside the text — exactly what the demonstration above used. Three rules regardless of route. Budget max_tokens generously: a thinking model reasons about the image before answering, and a tight cap returns an empty answer with the thinking complete (measured; §11's client contract applies to vision too). Clear old screenshots between iterations; each one keeps its thousands of tokens in the window. And on Windows, do not send image data-URIs from PowerShell 5.1 — Invoke-RestMethod silently fails to POST the roughly 261 KB body a 1440p PNG produces, and the signature is that the server log shows no task at all (§11). Python, curl, or PowerShell 7 and later all work.
13
When 27B is enough, and when it is not
A well-tuned local 27B is genuinely good at bounded work: answering factual questions, writing and explaining ordinary code, summarizing, and completing a self-contained task you can check when it finishes. It gets less reliable as a task gets longer, as the knowledge gets rarer, as the number of simultaneous requirements grows, and as the number of unchecked steps between you and the result increases. The decline is gradual rather than sudden, and much larger models are further along the same decline rather than free of it. The practical rule: give this model work you will look at, and escalate work that has to be right without you looking. This section is the evidence for that rule — and unlike the rest of this page, most of it is argued from published literature rather than measured here.
What this page did measure about capability: a five-set, 175-prompt benchmark suite at three effort levels, composite index 82.1 / 80.5 / 81.3 with GSM8K saturated at 100 and code execution pass rates of 92–100% on HumanEval and 84–92% on MBPP (§09). That is a real capability measurement on this file and this machine, and it is not comparable to the vendor's published figures for this model: different precision (4-bit against full), different harness, different sampling, different prompt formatting. Nor does it speak to the five territories below, all of which are about behaviour over long horizons that a 25-question set cannot reach. One agentic pipeline was also built and validated end to end on this machine — 22 model calls, 9 minutes 36 seconds of wall clock on one sandboxed repair task against real code — and it is worth reporting its outcome rather than only its existence: 0 of 96 target tests turned green, while all 561 previously-passing tests kept passing. The harness worked; the fix did not. That is one sample, and it proved the plumbing and the cost rather than the capability; the 10-task, three-effort sweep that would have measured capability was cut against a four-hour time gate on an eight-and-a-half-hour projection.
A model does not need trillions of parameters to act as an encyclopedia. High-frequency knowledge — geography, mainstream history, standard biology, common coding syntax — compresses comfortably into 27B (a model stores on the order of 2 bits of factual knowledge per parameter), which is why local factual question-answering feels frontier-grade. What frontier models buy with their scale looks like a difference in kind but is built from differences in degree: per-step reliability improves smoothly with capability, and because a long task succeeds only if every step does, those smooth gains compound into cliff-like gaps on exactly the work below — as reasoning engines, software architects, and autonomous agents. The boundary runs through five territories:
- Long-horizon reasoning. Direct questions are effortless at 27B; what degrades is the 30-step deduction, the formal proof, the architecture analysis — work where one early flaw silently poisons everything downstream: agent success decays exponentially with task length, and transformers solve deep compositional problems by pattern-matching that collapses as the reasoning graph grows — a collapse that reaches even frontier models at sufficient depth, so the frontier extends the horizon rather than escaping the regime. This page's judged runs located the local edge precisely: at
xhigh, the 27B delivered 100% of a complex single-file specification, twice. Multi-file systems, where the flaws hide in the seams between files, are where the footing gives — and note from §09 that no read-then-edit task against existing code was ever run here, so that boundary is argued rather than measured. - The long tail of knowledge. 27B keeps what training data repeated often — fact recall tracks how many times a fact appeared in pre-training — and frontier models retain far more of the rare: the niche legal precedent, the legacy framework, the unusual drug interaction (the studies measure trivia-style entity facts; these are the practical analogue). But the tail deficit shrinks rather than disappears — scaling fails to appreciably improve memorization of tail facts, and closing the gap by scale alone would take many orders of magnitude more model. The danger is the failure mode: on obscure ground no model reliably says "I don't know" — training and benchmarks reward confident guessing over abstention — and a small model simply has more gaps for the fluent guess to fill. Treat confident answers about rare domains as unverified drafts, or ground them with retrieval: retrieval-augmented small models beat plain models orders of magnitude larger on exactly these facts.
- Ultra-long recall. A whole codebase or a 500-page audit in one prompt is two problems wearing one coat: holding the text, which this card's 122k window does (§05), and keeping every cross-document dependency straight, which improves with model scale — in RULER's within-family test, a 34B degrades far less than a 6B trained on the same context — but scale mitigates rather than cures: only about half of models claiming 32K or more actually handle 32K, information buried mid-context is used far worse than at the edges, and once retrieval requires inference instead of literal matching, most long-context models fall below half their short-context accuracy by 32K.
- Multi-constraint instruction following. "Five custom interface rules, strict typing, an exact schema, and ten edge-case tests — simultaneously." Rule stacks like that degrade small models first, and quietly: the dominant failure at high instruction density is silent omission — the output stays a coherent, compliant-looking document — and violations are hard enough to spot that benchmarks need model judges to find them. Every model declines as constraints stack, so fidelity under heavy constraints is one of the clearest capability dividends — frontier reasoning models hold near-perfect through about 150 simultaneous instructions where small models decay exponentially — but it buys later, slower degradation, not immunity: even the best fall to about 68% at 500 instructions.
- Autonomous agentic reliability. An agent that calls tools, reads its own error logs, and self-corrects across dozens of steps multiplies its per-step error rate through every step — measured agent success declines exponentially with task length, like a half-life — and worse than independent errors predict, because models self-condition on their own earlier mistakes, amplifying drift into failure. Frontier agents push the horizon from minutes toward hours, but at 50% success — at their quoted horizon they still fail half the time. This is why the setups in §14 shine for supervised sessions and wobble on multi-hour autonomous loops, and why §09's defect risk matters most in exactly that mode, where nobody is watching the output.
| Task type | 27B local | Frontier model |
|---|---|---|
| Factual question-answering and summarization | excellent | excellent |
| Standard code and prose | great | exceptional |
| Multi-file system architecture | moderate — struggles | exceptional |
| 50-step autonomous tool loops | high drift / failure | much longer horizons — still ~50% at their own limit |
| Obscure or niche domain expertise | prone to confident guessing | better, still unreliable — verify or retrieve |
This table is a summary of the literature above plus this page's own judged runs; no cell in it is a measurement of a frontier model on this machine.
A well-tuned local 27B is a fast, private, tireless colleague for the bounded work that fills most of a day — and this page exists to make it as good at that as the hardware allows. Frontier models exist for the multi-step, high-stakes cognitive work where accuracy must decay as slowly as possible — and even they decay. The two compose: route the bounded work locally and escalate the architecture, which the planner-and-executor split in §14's DeepSeek Harness setup expresses as a configuration file. Knowing which task you are holding is itself a setting — the one no flag can fix.
14
Coding agent setups
Every recipe in §03 ends with an OpenAI-compatible server, so any OpenAI-compatible agent can connect with the same three facts: the address http://localhost:1234/v1, the model name the server reports (qwen/qwen3.8-27b, set by --alias), and the context and output limits from your -c setting. Two things go wrong often enough to state once for all five agents below. Tell the client the model's real limits — if its context setting and the server's -c drift apart, generations get cut off silently. And override the temperature — several clients send 0 by default, which puts this model into repetition loops; it wants 1.0 with top-p 0.95. The five setups below are ordered by popularity, and each was installed, configured exactly as printed, and run through a live request against the reference server on 2026-08-22.
OpenCode · the most-starred coding agent
# install: npm i -g opencode-ai (or the curl installer / brew) # config: ~/.config/opencode/opencode.json (or opencode.json in the project) { "$schema": "https://opencode.ai/config.json", "provider": { "llamacpp": { "npm": "@ai-sdk/openai-compatible", "name": "llama-server (local)", "options": { "baseURL": "http://localhost:1234/v1", "apiKey": "dummy" }, "models": { "qwen/qwen3.8-27b": { "name": "Qwen3.8 27B (local)", "attachment": true, "modalities": { "input": ["text", "image"], "output": ["text"] }, "limit": { "context": 122880, "output": 106496 } } } } } } # run: opencode --model llamacpp/qwen/qwen3.8-27b (or pick via /models) # the attachment/modalities pair is REQUIRED for images: without it OpenCode # silently reads a referenced PNG as text bytes (section 12's agent table) # # THE TWO NUMBERS ABOVE ARE TIED TO YOUR SERVER'S -c. Derive them, do not copy: # context = the server's -c value exactly # output = -c minus your longest expected prompt; 106496 = 122880 - 16384. # For recipe [2] at -c 180224 that is context 180224, output 163840; # for recipe [3] at -c 262144, context 262144, output 245760.
Two gotchas from the documentation: baseURL must sit inside options, not at the provider root, and the provider identifier must be a custom name ("llamacpp", "local") — reusing a built-in provider name breaks resolution. (llama.cpp example)
aider · the git-native pair-programmer
The most detailed walkthrough of the five — its metadata-and-limits pattern is the template every other client repeats in its own configuration format.
Step 1 · Install aider
# needs Python 3.9+ — any one of these: python -m pip install aider-install && aider-install # official installer (recommended) pip install aider-chat # plain pip uv tool install --force --python python3.12 aider-chat # uv, isolated environment
Step 2 · Point aider at the server
Aider's openai/ model prefix means "any OpenAI-compatible server" — llama-server and OpenVINO Model Server both qualify. Set two environment variables (the key is required by aider; make it match the server's --api-key if you set one, otherwise any placeholder works):
# Windows (cmd / batch) # Linux / macOS
set OPENAI_API_BASE=http://localhost:1234/v1 export OPENAI_API_BASE=http://localhost:1234/v1
set OPENAI_API_KEY=dummy export OPENAI_API_KEY=dummy
Step 3 · Tell aider the model's limits
Aider knows nothing about a local model until you describe it. Create .aider.model.metadata.json in the directory you launch aider from, or in your home directory. Match max_input_tokens to the server's -c value — these two numbers drifting apart is the classic silent failure:
{
"openai/qwen/qwen3.8-27b": {
"max_input_tokens": 122880, // = the server's -c, exactly
"max_tokens": 106496, // = -c minus your longest prompt (122880 - 16384)
"input_cost_per_token": 0,
"output_cost_per_token": 0,
"supports_vision": true
}
}
The model name after openai/ must match what the server reports at /v1/models — llama-server sets it with --alias, so every recipe in §03 serves as qwen/qwen3.8-27b.
Step 4 · Fix aider's defaults for a thinking model
Aider sends temperature 0 by default, which the model card advises against for thinking mode and which is a known repetition risk on reasoning models — this page's own audit of ten long greedy transcripts found 10 of 10 clean, so treat it as a risk worth avoiding rather than a certainty. It also needs to know to strip thinking blocks before parsing edits. Create .aider.model.settings.yml next to the metadata file:
- name: openai/qwen/qwen3.8-27b
use_temperature: 1.0
reasoning_tag: think
extra_params:
top_p: 0.95
Step 5 · Launch
aider --model openai/qwen/qwen3.8-27b --edit-format diff --timeout 3600
--edit-format diff— the model writes only changed sections instead of whole files. At local speeds this is the biggest quality-of-life setting there is: a 60 KB file edit drops from minutes to seconds. Diff mode also degrades gracefully on cut-offs, because complete sections still apply.--timeout 3600— long thinking runs outlive default client timeouts; the server log shows a timeout asClient disconnected. Stopping generation.- First-run sanity checks: the model-warning banner should be gone (the metadata file was found), and the banner should read diff edit format. If edits arrive mangled, the fallback is removing
--edit-format diffto use whole-file mode. - Vision: with
supports_vision: trueand the server launched with--mmproj,/add screenshot.pngworks and the local model will see it (§12).
Qwen Code CLI · Qwen's own terminal agent
# install (Node.js 20+) npm install -g @qwen-code/qwen-code # configure via environment variables — put them in ~/.qwen/.env to persist: OPENAI_BASE_URL=http://localhost:1234/v1 OPENAI_API_KEY=dummy OPENAI_MODEL=qwen/qwen3.8-27b # then just run: qwen
The key must be present even for a local server ("Missing credentials" otherwise), and set the context window explicitly in its settings if offered — Qwen Code's built-in defaults assume the cloud model's limits, not your -c value. For images, declare the model in settings.json under modelProviders with capabilities: {vision: true}; an environment-variables-only setup silently drops them (§12). (setup notes)
Pi coding agent
# custom providers live in ~/.pi/agent/models.json: { "providers": { "llamacpp": { "baseUrl": "http://localhost:1234/v1", "api": "openai-completions", "apiKey": "dummy", "compat": { "supportsDeveloperRole": false, "supportsReasoningEffort": false }, "models": [ { "id": "qwen/qwen3.8-27b", "name": "Qwen3.8 27B (local)", "contextWindow": 122880, "maxTokens": 106496, "reasoning": true, "input": ["text", "image"] } ] } } } # one-shot from the command line, no interactive session: # pi --model llamacpp/qwen/qwen3.8-27b --no-session -p "task" # contextWindow = the server's -c; maxTokens = -c minus your longest prompt.
The compat flags matter: they tell Pi to send a plain system role and to skip the reasoning_effort request parameter, which is the dialect llama-server expects — effort is a server-side setting here (§09). Pi enforces maxTokens on the client side, so size it like aider's; this is the client whose default cap causes the mid-generation cut-offs described in §11. (documentation)
DeepSeek Harness · plugin-based agent framework (developer preview)
# install globally, then run the web UI (an npx cold cache can hang for minutes): npm i -g @deepseek-ai/dsh dsh web # custom OpenAI-compatible providers live in $DSH_HOME/settings.yaml. # BOTH blocks are required: registering the provider is not enough — without the # routing block the harness silently keeps its cloud default and fails with # MISSING_CREDENTIAL (verified end-to-end on the reference machine, 2026-08-22): agent-default-model: provider: llamacpp model: qwen/qwen3.8-27b llm-pi-ai: providers: llamacpp: api: openai-completions baseURL: http://localhost:1234/v1 apiKeyEnv: LLAMA_API_KEY # set LLAMA_API_KEY=dummy to match --api-key compat: supportsDeveloperRole: false # llama-server wants a plain system role maxTokensField: max_tokens models: - id: qwen/qwen3.8-27b input: [text, image] # vision works when the server runs with --mmproj contextWindow: 122880 # match the server's -c value exactly maxTokens: 106496 # REQUIRED: dsh enforces this client-side (same # schema as Pi). Without it the small default cap # cuts long thinking replies mid-run — the symptom # is "Output token limit reached ... send continue" # on any xhigh answer (section 11)
The harness's headline trick is routing: it plans with heavy reasoning, then fans the plan out to sub-agents that execute with light reasoning. Two things to know when backing it with llama-server. First, llama-server ignores a per-request reasoning_effort field (the same dialect issue as Pi above) — effort is fixed server-side with --chat-template-kwargs, so either run one server at xhigh (this page's quality-first default), or run two llama-server instances — one xhigh planner on one port, one low executor on another — and register both as separate models in settings.yaml so the harness can route between them. Second, its parallel sub-agents each hold their own context: with --parallel 1 (every recipe here) requests queue rather than run concurrently, which is the right default on a single 24 GB card for memory reasons — raising --parallel multiplies the KV cache and walks you off §05's cliff. It is now also the right default for latency reasons, which was an open question in earlier editions and is not any more: the matched pair of 2026-08-25 measured two concurrent slots at +22.0% aggregate throughput and −35.4% per slot with the drafter on (§03), so a harness whose sub-agents you are waiting on finishes each one faster with a single slot. For non-thinking "fast mode" runs, Qwen's official instruct sampling is temperature 0.7 · top-p 0.8 · presence penalty 1.5.
Whatever the agent, its context and maximum-output settings must match the server's -c value — the client-side limits are the only thing preventing silent mid-generation cut-offs, because the model itself cannot budget its own tokens (§11).
15
Sources, instruments, and what was not measured
This section exists so that a reader who distrusts one number can find out exactly what produced it. It has four parts: which instrument read each class of number; where every external fact came from and how strongly it is sourced; a numbered list of the places this page corrected itself; and a register of what it never measured at all, so that nobody mistakes an omission for a result. It ends with one command you can run to check whether your own setup matches this machine.
Which instrument produced which number
| Class of number | Instrument | Notes and known biases |
|---|---|---|
| Decode and prefill speed | llama-server's own timings block — prompt eval time … / N tokens and eval time … / M tokens | Token-weighted. Never tokens ÷ wall-clock, which averages prefill into decode and reads 5–20× low at depth. A benchmark harness's own mean over requests is unweighted and read up to 3.8 t/s higher than the server's on these arms (42.2 against 38.4); the two are never mixed in one table here |
| Draft acceptance, mean draft length | the same timings block: draft_n and draft_n_accepted | Accepted per target pass is derived as draft_n_accepted ÷ (predicted_n − draft_n_accepted) |
| VRAM, server process | llama-server's own dedicated-VRAM report at load | This is what the budget model in §05 is fitted to, and it is stable to 127 MiB across 17 loads |
| VRAM, board total and spill | nvidia-smi --query-gpu=memory.used for dedicated; Windows performance counters \GPU Process Memory(*)\Shared Usage and \Dedicated Usage for the spill and per-process split | nvidia-smi on Windows cannot see the spill and shows [N/A] per process. The board VRAM total includes the desktop, which measured 1,179–1,669 MiB. The three counters and what each of them can and cannot see are in §02 |
| Board power and all energy figures | NVML through nvidia-smi --query-gpu=power.draw, logged at 1 Hz (the 08-22 effort arms) or 500 ms (the 08-23 sweep and matrix), integrated by scripts/power/attribute-power.py (trapezoid rule, edges linearly interpolated, gaps over 2 s excluded) | In-band tier — read from the card's own sensor, not from a meter at the wall. Inside the number: the graphics chip, its memory, board regulators and fans. Excluded and never measured: the power supply's conversion loss, the processor, system memory, drives, the display, and any facility overhead. The integrator's self-test passed 7 groups and 27 assertions on 2026-08-23. Every joule is joined to the server's own per-request timings, so prefill and decode are attributed separately |
| Open-ended writing quality (ALPACA, MT-Bench) | a blind panel of three Claude Opus 5 seats over the kept transcripts, standard 1–10 single-answer rubric, normalized (r−1)/9×100; scripts/bench/judge-panel.py | The most subjective instrument on this page, and the one with a named conflict. Qwen wrote the answers and Claude read them, so no model graded its own output — but the judge and this page's author are both Claude models, which is a correlated instrument, not an independent one. Mitigations, all of them partial: three seats rather than one; blinding by opaque salted identifier with the arm mapping sealed; a different shuffle seed per seat; seat spread published beside every mean (0.28–0.92 rating points); and arm-against-arm conclusions drawn from a paired bootstrap over the same prompts rather than from raw means. Known residual bias: judges tend to reward length, and these seats were instructed against it and reported correcting for it — which is self-report, not proof. A second-vendor or human judge is an open item in §15 |
| Perplexity | llama-perplexity over the full wikitext-2-raw test set at -c 8192 -fa on -ngl 99 | 36 chunks × 8,192 = 294,912 scored positions; the corpus file was hash-verified identical across the cache-precision comparison. The count is tokenizer-bound (§08) |
| Benchmark scores, 2026-08-23 | measured-inference/scripts/bench, suite hash 1cdf54f8eb9d3f8f, 175 prompts, greedy, seed 42 | Three grader bugs were found and fixed symmetrically during this run — applied to prediction and reference alike, so a fix can only make one value written two ways compare equal, never make two different values match. All arms were re-graded offline from kept transcripts with the final grader; the self-test stayed at 78 passed, 0 failed |
| Benchmark scores, 2026-08-21 and 2026-08-22 | chinkeong/benchmark — final-number / boxed-answer exact match, truncations counted wrong | A different grader from the one above, and unaudited against the three bugs it found. Every score from these two days carries that caveat: §08's 20-question table and its GSM8K n=200 comparison |
| Whether the emitted code runs | measured-inference/scripts/bench/execute-probe.py — each file's probe-A output rebuilt with the prompt's own prefix and executed under node v24.15.0, 15 s timeout, greedy (temperature 0, top_k 1), n=1 per file | One task, one language, one sample per file — a threshold, not a rate, and it must not be read as a pass rate. Probe A is a continuation task, so an output beginning mid-expression is correct rather than truncated; the prefix is prepended per output according to whether the file continued the program, restarted it, or wrapped it in the markdown fences the prompt forbade. The demo graph is model-generated, so there is no fixed correct answer to check — the test is that the program parses, runs to completion and prints the path and cost it promised. Zero GPU time: it reads files already on disk. Added 2026-08-25 after the judge panel and this probe both found degeneration that the four lexical detectors (§02) cannot see — a missing closing brace contains no repeated n‑grams, so no threshold on them could ever catch it |
| Image token counts | the server's own prompt_n differential with and without the image | Confirmed the arithmetic law to within 2 tokens at two resolutions |
| Clock, temperature and power state | nvidia-smi --query-gpu=clocks.sm,clocks.mem,temperature.gpu,pstate,utilization.gpu | Present in the 08-23 logs and absent from the 08-22 ones, which is why the earlier logs cannot prove whether a low power sample was a ramping board or an efficient one |
External sources, graded
Each bullet is marked [P] primary document, [A] arithmetic on published specifications, or [C] carried from a dated prior fact-check on this page.
- [P] §13 task boundaries — long-tail knowledge: Kandpal et al., ICML 2023, PopQA / Mallen et al., ACL 2023, knowledge-capacity scaling laws, Why Language Models Hallucinate · long-horizon reasoning: Faith and Fate (NeurIPS 2023), emergent-abilities mirage (NeurIPS 2023), METR time horizons, agent half-life (Ord, 2025) · long context: RULER (COLM 2024), Lost in the Middle (TACL 2024), NoLiMa (ICML 2025) · constraint stacking: FollowBench (ACL 2024), IFScale, ComplexBench (NeurIPS 2024) · agentic reliability: METR blog, self-conditioning in long-horizon execution
- [P] Arc Pro B70 — 32 GB, 256-bit, 608 GB/s: Intel datasheet, corroborated by Tom's Hardware and Puget Systems
- [P] RTX 50-series bandwidth (5090 1792, 5080 960, 5060 Ti 448 GB/s) — GamersNexus, TechSpot; the RTX 3090's 936 GB/s is NVIDIA's own GA102 whitepaper figure (384-bit × 19.5 Gbps)
- [P] DGX Spark GB10: 128 GB unified LPDDR5X, 273 GB/s — Tom's Hardware review, LMSYS
- [P] Arc B390-class iGPU (Core Ultra): dual-channel LPDDR5X-9600 mandate, 153.6 GB/s, ≤96 GB shared, shipping in Core Ultra X9 388H laptops — Intel ARK, TechPowerUp, Tom's Hardware, HotHardware, shipping laptop list
- [C] NVFP4 on Blackwell: vLLM-only, about 1.5× BF16, 92–97% accuracy, FP8 KV — Unsloth documentation, NVIDIA forums, Kaitchup. Vendor claims, unmeasured here
- [C] Quantization quality claims (unsloth Dynamic v3, Qwen's AD-IQ3_S, the lmstudio Q6_K comparison) — HF discussion #65, kingy.ai
- [C] The
xhighoverthinking anecdote (22,276 thinking tokens, 21 minutes, for one simple drawing) — Simon Willison - [P] Intel software stack: IPEX-LLM archived January 2026 — the repository itself; Vulkan beats SYCL on Battlemage — llama.cpp #22413, B70 benchmark notes, Phoronix
- [P] Model card, OpenVINO builds, drafting support — Qwen/Qwen3.8-27B, int4-ov, int8-ov, qwen38-mtp recipe and benchmarks, llama.cpp discussion 27164, NVFP4-MTP-GGUF card (whose Blackwell-only claim this page's Ampere measurements contradict)
- [P] DFlash2 block drafting and path selection — inco.ai announcement, llama.cpp PR 27342, drafter GGUFs
- [A] Every non-3090 row of §07: bandwidth ÷ file size × the format constant (0.70 K-quant, 0.65 IQ-quant), using the bandwidth figures above
This page's own measurement runs
- 2026-08-21 — llama.cpp build 5ecbe1ac1 (PR 27342), CUDA 12.8, RTX 3090: the drafting comparison (baseline / built-in head / DFlash2 / NVFP4 on Ampere) and the 18-configuration n-max and p-min sweep, all temperature-0 600-token code generation, plus llama-bench prompt-processing and generation figures. Also the FP4-against-INT4 scored comparison: GSM8K and MATH-500, 20 each, greedy, 4,096-token budget, llama-server build 10502, graded with chinkeong/benchmark
- 2026-08-22 — RTX 3090: the context-ceiling probe (stepped
-cwith short temperature-0 probes), the reasoning-effort comparison at temperature 1.0 with blind judging, the acceptance demonstration at n-max 10 / p-min 0.5, the wikitext-2 perplexity table over six files, the GSM8K n=200 runs per effort level and per file, the 10-configuration drafting re-sweep, the two half-speed failure modes, and the vision demonstration. Reference recipe: Q4_K_M,q8_0KV, drafter on, projector loaded - 2026-08-23, independent re-measurement round (00:19–02:19, ~2 h GPU) — the source of most corrections dated 08-23 in §04–§12: the drafter's VRAM bill, the two-constant VRAM model, the measured window slope, the deep-fill collapse, the answer-against-reasoning token regimes, depth on the shipped recipe, the vision token law, per-effort energy, appetite as a distribution, chat-template forensics, the perplexity position count, the PowerShell image-POST trap, and the bandwidth citation fixes. Same RTX 3090, driver 596.36, Windows 11 Pro 26200, i5-13600KF, 31.8 GiB DDR4-3200 dual-channel; llama.cpp build 10502 (version 0.1.2-dev, commit
0adcc3bb5, Clang 20.1.8, binary dated 2026-08-19); Qwen3.8-27B-UD-IQ4_XS.gguf (14,252,845,984 B) withmmproj-Qwen3.8-27B-BF16.gguf;--load-mode mmap, port 1235,--parallel 1,q8_0KV,-ngl 99. Run without sight of this page, so where its numbers differ from an 08-21 or 08-22 row, the difference is either a genuine condition change or a correction, and each is labelled as one or the other where it lands - 2026-08-23, follow-up round (02:48–04:29, ~1.7 h GPU) — same machine and build. Four measurements: (1) the matched drafter-flag sweep — 7 configurations × 2 files × 2 probes,
-c 32768, thinking off, temperature 0 / top-k 1, 700 tokens, fresh server per configuration, on a novel 149-token rate-limiter prompt chosen to avoid the acceptance inflation a textbook algorithm causes (§06). (2) the projector-at-depth pair — byte-identical 90,862-token prompts, loads ordered A-B-B-A, post-prefill probe discarded, 45 s cooldown, 5 probes per load with the drafter off and 4 with it on, plus the cooled depth ladder and the thinking-on-against-off isolation that produced the mean-draft-length result. (3)llama-perplexityat-ctk q4_0 -ctv q4_0, flags identical to the campaign's cache series, on a corpus file hash-verified identical to the fp16 andq8_0baselines'. (4) a no-GPU repetition audit of 10 long greedy transcripts: 10 of 10 clean, the single detector flag a false positive, and the one incomplete file cut by its 65,536-token budget rather than looping - 2026-08-23, effort benchmark sweep (04:40–13:19, 8.48 h GPU) — the evidence behind §09 and §10. Suite hash
1cdf54f8eb9d3f8f, 175 prompts (7 sets × 25) × 3 effort arms = 525 generations, greedy at temperature 0 / top-k 1 / seed 42,--max-tokens 16384with a rule-driven re-run of only the truncating datasets at 32,768 and-c 65536, no speculative decoding,q8_0KV. ALPACA and MT-Bench ran unscored at the time — transcripts kept, and scored on 2026-08-24 by the judge panel below. Includes the determinism check (139 of 139 byte-identical) and the three grader fixes - 2026-08-24, judge panel (zero GPU) — the evidence behind §09. The kept ALPACA and MT-Bench transcripts from the three effort arms, 150 answers, read blind by three independent Claude Opus 5 seats for 450 ratings with none missing or partial. Opaque salted identifiers; the identifier-to-arm mapping sealed in a key file no seat reads; a separate shuffle seed per seat; the standard 1–10 single-answer rubric; ratings normalized (r−1)/9×100. Arm comparison by 20,000-resample paired bootstrap over the per-prompt differences, seed 42, six comparisons. Builder, scorer and paired test:
scripts/bench/judge-panel.py. Instrument limit, stated with the result: the answers are Qwen's and the judge is Claude, so no model grades its own output — but the judge and this page's author are both Claude models, a correlated instrument, which is why three seats run and the seat spread is published - 2026-08-23, energy joins (zero GPU) — two analyses computed entirely from logs already on disk, using
scripts/power/attribute-power.py. The first joined a 500 ms power log (17,715 samples, 10:48–13:19) to the sweep's per-request timings, counting only the 115 of 150 requests whose whole window lay inside the log, and produced the drafter-off reference of 7.884 ± 0.307 J per decode token plus the sustained-power drift anomaly. The second re-integrated the 08-22 effort arms' own 1 Hz logs and reproduced all four published Wh figures to within 0.05%, adding joules per token, tokens per kWh and energy-delay product per level - 2026-08-23, power matrix (13:21:49–14:05:20, 43.5 min of logged span, 19 arms, 100% sample coverage) — thirteen of those arms carry full per-arm energy tables (the drafter ladder, three quantizations, two cache precisions, both token regimes, three depths); the other six are the two idle arms, the two concurrency arms, and the two power-cap arms that were skipped. The source of every figure in §03 and §03: settled idle with and without a server, the three-point drafter ladder, three quantizations, two cache precisions, both token regimes, three depths, and the concurrency pair. In-band NVML board power at 500 ms, three requests per arm, 700 generated tokens each. The two power-cap arms were skipped for lack of an elevated shell; the runner's own net columns used a 31.0 W idle constant rather than the measured 34.1 W loaded idle, a uniform shift of about 1% that changes no ranking, and the net columns published here are re-derived at 34.1 W
- 2026-08-23 to 08-24, quantization ladder (14:13 → 12:35, ~6 h GPU) — complete. Nine files of this model ranked on perplexity over 294,912 scored token positions each, under one identical condition set, with functional detectors beside every rung. Its gate re-ran the UD-IQ4_XS reference file and reproduced 6.5956 bit-identically twice at 0.000% drift before starting, which is what pins §08's error bars, and a same-size pair resolved a 4.27% quality gap against ±0.045 error bars — so the rig both reproduces and still discriminates. It also recorded that the same corpus tokenizes to 297,193 tokens under Qwen's tokenizer and 295,216 under Gemma's, which is the source of the tokenizer caveat. The full treatment is now in §08: the eight-rung table, the point where the perplexity curve turns, the accuracy ladder, the empty-answer audit and the paired tests. This entry's earlier "not yet on this page" caveat is retired by the two rounds below
- 2026-08-24 into 08-25, accuracy ladder over the same files — eight arms run across the day; most took under thirty minutes each (1,705 s and 1,766 s on two of them) and the bottom rung took 8,965.8 s — two and a half hours, more than four times the next-longest arm, because at 1.835 bits per weight the model stops terminating. The evidence behind §08's Mean, empty and truncation columns and behind every paired McNemar row. Frozen suite
1cdf54f8eb9d3f8f— GSM8K + HumanEval + MBPP, 25 questions each, 75 items per file — greedy at temperature 0 / seed 42,--max-tokens 16384,-c 32768,q8_0KV,reasoning_effort=low, no speculative decoding, eight files scored on the identical items so that every arm-against-arm claim is paired. Two instrument failures are recorded rather than quietly fixed: a grader crashed on a bare####with an empty tail, a shape no file above 2.15 bits had ever produced (fixed to grade it wrong; no recorded score could change, because any earlier run reaching that path would have crashed and none did); and an empty-answer count was first published from the truncation counter and was wrong — UD-IQ1_M reported zero truncations while returning nothing to five questions. Both columns are now read from the saved artifacts bywork/ladder-repcheck.py, which is why they are separate columns. A cross-model arm on the identical suite — gemma-4-12B-it-QAT-Q4_0 at 6.497 GiB — ran the same day under disclosed condition asymmetry - 2026-08-25, the four rounds that turned the ladder into a recommendation — (1) the drafter on/off pair on both candidate files: 700-token novel-code generations,
-c 32768,q8_0KV, thinking off, temperature 0, three settled probes per arm with the first post-prefill probe discarded, which found the ranking inverting between drafter-off and the shipped drafter-on recipe. (2) the full-native-window trial: three arms at-c 262144, each loaded, VRAM read as a drafter on/off pair, then deep-filled to 218,233 real tokens — 83% of the window, 424 s of prefill at 514.7 t/s — and probed with the first post-prefill probe discarded. A self-caught instrument error belongs with it: the probe script averagedprompt_nacross settled probes, which read 4 because probes 2 and 3 hit the server's prompt cache; the real fill is in the server's ownprompt eval timeline and the decode figures are at depth as published. (3) the requirement sweep behind §08's memory-and-speed table and §08's depth series: three files — UD-IQ4_XS, UD-Q3_K_XL, UD-Q2_K_XL — × three windows (-c 32768,65536,131072) × the drafter on and off, plus a fourth window at-c 98304drafter-on, each window deep-filled to about 90% of itself with wikitext-2-raw prose, board VRAM read at that depth, first post-prefill probe discarded and two settled probes taken. At-c 131072every file was loaded twice from scratch with five settled probes per load, and at-c 98304UD-Q3_K_XL was loaded twice, which is what settled the ordering after it had been published, withdrawn and reinstated (the episode is recorded at §08 rather than only in a log). The condition on that round is stated wherever its numbers appear: the fill is prose, the drafter's confidence gate is content-dependent, and the ordering is therefore published as a measurement on prose rather than as a property of the files. (4) the execute probe (§08), which used no GPU at all: every file's probe-A JavaScript, already on disk, rebuilt with the prompt's own prefix and run undernode v24.15.0at a 15-second timeout. Six of ten programs run; the three files below 2.481 bits per weight do not, and neither does the cross-model gemma arm at 4.651. It moved a published number — the “functional floor” had been read off four detectors that never executed anything — and it is the cheapest round in the whole campaign. The same day's matched--parallelpair closed the concurrency question (§03) - Methodology version. This page follows the campaign methodology as of 2026-08-25: 26 numbered rules plus four review gates, a two-voice writing law, a recipes-first structure, and a standardized-energy-metrics mandate. Two rules were amended by the work in this revision and the amendments are visible in §08's tables rather than only in a log. Rule 20 now requires empty answers and truncations as separate columns, because the truncation counter is structurally blind to an answer that terminates normally and returns nothing — a blindness that produced a false statement in an earlier draft of that very table. Rule 25 now requires a sweep to be run at the recipe that ships: the whole quantization ladder was measured with the drafter off and gave the wrong file ordering for the configuration every recipe on this page uses. The methodology is pinned here for the same reason the build is pinned to commit
0adcc3bb5: a page that names a rule number without naming the ruleset cannot be checked later
The twenty-nine places this page corrected itself
Each entry is a claim an earlier edition of this page made and this edition retires, with the measurement that forced it. They are kept visible rather than quietly deleted, because a page that hides its corrections teaches its readers to trust the wrong things.
- "ALPACA and MT-Bench are unscored by design." They were unscored by circumstance — no judge was available — and calling that a design choice dressed a gap up as a decision. A blind three-seat panel scored all 150 kept answers on 2026-08-24, the headline index became a seven-set index, and the pair turned out to be the only instrument in the campaign that separates the effort levels at all:
xhighlost tomediumon 14 of 25 MT-Bench prompts. The gap that remains is real and now stated as a gap — the judge and this page's author are both Claude models. §09, §15 - "The built-in draft head costs no VRAM." It costs 1,008 MiB fixed, plus 5,120 B per window token, plus another 898 MiB at n-max 10 — about 1.8 GiB at a 163,840-token window. §05, §06
- The vision projector was budgeted at its file size (~0.9 GB). It occupies 1,138 MiB resident, measured four times. §04, §05
- Large windows were declared safe on the strength of short probes — 196,608 shipped with a claimed ~3 GiB of slack, and 262,144 shipped as verified. Filled properly, 196,608 leaves 415 MiB, the real text-only ceiling with the drafter is 180,224, and 262,144 needs the drafter off. §05, §08
- "Past the resident ceiling, speed degrades progressively." It collapses, and only when real tokens land there: 3.82× at a 91,000-token fill, from 30.6 to 8.0 t/s. §05
- The projector was thought to cost about 15% of decode at depth. Paired probes measured 0.04–0.09%: it costs memory and nothing else. §06
- One universal 0.7 bandwidth-efficiency constant. It is format-specific: 0.70 for K-quants, 0.649 measured for UD-IQ4_XS. §04, §07
- Perplexity was said to score "~330k token positions". It scores 294,912 — 36 chunks × 8,192 — and that count is tokenizer-bound, so it is a count of scored windows rather than of corpus tokens. §08, §10
- 1080p and 1440p image costs were interpolated and ran 28–30% high. The measured law is tokens = pixels ÷ 1024 + 2; a 1440p shot costs 3,602, not "~4.7k". §12
- 81.7 t/s was demoted to "best case, not reproducible on real work". It reproduces at 81.71 on a deliberately novel code prompt; the demotion was an unlabelled token regime. §06
- The drafter flags were split per file. A matched sweep across both files ranks n-max 10 / p-min 0.5 first on each; acceptance lands within 1.6 points at six of the seven configurations and 3.7 at the seventh, so it is a property of the draft head and not of the quantization. §06
- Configurations were ranked by draft acceptance. Across configurations acceptance inverts: the 96.5%-acceptance setting is the slowest speculating one. Mean draft length — strictly, accepted tokens per verification pass — is what ranks them. §06
- "Acceptance never moves with depth." It rises, 0.80 → 0.92 across a 60× span, in two independent series. §06
- The agent-depth band was published as 34–44 t/s. That was the reasoning stream; the tokens an agent receives run 65–70 t/s at the same depth. §06, §03
- "The board pulls a flat ~344 W." Sustained draw is a range — 306–341 W drafter-off — drifting 11.7% with its own SM clock at constant throughput, constant temperature and constant memory clock. §03
- Idle was published as 33.2 W with no server against 30.7–31.1 W loaded — physically backwards, and the reason was a board still cooling. Settled with little on screen: 29.9 W with no server, 34.1 W with the model resident. §03
- "Never use a
q4_0KV cache." It costs +0.693% perplexity against fp16 — real, super-linear in bits, and smaller than the gap between two respectable 4-bit weight files. A knowing trade, not a prohibition. §06, §08 - "16,384 proved sufficient for every
xhighthought." True on GSM8K only. On a seven-set suite the same cap truncatedxhigheight times, and three prompts exceeded even 32,768. §09, §10 - An effort sweep read as "quality falls as effort rises" (81.3 / 80.5 / 77.3). The entire penalty was the token cap; at 32,768 the arms read 82.1 / 80.5 / 81.3, a tie. §09, §10
- "
lowships broken code, 2 of 2 runs" was printed as a flat claim. A second campaign saw the fatal defect atmediuminstead, and on the 175-prompt suitelowwas the highest-scoring arm. What is solid is that a fatal defect is a live risk at every level at n=1. §09 - "Two concurrent requests gain +60.3% aggregate throughput." That arm ran with the drafter off, and this page quarantined it against an older ~+11% drafter-on figure. The matched drafter-on pair ran on 2026-08-25 and the quarantine is lifted: +22.0% aggregate and −35.4% per slot (82.98 → 101.25 t/s aggregate, 85.79 → 55.41 per slot), acceptance unmoved at 0.618 → 0.620. The older figure was much closer to right, for the reason the quarantine named — the drafter has already saved most of the repeated reading of the model that batching would otherwise save. Every recipe still ships
--parallel 1, now on evidence rather than on caution. §03, §06 - "Quality falls about 1% of perplexity per gigabyte down to 2.9 bits, then five times that — so stop at about 9 GiB." The advice stands and the reasoning behind it did not. The per-gigabyte figures were right, but 2.9 bits per weight is not the point after which quality degrades — it is the last rung that still performs like the full-quality file. Accuracy on 75 paired items is flat from 4.22 bits through 2.91 and still a tie at 2.48, and empty answers are exactly zero down to 2.91. §08
- The whole quantization ladder was swept with the drafter off, and it gave the wrong ordering for the configuration every recipe ships. Drafter off, UD-Q2_K_XL is 7.8% faster than UD-IQ4_XS and the ladder reads "swap your daily file". Drafter on, UD-IQ4_XS is 12.9% faster despite being 45% larger, because the draft head degrades with bit width. Sweep at the recipe you ship. §08
- Full native context on UD-IQ4_XS was said, from arithmetic alone, to need the card free of any graphical session. It is now measured: deep-filled to 218,233 real tokens the 4-bit file reaches 23,821 MiB with speculation already off, leaving 755 MiB — 1,041 MiB short of the 1,796 MiB this page reserves for a desktop. UD-Q2_K_XL holds the same window with the drafter at 1,717 MiB of slack and decodes 34% faster there. §08, §05
- "Board watts fell as the drafter got wider — the win compounds." They did not. That claim came from the whole-window mean watts (325.2 → 308.4 → 302.4), which fall because each request also contains a short low-power prefill and setup segment drawing 101–139 W, and that segment is a larger share of a shorter run. Power while actually decoding is flat: 344.6 → 341.7 → 341.0 W, a 1.0% change. The 2.52× energy saving is the 2.50× throughput gain and nothing else — a duller mechanism and the true one, and the same distinction is why joules per token divide by the decode watts, not by the column headed mean load. §03
- The word "cliff" named two unrelated things. It is this page's term for the collapse that happens when a window overflows the card — a real discontinuity, 3.8× at a 91k fill — and it was also being used for the quality step between
lowandmedium, which is a step in a small sample and not a discontinuity in anything. The second use is gone. §05, §09 - The short-context speed band was published as 71–84 t/s. Its lower endpoint came from one live check of the vision recipe whose token regime was never recorded — the one thing this page says no speed number may lack. Rebuilt from two measurements that do carry their regime: 83.5 t/s on a novel code prompt at
-c 32768and 86.3 at a 1,458-token fill under the cooled protocol, both answer tokens at the shipped flags. The band is now 83–86 t/s, and the 71.3 figure is kept only as a labelled loose end. §03, §12 - The drafter-off floor was printed as 39.7–43.8 t/s, and prose speculation was credited with 1.16×. Both were inherited from a summary bullet in the campaign log that this page checked against the log's own table and found wrong. 43.80 t/s is a drafter-ON prose figure, so it cannot be a floor endpoint; the measured drafter-off floor for this file is 41.46–42.97 across five contents and both token regimes. And 43.80 ÷ 41.55 is 1.05×, not 1.16× — the 1.16× belongs to
n4/p0.75's 48.35 t/s on the same content, which is what the log's own "best vs floor" column says. §06 - Three memory budgets converted MiB to GiB by dividing by 1000. The 262,144-token window costs 9,984 MiB (9.75 GiB, not "9.98"), that configuration totals 23,216 MiB (22.7 GiB, not "23.2"), and the worked vision example totals 21,556 MiB (21.05 GiB, not "21.6"). The slack figures those pages quoted were right; the GiB labels were not. §05
What was not measured
Stated plainly so that nobody mistakes an omission for a result. Each entry names the measurement that would close it and roughly what it would cost.
| Gap | Why it is missing | What would close it, and its price |
|---|---|---|
| Power capping | The probe returned exit 4 Insufficient Permissions; the cap was never changed and remains the stock 350 W | nvidia-smi -pl 250 and -pl 300 from an elevated shell, then re-run three decode arms. About twenty minutes. This is the strongest untested hypothesis on this machine, because the board demonstrably spent 6–12% more power at clocks that bought zero extra tokens |
| Two concurrent requests with the drafter on — CLOSED 2026-08-25 | The one concurrency arm that had ever run used the drafter off, and it contradicted an earlier measurement, so both were quarantined | The matched pair ran: UD-IQ4_XS at n10/p0.5, three reps each, +22.0% aggregate and −35.4% per slot with acceptance unmoved (§06, §03). What is still open is the energy half — the pair measured throughput, not board power, so the retired 5.19 J/token figure has no drafter-on replacement. One re-run of the two arms under the power logger would close it. Minutes |
| Speed of any sub-4-bit file on a 16 or 12 GB card — narrowed 2026-08-25 | Quality was measured on every rung and does not depend on the card. The VRAM requirement is now measured too, at three files × three windows × the drafter on and off, deep-filled (§08) — and because a requirement is a property of the model, the window and the flags, it transfers to any card. So the fit half of this gap is closed and only the arithmetic of card-total-minus-reserve is derived. What remains open is speed, which belongs to a card's memory bandwidth and to nothing else | One 16 GB card and one 12 GB card, loading UD-Q3_K_XL and UD-Q2_K_XL at a stated window and timing decode at depth. Nothing on this machine can close it — it owns one 24 GB card (§08) |
| Board power for any file below 4 bits | The 19-arm power matrix predates the ladder; its three quantization arms are all 4-bit. Nothing on this page prices the energy of UD-Q2_K_XL or anything smaller | Re-run the drafter ladder's three arms on UD-Q2_K_XL under the same in-band NVML protocol. About twenty minutes, and it would say whether the file that is 34% faster at the full window is also cheaper per token |
| A read-then-edit task against existing code | Every scored generative task here is greenfield. One real repair task did run, inside an agentic pipeline validation: 22 model calls, 9 m 36 s, 0 of 96 target tests fixed, all 561 regression tests still passing. The 10-task × 3-effort sweep that would have made it a measurement was cut on a ~4-hour time gate against a ~8.5-hour projection | That cut sweep, or a smaller suite of repair tasks against real files judged blind at n≥10 per effort level. Several hours, and it is the measurement that would let this page speak about the work it recommends the model for |
| Blind, repeated quality judging | The 08-22 pages were graded blind but only n=2 per level; the 08-23 single runs were n=1 and not blind. Neither is enough | n≥10 per level on one task with blind judging. Separating a real effort-to-defect effect from sampling noise needs that and nothing less |
| The vision critique loop's control | The same question was never asked with the image withheld | One extra request per demonstration. Minutes — and without it, no claim on this page about the model "seeing" a defect is more than plumbing |
The quality cost of --image-max-tokens 1024 | The saving was measured (3.6×); the consequence was not | The same image at both budgets, asked a question whose answer depends on fine detail. Minutes |
Long-context retrieval with a q4_0 cache | Perplexity at -c 8192 cannot see a retrieval failure at 200,000 tokens; that is the actual argument against a 4-bit cache and it remains untested | A needle-retrieval test at 200k with both cache precisions. An evening |
| System RAM under either load mode | Never measured on this machine at all. The 15 GB saving attributed to --load-mode none is an inherited figure, not a reading taken here | One working-set reading per load mode. Minutes |
| Any drafter but the built-in head | DFlash2 was measured and lost; no other external drafter was tried | Out of scope rather than untested-by-accident. Read the absence as untested |
| Any machine but this one | Every measured row on this page is the same RTX 3090 on the same driver | Nothing here can close it. The non-3090 rows in §07 are arithmetic and are labelled as such |
| ALPACA and MT-Bench scores — CLOSED 2026-08-24 | Both sat unscored for want of a judge. A blind three-seat Claude Opus 5 panel read all 150 kept answers on 2026-08-24; the scores are in §09 and the suite now reports a seven-set index | What remains open is stronger independence: the judge and this page's author are both Claude models. Closing that needs a judge from a different vendor, or human raters. The transcripts, the sealed key and all 450 ratings are kept, so any other judge can be run over exactly the same answers and compared |
| A second-vendor or human judge | Opened by the line above. One judge family has read these answers; nothing tests whether a different family would rank the arms the same way | Re-run judge-panel.py's packets through another model, or a small human panel. The packets are built and blinded already, so the cost is the judging itself |
| Wall power, cost per answer, and carbon | Only in-band board power was ever measured. The power supply, processor, memory, drives and display are excluded | A plug meter. Until then, no figure on this page may be called system power or divided into an electricity bill |
| Energy for four of the seven benchmark sets | GSM8K, ALPACA, MeetingBank and MT-Bench ran entirely outside the window the power logger covered | Re-run those sets with the logger active. About an hour |
| An image-attachment matrix at scale | Each agent got one image and one question — enough to say the plumbing works, not enough to rank them | Several resolutions and multi-image requests per agent. An afternoon |
| Quality ranking between quantizations at usable sample size | §10's own arithmetic says it needs thousands of questions per file | About 12–60 hours and 4–20 kWh per file. This is why the page ranks files by perplexity instead |
The reproduction check
One command, one number, one pass band. If your machine returns something inside the band, your build and your card behave like the one every measurement on this page came from. If it returns something below the band, the diagnostic table in §11 tells you which of the two half-speed failures you have.
:: 1. start the server with the drafter OFF — this measures the bandwidth floor, :: which is the one number that does not depend on your content: llama-server.exe -m Qwen3.8-27B-UD-IQ4_XS.gguf -c 32768 -ngl 99 --parallel 1 ^ -ctk q8_0 -ctv q8_0 --spec-type none --jinja --host 127.0.0.1 --port 1235 :: 2. send TWO identical requests and discard the first (it inherits the cold-clock :: state that section 11 measures at up to 25%). Each request: 700 tokens, :: temperature 0, top-k 1, thinking OFF, a short novel code prompt. :: ONE LINE - a line break inside the -d body will not survive cmd: curl -s http://127.0.0.1:1235/v1/chat/completions -H "Content-Type: application/json" -d "{\"model\":\"qwen/qwen3.8-27b\",\"max_tokens\":700,\"temperature\":0,\"top_k\":1,\"chat_template_kwargs\":{\"enable_thinking\":false},\"messages\":[{\"role\":\"user\",\"content\":\"Write a sliding-window rate limiter class in Python.\"}]}" :: 3. read the SECOND run's decode speed from the server's own log line: :: eval time = ..... ms / 700 tokens ( ... ms per token, NN.NN tokens per second) :: Use that number. Do NOT use tokens divided by wall-clock (section 10).
Expected: 42.97 t/s (measured 2026-08-23, RTX 3090, driver 596.36, llama.cpp build 10502, UD-IQ4_XS). Pass band: 41.7 to 44.3 t/s — that is ±3%, and the 3% is not arbitrary: it is this campaign's measured spread of the drafter-off decode floor across five different contents and both token regimes (§01). One clarification, because the two numbers look like they disagree: the floor across contents runs 41.46–42.97 t/s, and the low end of that belongs to a different prompt. This check pins the prompt, so the band around 42.97 is the right one to judge your own run by. Inside it, your setup matches. Between about 34 and 41 t/s, suspect a different file, a different build, or fp16 rather than 8-bit cache. Below about 35 t/s something is wrong, and there are only two common causes: the -ngl off-by-one reads 29.8 t/s on this file, and a memory spill reads 20–35 (§11). Above the band, check that the drafter really is off — --spec-type none makes the server log unused tensor blk.64.*, and if that line is missing you are measuring a speculating server, which starts at about 73 t/s.