Jose Romero
← All letters
Prefer email? Read this on Substack

The 29 GB version lost to the 17 GB one

Yesterday the new Qwen 3.8-27B swept my benchmark at full precision but was three times too slow to daily-drive. So I did the obvious next thing: I shrank it. Four quantization levels, from the full 55 GB down to 11 GB, plus the exact build most people actually run, the one Ollama gives you by default. Same four one-shot tests, thinking off, tools off, no retries. This letter has the complete data, so you can check every claim.

See it for yourself

Everything in this video is checkable. Open the full results page → every prompt, all 20 playable outputs across the five quants, and the numbers.

What a quant is, in one breath

A model is a giant list of numbers, 27 billion of them here. Quantization just changes how carefully each number is written down. Full precision (BF16) writes pi as 3.14159265. Q8 writes 3.1416. Q4 writes 3.14. Q2 writes "3-ish". Same photo, lower JPEG quality each time. The question is whether you can see the compression.

Ground rules, so the comparison is fair

The tale of the tape

ContestantAvg speedWhole examLoadPeak memory
BF16 (54.7 GB)8.8 tok/s57m 50s54 s71 GB
Q8_0 (29 GB)16.8 tok/s28m 07s34 s48 GB
Q4 Unsloth (17.9 GB)20.2 tok/s24m 23s14 s39 GB
Q4_K_M Ollama (17 GB)18.8 tok/s26m 32s6.5 s33 GB
Q2 Unsloth (10.7 GB)27.9 tok/s17m 27s14 s32 GB

Peak memory is total system RAM at the highest point of the run (my rig shares one 128 GB unified pool between CPU and GPU, so there is no separate VRAM number; the ~9 GB OS baseline is included).

The speed ladder behaved perfectly. Local speed is memory bandwidth: every token reads all the model's weights from memory, and my chip's bus peaks at 256 GB/s on paper (about 215 measured). Halve the gigabytes, roughly double the speed. For scale, the average adult reads at about 240 words per minute, roughly 5 tokens per second: full precision is barely faster than you read.

SpeedFeels likeAn 8,000-token answer takes
~5 tok/sAverage human reading speed, the floor of usable~25 min
8.8 tok/s (BF16)Just above reading pace, painful for long code15m 09s
16.8 tok/s (Q8)Waiting, not reading7m 56s
19-20 tok/s (both Q4s)Same league, a notch quicker~7m
27.9 tok/s (Q2)The daily-driver zone, same speed as my Qwen 3.64m 47s
~70 tok/s (cloud)The frontier-API experience~2m

Quality did NOT behave. I expected a clean staircase, each smaller quant a little worse. That is not what happened.

Test 1: Blockfall (a one-shot Tetris clone)

Only full precision shipped a working game. Q8's board broke on the first hard drop, Q4 rendered no shapes at all, Ollama's rotation was scrambled, and Q2's game never started. Winner: BF16.

ModelTimePrompt tokOutput tokSpeed
BF16 🏆14m 54s2147,9738.92 tok/s
Q8_05m 23s2145,46216.89 tok/s
Q4 Unsloth6m 13s2147,51020.12 tok/s
Q4_K_M Ollama6m 55s2147,95719.2 tok/s
Q2 Unsloth5m 08s2148,57627.76 tok/s

Test 2: Eruption (a volcano physics sim)

The upset begins. BF16 looked the nicest but its lava ran uphill. Q8 was broken outright. Unsloth's Q4 was clickable but had the same uphill-lava physics. Ollama's Q4 put ten thousand particles on screen with lava that actually flowed downhill, wind that responded, and a live FPS counter. Winner: Q4_K_M Ollama.

ModelTimePrompt tokOutput tokSpeed
BF1617m 31s2389,2408.79 tok/s
Q8_08m 45s2388,70716.58 tok/s
Q4 Unsloth6m 47s2388,21320.16 tok/s
Q4_K_M Ollama 🏆7m 15s2387,80717.98 tok/s
Q2 Unsloth5m 07s2388,60227.97 tok/s

Test 3: The Ledger (a messy CSV that must be cleaned correctly)

The trap: duplicate order IDs, three date formats, refunds, N/A rows. BF16 got everything right, with charts. Q8, twice its 4-bit siblings' size, shipped a page with console errors. Unsloth's Q4 got the numbers right but its chart was broken. Ollama's Q4 got every number right AND shipped the nicest page, in under 8 minutes instead of BF16's 16. Q2 died on a syntax error. Winner: Q4_K_M Ollama, over BF16 on speed at equal correctness.

ModelTimePrompt tokOutput tokSpeed
BF1616m 12s1,1628,6688.92 tok/s
Q8_09m 27s1,1629,71117.12 tok/s
Q4 Unsloth6m 43s1,1628,24320.42 tok/s
Q4_K_M Ollama 🏆7m 40s1,1628,63618.93 tok/s
Q2 Unsloth5m 01s1,1628,34727.7 tok/s

Test 4: Blind Artist (draw a scene in raw SVG, then sign your own name)

Unsloth's Q4 drew the best forklift dog of the whole board AND was the only model that signed its own name. Which brings us to the saga: the prompt asks each model to sign the artwork with its own model name, and four out of five wrote "Claude". This time I opened the raw output file on camera: it is right there in the SVG, generated entirely on my machine, no cloud involved. It still points to Claude-generated text somewhere in the training data, and it still makes me laugh. Winner: Q4 Unsloth.

ModelTimePrompt tokOutput tokSpeedSigned
BF169m 12s1924,8498.77 tok/s“Claude”
Q8_04m 31s1924,53216.67 tok/s“Claude”
Q4 Unsloth 🏆4m 38s1925,65520.28 tok/s“Qwen” ✓
Q4_K_M Ollama4m 41s1925,38519.21 tok/s“Claude”
Q2 Unsloth2m 09s1923,65528.22 tok/s“Claude”

The final board

ContestantTest wins
BF16 (54.7 GB)1
Q8_0 (29 GB)0
Q4 Unsloth (17.9 GB)1
Q4_K_M Ollama (17 GB)2
Q2 Unsloth (10.7 GB)0

Read that Q8 row again. The 29 GB quant, double the size of the 4-bits, won nothing and broke tests both 17 GB builds passed. Quantization is not a straight line, and on this exam Ollama's default download was the best value on the whole board. My verdict: I'm switching my daily driver experiment to Q4.

Will it fit on YOUR GPU?

Rule of thumb: the model file plus roughly 2 to 3 GB (context cache and overhead) must fit in VRAM for full speed. If it does not fit, Ollama and llama.cpp split the layers between GPU and system RAM, every token still walks every layer, and the CPU-resident slice sets the pace. It still runs, it just quietly stops being fast.

QuantVRAM wantedFits fully on (examples)
Q2 (10.7 GB)~13-14 GB16 GB cards (4060 Ti 16GB, 5070 Ti); tight on 12 GB
Both Q4s (17-17.9 GB)~20-21 GB24 GB cards (RTX 3090, 4090, RX 7900 XTX)
Q8 (29 GB)~32 GB32 GB cards only (RTX 5090)
BF16 (54.7 GB)~58+ GBNo single consumer GPU, unified-memory territory

And past gaming cards there are exactly two roads: pro GPUs keep the bandwidth and charge used-car money for it (48 to 96 GB at 800 to 1,800 GB/s), or unified memory buys huge capacity at a fraction of the bandwidth (my 128 GB box at 256 GB/s spec, Apple silicon up to 512 GB). Everything fits, it just runs slower per token. Bandwidth is tok/s. That trade is this whole channel's rig story.

Tomorrow: thinking mode, done properly

Everything above ran with thinking off, which means I deliberately turned OFF Qwen 3.8's default mode. I got a sneak peek of Q4 with thinking enabled and the result was so much better it made the whole ladder look like a warm-up. Next: the same exam with thinking at official settings, across multiple reasoning levels. Day 30.

My rig

Gear in this one

Some links on this page are affiliate links. If you buy through them I may earn a small commission at no extra cost to you. It helps support the work. As an Amazon Associate I earn from qualifying purchases.

Sources

I am documenting a life of interest and curiosity, tech, money, health, and family included. New video daily.

Get the letters

Every post goes out free on Substack. What I build, break, and fix, written up in plain English. No spam, unsubscribe anytime.

Subscribe on Substack