The 29 GB version lost to the 17 GB one
Yesterday the new Qwen 3.8-27B swept my benchmark at full precision but was three times too slow to daily-drive. So I did the obvious next thing: I shrank it. Four quantization levels, from the full 55 GB down to 11 GB, plus the exact build most people actually run, the one Ollama gives you by default. Same four one-shot tests, thinking off, tools off, no retries. This letter has the complete data, so you can check every claim.
See it for yourself
Everything in this video is checkable. Open the full results page → every prompt, all 20 playable outputs across the five quants, and the numbers.
What a quant is, in one breath
A model is a giant list of numbers, 27 billion of them here. Quantization just changes how carefully each number is written down. Full precision (BF16) writes pi as 3.14159265. Q8 writes 3.1416. Q4 writes 3.14. Q2 writes "3-ish". Same photo, lower JPEG quality each time. The question is whether you can see the compression.
Ground rules, so the comparison is fair
- Instruct mode for all five, thinking OFF. Not just configured: verified, 0 thinking tokens in every output. (Thinking is actually Qwen 3.8's default mode; it gets its own video next, at the official thinking settings.)
- Identical samplers everywhere, the official card's recommended instruct settings (temperature 0.7, top_p 0.8, top_k 20, presence 1.5). Ollama ships different defaults; I overrode them so the only variable is the build and the quant.
- Tools disabled on both stacks. One prompt, one shot, no retries, seed 42, 60-minute cap per test.
- One caveat on the Ollama leg: its default pull is Q4_K_M (17 GB) while my Unsloth 4-bit is UD-Q4_K_XL (17.9 GB). Different 4-bit recipes, so where outputs differ, the quant format itself can be the cause, not the runtime.
The tale of the tape
| Contestant | Avg speed | Whole exam | Load | Peak memory |
|---|---|---|---|---|
| BF16 (54.7 GB) | 8.8 tok/s | 57m 50s | 54 s | 71 GB |
| Q8_0 (29 GB) | 16.8 tok/s | 28m 07s | 34 s | 48 GB |
| Q4 Unsloth (17.9 GB) | 20.2 tok/s | 24m 23s | 14 s | 39 GB |
| Q4_K_M Ollama (17 GB) | 18.8 tok/s | 26m 32s | 6.5 s | 33 GB |
| Q2 Unsloth (10.7 GB) | 27.9 tok/s | 17m 27s | 14 s | 32 GB |
Peak memory is total system RAM at the highest point of the run (my rig shares one 128 GB unified pool between CPU and GPU, so there is no separate VRAM number; the ~9 GB OS baseline is included).
The speed ladder behaved perfectly. Local speed is memory bandwidth: every token reads all the model's weights from memory, and my chip's bus peaks at 256 GB/s on paper (about 215 measured). Halve the gigabytes, roughly double the speed. For scale, the average adult reads at about 240 words per minute, roughly 5 tokens per second: full precision is barely faster than you read.
| Speed | Feels like | An 8,000-token answer takes |
|---|---|---|
| ~5 tok/s | Average human reading speed, the floor of usable | ~25 min |
| 8.8 tok/s (BF16) | Just above reading pace, painful for long code | 15m 09s |
| 16.8 tok/s (Q8) | Waiting, not reading | 7m 56s |
| 19-20 tok/s (both Q4s) | Same league, a notch quicker | ~7m |
| 27.9 tok/s (Q2) | The daily-driver zone, same speed as my Qwen 3.6 | 4m 47s |
| ~70 tok/s (cloud) | The frontier-API experience | ~2m |
Quality did NOT behave. I expected a clean staircase, each smaller quant a little worse. That is not what happened.
Test 1: Blockfall (a one-shot Tetris clone)
Only full precision shipped a working game. Q8's board broke on the first hard drop, Q4 rendered no shapes at all, Ollama's rotation was scrambled, and Q2's game never started. Winner: BF16.
| Model | Time | Prompt tok | Output tok | Speed |
|---|---|---|---|---|
| BF16 🏆 | 14m 54s | 214 | 7,973 | 8.92 tok/s |
| Q8_0 | 5m 23s | 214 | 5,462 | 16.89 tok/s |
| Q4 Unsloth | 6m 13s | 214 | 7,510 | 20.12 tok/s |
| Q4_K_M Ollama | 6m 55s | 214 | 7,957 | 19.2 tok/s |
| Q2 Unsloth | 5m 08s | 214 | 8,576 | 27.76 tok/s |
Test 2: Eruption (a volcano physics sim)
The upset begins. BF16 looked the nicest but its lava ran uphill. Q8 was broken outright. Unsloth's Q4 was clickable but had the same uphill-lava physics. Ollama's Q4 put ten thousand particles on screen with lava that actually flowed downhill, wind that responded, and a live FPS counter. Winner: Q4_K_M Ollama.
| Model | Time | Prompt tok | Output tok | Speed |
|---|---|---|---|---|
| BF16 | 17m 31s | 238 | 9,240 | 8.79 tok/s |
| Q8_0 | 8m 45s | 238 | 8,707 | 16.58 tok/s |
| Q4 Unsloth | 6m 47s | 238 | 8,213 | 20.16 tok/s |
| Q4_K_M Ollama 🏆 | 7m 15s | 238 | 7,807 | 17.98 tok/s |
| Q2 Unsloth | 5m 07s | 238 | 8,602 | 27.97 tok/s |
Test 3: The Ledger (a messy CSV that must be cleaned correctly)
The trap: duplicate order IDs, three date formats, refunds, N/A rows. BF16 got everything right, with charts. Q8, twice its 4-bit siblings' size, shipped a page with console errors. Unsloth's Q4 got the numbers right but its chart was broken. Ollama's Q4 got every number right AND shipped the nicest page, in under 8 minutes instead of BF16's 16. Q2 died on a syntax error. Winner: Q4_K_M Ollama, over BF16 on speed at equal correctness.
| Model | Time | Prompt tok | Output tok | Speed |
|---|---|---|---|---|
| BF16 | 16m 12s | 1,162 | 8,668 | 8.92 tok/s |
| Q8_0 | 9m 27s | 1,162 | 9,711 | 17.12 tok/s |
| Q4 Unsloth | 6m 43s | 1,162 | 8,243 | 20.42 tok/s |
| Q4_K_M Ollama 🏆 | 7m 40s | 1,162 | 8,636 | 18.93 tok/s |
| Q2 Unsloth | 5m 01s | 1,162 | 8,347 | 27.7 tok/s |
Test 4: Blind Artist (draw a scene in raw SVG, then sign your own name)
Unsloth's Q4 drew the best forklift dog of the whole board AND was the only model that signed its own name. Which brings us to the saga: the prompt asks each model to sign the artwork with its own model name, and four out of five wrote "Claude". This time I opened the raw output file on camera: it is right there in the SVG, generated entirely on my machine, no cloud involved. It still points to Claude-generated text somewhere in the training data, and it still makes me laugh. Winner: Q4 Unsloth.
| Model | Time | Prompt tok | Output tok | Speed | Signed |
|---|---|---|---|---|---|
| BF16 | 9m 12s | 192 | 4,849 | 8.77 tok/s | “Claude” |
| Q8_0 | 4m 31s | 192 | 4,532 | 16.67 tok/s | “Claude” |
| Q4 Unsloth 🏆 | 4m 38s | 192 | 5,655 | 20.28 tok/s | “Qwen” ✓ |
| Q4_K_M Ollama | 4m 41s | 192 | 5,385 | 19.21 tok/s | “Claude” |
| Q2 Unsloth | 2m 09s | 192 | 3,655 | 28.22 tok/s | “Claude” |
The final board
| Contestant | Test wins |
|---|---|
| BF16 (54.7 GB) | 1 |
| Q8_0 (29 GB) | 0 |
| Q4 Unsloth (17.9 GB) | 1 |
| Q4_K_M Ollama (17 GB) | 2 |
| Q2 Unsloth (10.7 GB) | 0 |
Read that Q8 row again. The 29 GB quant, double the size of the 4-bits, won nothing and broke tests both 17 GB builds passed. Quantization is not a straight line, and on this exam Ollama's default download was the best value on the whole board. My verdict: I'm switching my daily driver experiment to Q4.
Will it fit on YOUR GPU?
Rule of thumb: the model file plus roughly 2 to 3 GB (context cache and overhead) must fit in VRAM for full speed. If it does not fit, Ollama and llama.cpp split the layers between GPU and system RAM, every token still walks every layer, and the CPU-resident slice sets the pace. It still runs, it just quietly stops being fast.
| Quant | VRAM wanted | Fits fully on (examples) |
|---|---|---|
| Q2 (10.7 GB) | ~13-14 GB | 16 GB cards (4060 Ti 16GB, 5070 Ti); tight on 12 GB |
| Both Q4s (17-17.9 GB) | ~20-21 GB | 24 GB cards (RTX 3090, 4090, RX 7900 XTX) |
| Q8 (29 GB) | ~32 GB | 32 GB cards only (RTX 5090) |
| BF16 (54.7 GB) | ~58+ GB | No single consumer GPU, unified-memory territory |
And past gaming cards there are exactly two roads: pro GPUs keep the bandwidth and charge used-car money for it (48 to 96 GB at 800 to 1,800 GB/s), or unified memory buys huge capacity at a fraction of the bandwidth (my 128 GB box at 256 GB/s spec, Apple silicon up to 512 GB). Everything fits, it just runs slower per token. Bandwidth is tok/s. That trade is this whole channel's rig story.
Tomorrow: thinking mode, done properly
Everything above ran with thinking off, which means I deliberately turned OFF Qwen 3.8's default mode. I got a sneak peek of Q4 with thinking enabled and the result was so much better it made the whole ladder look like a warm-up. Next: the same exam with thinking at official settings, across multiple reasoning levels. Day 30.
My rig
Gear in this one
- GMKtec EVO-X2 AI Mini PC, Ryzen AI Max+ 395, 128GB (my exact build)
- GMKtec EVO-X3, the newer tower model with OCuLink for eGPUs
Some links on this page are affiliate links. If you buy through them I may earn a small commission at no extra cost to you. It helps support the work. As an Amazon Associate I earn from qualifying purchases.
Sources
- The quants: Qwen 3.8-27B GGUF (Unsloth) · qwen3.8:27b on Ollama
- The modes and recommended settings: official Qwen 3.8 model card
- The baseline fight: yesterday's BF16 showdown
I am documenting a life of interest and curiosity, tech, money, health, and family included. New video daily.
