Jose Romero
← All letters
Prefer email? Read this on Substack

Qwen 3.8 just dropped and it can SEE

Two weeks ago I covered a 400-gigabyte AI file almost nobody on Earth can run. Today its little sibling arrived, and this one is for the rest of us: Qwen 3.8-27B, free, Apache 2.0, small enough for a gaming GPU, and, for the first time in my local lineup, able to see.

The drop

The announcement post went up this morning: "We promised open weights for Qwen3.8. Now, time to meet them." The weights hit Hugging Face and ModelScope within the hour. It's a 27 billion parameter multimodal dense model with 262K native context (extensible to 1M via YaRN), and the 4-bit build runs on a 17 to 19 GB setup: an RTX 5080, a 4090, or a Mac with 24 GB of RAM.

The ecosystem moved the same day. When I first checked the quantized variants on Hugging Face there were around 40. When I refreshed live on camera: 349. The most popular are Unsloth's GGUF builds, a GGUF being the MP3 of AI models, the ready-to-play file, with sizes from roughly 58 GB at 16-bit down to the 17-19 GB 4-bit sweet spot. Blackwell GPU owners even get a dedicated NVFP4 build tuned for extra speed.

The headline: it has eyes

Every popular local model I've tested this month, including the beloved Qwen 3.6, was blind. There were ways to hack vision onto it, but nothing native. Qwen 3.8-27B ships with image and video understanding out of the box: STEM diagrams, documents, screenshots, hour-scale videos. Fully offline, so your files never leave your machine.

Two more quality-of-life wins from the model card: thinking effort levels (none, low, medium, extra high, where 3.6 ran everything at one effort), and honest guidance from Qwen about the 1M context mode, which costs tokens per second and should be saved for genuinely big jobs.

The benchmarks, with the grain of salt attached

The launch numbers are published by Alibaba, so treat them accordingly until independent testing lands. What they claim: a significant jump over Qwen 3.6-27B on nearly every score (their own software-engineering benchmark goes from 49 to 79), and, the part making headlines, this free 27B trades wins with Opus 4.6 Max, a frontier model, on several agent and coding benchmarks. It doesn't win everywhere. That it's close at all, on a model with unlimited tokens and zero data leaving your house, is the story, and it asks an uncomfortable question of every per-token subscription.

Where I think most people will actually use it: agentic work. Tool calling is what Qwen highlighted hardest, running inside coding agents and automation without a human interfering at every step.

What happens next

The download is already running on my machine through Ollama, plus a few of Unsloth's quantization levels to compare intelligence at different compression levels. My own tests, on my own hardware, with my frozen exam, are next; the champion it has to beat lives in my local AI Olympics letter.

My rig

Gear in this one

Some links on this page are affiliate links. If you buy through them I may earn a small commission at no extra cost to you. It helps support the work. As an Amazon Associate I earn from qualifying purchases.

Sources

I am documenting a life of interest and curiosity, tech, money, health, and family included. New video daily.

Get the letters

Every post goes out free on Substack. What I build, break, and fix, written up in plain English. No spam, unsubscribe anytime.

Subscribe on Substack