
Letters
41 letters, page 1 of 5
Can a 128 GB desktop run GLM-5.3-Flash, Z.ai's 320B open-weights model? The smallest Unsloth quant (93 GB) on a branch build of llama.cpp: one crash in Unsloth Desktop, then 9 tokens a second, four one-shot tests finished twice, vision working at 1-bit, and the one bug a harness would have fixed. Every output and number on the results page.
Local AI · Benchmarks
- How Claude Code makes my YouTube videos2026-09-05
The channel at 90 days: 76,000 views, 462 subscribers, and the workflow that gets every video out. Ten stages, one folder per video, the Claude Code skills, the one post-processing command, what is still friction, and the gear on the desk.
Workflow
OpenAI released GPT-6 Astra. I read the full announcement on camera: the 98% FrontierMath, 99.9% ARC-AGI-3 and 100% ExploitBench claims, the Hugging Face honeypot eval (48% to 0%), the computer-use demos, DeepSWE and Artificial Analysis, and Matthew Berman's early-access review. A first impression, and why the hype still points back to local AI.
AI news
A first impression of Claude Fable 5.1 and Mythos 5.1: what changed, who gets which model, the pricing read off the docs, and then two days of Claude Code in auto mode on my own local coding-agent benchmark. $165 on the estimator, a harder suite, and one honest caveat.
AI news · Benchmarks
OpenAI is ending its partnership with Cursor after SpaceX's $60 billion acquisition, with a proposed shutoff of November 12. What OpenAI actually said, the Musk history it points at, the disputed 5% number, what still works after the cutoff, and why the answer is owning your tools.
AI news
Nvidia has reportedly agreed to buy Hugging Face for $12.9 billion, about 86 times its self-reported revenue. Reported, not signed. The numbers, the offer Hugging Face turned down nine months ago, the open-weights letter both companies signed, and why a chip company owning the open-weights hub is not the worst outcome for people who run models at home.
AI news
Qwen's new 125B mixture-of-experts (6B active) on a 128 GB mini PC, the night it became runnable: llama.cpp built from the support PR, four Unsloth quants from 1-bit to IQ4_XS through the same four one-shot tests, my Qwen 3.8-27B as the reference. Every quant ran at about 20 tokens per second. Then the honest question: would I actually use it?
Local AI · Benchmarks
Qwen's new 125B mixture-of-experts model with 6B active parameters dropped today. Before running a single benchmark I read the Unsloth guide and the Qwen model card on camera to answer one question: can a 128 GB unified-memory mini PC even hold it? The smallest quant is 75 GB, they recommend 96, and BF16 is 355 GB.
Local AI
MTP had been silently on in every benchmark I published and I could not explain it. So I measured multi-token prediction on Qwen 3.8-27B at BF16: off versus depths 2, 3 and 4, three prompts, two protocols, and the baseline run twice. 4.2 to 12.9 tokens per second on the same hardware, and every raw output hashed.
Local AI · Benchmarks
- I wired DeepSeek Harness to my local AI2026-08-20
DeepSeek's new open-source agent harness passed 175,000 GitHub stars in about a week. I installed it from source and wired it to local AI two ways, Ollama and Unsloth, both running Qwen 3.8-27B. Every command, both live errors, and a working app in 30 minutes.
Local AI · Workflow
Get the letters
Every post goes out free on Substack. What I build, break, and fix, written up in plain English. No spam, unsubscribe anytime.
Subscribe on Substack