Jose Romero
← All letters
Prefer email? Read this on Substack

How Claude Code makes my YouTube videos

A different kind of letter. I opened my YouTube Studio analytics on camera and walked through how the channel got here, then showed the workflow that gets every video out the door. This is the written version, with the parts I skipped on camera because the video was already ten minutes in.

Source files for this video

Everything below is built from these two pages. Open them alongside the letter.

  1. The slide deck: 29 slides, the ten stages, the edit chain, the publish order and the channel numbers. Fullscreen, arrow keys to move.
  2. The workflow and skill files: the stage map, the tools, and all ten Claude Code skill files sanitized of anything personal, each with a copy button and a raw .md link. A snapshot from this video, not a maintained page.

The numbers, honestly

I started posting a video every day on July 20 and kept that up for 30 days, through the middle of August. Since then I have posted every other day or so, based on whatever I was into that week. Ninety days in, the channel sits at about 76,000 views and 462 subscribers.

Most of that came from one topic. Qwen 3.8 has been the biggest lifter by far: the video on running Qwen 3.8-Flash-Next at 1-bit on my 128 GB rig is at 17,000 views, and the thinking-levels and quant-ladder videos are right behind it. The surprise was a one-off about switching to a Corne split keyboard because of wrist pain. That one is at 10,000 views and I did not expect it. It is the only hardware video on the channel, so I may go back to that lane.

Local AI is my niche for now. I am not married to it, but it is what I spend most of my free time on, so it is what I make videos about. Monetization is not on yet: memberships need more subscribers than I have, ads need 1,000, and I am about halfway there.

The workflow, before the camera turns on

Everything runs out of one folder with Claude Code, and every video gets its own numbered subfolder. A handful of slash commands do the prep:

The ten stages of the pipeline in two rows: idea, research, packaging first, beat map, slides before the camera; record, edit, copy and check, publish, retro after it. Cyan stages are AI drafted, violet stages are human only.
Ten stages. Cyan is drafted by the AI, violet is human only.

There is also a beat map step that writes the hook and the key talking points. I will be honest: I do not really use it. Sometimes I record part of the video first and build the slides from what I said.

After recording: one command

I record in OBS with a Sony camera and an SM7B, do a take or several, and drop the file on the desktop. Then one post-processing command does the rest.

The edit chain: archive the raw into the folder, A/V sync, word-timed cuts, silence cut with auto-editor at 256k audio, captions with faster-whisper, a 142-rule spelling dictionary, a loudness pass to -14 LUFS, and the final files.
The edit chain, mostly unattended.
  1. Copies the raw into the video folder so it is never lost when I clean the desktop.
  2. Fixes the delay between my camera and my microphone. OBS drifts depending on what else the machine is doing, so the AI measures from my clap at the start and picks the offset instead of me guessing.
  3. Cuts the dead air, the over-breathing, the ums, and the flubs. If I mess up a section I pause, clap, and restart, and it knows to cut back to that point.
  4. Runs privacy checks on the screen share so I am not showing more personal information than I mean to.
  5. Writes the captions with faster-whisper. YouTube's automatic captions get AI model names wrong constantly, so my transcription runs against a dictionary of the names I actually say.
  6. Renders a second version with a loudness pass for the nights I get tired and my voice drifts low.

That command is the single reason I can put out this many videos. This part used to take me hours, and that friction is what made me stop making videos in the past. It is not perfect, but the small mistake here and there is not enough to make me not publish.

Where it still hurts

The workflow on paper is ten stages. In practice I am running about six of them. The retro step at the end has never run once. The friction that is left is all the copy-paste: getting the article into Substack, listening to both audio versions to pick one, updating the YouTube description after the links exist.

Publish order: upload to YouTube unlisted with the caption file, publish on Substack and get the link, finish the description and flip the video to public, then push the site letter with the Substack link.
Publishing is manual, in this order. Substack first.

What is next

I am migrating the annoying parts to Codex. The plan is to record on the Mac, funnel the files to the Strix Halo rig for all the post-processing, and let an agent handle the file shuffling, descriptions, and posting with light oversight. Not overnight on full auto, but something I can run in a side panel while I work on something else.

What is on the desk

The current setup, since a few of you asked. A proper setup-upgrade video is coming later; this is what the videos are recorded on today.

Gear in this one

Some links on this page are affiliate links. If you buy through them I may earn a small commission at no extra cost to you. It helps support the work. As an Amazon Associate I earn from qualifying purchases.

Sources

The slides from this videoHow a video gets made here · 29 slides · opens fullscreen, arrow keys to navigate

Get the letters

Every post goes out free on Substack. What I build, break, and fix, written up in plain English. No spam, unsubscribe anytime.

Subscribe on Substack