Packaging one video used to take me two hours
10 August 2026 · 5 min read
Before packaged existed I ran a faceless YouTube channel. The videos themselves were the easy part. The part that reliably ate my evening was everything that happened after the video was finished.
Two hours per upload, give or take. Here is where they went.
The two chat windows
I had ChatGPT open in one tab and Claude in the other, and I would paste the same transcript into both. Not because I had a clever system — because neither one gave me something I could use on the first try, and comparing two mediocre answers felt more productive than staring at one.
Then the loop started:
- Ask for titles. Get ten. Eight are the same title. One has a colon in a place no human would put a colon.
- Ask for a description. Get four paragraphs of throat-clearing before anything a viewer would care about, with the actual hook buried in paragraph three where the search preview will never show it.
- Ask for tags. Get a comma-separated wall, some of it about a topic the video does not cover, and no idea whether it fits the character limit.
- Ask for chapters. Get chapter names that sound right and correspond to nothing in the transcript.
- Paste it into Studio. Find the formatting arrived as literal asterisks, because YouTube descriptions do not render markdown.
Then thumbnail text. Then the thumbnail concept itself — what is on it, where the words sit, whether the words survive being shrunk to the size of a fingernail on a phone.
What made it two hours instead of twenty minutes
Not the generating. The generating was fast. The two hours were iteration and re-explaining.
Every session started from nothing. The model did not know my channel, my niche, the phrases I use, the phrases I refuse to use, or that I had already rejected the curiosity-gap angle on the last four videos. So I re-explained it, every time, in a slightly different way, and got a slightly different flavour of the same output.
That is the actual problem. Not “AI can't write a title” — it can. It is that a blank chat window has no memory of you, no opinion about what good looks like on YouTube specifically, and no reason to give you the same quality twice.
What packaged is, in one sentence
It is that two hours, collapsed into one paste, with the opinions already made.
You drop the transcript. It returns every field Studio asks for, in the order Studio asks for them, each with its own copy button and a line explaining why it is written that way. No prompt to compose. No iteration loop unless you want one — and if you do, a regeneration is required to come back with a genuinely different angle rather than the same idea reworded.
And it remembers your channel, so the second video is better than the first for reasons you did not have to type out again.
The part I did not expect
Building this, the thing that turned out to matter most was not the writing quality. It was refusing to fake anything.
If there is no caption file, packaged does not invent timestamps — it gives you chapter titles and says plainly that it cannot align them. If a generation fails, it does not charge you. If the transcript is too messy for chapters, the chapters bail and every other block still ships.
Those are all small. They are also the difference between a tool you trust with the last step before you hit publish, and one you have to check behind.