← All posts
Comparison

packaged vs just using ChatGPT

15 August 2026 · 6 min read

ChatGPT writes good YouTube metadata. We tested it and it does. This post is about the gap between a good answer and a published video.

What we did, exactly

On 15 August 2026 we gave GPT-5.6 a real 1,600-word transcript and asked for a title, description, tags and chapters, then asked it to list what a creator would still have to fix by hand.

The model knew its output would be compared to packaged. We told it so. We asked it not to polish and not to add caveats, but we cannot prove it complied, and a model that knows it is being measured may behave differently from one that does not.

This is one transcript, one run, one model, on one day. Everything below is what happened in that run. None of it establishes how ChatGPT behaves generally, and we will not pretend otherwise.

Most of what we expected to criticise, we could not

  • The tags were fine. 458 characters across 30 tags, inside the 500 the creator tooling community treats as the ceiling. Google does not publish that number, so we will not call it YouTube's limit.
  • The hook was well placed. Its strongest line landed in the first 125 characters, where the search preview cuts.
  • The description body was clean. No stray formatting characters inside it.

We had expected to find problems in all three. We did not, and saying so is the only reason to trust the rest of this post.

One thing it produced that a creator could not check

It supplied eight chapter timestamps — 0:00, 1:40, 3:35, and so on through the video.

There was no finished video. The transcript carried no timing data, so those numbers were estimated from how long each section of text was.

When we asked it to audit its own output, it said so plainly: not timestamps measured against the finished voice-over and edit. A creator who does not ask for that audit does not get that sentence. They get eight numbers that look exactly like measured ones, and the mistake stays invisible until a viewer clicks a chapter and lands in the wrong place.

They are checkable. You open the finished edit and verify all eight by hand. That manual check is the cost, and it arrives at the end of the job, when you are least likely to do it.

With no caption file, packaged returns chapter titles and tells you the timings are unavailable. A timestamp not derived from real timing data is rejected by the validator rather than discouraged in a prompt — a distinction worth naming, because a prompt alone leaves it to the model.

The rest is the loop, not the writing

In this run, the answer arrived as one document. Three of the five manual steps it listed were placing or cleaning fields: strip the scaffolding, move the chapters into the description, separate the fields. The other two were correcting the timestamps and a final editorial read. packaged returns each field separately with its own copy button, in the order YouTube's upload page asks for them.

A capable chat can vary its angles when you supply the history. Tell it which angles you have already seen and it will avoid them. packaged tracks that for you, refuses a repeat by default, and says so when the angles run out rather than charging you for the same idea again.

A configured chat can carry your channel context too — through memory, projects, or a prompt you maintain. The difference is who maintains it. packaged carries your channel profile by default, on the first package and on every retake after it.

A hard failure does not consume credits. The sentence you see is the one in the code, not a support policy you have to invoke.

And packaged does not require YouTube account authorization. It cannot sign in to your channel, because no such capability exists in the product — a check runs on every build and fails it if one ever appears. The one thing that does leave our server: the page title of a link you paste into your own description, so the call to action can name what it points at.

So which should you use

If you package one video a month, use ChatGPT. It will do a good job. We would rather say that than sell you something you do not need.

If you package videos every week, the cost is not the writing. It is the disassembly, the context you rebuild every session, the angles you police yourself, and the timestamps you verify at the end of a long job.

Publishing weekly? Drop in one transcript and compare the finished workflow against what you did by hand. 100 credits, no card. Start free.

What this post does not claim

  • We do not claim ChatGPT writes bad metadata. We measured it and it does not.
  • We do not claim this is how ChatGPT usually behaves. One run cannot show that.
  • We do not claim a ranking, a view count or a result. Metadata does not guarantee views; performance depends on viewers and on factors no tool controls.
  • We do not claim packaged is smarter. The difference is the workflow around the answer, which is the only difference we can show you.