
To test an AI agent on a small sample first, have Codex or Claude Code do only the first piece of the job, like five minutes of a 45-minute video or five pages of a deck, and review it before the agent runs the rest. Skip the test and one wrong guess costs you a full run plus the redo. I have been creating AI tutorials for business owners since January 2023, and I help established business owners build agents inside their companies. One of the biggest wastes I see, and one I caught in my own weekly usage report, is letting an agent run a whole job before anyone has checked a sample. If you have to edit something that's 45 minutes, you want to have it prove its rules and prove itself with the first five minutes.
In short: - Have Codex or Claude Code prove its rules on a small piece before the full job: the first five minutes of a video, the first five pages of a deck, or one thumbnail before a batch. - Check the things that repeat across the whole job, such as branding, messaging, caption style, or the fields and sales stages in a CRM. - Correct the sample once, then have the agent apply those corrections to everything that follows. - Shanee's own weekly usage report flagged about 20 thumbnail regenerations and told her to show one thumbnail before making the full batch.
Part of the series: How to Reduce Codex and Claude Code Token Usage: 20 Ways. This is way 17. Previous: Fix the root cause of agent mistakes · Next: Turn repeat agent work into a script
1. Test it on a small piece before the agent does the whole job (the AI video editing example)
A great example of this is editing a video. I've seen people say, "Okay, edit the video, do the captions, whatever." Even in the event of a five-minute video, if it has to do the whole five minutes and then submit a result back to you, it's going to take a long time. And if it messes up even just in the first 30 seconds, now it's generated an extra 4 minutes and 30 seconds. That is a complete waste.
So you want to do things in chunks.

What does it cost your business to skip the sample?
If you have to have it do a 100-page slide deck, you have to have it prove itself with maybe the first five pages, so that it could prove that it has the branding, the messaging, all of that, how you want it. Because if not, it's going to try to produce the whole thing every single time. It'll take an hour, maybe longer, and then you look at it after an hour and you're disappointed or deflated or frustrated because it's not better. It didn't understand what you were doing, and you've just waited an hour for nothing.
So this is actually pretty critical. You pay once for the wrong full run, and then you pay again for the redo.
2. How to pilot test an AI agent: pick the small piece for the job
The small piece is whatever lets you check the agent's rules before it repeats them: the first five minutes of a video, the first five pages of a deck, one thumbnail before the batch. Shanee's field guide for this series lays out three before-and-after examples side by side.

| Job | Expensive way to start | Small piece first |
|---|---|---|
| Editing a 45-minute video | Edit all 45 minutes, review, find the wrong caption style, redo all 45 minutes | Edit the first 5 minutes, review, fix the style once, then edit the other 40 minutes |
| Building a CRM (5,000 contacts and three years of sales out of spreadsheets) | Set up every field and sales stage, import 5,000 contacts, find wrong stages and duplicates, rebuild and re-import | One set of sales stages and 50 contacts, review, fix the fields and double-entry rules, then import the other 4,950 |
| Rebuilding a 30-page website | Build all 30 pages, review, find the words and look are wrong, rewrite all 30 pages | Build the homepage first, review, lock in the look and the words, then build the other 29 pages |
| A 100-page slide deck | Produce the whole deck and wait an hour | Prove the branding and messaging on the first five pages |


The same idea works for any big job: one email before all twelve, one slide before the whole deck.
How to test an AI agent on a small sample in Codex or Claude Code
Step 1. Name the small piece in the request, not the whole job.
Step 2. Ask the agent to stop and show you what it did and what it guessed.
Step 3. Correct the sample once, then let it apply your corrections to the rest.
This prompt from Shanee's field guide does all three:
Before you do the full job, prove the approach on a small piece. Do only [the first 5 minutes of the video / the first 50 records / one email / one slide / one page]. Then stop and show me: - What you did, and how you'd do the rest the same way - Anything you had to guess about - Your estimate for the full job Don't start the rest until I approve the sample. Once I do, apply my corrections to everything that follows.
3. Shanee's own example: one thumbnail before the full batch
This one showed up in Shanee's own weekly usage report. She was creating article thumbnails for a client's site migration, and to get them right took about 20 times, because she didn't have a defined reference photo and didn't give the agent an example from the start. One of the three habits the report told her to test was show one thumbnail before making the full batch.
It's the same lesson she gives for standards in Part 1. When you save something as a standard, you wanna say, "Prove it to me by creating an example deck now." And then whatever it gets wrong, you tweak and you refine. You don't just want to blindly trust it.
Frequently asked questions
How do I test an AI agent before it does the whole job?
Ask Codex or Claude Code to do only a small piece, then stop and show you what it did, what it guessed about, and its estimate for the full job. Correct the sample once, and only then let it apply your corrections to the rest. The prove-it-first prompt above does all three steps.
How small should the test sample be?
Small enough to check in a few minutes, big enough to show the agent's rules. Shanee's examples are the first five minutes of a 45-minute video, the first five pages of a 100-page deck, and one thumbnail before a batch (18:50).
How should I start AI video editing with Claude Code or Codex?
Start with the first five minutes, not the whole video. Have Claude Code or Codex edit that piece, review it, and fix the caption style once before it edits the other 40 minutes. If it messes up in the first 30 seconds of a full run, everything it produced after that is wasted (18:38).
Is testing a small sample worth it for a small business?
Yes, because it adds one short review and protects you from paying for a full run that has to be redone. On a 100-page deck, Shanee says the full run can take an hour or longer, and if the agent misunderstood you, you have waited an hour for nothing (19:27).
What if the sample is wrong twice?
Stop and give the agent a real example instead of vague feedback. That is way 15 in this series.
Related ways in this series
- Way 9: Make great agent output the standard, including "prove it to me" before you trust a new standard.
- Way 14: Have the agent look before it fixes, for big jobs where the plan itself is uncertain.
- Way 15: Stop when an AI agent fails twice, for when the sample keeps coming back wrong.
- Way 20: Build a weekly token usage audit agent, where the "one thumbnail first" habit came from.
- Back to the hub: How to Reduce Codex and Claude Code Token Usage: 20 Ways
Sources
Official / Primary Sources
- Shanee Moret, "20 Ways to Optimize How You Use Codex and Claude Code Part 2" (YouTube, 2026-10-03) , the small-piece rule, the video and slide deck examples, and her own thumbnail report.
- Shanee Moret, "20 Ways to Use Codex and Claude Code Without Burning Your Usage (Part 1)" (YouTube, 2026-10-02) , "prove it to me" when creating a standard.
- Shanee Moret, "20 Ways to Optimize Your Token Usage in Codex and Claude Code" field guide (GrowthAcademy.Global, 2026) , the CRM and 30-page website examples and the prove-it-first prompt.