AI Tutorials · Codex and Claude Code

Test Your AI Agent on a Small Sample First: Codex and Claude Code for Small Teams

Small business owners: have Codex or Claude Code prove its rules on 5 minutes of video or 5 slides before the full job, so one bad guess doesn't cost an hour.

Shanee Moret thumbnail with the words TEST SMALL SAMPLES FIRST

To test an AI agent on a small sample first, have Codex or Claude Code do only the first piece of the job, like five minutes of a 45-minute video or five pages of a deck, and review it before the agent runs the rest. Skip the test and one wrong guess costs you a full run plus the redo. I have been creating AI tutorials for business owners since January 2023, and I help established business owners build agents inside their companies. One of the biggest wastes I see, and one I caught in my own weekly usage report, is letting an agent run a whole job before anyone has checked a sample. If you have to edit something that's 45 minutes, you want to have it prove its rules and prove itself with the first five minutes.

20 Ways to Optimize How You Use Codex and Claude Code Part 2. Watch the full video on YouTube (published 2026-10-03).

In short: - Have Codex or Claude Code prove its rules on a small piece before the full job: the first five minutes of a video, the first five pages of a deck, or one thumbnail before a batch. - Check the things that repeat across the whole job, such as branding, messaging, caption style, or the fields and sales stages in a CRM. - Correct the sample once, then have the agent apply those corrections to everything that follows. - Shanee's own weekly usage report flagged about 20 thumbnail regenerations and told her to show one thumbnail before making the full batch.

Part of the series: How to Reduce Codex and Claude Code Token Usage: 20 Ways. This is way 17. Previous: Fix the root cause of agent mistakes · Next: Turn repeat agent work into a script

1. Test it on a small piece before the agent does the whole job (the AI video editing example)

A great example of this is editing a video. I've seen people say, "Okay, edit the video, do the captions, whatever." Even in the event of a five-minute video, if it has to do the whole five minutes and then submit a result back to you, it's going to take a long time. And if it messes up even just in the first 30 seconds, now it's generated an extra 4 minutes and 30 seconds. That is a complete waste.

So you want to do things in chunks.

Way 17 card titled "Test it on a small piece before the agent does the whole job" with what it is, why it matters, and skip-it and do-it examples
Shanee's card for way 17, including the 45-minute video example: edit the first five minutes, fix the captions once, use that fix for the rest. From Shanee's field guide on the 20 ways. Shown in the video at 18:50.

What does it cost your business to skip the sample?

If you have to have it do a 100-page slide deck, you have to have it prove itself with maybe the first five pages, so that it could prove that it has the branding, the messaging, all of that, how you want it. Because if not, it's going to try to produce the whole thing every single time. It'll take an hour, maybe longer, and then you look at it after an hour and you're disappointed or deflated or frustrated because it's not better. It didn't understand what you were doing, and you've just waited an hour for nothing.

So this is actually pretty critical. You pay once for the wrong full run, and then you pay again for the redo.

2. How to pilot test an AI agent: pick the small piece for the job

The small piece is whatever lets you check the agent's rules before it repeats them: the first five minutes of a video, the first five pages of a deck, one thumbnail before the batch. Shanee's field guide for this series lays out three before-and-after examples side by side.

Before and after timeline: editing all 45 minutes stalls at the wrong caption style, while editing the first 5 minutes lets you fix the style once
The 45-minute video example. The "before" path stalls at the wrong caption style and has to redo all 45 minutes. From Shanee's field guide on the 20 ways. Shown in the video at 19:30.
Job Expensive way to start Small piece first
Editing a 45-minute video Edit all 45 minutes, review, find the wrong caption style, redo all 45 minutes Edit the first 5 minutes, review, fix the style once, then edit the other 40 minutes
Building a CRM (5,000 contacts and three years of sales out of spreadsheets) Set up every field and sales stage, import 5,000 contacts, find wrong stages and duplicates, rebuild and re-import One set of sales stages and 50 contacts, review, fix the fields and double-entry rules, then import the other 4,950
Rebuilding a 30-page website Build all 30 pages, review, find the words and look are wrong, rewrite all 30 pages Build the homepage first, review, lock in the look and the words, then build the other 29 pages
A 100-page slide deck Produce the whole deck and wait an hour Prove the branding and messaging on the first five pages
Before and after rails for building a CRM: importing all 5,000 contacts stalls at wrong stages and duplicates, while starting with 50 contacts lets you fix the fields before importing the other 4,950
The CRM example: moving 5,000 contacts and three years of sales out of spreadsheets. Prove one set of sales stages on 50 contacts, then import the other 4,950. From Shanee's field guide on the 20 ways. Shown in the video at 20:00.
Before and after rails for rebuilding a 30-page website: building all 30 pages stalls when the words and the look are wrong, while building the homepage first locks them in before the other 29 pages
The website example: build the homepage first, lock in the look and the words, then build the other 29 pages. From Shanee's field guide on the 20 ways.

The same idea works for any big job: one email before all twelve, one slide before the whole deck.

How to test an AI agent on a small sample in Codex or Claude Code

Step 1. Name the small piece in the request, not the whole job.

Step 2. Ask the agent to stop and show you what it did and what it guessed.

Step 3. Correct the sample once, then let it apply your corrections to the rest.

This prompt from Shanee's field guide does all three:

Copy-ready
Before you do the full job, prove the approach on a small piece.

Do only [the first 5 minutes of the video / the first 50 records /
one email / one slide / one page]. Then stop and show me:
- What you did, and how you'd do the rest the same way
- Anything you had to guess about
- Your estimate for the full job

Don't start the rest until I approve the sample.
Once I do, apply my corrections to everything that follows.

3. Shanee's own example: one thumbnail before the full batch

This one showed up in Shanee's own weekly usage report. She was creating article thumbnails for a client's site migration, and to get them right took about 20 times, because she didn't have a defined reference photo and didn't give the agent an example from the start. One of the three habits the report told her to test was show one thumbnail before making the full batch.

It's the same lesson she gives for standards in Part 1. When you save something as a standard, you wanna say, "Prove it to me by creating an example deck now." And then whatever it gets wrong, you tweak and you refine. You don't just want to blindly trust it.

Frequently asked questions

How do I test an AI agent before it does the whole job?

Ask Codex or Claude Code to do only a small piece, then stop and show you what it did, what it guessed about, and its estimate for the full job. Correct the sample once, and only then let it apply your corrections to the rest. The prove-it-first prompt above does all three steps.

How small should the test sample be?

Small enough to check in a few minutes, big enough to show the agent's rules. Shanee's examples are the first five minutes of a 45-minute video, the first five pages of a 100-page deck, and one thumbnail before a batch (18:50).

How should I start AI video editing with Claude Code or Codex?

Start with the first five minutes, not the whole video. Have Claude Code or Codex edit that piece, review it, and fix the caption style once before it edits the other 40 minutes. If it messes up in the first 30 seconds of a full run, everything it produced after that is wasted (18:38).

Is testing a small sample worth it for a small business?

Yes, because it adds one short review and protects you from paying for a full run that has to be redone. On a 100-page deck, Shanee says the full run can take an hour or longer, and if the agent misunderstood you, you have waited an hour for nothing (19:27).

What if the sample is wrong twice?

Stop and give the agent a real example instead of vague feedback. That is way 15 in this series.

Sources

Official / Primary Sources