Age and life stage
The single field that moves performance most, because it decides who the viewer thinks the ad is about. A claim about waking at three in the morning lands differently at 26 and at 45.
How to describe the person on camera so the ad reads as a person, which fields actually change the result, and where a generated performer still gives itself away.
{
"creator": {
"age": "mid thirties",
"look": "no makeup, hair tied back",
"wardrobe": "grey sweatshirt, one size too big",
"setting": "kitchen, morning light, clutter on the worktop",
"delivery": "unrehearsed, half smile, one small stumble",
"camera": "handheld at arm length, slight drift"
},
"aspect": "9:16"
}The three fields that decide realism are the last three, not the first two.
A synthetic performer built from a written brief, standing in for the person you would otherwise book.
An AI UGC creator is a face, a voice and a delivery generated for a video ad. No casting call, no shipping the product to someone, no waiting a week for a file that turns out to be filmed in the wrong orientation. You write who the person is and what they say, and a vertical take comes back.
You will see the same idea sold under three names. AI UGC actors usually means a performer generated per brief. A UGC AI avatar usually means a fixed identity from a library, reused across many videos so a campaign has one recognisable face. Creator gets used for both. The distinction that matters when you are choosing is whether you can describe the person you want or only pick one that already exists.
Libraries are faster to start with and better for continuity, because the same avatar in ad four looks like the one in ad one. Described creators are better for testing, because the person is a variable you can move. Most accounts end up using both: a library face for the campaign that is working, and described ones while hunting for the next angle. What both produce is the same kind of file, covered on AI UGC videos.
Most briefs spend all their detail on the face. The face is the field that matters least, because every model already renders a plausible one.
The single field that moves performance most, because it decides who the viewer thinks the ad is about. A claim about waking at three in the morning lands differently at 26 and at 45.
Rooms carry more signal than faces. A kitchen with things on the worktop reads as a home, and an empty room with a plain wall reads as a set.
Slightly wrong clothes make a person look real. Anything crisp, matched or freshly pressed drifts back toward a brand shoot.
Ask for unrehearsed rather than confident. Confident is the default and the default is what people have learned to scroll past.
Held at arm length, with drift and a reframe, is the tell that says phone. A locked static frame says tripod, and tripods belong to brands.
Pace, accent and breath decide whether the lip sync survives a second viewing. A slightly slower read gives the mouth easier work to do.
A generated performer will do exactly what the brief says, including the parts you left on default. Defaults are what make an ad look generated.
Read the script out loud and cut anything you stumble on. Contractions, short clauses and one filler word do more for believability than any render setting.
A person doing nothing with their hands looks like a talking head. Pick the product up once, on a specific line, and the whole take stops feeling like a presentation.
A half second of dead air before the close, a glance away, an unbalanced frame. Ask for these directly, because no default setting will give them to you.
notes for take 2
- drop the word "amazing", she would not say it
- start talking before the shot settles
- one physical beat: pick the sachet up on
the line about the change
- eyes to the lens on the close, not before
- leave the pause after "three in the morning"
- no music, room tone onlyNotes for a second take are just text, so they can be replayed on the next product.
Worth knowing before you spend, because most of it can be designed around rather than fixed.
Picking a creator from a grid of thumbnails is pleasant once. It is the slowest part of the job by the fortieth ad.
Wireflow is our own product, and it exposes its generation workflows over a hosted MCP endpoint for exactly this reason. Listing the workflows and the models is read only, so an agent can look before anything is spent. The same approach applied across a client roster is on AI UGC for agencies.
Same sleep tea script, three different creators: late twenties, mid thirties, mid fifties. Kitchen, handheld, unrehearsed.
I will keep the script, the setting and the camera note identical and change only the age band, so the comparison is clean.
Three takes. The fifty five read is the calmest and holds the pause before the close. The late twenties one rushes the second line, so I can slow that read and rerun just that one.
list_models, run_workflow and get_execution are real tools on the server. Only run_workflow spends credits.
Once the person on camera is settled, the next decisions are about the script and the buy.
It is a synthetic performer generated for creator style video ads: a face, a voice and a delivery, produced from a written brief instead of booked for a shoot. Some tools call the same thing an AI UGC actor or a UGC AI avatar. The practical difference between them is whether you pick from a fixed library or describe the person you want.
Close enough in everyday use, though the words hint at different products. Avatar usually means a fixed identity from a library that you reuse across many ads. Actor or creator usually means a performer generated per brief, which gives you more range and less continuity between videos.
Yes, and you should when the ads are meant to feel like one person. Reuse the same avatar identity or the same seed and casting fields across the batch. Keep the description in a file so a run six weeks later starts from exactly the same person rather than a close cousin.
Hands and product handling are the weak spot, along with long unbroken takes, over even eye contact and any packaging text the model has to render. Short takes, one physical action per shot and product inserts filmed separately hide most of it. If a shot needs a label read on camera, film that shot or composite it.
Not for a generated performer, because there is no real person to license. That is also the constraint: if your ad depends on a specific known face, this is the wrong category of tool. Use generated creators for volume testing and book real people for the ads that trade on who is saying it.
Connect the MCP endpoint to Claude, Cursor or n8n, then describe the person and the read in plain language. The agent calls the generation workflow with those fields filled in, polls the run, and returns the video URL. Notes for a second take are just another message in the same conversation.
Connect the endpoint, write the casting brief in a sentence, and read the result in the same window. The read only tools cost nothing, so you can look before you generate.