Higgsfield AI
36 min video
3 min read
Make Cinematic AI Ads in 3 Steps
You just saved 33 min.
The big takeaway
A complete workflow for generating ultra-realistic AI commercials using Soul Cinema, GPT Image 2.0, and Claude to build locked assets, write connected prompts, and iterate through Citas 2.0 to produce broadcast-quality ads.
Stage 1: Building & Testing Assets
Organize Everything Before Generating
Create a centralized workspace (canvas) to lay out all generated images side by side. This prevents wasting half your day digging through scattered files during editing and keeps all candidates visible for comparison.
Product Sheets from Every Angle
Use GPT Image 2.0 to generate multi-view product sheets (front, 3/4 views) so the model sees the product from all angles. A single image causes hallucination halfway through scenes; multiple views lock consistency.
Character Sheets with Close-Up + Full Body
Use Claude to write detailed prompts for character sheets with two panels: close-up face and full body (front and back) on a gray background. The close-up locks the exact face; the full body shows height and build. Gray background increases usability rate because it eliminates background clutter competing with the character.
Test Multiple Candidates, Not Just One
Generate several character and location options and test them in motion before locking. A face that looks great as a still may fall apart the moment it moves; only actual footage reveals which holds up.
Three-Quarter Angle for Locations
Generate locations at a three-quarter angle rather than head-on. This gives the room more depth for the model to hold onto when the camera moves, resulting in a much higher win rate than flat shots.
Change One Variable at a Time
When testing, keep the prompt identical and swap only the hero or only the kitchen. This isolates what's making the difference and prevents confusing results.
Edit Locations with GPT Image 2.0
Once a location is chosen, refine it by requesting specific edits (e.g., clear the island, add a stove, remove the TV, add a door). This locks the exact layout needed for animation.
Erase Extra Faces in Character Sheets
If a character sheet has multiple faces in one image, the video model won't know which to grab during animation and the face will drift. Use GPT Image 2.0 to erase all but one face per panel.
Generate Outfit Variations, Then Combine
Ask Claude for 10 outfit prompts for your character, generate all in GPT Image 2.0, then mix and match. If you love the shirt from one and jeans from another, ask Claude to combine them into a single prompt and regenerate.
Preserve Quality When Editing Assets
Every edit in GPT Image 2.0 softens the original quality. To keep the crisp, real look: layer the original high-detail Soul Cinema shot on top of the edited GPT version in a photo editor, mask out the old outfit, and let the new one show through. Face and skin stay from the original; only the outfit comes from the edit.
Props Get Clean Reference Sheets
For repeating props (sky dancer, sneakers, backpack), generate clean reference sheets in GPT Image 2.0. Props don't need motion tests like characters do; the sheet is sufficient.
Stage 2: The Shot List & Prompt System
Use a Claude Skill to Write Prompts
Load a custom Claude skill (provided in description) that writes shot lists tailored for Citas 2.0. This skill encodes everything learned about what holds character consistency, how Citas wants shots described, and all best practices—so Claude stops guessing and writes prompts like a professional.
Feed Claude Your Script + All Locked Assets
Give Claude three things: the full script, every locked asset uploaded as images (not just descriptions), and names for each asset. The more it can actually see, the better it writes prompts. Generic descriptions produce generic prompts.
Shot List as One Connected Document
Claude produces a single connected shot list, not loose separate prompts. It has a style prefix (lighting, camera, color) glued to every prompt, and each prompt has a name. Change the prefix once and it updates everywhere; edit 'prompt 1A' and only that one changes.
Auto-Attach Elements in Citas 2.0
Save all locked assets as named elements in your Citas project using the same names as in Claude. When you paste a prompt into Citas, all required images attach automatically—no manual work.
Stage 3: Scene Generation & Iteration
Fix Lighting Globally via Style Prefix
If the first generation has wrong lighting across all scenes, edit the style prefix in Claude (e.g., 'kill contra jour, go soft even daylight, bright and clean') and apply it to every prompt. This fixes lighting everywhere at once without rewriting individual prompts.
Override Prefix for Single Scenes
When one scene needs its own look (e.g., stadium at midday needs harsh sun and hard shadows, not the soft morning light of the kitchen), override the prefix for just that scene. Tell Claude to apply the override only to that prompt.
Choreograph Motion Move by Move
Never just say 'he dances.' The model interprets that as random flailing. Instead, spell out every move: 'two head nods, shoulder rolling one at a time, a knee dip, a finger snap, quarter spin at the door.' Specific choreography produces specific motion.
Use Music Track as Input for Beat Sync
Drop your actual music track into the Citas prompt as an input and tell the model to have the character dance in time with its beat. Every move then lands on the beat, creating perfect sync without manual timing.
Build Multiple Character Versions for State Changes
If a character must change mid-scene (e.g., dry before a run, sweaty after), don't fight one reference. Build a second sheet now. For example, create both 'S hero dry' and 'S hero wet,' then tell Claude when to use each. Images are cheap; videos aren't.
Use Schematics to Lock Locations & Props
When props or characters drift between takes, create a schematic map in GPT Image 2.0 marking exactly where things sit and their relative sizes (e.g., 'fire hydrant here, sky dancer two times a person's height to its right'). Feed this map to Claude to rewrite the prompt around it. The model then gets it right take after take, eliminating the lottery of 20 generations.
Match Cuts Between Scenes
To stitch scenes together smoothly, match the opening cut of the next scene to the closing cut of the previous one. Ask Claude to 'match the opening tap to the closing tap of scene 1C exactly—same hand, same motion.' This creates clean transitions.
Lock Props Frame to Frame
If a prop appears in multiple cuts (backpack, headphones), explicitly lock it in the prompt: 'lock the backpack on both shoulders in every single cut.' Without this, some batches include it and some don't, breaking continuity.
Iterate on One Prompt, Don't Rewrite
Instead of writing three separate prompts, tweak one shot until it's perfect. Tell Claude 'edit prompt 1A' and describe the fix. This keeps the system connected and prevents losing track of what's working.
Pull Best Seconds from Multiple Generations
The finished ad is the best few seconds out of 100 tries cut together. Run several batches of each prompt, then cherry-pick the keeper faces: the spoon from one video, the flame from another, the pour from a third. Iteration is the skill.
Use Body Rig Shots for Product Focus
For a product shot that demands the headphones stay dead center while the world moves around them, ask for a 'body rig snorricam locked on the right ear cup—camera bolted to his body, headphones rock solid dead center, everything behind whipping past in motion blur.' This creates a cinematic product moment without a real camera rig.
Workflow Metrics & Results
Full Commercial Production Breakdown
The complete headphones ad was built across five scenes using multiple prompts per scene, with the final cut assembled from the best seconds of many generations. Scene 1 (kitchen) used 3 prompts and 9 generations; the stadium, street, and office scenes each required similar iteration. The entire workflow—from asset generation through final edit—demonstrates how iteration and specificity replace brute-force generation.
Worth quoting
"The finished ad is just the best few seconds out of 100 tries cut together. That's the whole game."
— Adele, at [34:51]
"A face that looks great as a still might fall apart the moment it moves."
— Adele, at [3:54]
"Images are cheap; videos aren't."
— Adele, at [22:56]
Try this
Download the Claude skill from the video description and load it into a new Claude chat to begin writing shot lists.
Create a centralized canvas workspace (Figma, Notion, or similar) and organize all generated assets by scene and type (characters, locations, props).
Generate a product sheet with multiple angles (front, 3/4 views) in GPT Image 2.0 to ensure model consistency.
Build character sheets with close-up and full-body panels on gray background; test multiple options in motion before locking.
Generate locations at three-quarter angles and test them with your hero character before finalizing.
For any character that changes state mid-scene, create two versions now (e.g., dry and wet) rather than trying to edit one.
Create a schematic map in GPT Image 2.0 for complex scenes with multiple props or characters to lock their positions and sizes.
Write out choreography move-by-move (head nods, shoulder rolls, spins) instead of generic descriptions like 'he dances.'
Drop your music track into Citas 2.0 as an input and ask the model to sync motion to the beat.
Batch generate multiple versions of each prompt, then cherry-pick the best seconds from different videos to assemble the final cut.
Use match cuts between scenes by ensuring the opening of the next scene mirrors the closing pose of the previous one.
Preserve original asset quality by layering the high-detail Soul Cinema shot over edited GPT versions in a photo editor, masking only the changed element.
Made with Glimpse by Wozart
glimpse.wozart.com/v/xxzomazb
Share this infographic
Read this infographic as text

Make Cinematic AI Ads in 3 Steps

Summary of the video “3-Step Workflow To Make Ultra-Realistic AI Ads by Higgsfield AI.

A complete workflow for generating ultra-realistic AI commercials using Soul Cinema, GPT Image 2.0, and Claude to build locked assets, write connected prompts, and iterate through Citas 2.0 to produce broadcast-quality ads.

Stage 1: Building & Testing Assets

Organize Everything Before Generating

Create a centralized workspace (canvas) to lay out all generated images side by side. This prevents wasting half your day digging through scattered files during editing and keeps all candidates visible for comparison.

Product Sheets from Every Angle

Use GPT Image 2.0 to generate multi-view product sheets (front, 3/4 views) so the model sees the product from all angles. A single image causes hallucination halfway through scenes; multiple views lock consistency.

Character Sheets with Close-Up + Full Body

Use Claude to write detailed prompts for character sheets with two panels: close-up face and full body (front and back) on a gray background. The close-up locks the exact face; the full body shows height and build. Gray background increases usability rate because it eliminates background clutter competing with the character.

Test Multiple Candidates, Not Just One

Generate several character and location options and test them in motion before locking. A face that looks great as a still may fall apart the moment it moves; only actual footage reveals which holds up.

Three-Quarter Angle for Locations

Generate locations at a three-quarter angle rather than head-on. This gives the room more depth for the model to hold onto when the camera moves, resulting in a much higher win rate than flat shots.

Change One Variable at a Time

When testing, keep the prompt identical and swap only the hero or only the kitchen. This isolates what's making the difference and prevents confusing results.

Edit Locations with GPT Image 2.0

Once a location is chosen, refine it by requesting specific edits (e.g., clear the island, add a stove, remove the TV, add a door). This locks the exact layout needed for animation.

Erase Extra Faces in Character Sheets

If a character sheet has multiple faces in one image, the video model won't know which to grab during animation and the face will drift. Use GPT Image 2.0 to erase all but one face per panel.

Generate Outfit Variations, Then Combine

Ask Claude for 10 outfit prompts for your character, generate all in GPT Image 2.0, then mix and match. If you love the shirt from one and jeans from another, ask Claude to combine them into a single prompt and regenerate.

Preserve Quality When Editing Assets

Every edit in GPT Image 2.0 softens the original quality. To keep the crisp, real look: layer the original high-detail Soul Cinema shot on top of the edited GPT version in a photo editor, mask out the old outfit, and let the new one show through. Face and skin stay from the original; only the outfit comes from the edit.

Props Get Clean Reference Sheets

For repeating props (sky dancer, sneakers, backpack), generate clean reference sheets in GPT Image 2.0. Props don't need motion tests like characters do; the sheet is sufficient.

Stage 2: The Shot List & Prompt System

Use a Claude Skill to Write Prompts

Load a custom Claude skill (provided in description) that writes shot lists tailored for Citas 2.0. This skill encodes everything learned about what holds character consistency, how Citas wants shots described, and all best practices—so Claude stops guessing and writes prompts like a professional.

Feed Claude Your Script + All Locked Assets

Give Claude three things: the full script, every locked asset uploaded as images (not just descriptions), and names for each asset. The more it can actually see, the better it writes prompts. Generic descriptions produce generic prompts.

Shot List as One Connected Document

Claude produces a single connected shot list, not loose separate prompts. It has a style prefix (lighting, camera, color) glued to every prompt, and each prompt has a name. Change the prefix once and it updates everywhere; edit 'prompt 1A' and only that one changes.

Auto-Attach Elements in Citas 2.0

Save all locked assets as named elements in your Citas project using the same names as in Claude. When you paste a prompt into Citas, all required images attach automatically—no manual work.

Stage 3: Scene Generation & Iteration

Fix Lighting Globally via Style Prefix

If the first generation has wrong lighting across all scenes, edit the style prefix in Claude (e.g., 'kill contra jour, go soft even daylight, bright and clean') and apply it to every prompt. This fixes lighting everywhere at once without rewriting individual prompts.

Override Prefix for Single Scenes

When one scene needs its own look (e.g., stadium at midday needs harsh sun and hard shadows, not the soft morning light of the kitchen), override the prefix for just that scene. Tell Claude to apply the override only to that prompt.

Choreograph Motion Move by Move

Never just say 'he dances.' The model interprets that as random flailing. Instead, spell out every move: 'two head nods, shoulder rolling one at a time, a knee dip, a finger snap, quarter spin at the door.' Specific choreography produces specific motion.

Use Music Track as Input for Beat Sync

Drop your actual music track into the Citas prompt as an input and tell the model to have the character dance in time with its beat. Every move then lands on the beat, creating perfect sync without manual timing.

Build Multiple Character Versions for State Changes

If a character must change mid-scene (e.g., dry before a run, sweaty after), don't fight one reference. Build a second sheet now. For example, create both 'S hero dry' and 'S hero wet,' then tell Claude when to use each. Images are cheap; videos aren't.

Use Schematics to Lock Locations & Props

When props or characters drift between takes, create a schematic map in GPT Image 2.0 marking exactly where things sit and their relative sizes (e.g., 'fire hydrant here, sky dancer two times a person's height to its right'). Feed this map to Claude to rewrite the prompt around it. The model then gets it right take after take, eliminating the lottery of 20 generations.

Match Cuts Between Scenes

To stitch scenes together smoothly, match the opening cut of the next scene to the closing cut of the previous one. Ask Claude to 'match the opening tap to the closing tap of scene 1C exactly—same hand, same motion.' This creates clean transitions.

Lock Props Frame to Frame

If a prop appears in multiple cuts (backpack, headphones), explicitly lock it in the prompt: 'lock the backpack on both shoulders in every single cut.' Without this, some batches include it and some don't, breaking continuity.

Iterate on One Prompt, Don't Rewrite

Instead of writing three separate prompts, tweak one shot until it's perfect. Tell Claude 'edit prompt 1A' and describe the fix. This keeps the system connected and prevents losing track of what's working.

Pull Best Seconds from Multiple Generations

The finished ad is the best few seconds out of 100 tries cut together. Run several batches of each prompt, then cherry-pick the keeper faces: the spoon from one video, the flame from another, the pour from a third. Iteration is the skill.

Use Body Rig Shots for Product Focus

For a product shot that demands the headphones stay dead center while the world moves around them, ask for a 'body rig snorricam locked on the right ear cup—camera bolted to his body, headphones rock solid dead center, everything behind whipping past in motion blur.' This creates a cinematic product moment without a real camera rig.

Workflow Metrics & Results

Full Commercial Production Breakdown

The complete headphones ad was built across five scenes using multiple prompts per scene, with the final cut assembled from the best seconds of many generations. Scene 1 (kitchen) used 3 prompts and 9 generations; the stadium, street, and office scenes each required similar iteration. The entire workflow—from asset generation through final edit—demonstrates how iteration and specificity replace brute-force generation.

Notable quotes

The finished ad is just the best few seconds out of 100 tries cut together. That's the whole game. — Adele
A face that looks great as a still might fall apart the moment it moves. — Adele
Images are cheap; videos aren't. — Adele

Action items

  • Download the Claude skill from the video description and load it into a new Claude chat to begin writing shot lists.
  • Create a centralized canvas workspace (Figma, Notion, or similar) and organize all generated assets by scene and type (characters, locations, props).
  • Generate a product sheet with multiple angles (front, 3/4 views) in GPT Image 2.0 to ensure model consistency.
  • Build character sheets with close-up and full-body panels on gray background; test multiple options in motion before locking.
  • Generate locations at three-quarter angles and test them with your hero character before finalizing.
  • For any character that changes state mid-scene, create two versions now (e.g., dry and wet) rather than trying to edit one.
  • Create a schematic map in GPT Image 2.0 for complex scenes with multiple props or characters to lock their positions and sizes.
  • Write out choreography move-by-move (head nods, shoulder rolls, spins) instead of generic descriptions like 'he dances.'
  • Drop your music track into Citas 2.0 as an input and ask the model to sync motion to the beat.
  • Batch generate multiple versions of each prompt, then cherry-pick the best seconds from different videos to assemble the final cut.
  • Use match cuts between scenes by ensuring the opening of the next scene mirrors the closing pose of the previous one.
  • Preserve original asset quality by layering the high-detail Soul Cinema shot over edited GPT versions in a photo editor, masking only the changed element.

More like this