33 min readMarkdown
Seedance 2.5: the complete prompting guide
Thirty seconds of video with sound from one written brief. Seven rules, three pipelines and the failures worth knowing about before you spend anything.
Seedance 2.5 makes up to thirty seconds of video in a single run, with the sound made at the same time as the picture: speech, sound effects and music. No joining clips, no separate voice step, no second render to add sound.
That single run is what makes it different from a model that gives you five seconds at a time. A thirty-second One run of the model: a single generated clip, start to finish. Running the same brief again gives you another take, and you pay for it again. has no seams to hide because there are none. It also means every generation is all or nothing: you cannot fix second nineteen without re-running seconds one to thirty.
So the work moves earlier. Almost everything that decides whether a clip is good happens before you press generate.
Treat it as a shoot you are directing, not a slot machine you are pulling.
Overview
The short version
Four decisions do most of the work. If you read nothing else, read these. Each one is written out in full further down.
-
Get the picture right before you make it move
Video is priced by the second, so every mistake you find in a thirty-second take costs a whole take to fix. One image is the cheap place to be wrong, and the one place you can look at the result before anything moves.
Settle the room, the clothes, the light, the framing and any words on screen as a single image, then hand that finished image to the video model.
-
Describe what the camera sees, not what a film crew would call it
The model does what it can picture and guesses at everything else. A camera term is a direction a person understands and a model only half-follows; a list of what you can actually see in the frame gets followed. Nearly every wording failure we had was this one.
Instead of
Three-quarters from behind and to her left
Write
We see the back of her head, her claw clip, the back of her neck
-
Tell every image you upload what it is for, and what it is not
An uploaded picture with no instructions gives away everything in it: its background, its lighting, its framing, its pose. Name the one thing you want from it and rule out the rest, or the studio backdrop behind your product quietly becomes the backdrop of your film.
Instead of
[Image1] is Ana.
Write
[Image1] is Ana. Use it for her face, build and hair colour only. Do not use its background, lighting, framing or pose.
-
A photoreal person has to be drawn, or made in Seedream
Hand the model a photo of a real person's face and the generation is refused, seven times out of seven in our tests, whether you supply it as the A picture you supply that the clip opens on exactly. It becomes frame one, and the film takes its shape from it rather than from your aspect ratio setting. or as a reference. A photoreal character you generated yourself is refused the same way — inventing the face in GPT Image 2 or Nano Banana buys you nothing, because what is inspected is the picture. Seedream is the exception: photoreal people made there are accepted.
Make the cast in Seedream and reference those images. Or cast an illustrated, anime or 3D character, which passes whatever drew it: the same brief that was refused with a real face rendered first try with a drawn one. Or supply no picture of a person at all and describe them in words: text to video invents a face rather than copying one.
What it accepts
| Input | Limit | Notes |
|---|---|---|
| Prompt | 4,000 chars | Room for every block a brief needs. Only a long, talky thirty seconds comes close. |
| Duration | 4–30 seconds | Or automatic, where the model picks the length. |
| Resolution | 480p · 720p | Test at 480p, finish at 720p. Upscale afterwards. |
| Reference images | up to 30 | Characters, products, locations, styles. |
| Reference videos | up to 10 | Motion, pacing, camera behaviour. 30s combined. |
| Reference audio | up to 10 | Voices and ambience. 30s combined, one voice per track. |
| First frame | 1 image | Cannot be combined with reference images. |
What it costs
Video is priced by the second, so length and resolution are the only two decisions that cost you money. Sound is not a third: speech, effects and music come out of the same per-second rate whether you write for them or not, unlike Veo 3.1, where the audio switch doubles what a take costs.
| Length | 480p | 720p | What it’s for |
|---|---|---|---|
| 5s | ~43 | ~97 | Testing: does the look hold, does it pass, do the inputs work |
| 15s | ~129 | ~289 | A complete short piece: a User-generated content: an ad shot to look like a real customer filmed it on their phone, rather than like an ad. ad, a vlog, a product spot |
| 30s | ~257 | ~578 | A film with a beginning, a middle and a reveal |
A thirty-second take costs too much to get wrong, so do not start one cold. Begin from a first frame, or from a set of references, and be sure of your prompt before you press generate.
We still work in much smaller pieces most of the time, 4 to 15 seconds at a go.
The rules
Seven things that are true whatever you are making. Each one cost us money to learn, and they are ordered by how much they will save you.
Name the pixel, not the idea
The model obeys what it can picture, and guesses at everything else.
An angle, a lighting theory, a rule: these get interpreted. A description of what should be visible in the frame gets followed.
Every wording failure we had was this one. “Three-quarters from behind” is a direction a person understands and a model only half-follows:
- As an angle, “three-quarters from behind and to her left”, it produced the intended framing once in five attempts.
- As a view, “we see the back of her head, her claw clip, the back of her neck”, it worked twice in two, first try.
The same rule fixed two other things in the same run. A monitor that kept coming out dark: “the monitor is the The main light on a subject: the one that decides where the shadows fall. Everything else fills in around it.” became “its glow spills onto the wall in front of her”. An icon that kept getting a label on it: “no other text” became “the rectangle is empty”.
A failure is a prompt bug until proven otherwise
Several outputs wrong in the same way are not bad luck.
When a generation comes back wrong, the instinct is to blame the dice and run it again. Usually the sentence is at fault, and re-rolling a sentence bug is paying again and again for the same mistake.
We spent 44 re-rolling a framing we thought was random. It was one unclear sentence, and every re-roll had failed in exactly the same way. Rewriting it fixed the shot on the next attempt.
Settle it at the cheap layer
Frames before motion. Text before video.
Anything that can be decided in one image should be. Make the image, look at it, fix it, look again, then hand the finished image to the video model.
This matters most for text. An image model spells a headline or a screen correctly; a video model asked to invent the same text at 720p only gets close.
- Four image models, three bits of text each: twelve for twelve spelled correctly.
- The same screen asked for inside a video: believable, readable, and nothing like the real brand.
Test at 480p before you pay for 720p
257 against 578. Same prompt, same references, same length, same model, at a little under half the price.
Dropping to 480p changes nothing a prompt controls. You get the same shots, the same timing, the same performance, the same voices, the same character. What you give up is sharpness, and sharpness is the one thing that was never in doubt.
So run the take at 480p first and read it honestly. It answers everything a brief can get wrong: whether the framing is what you pictured, whether the One beat is one thing happening: a single moment or action in the film. A fifteen-second clip is usually four to six of them in a row. land where you asked, whether the character holds, whether your inputs are accepted. Once it is right you re-run the same brief at 720p for the version you keep.
Going shorter stacks on top of it: five seconds at 480p is 43. What wastes the most money is a full-length 720p take of a brief nobody has looked at yet.
Know which slot you are filling
A first frame is frame zero. References are materials. They cannot be used together.
A first frame means the film begins on that exact image and takes its The shape of the frame: its width against its height. 16:9 is the wide shape a television is; 9:16 is a phone held upright. from it. Reference images are things the model builds a new frame from.
- Want the clip to start on a picture you already have? That’s a first frame.
- Want the clip to contain a person, product or place you already have? Those are references.
Picking the wrong one gives results that look like the model failed: a shot you did not ask for, an aspect ratio you did not set, an image redrawn when you wanted it copied.
Timestamps share out time, they do not cut it
Close enough to structure a film. Not close enough to cut to.
Writing at 12s tells the model roughly how much of the clip that beat should
take up. It is a budget, not an edit point.
- A six-beat sheet produced cuts within 0.83 seconds of every boundary asked for, six for six, five of them inside six tenths.
- Seven timed spoken lines in another take landed zero to two seconds early, never late, with the closing line exact.
Use ranges to set the pace of a sequence and save exact marks for the two or three moments that matter. Pinning down every second makes the pacing worse: too much packed into a window plays rushed, too little plays like slow motion.
The only photoreal person it will animate is one Seedream made
The refusal is about your materials, not your words.
Hand it a photoreal human face, as a first frame or as a reference, and the generation is refused. A photograph of a real person, a photoreal character you made in GPT Image 2, Nano Banana or Flux: inventing the person changes nothing, because what gets inspected is the picture.
Seedream is the exception. Photoreal people made there are accepted, so a photoreal cast is a decision about where the cast gets made, taken before you write a line of the brief.
The refusal arrives before anything renders, so it costs you nothing but the attempt — and it reads pictures, so no rewording gets past it.
Drawn characters are accepted whatever drew them: anime, illustration, 3D. The same brief with a drawn character rendered first time, from a sheet an image model outside the Seed family had made.
So if your piece needs someone who speaks to camera, you have three ways round it. Make them in Seedream, cast a drawn character, or supply no picture of a person at all and describe them in words, since text to video hands the check nothing to inspect and invents a face instead of copying one. That last route costs you the likeness: you get a convincing person, not that person, and nothing holds them still across a second take.
The pipelines
Everything you make starts in one of three states: you have nothing but an idea, you have one picture, or you have a set of materials.
The blocks every brief uses
Whichever pipeline you are in, the prompt is built from the same parts, in the order the model reads most reliably:
[Goal] one sentence: what this video is and what happens in it
[Materials] what each supplied image or audio track is FOR, and what it is not
[Subject] who or what must stay identical from first frame to last
[Look] the colour, the light, the texture, skipped if a frame carries it
[Camera] where it starts, how it moves, how fast, where it ends
[Beats] the sequence, with time ranges and any spoken lines
[Audio] music, effects, voices
[Keep/Avoid] what must not change, and what must never appear
At 4,000 characters all eight blocks fit, written out properly. Only a thirty-second film with several stages and a lot of dialogue comes near the ceiling, so the question is no longer what to leave out. It is what is worth saying.
Delete anything a supplied image already carries. If you gave the model a first frame of a room, do not describe the room. Describe what happens in it. Having the room in the budget is not a reason to spend it: words that repeat a picture add nothing, and words that contradict one make something uncertain that was already settled.
How to write the audio block
Music, effects and speech each have their own brackets. You can write everything in plain sentences instead, but the brackets leave no room for doubt and the model reads them cleanly.
( ) music and background sound (No music until the last beat.)
< > sound effects <rain on glass, a mug set down>
{ } spoken lines She says in French: {Je lis.}
Name the language immediately before each spoken line, not once at the top of
the brief. She says in French: {Je lis.} holds; one rule at the end drifts,
especially when more than one language is in play.
Pipeline A: start from text
Nothing but a written brief.
What it is for
- Best for: mood pieces, Footage with nobody speaking in it, like scenery, hands, a street or a detail, cut in around the main shots to cover a voice-over or to let a scene breathe., The wide shot that opens a scene and tells you where you are before anyone in it does anything., montages, anything with a strong style.
- Costs you: control. You get a good version of your description, not your specific person or product.
- Watch for: faces and objects changing as the clip runs. Nothing is holding them in place.
How to write it
Because nothing is supplied, every block has to work. Put the detail into the subject and the beats, and describe the look in camera words rather than mood words: a named lens, a stated light direction, a grain. “Cinematic” and “beautiful” buy nothing.
Describe a character who comes back once, in detail, then point back to that. Describing them again invites the model to redraw them.
Worked example: A film that is its own brief: thirty seconds, one take, a whole apartment and the woman in it from words.
Pipeline B: start from text with a first frame
You have one picture, and the clip begins on it exactly.
What it is for
- Best for: bringing a picture to life: a product shot, a poster, a character, a screenshot.
- Costs you: reference images. It is one road or the other, never both.
- Watch for: aspect ratio comes from the image, not your setting.
How to write it
This is the pipeline where less prompt is better.
The frame already carries the framing, the clothes, the room, the light and the colour. Describing them again buys nothing and risks contradicting the picture, and if your words and the picture disagree, you have made something uncertain that was already settled. The room in the budget is the trap here: there is space to restate the whole frame, and doing it is how you lose it.
Say “exactly as the first frame” once, then write only what changes: the motion, the performance, the speech, the camera.
The one known limit
Small text falls apart over the length of a clip. In a fifteen-second test, the headline and button on a phone screen stayed perfectly readable while the small address bar turned into something almost, but not quite, itself by the final second.
Anything that must stay readable should be large in frame. Treat small text as texture, not words.
Worked example: Your product, in their hands: fifteen seconds from one still, with the decay visible on the phone.
Pipeline C: start from text with references
You have materials (a cast, a product, a location) and they must survive the clip.
What it is for
- Best for: characters who come back, real products, anything that has to match something else.
- Costs you: the first frame. And a photoreal face is refused here as everywhere unless Seedream made it.
- Watch for: a reference with no job given leaks. A studio backdrop becomes the film’s backdrop.
Every reference needs a job and a boundary
This is the whole skill of the pipeline. A reference with no instructions hands over everything in it: its background, its lighting, its framing, its mood.
[Image1] is <subject>. Use it for <face, build, colour> only.
Do not use its background, lighting, framing or pose.
For characters, supply several separate angles rather than one sheet with them all on it. With thirty slots there is no reason to save space, and a full-size side view carries far more of a face than the same view shrunk into one cell of a grid.
Where a photoreal cast has to come from
Make photoreal people in Seedream. A reference that shows a photoreal human face is refused unless it came from there — a photograph of a real person and a photoreal character you invented in another image model fail the same check, and the run is rejected before it costs you anything. Drawn characters are exempt, whatever made them, which is why the worked example below runs on a sheet from an image model outside the Seed family.
Decide this before you build the cast, because it is not something a prompt can recover from. The rule in full is above.
Voices
Audio references copy a voice. Ten to fifteen seconds is enough, and the model copies how the words are said as well as the sound of the voice, so record the reference in character.
- One voice per track.
- Never music in a voice track.
- Thirty seconds combined, across all audio.
Worked example: A creator who does not exist: one character sheet, thirty seconds, sixteen cuts she survives.
When it breaks
“The generation was refused as sensitive”
Almost always a photoreal human face in what you supplied — either a photograph of a real person, or a photoreal character another image model made for you. It looks at your images, not your words, so a harmless prompt is refused just as fast as a detailed one.
Fix: remake the character in Seedream, whose photoreal people are accepted; use a drawn character, which passes whatever drew it; describe the person in words and supply no picture of them; or set the shot up so no face is visible: from behind, hands only, over the shoulder. Rewriting the prompt will not help.
“My prompt was rejected before anything rendered”
Over 4,000 characters.
Fix: cut what your images already carry. If you supplied a first frame, delete every sentence describing the room, the clothes, the light and the colour. That alone usually wins back a third of the budget. Then cut repeats before content: a rule stated once in the right block beats the same rule stated three times.
“The camera is in the wrong place”
You named an angle instead of a view.
Fix: write what is visible in the frame, and what is not. If two renders get it wrong the same way, that is a sentence problem, not bad luck, so do not re-roll it.
“The aspect ratio or the length isn’t what I asked for”
These are settings, not prompt text. Writing “16:9” in the brief does nothing, and a first frame makes the clip inherit that image’s ratio regardless of the setting.
Fix: set them in the controls, and make the first frame in the ratio you need.
“Text in the frame turned to mush partway through”
The model holds large type and re-invents small type as the clip runs.
Fix: make text large, or accept it as texture. Check the last second, not the first: frame one always looks perfect, because it often is your supplied image.
“The face or product drifts across the clip”
Not enough holding it in place. One reference gives the model less to hold onto than it needs over fifteen seconds.
Fix: supply several separate full-size angles, say that they are all the same subject, and add a clear instruction that the subject never changes or gets redrawn.
“A character spoke the wrong language”
The language was implied, or named once at the end instead of at each line.
Fix: name it immediately before every line. If you need a specific voice, supply an audio reference rather than describing one.
Worked examples
Three finished pieces, one per pipeline, each with the brief that made it. Every one is a single generation, and nothing here is several takes cut together.
A film that is its own brief
Pipeline A: the picture from words alone.
- Length
- 30s, one take
- Frame
- 9:16, 720p
- Supplied
- A logo, two voices
- Cost
- 578
A woman reads a document at one in the morning while a voice behind the camera asks her what it is. At twenty-one seconds the camera arrives at her monitor, and the document turns out to be the brief for the film you are watching.
Nothing you can see was supplied. The woman, the flat, the rain, the light at 1am are all description. A photograph of her face would have been refused, and text to video invents one instead of copying one. What was handed over is not a picture of the scene at all: our logo, so the player in the last shot carries the real wordmark rather than a plausible one, and two voice recordings, one French and one Japanese.
Seven lines were asked for at 2.5, 5, 6.5, 10, 12, 15 and 23 seconds. They came back in order, zero to two seconds early, never late, and the closing line landed on its mark.
The captions are not the model’s. The brief asked for subtitles on screen and got none: every line was spoken and nothing was printed. They were burned in afterwards from a transcript, which is the dependable way round it: write the speech, caption the result.
The brief, in full
Eight thousand characters, twice the ceiling. It was written before we hit the limit and run where there is not one; the version that runs here is the same film compressed to fit, which is what the cutting rule above is for.
[Generation Goal]
Generate a 30-second single-take live-action film, one continuous handheld take with no cuts. The central subject is Camille; the primary event is that she reads a document at 1am while an unseen woman questions her from behind the camera, until the camera arrives at her monitor and the document turns out to be the brief for this film.
[Materials]
[Image1] is the justpictur.ing logo — a rounded dark tile holding a small mark of coloured dots, followed by a lowercase wordmark. It is a graphic asset, not a scene: never place it anywhere in the room and never show it before 21 seconds. From 21s it appears exactly once, reproduced faithfully, as the small logo on the chrome of the video player on her monitor, above the picture area at the left. Keep its shape, its colours and its spelling exactly.
[Audio1] is Camille's voice: French, adult female, low and close.
[Audio2] is the unseen woman's voice: Japanese, adult female, low and close.
[Character]
CAMILLE, a woman in her early thirties of mixed East Asian and European heritage. An oval face with high cheekbones, a straight nose, straight dark brows and dark brown eyes. Warm light-olive skin with visible pores, a light scatter of freckles across her nose and upper cheeks, and natural asymmetry; no makeup. Thick dark brown, almost black, slightly wavy hair worn up in a tortoiseshell claw clip, with loose strands escaping at her temples and the nape of her neck. Slim, average height. She wears an oversized oatmeal wool cardigan over a white t-shirt, sleeves pushed back over her wrists. She is tired and unperformed. She never looks at the camera and never smiles for it.
[Scene]
The working corner of a Paris apartment at one in the morning. A plain plaster wall behind her carries four or five small photographs pinned in a loose group. At the far left a tall window with an off-white curtain half drawn, rain running down the glass, a wrought-iron balcony rail and the blur of the street beyond. She sits at a plain wooden desk in a dark chair, facing a large monitor that stands at the right of frame with its screen turned toward her. A small warm desk lamp sits low behind her right shoulder. On the desk: a ceramic mug, a keyboard, a few loose papers, a notebook. A plant at the top right. Nothing else.
[Look]
Photoreal, 35mm, practical light only. The monitor is the key light and it is cold blue-white — it lights the left side of her face, her hands at the keyboard and the wall in front of her, and it is the brightest thing in the room. The small lamp behind her is the only warm source and reads as a small accent, giving a warm edge to her hair and shoulder. The ceiling light is off. The room reads cool overall with that one warm accent, the background about one and a half stops darker than she is. Real skin with visible pores, fine sensor noise in the darker wall, neutral white balance, low saturation, no teal-and-orange grade, no lens flare, no bloom. Footage of a real room at 1am, not a commercial.
[Camera]
One unbroken handheld take at standing height, a couple of metres from her, seeing her from her left in near profile with the monitor at the right of frame — slight drift, small reframes, brief focus hunting. From 3s it makes one slow continuous arc forward and to the right at 10-15 cm/s, never faster, arriving square-on to her monitor at 21s and holding there, breathing, to the end. It never turns toward the voice, never pans off her, and never stops moving until it arrives.
[Stage 1] 0-10 seconds
Initial state: Camille seen from her left in near profile, her face and hands lit cold by the screen, the lamp warm behind her. She is reading a document on the monitor — reading, not typing. Rain runs down the window at the far left.
Primary event: the camera begins its arc while she is questioned and never turns toward it.
At 2.5s the unseen woman says in Japanese: {まだ来ないの?} 【Are you coming to bed?】
At 5s she says in Japanese: {まだ書いてるの?} 【Still writing?】
At 6.5s Camille, without looking away from the screen, says in French: {Je lis.} 【I'm reading.】
She lifts her mug an inch without looking, finds it cold, sets it back down, then pulls her cuffs over her hands.
End state: the camera is closer and further round to her right, still moving; she is still reading; the mug is back in the same place.
[Stage 2] 10-21 seconds
Continue from Stage 1: same woman, same wardrobe, same room, same seated position, camera still arcing.
Primary event: the face of the monitor becomes visible for the first time and the document proves endless — a production brief of dense monospace blocks and bracketed lines, too far to read, its scrollbar a sliver. She flicks the scroll wheel three times; it barely moves.
At 10s the unseen woman says in Japanese: {何読んでるの?} 【Reading what?】
At 12s Camille says in French: {Le brief.} 【The brief.】
At 15s the unseen woman says in Japanese: {何のブリーフ?} 【The brief for what?】
She does not answer. Nobody speaks for four seconds. She clicks the mouse once.
End state: the camera is almost square to the monitor, close but not yet close enough to read it.
[Stage 3] 21-30 seconds
Continue from Stage 2: same woman, same wardrobe, same room, same seated position.
Primary event: at 21s the camera arrives square-on and the monitor fills the frame, sharp, showing a video player on a flat near-black interface with the [Image1] logo on its chrome above the picture area at the left. Inside the player is this same room and this same woman at this same desk, seen from where this film opened. She clicks again and the picture inside the player begins to move: the same slow drift, rain on the same window inside the frame.
At 23s the unseen woman says in Japanese, quieter: {…待って。それ、今?} 【…Wait. Is that now?】
End state: the monitor full-frame, the player playing, the camera still handheld and breathing. The take ends on motion, not on a cut and not on a black frame.
[Audio]
(No music before 21 seconds. One low sustained synth note enters under the arrival at 21 seconds and holds to the end. Nothing else.)
<steady rain falling on the window glass, close and clearly audible from the first frame to the last, the constant bed the whole film sits on; a fridge compressor in another room; a ceramic mug set down on wood; a chair castor rolling an inch; three scroll-wheel clicks; two single mouse clicks>
The rain is always present and always audible. It is never silenced, never faded under the dialogue and never replaced by music — the voices and the synth note sit on top of it, and between the spoken lines the rain is the loudest thing in the film.
Camille speaks only French, in the voice of [Audio1]. The unseen woman speaks only Japanese, in the voice of [Audio2], further away and off-axis, from a doorway behind the lens. Both unprocessed and quiet — it is 1am. Seven spoken lines only; nobody speaks over anybody.
Every on-screen subtitle is the English translation of the line spoken at that moment. The dialogue channel carries only French and Japanese; the subtitle channel carries only English.
[Maintain Consistency]
Camille is the same woman from the first frame to the last — the same face, the same freckles, the same claw clip and escaping strands, the same oatmeal cardigan over a white t-shirt. The room and every object in it stay exactly where they are. One person on screen; the unseen woman is never seen, never reflected and never lit, and the camera never turns toward her. The rain never stops or changes. The monitor stays the cold key light and the lamp stays the only warm accent. No cuts, dissolves, fades or black frame, including the last. No music before 21 seconds. Camille never looks at the camera and never smiles for it. No on-screen text other than the subtitles and the [Image1] logo. No English speech, and no French or Japanese subtitles.Your product, in their hands
Pipeline B: one first frame.
- Length
- 15s, one take
- Frame
- 9:16, 720p
- Supplied
- One first frame
- Cost
- 11 + 289
The first frame, made first, with a real screenshot as its reference.
An illustrated creator holds a phone to camera and talks through the product in six beats, which is the shape of nearly every ad made to look like a customer filmed it. The still was made first, with a real screenshot bound in as a reference, so the page on the phone was settled before anything moved.
She is drawn rather than photographed because the identical brief with a photoreal person was refused. That is the whole reason for the choice, and it is not a style decision. The other road out is Seedream, whose photoreal people are accepted; drawn was the one taken here, and a drawn character clears the check whatever made it.
The limit is visible in the clip if you watch the phone. The headline and both buttons hold for the full fifteen seconds; the address bar and the small copy under the headline drift into something that is almost, but not quite, English. Large type survives a video model. Small type is texture.
The brief, in full
Style: 2D anime, hand-drawn cel animation matching the first frame exactly — clean line art, flat colour, soft cel shading, painted background. Not photoreal, not 3D, no live action.
Camera: locked off at her eye level. It never moves — no pan, tilt, zoom or cut — only a slight breathe, for all 15 seconds.
Subject: the animated character in the first frame, identical throughout — face, hair and sweatshirt never drift or re-roll. She speaks warm and quick, like someone showing a friend something good, never like an advert.
Screen: she holds a black phone in her right hand in the left third of frame, flat to the viewer, bright and sharp. Its page is frozen exactly as in the first frame and stays legible — it never scrolls, changes, dims or reflects.
0.0-2.5s She tips her head toward the phone, raises her eyebrows and grins at the viewer. She says, "Okay, you have to see this."
2.5-5.0s She turns the phone a few degrees and back, flat again, and nods once. She says, "It's called Just Picturing."
5.0-7.5s She lifts her left hand into frame and points up at the screen. She says, "You describe it, and it makes the picture."
7.5-10.0s Her hand drops away and she brings the phone closer, still flat. She says, "Then it turns that picture into a clip."
10.0-12.5s She flicks three fingers up, counting them off. She says, "Clips, music, whole videos. One app."
12.5-15.0s She pushes the phone closer, taps it twice with her thumb and laughs. She says, "Free to start. That's the site right there."
Audio: her voice only, close and clear in a small room, natural breaths, a short laugh at the end. Two soft thumb taps on glass. No music, no second voice.
Avoid: music, narrator, a second character, captions, subtitles, on-screen text, watermark or logo, any cut or scene change, screen reflections, scrolling or changing the page, any change of room, clothing or light, slow motion, any shift toward photorealism, and any drifting of her face or hair.A creator who does not exist
Pipeline C: one reference.
- Length
- 30s, 16 cuts
- Frame
- 16:9, 720p
- Supplied
- One character sheet
- Cost
- 11 + 578
The one reference, from a seven-word prompt.
A surfer paddles out, catches a wave and comes through the barrel. She is pinned down once in a character sheet and the film is that sheet turned into thirty seconds of footage: two generations, ten minutes apart, the first an input to the second, and both of them first attempts.
The film’s brief is not a mood board, it is a schedule: six timed beats, each naming its own shots. Where a beat lists five images, five cuts appear. Against the six boundaries it asked for, the cuts landed within 0.83 seconds, six for six, and the character holds across sixteen of them.
Note how little the reference is asked to do. It has one job, it is the only input, and no sentence in the brief describes her face.
The sheet was made in GPT Image 2, which is worth saying out loud: she gets through because she is drawn. The identical sheet rendered photoreal there would have been refused, and a photoreal version of this film would have had to start in Seedream.
Both briefs, in full
The character sheet came from the first of these, through the Character Sheet recipe. The film came from the second, with that sheet as its only reference, reproduced verbatim, including the typo at 0–4s.
LANI - a cute surfer girl in bikini0–4s — Quiet hook
Extreme close-up of Lani's eye. Ocean reflected in it. Wind moves loose
strands of hair. She hears a wave building and turns. We her from behind,
focusing on her bottom.
Cut wide: she's standing on the beach with the pink board under her arm.
Almost no dialogue.
4–9s — Preparation montage
Quick, beautiful anime cuts:
fingers tightening the board leash / bracelets moving on her wrist / bare
feet stepping into wet sand / hand brushing hair away from her face / pink
board hitting the water
This establishes all her recognizable details naturally.
9–14s — Paddle out
Camera from water level as she paddles toward an incoming swell.
Then underwater shot: her silhouette and pink board passing above, sunlight
breaking through the surface. Music starts building here.
14–18s — The wave arrives
She looks over her shoulder. Huge turquoise wall forming behind her.
Close-up: tiny confident smile. She turns and paddles hard.
18–25s — PAYOFF
She catches it. Fast sequence: low angle alongside the board / water
spraying across frame / aerial shot of her carving / close-up of her face /
camera goes inside the barrel / she appears through the translucent wave
This is where you spend most of the animation quality.
25–28s — Signature shot
Camera inside the wave, looking outward. Lani passes through the barrel with
sunlight behind her, hair and necklace moving naturally. For maybe half a
second she looks toward camera. Not posing — just part of the motion.
28–30s — Ending
She exits the wave. Cut to very wide shot. Tiny Lani gliding across a huge
ocean as the music drops out and you only hear water. Then cut to black.Everything here runs in the studio.
Every model named on this page is one you can pick from the composer. New accounts start with credits to spend on exactly this.
Start picturing