Prompt systems for ads
The route frames commercial video creation around usable prompts, platform formats, and production outcomes.
Commercial Guide
Read practical guides for ads, product demos, creator monetization, and copy-ready video prompt workflows.
The route focuses on practical ad, product, and monetization workflows.
Prompt blocks and media steps are organized for repeatable output.
Public guide routes help Seedance capture business-intent searches.
ByteDance has released its next-generation video generation model **Seedance 2.5**. For anyone making videos, the change goes beyond image quality alone: a single video can now contain an entire scene, dozens of reference assets can be directed at once, and edits can be pinpointed down to a specific second.
It supports pure text generation, and also generation with simultaneous reference to images, video, and audio, as well as editing of existing videos.
ByteDance has published an official prompt guide: from one-sentence generation, to how to direct 50 reference assets, how to break a 30-second long film into segments, and how to change only one spot when editing a video — every type of task comes with templates that can be copied and used directly.
---
Prompts can be flexibly combined according to the formula below:
**FORMULA: Subject + Action or Event + Scene and Environment (optional) + Visual Style (optional) + Camera Movement or Cut (optional) + Sound (optional)**
| Component | Description |
|------|------|
| **Subject + Action or Event** | States "who or what" is "doing what" — this is the foundation of the video content. The action should prioritize summarizing the main process; write specific details only for key actions, and avoid describing the same action repeatedly. |
| **Scene and Environment** | States the location, time, weather, spatial relationships, and background state. |
| **Visual Style** | States lighting, color, materials, image texture, or the overall atmosphere. |
| **Camera Movement or Cut** | States the shot size, camera position, camera movement, focus target, and shot continuity. |
| **Sound** | States dialogue, voice quality, ambient sound, sound effects, and music. |
<Subject> performs <main action or event> in <scene and environment>. The frame presents <visual style>. The camera uses <shot size, camera position, camera movement or cut>. The sound includes <dialogue, ambient sound, sound effects, or music>.A ceramic artist completes a light-blue ceramic cup in a morning studio, takes it off the wheel, and places it in the center of a wooden shelf. Soft morning light streams in through the window; the damp clay shows a delicate sheen, and the worktable stays tidy. The camera first records the throwing motion in a medium shot, then slowly pushes in toward the texture on the cup's surface, and finally cuts to a frontal view of the wooden shelf. The low hum of the spinning wheel, the sound of clay rubbing, and subtle indoor ambient sound are kept.
**Note**: Parts you don't need can be omitted. Generation parameters don't need to be written into the prompt; adjustable parameters are set on the generation page or interface.
---
Up to **50 reference assets** can be combined and used.
| Asset type | Input range | Recommended range |
|----------|----------|----------|
| **Images** | Up to 30, no single image exceeding 4K | 1 to 8 subjects for subject images is ideal |
| **Videos** | Up to 10 clips, with the total duration of all videos not exceeding 30 seconds | 1 to 5 subjects and 5 to 10 seconds per clip for subject audio/video is ideal |
| **Audio** | Up to 10 clips, with the total duration of all audio not exceeding 30 seconds | Keep dialogue, voice quality, ambient sound, or music directly relevant to the task |
| **Video editing** | Video and reference images can be combined for editing | The original video within 20 seconds and 1 to 5 reference images is ideal |
**Note**: Going beyond the recommended range is still possible, but the more assets there are, the easier it is for stability to drop.
If a subject image has more than 5 subjects and multiple viewpoints are still needed, it is recommended to split the different viewpoints into multiple images; several separate viewpoint images are usually more stable than combining multiple viewpoints into a single image.
After uploading reference assets, you need to state exactly what each asset provides. When people, backgrounds, or compositions in an asset are likely to be mistakenly carried into the finished video, also specify what is not to be used.
**The asset mapping relationships must be written into the prompt.** Don't rely only on text annotations inside the images, and don't let the model decide on its own which person, prop, or scene each of several assets corresponds to.
@Image1 is used for the <appearance, clothing, structure, or material> of <subject>.
@Video1 is used for <action, camera movement, or pacing>.
@Audio1 is used for the <voice quality, dialogue, ambient sound, or music> of <character or sound type>.
<Subject> completes <main action or event> in <scene>. The frame presents <visual style>, and the camera uses <camera expression>.@Image1 is used for the <ceramic artist>'s facial features, hairstyle, and dark green apron; the image background is not used. @Image2 is used for the <ceramic studio>'s wooden worktable, window placement, and morning light; the people in the image are not used. @Video1 is used for the action pacing of throwing clay with both hands, lifting the cup, and setting it down; the character identity, clothing, and scene in the video are not used. The <ceramic artist> completes a light-blue ceramic cup in the morning <ceramic studio>, takes it off the wheel, and places it in the center of a wooden shelf. The camera first records the throwing motion in a medium shot, then slowly pushes in toward the texture on the cup's surface; the sound of the spinning wheel, the friction of clay, and indoor ambient sound are kept.
@Image1 defines the front appearance of the same folding desk lamp. @Image2 defines the left-side structure of the same folding desk lamp. @Image3 defines the right-side structure of the same folding desk lamp. @Image4 defines the back structure of the same folding desk lamp. The four images together define the same folding desk lamp, and there is always only one folding desk lamp in the finished video.
**Tip**: When a reference video has already accurately provided the action, camera movement, and sequence, the prompt only needs to state which content is inherited, without reciting each action; repeated description may conflict with the asset itself.
---
Prompts can use natural language directly. When you need to further distinguish music, sound effects, dialogue, and subtitles, you can use the following special characters:
| Category | Symbol | Example |
|------|------|------|
| **Music** | `· ( )` | `·(Calm piano music plays in the background)` |
| **Sound effect** | `< >` | `<A bell tolls in the distance>` |
| **Dialogue** | `{ }` | `{Hello, welcome back}` |
| **Subtitle** | `[ ]` | `[Chapter One: Departure]` |
When you need to control subtitles or sound, you can state directly which sound categories are kept and which are not. For the visual content, positive descriptions should still be prioritized; only asset roles, editing scope, and content that is easily brought in by mistake need explicit restrictions.
**Examples:**
When the dialogue is not in Chinese, it is recommended to state the language clearly before the dialogue:
The girl speaks softly in Japanese: {もう大丈夫です}
**Recommended formula**:
Dialogue language + Regional variant or accent + Manner of speech + Speaker + {Dialogue content}**Examples**:
Dialogue language: American English. The girl says in natural, colloquial American English: {I thought you weren't coming.}
Dialogue language: authentic Los Angeles American English. The young man says in natural LA slang: {No way, you actually made it.}
---
Images, videos, and audio can be combined and used. When there are many assets, the focus of the prompt is not to write all the assets into a single sentence, but to clarify the correspondence between characters, props, scenes, actions, and sounds.
The following order of organization is recommended:
**ORDER: per-asset role → subject mapping → group by type → subject definitions → call assets per scene**
Different people, products, and props need to be bound to their own assets:
`<Person A> corresponds to @Image1, using only appearance, hairstyle, and clothing. <Person B> corresponds to @Image2, using only appearance, hairstyle, and clothing. <Prop A> corresponds to @Image3, using only structure, material, and color. <Scene A> references @Image4, using only spatial layout, architecture, and lighting, not the people in the image.`
**Don't write "images 1 through 4 each define one of four characters"** — this kind of phrasing does not clarify which image corresponds to which character.
[Characters]<Restorer> corresponds to @Image1, using only appearance, hairstyle, and clothing. <Recorder> corresponds to @Image2, using only appearance, hairstyle, and clothing. <Exhibition Setup Staff> corresponds to @Image3, using only appearance, hairstyle, and clothing. <Guide> corresponds to @Image4, using only appearance, hairstyle, and clothing. The four people's appearance, clothing, actions, positions, and dialogue are not swapped with one another.
[Props]<Sample Box> corresponds to @Image5, belonging only to <Restorer>. <Record Board> corresponds to @Image6, belonging only to <Recorder>.
[Scene]<Restoration Room> references @Image7, using only space, materials, and lighting. <Exhibition Hall> references @Image8, using only space, materials, and lighting.
[Action and Sound]@Video1 is used for <Restorer>'s action of opening the sample box, not the people and scene in the video. @Audio1 is used for <Guide>'s voice quality and specified dialogue.When the same person needs to use multiple assets across scenes, subject definitions can be added:
[Subject Definition: Restorer]Appearance and clothing: @Image1. Fixed prop: the <Sample Box> in @Image5. Places of appearance: <Restoration Room> and <Exhibition Hall>. Action references: @Video1's box-opening action and @Video2's sample-placing action. Not used: other people's clothing; does not hold the <Record Board> or explanation devices.**Scene One | Restoration Room Inspection** uses: <Restorer>, <Sample Box>, <Restoration Room>, and @Video1's box-opening action. Event: <Restorer> opens the sample box in front of the worktable and inspects the samples inside. Ending state: <Restorer> stops on the inner side of the worktable, and the sample box remains at <Restorer>'s own right hand, that is, on the left side of the frame.
**Scene Two | Exhibition Hall Registration** uses: <Recorder>, <Record Board>, and <Exhibition Hall>. Event: <Recorder> checks the numbers on the record board beside the display case. Ending state: the record board is still held by both of <Recorder>'s hands, and no other people enter the display case area.
**Goal**: The goal of multi-asset creation is for the model to choose the correct assets in the current scene; it does not require all assets to appear in the frame at the same time.
---
When there are many events, it is recommended to break the story into continuous stages. Each stage should arrange only one major state change, and clearly write out the state that can be directly seen in the frame when the stage ends.
[Generation Goal]Generate a <video type>. The core subject is <subject>, and the main event is <story summary>.
[Stage One]Beginning state: <initial state of people, props, and scene>. Main event: <one main action or event>. Ending state: <character position, prop ownership, or frame state>.
[Stage Two]Continuing from the previous stage: <state to maintain>. Main event: <one main action or event>. Ending state: <observable state>.
[Stage Three]Main event: <ending event>. Ending state: <final frame state>.
[Keep Consistent]Keep <character identity, number, clothing, prop ownership, spatial orientation, and sound relationships> stable.Generate an explanatory video of a flower shop order packaging process. The <Florist> and the <Shop Assistant> together complete the sorting, wrapping, and delivery of bouquets.
**Stage One**
- Beginning state: the <Florist> stands behind the worktable, with loose flower stems, scissors, and wrapping paper on the tabletop.
- Main event: the <Florist> sorts the flower stems and trims them to length.
- Ending state: the bouquet is held in the <Florist>'s left hand, and the scissors are placed back on the right side of the worktable.
**Stage Two**
- Continuing from the previous stage: the two keep the same identities and clothing, and the bouquet is still held by the <Florist>.
- Main event: the <Shop Assistant> unfolds the wrapping paper, and the <Florist> places the bouquet in and ties it with a green ribbon.
- Ending state: the wrapped bouquet lies flat in the center of the worktable, with the ribbon knot facing the camera.
**Stage Three**
- Main event: the <Shop Assistant> picks up the bouquet and places it on the pickup rack.
- Ending state: the bouquet is only in the center of the pickup rack, and the two stand behind the worktable inspecting the finished product.
**Keep Consistent**: keep the identities, clothing, worktable orientation, scissor position, and bouquet ownership of the <Florist> and the <Shop Assistant> stable.
For ordinary narratives, "stages" should be prioritized. Only use time expressions in 1-second units when key handovers, entrances and exits, transitions, or explicit beats need to be controlled.
**Examples**:
0-5 seconds: show an empty wooden display stand; a hand sets down a white ceramic plate; when it ends, the hand has left, and only the white ceramic plate remains in the center of the display stand.
5-10 seconds: remove the white ceramic plate, then set down a clear glass cup; when it ends, only the clear glass cup remains in the center of the display stand.
10-15 seconds: remove the clear glass cup, then set down a green ceramic bottle; when it ends, only the green ceramic bottle remains in the center of the display stand.
**Note**: Time ranges should be continuous and non-overlapping. Time ranges express the time budget for events, not precise edit points, so actions may run slightly early or late around the boundaries. Too little content in a time range gives the model more room to improvise, while too much content may cause excessive cutting or omitted plot.
**Don't use timestamps to control frequencies such as "complete three actions in one second."**
---
Video editing, first-frame or first-and-last-frame video generation, and video extension will automatically lock certain generation parameters based on the input assets. The specific rules are as follows:
| Task type | Aspect ratio | Duration |
|----------|----------|------|
| **Video editing** | Automatically keeps the input video's ratio, **not separately configurable** | Automatically roughly keeps the input video's duration, **not separately configurable**; there may be a difference of up to about 0.3 seconds |
| **First-frame or first-and-last-frame video generation** | Automatically uses the **ratio of the first-frame image**; the first frame and last frame should use the same aspect ratio | Configurable |
| **Video extension** | Automatically keeps the input video's ratio, **not separately configurable** | Configurable |
The parameters automatically locked in these tasks cannot be specified separately on the generation page or interface; all other parameters follow the currently available options.
---
When editing an existing video, first define the original video as the sole master, then state the editing target, scope of effect, target assets, and content to keep.
[Editing Goal]Edit @Video1, performing <addition, removal, replacement, or adjustment> on <on-screen object, region, or sound category> across <the entire video or a specific time range>.
[Original Video Role]@Video1 is the sole editing master, responsible for <people, scenes, actions, composition, shots, occlusion relationships, sound, and event order>.
[Target Asset Role]@Image1 or @Audio1 is used for <the specified attributes of the target object or sound>.
[Editing Scope]Only process <object, region, time range, or sound category>.
[Content to Keep]Keep <on-screen content, actions, sounds, and time relationships in @Video1 that should not change>.[Editing Goal]Edit @Video1, adjusting the cool blue lighting on the wall at the right of the frame to warm orange lighting only within 4-7 seconds.
[Original Video Role]@Video1 is the sole editing master, responsible for the people, room layout, actions, composition, camera movement, sound, and event order.
[Editing Scope]Only adjust the light color of the right wall and the area it illuminates; the character's skin tone changes naturally with the ambient light.
[Content to Keep]The character's identity, clothing, expression, position, actions, room structure, camera movement, dialogue, and ambient sound keep @Video1.
**Template**:
[Editing Goal]Edit @Video1, modifying only <original object> into <target object>.
[Original Video Role]@Video1 is the sole editing master, responsible for the original scene, camera position, camera movement, action trajectory, occlusion relationships, and event order.
[Target Asset Role]@Image1 is used for the <appearance, structure, or material> of <target object>, not <irrelevant background, people, or composition>.
[Editing Object and Scope]Only modify <explicit object and region>. The number of target objects across the entire video is <number>. Do not modify <content that must be kept>.
[Timeline Inheritance]<Target object> inherits the time points, durations, paths, and speed changes of <original object>'s every appearance, movement, occlusion, and exit. Apart from the explicitly modified objects or regions above, the other people, props, scene content, camera movements, camera cuts, and event order in @Video1 remain unchanged.**Example**:
[Editing Goal]Edit @Video1, replacing only the yellow folding desk lamp in the video with the white folding desk lamp in @Image1.
[Original Video Role]@Video1 is the sole editing master, responsible for the desk, books, hand actions, camera position, camera movement, occlusion relationships, and event order.
[Target Asset Role]@Image1 is used only for the white folding desk lamp's appearance, structure, and material, not the background, composition, or other objects in the image.
[Editing Object and Scope]There is always only one white folding desk lamp throughout the entire video. Only replace the original yellow folding desk lamp; do not modify the books, desk, hands, or background.
[Timeline Inheritance]The white folding desk lamp inherits the time points, paths, and speed changes of the original yellow folding desk lamp's every appearance, lamp-arm rotation, occlusion by the hand, and exit from the frame. Apart from the explicitly modified objects or regions above, the other people, props, scene content, camera movements, camera cuts, and event order in @Video1 remain unchanged.
**Template**:
[Editing Goal]Edit @Video1, replacing only <original background region> with <target environment> in @Image1.
[Original Video Role]@Video1 is the sole editing master, responsible for the people, foreground objects, actions, composition, camera movement, and event order.
[Target Asset Role]@Image1 is used only for the <target environment>'s spatial layout, materials, depth of field, environmental colors, and light direction, not the people or foreground objects in the image.
[Editing Object and Scope]Only modify <background regions outside the subject's outline>. Do not modify <subject identity, facial features, hairstyle, clothing, expression, position, size, and actions>.
[Timeline Inheritance]The character's actions and occlusion relationships keep @Video1. Apart from the explicitly modified objects or regions above, the other people, props, scene content, camera movements, camera cuts, and event order in @Video1 remain unchanged.**Example**:
@Video1 is the sole editing master, responsible for the people, actions, composition, shots, and event order. @Image1 provides only the daytime glass greenhouse's spatial layout, depth of field, environmental colors, and light direction, not the people in the image. Only replace the light-gray background outside the character's outline in @Video1 with the daytime glass greenhouse in @Image1. The character's identity, facial features, hairstyle, clothing, expression, position, size, and raising-hand action keep @Video1. Apart from the explicitly modified objects or regions above, the other people, props, scene content, camera movements, camera cuts, and event order in @Video1 remain unchanged.
Dialogue, language, voice quality, background music, and sound effects can be handled separately. The prompt needs to state the speaker or sound category, the target change, and whether other sounds are kept.
**Examples**:
Edit @Video1, removing only the original background music, keeping the character dialogue, lip movements, ambient sound, and action sound effects; the on-screen content, shots, and editing pacing keep @Video1.
Edit @Video1, adjusting the <guide>'s dialogue language to natural American English, keeping the dialogue content and the timing of speech unchanged; the other people's voices, background music, ambient sound, and picture keep @Video1.
---
Video extension continues creating beyond the boundaries of the original video. The output ratio automatically keeps the input video's ratio, and the extension duration is set by the user.
| Direction | Boundary requirement |
|------|----------|
| **Extending backward** | The first frame of the extended segment needs to continue from the original video's last frame |
| **Extending forward** | The last frame of the extended segment needs to connect with the original video's first frame |
In addition to the boundary frames, the people, props, background, motion trend, and sound must also remain continuous.
First describe the continuous state of the original video's last frame, then describe what happens after the last frame.
**Basic template**:
@Video1 is the original video to be extended backward. Extend @Video1 backward. The first frame of the extended segment continues directly from the last frame of @Video1: keep <subject pose and orientation>, <prop position>, <background and spatial relationships>, <camera position and composition>, <lighting>, <sound state>, and <motion trend> continuous. Then, <describe the new actions, events, shots, or sounds to be extended>. Throughout the extension, keep <character identity and clothing>, <key props>, <background layout>, <camera axis>, and <original sound environment> continuous. The same subject always remains a single continuous object, never duplicated or split; the number of character features or object parts stays stable.**Example**:
@Video1 is the original video to be extended backward. Extend @Video1 backward. The first frame of the extended segment continues directly from the last frame of @Video1: keep the same fixed medium shot, the position and orientation of the orange paper airplane, the classroom window-side background, the afternoon light, and the motion trend flying toward the right of the frame continuous. Then let the orange paper airplane continue gliding toward the right of the frame and fly out of the shot, with the white curtain by the window swaying slightly. The camera and classroom background keep the last-frame state of the original video.
First state one by one what the other assets are responsible for, then clarify that the original video is responsible for the extension boundary.
**Template**:
@Image1 defines <Person A>'s facial features. @Image2 defines <Person A>'s clothing. @Image3 defines the structure and material of <key prop>. @Video1 is the original video to be extended backward. Extend @Video1 backward. The first frame of the extended segment continues directly from the last frame of @Video1: keep <boundary frame and sound state> continuous. Then, <the new action or event completed by Person A using the key prop>. Throughout the extension, keep <character identity and clothing>, <key props>, <background layout>, <camera axis>, and <original sound environment> continuous. The same subject always remains a single continuous object, never duplicated or split; the number of character features or object parts stays stable.**Example**:
@Image1 defines the <gardener>'s facial features. @Image2 defines the <gardener>'s light-green work apron. @Image3 defines the structure and material of the <woven rattan flower basket>. @Video1 is the original video to be extended backward. Extend @Video1 backward. The first frame of the extended segment continues directly from the last frame of @Video1: keep the greenhouse worktable, the <gardener>'s position, and the position of the <woven rattan flower basket> continuous. Then, the <gardener> lifts the <woven rattan flower basket> with both hands and places it on the middle shelf of the rack behind. Throughout the extension, keep the <gardener>'s face, apron, greenhouse layout, and camera direction continuous.
First describe what happens before the original video begins, then write the original video's first frame as the explicit ending state of the extended segment.
**Basic template**:
@Video1 is the original video to be extended forward. Extend @Video1 forward. Before the original video begins, <describe the preceding actions, events, shots, or sounds>. The last frame of the extended segment naturally connects to the first frame of @Video1: <subject pose and orientation>, <prop position>, <background and spatial relationships>; keep <camera position and composition>, <lighting>, <sound state>, and <motion trend> consistent with the first frame of @Video1. Throughout the extension, keep <character identity and clothing>, <key props>, <background layout>, <camera axis>, and <original sound environment> continuous. The same subject always remains a single continuous object, never duplicated or split; the number of character features or object parts stays stable.**Example**:
@Video1 is the original video to be extended forward. Extend @Video1 forward. Before the original video begins, present an empty shot of the same glass greenhouse: morning mist slowly disperses close to the ground, the top shading blinds gradually rise, and there are temporarily no people in the frame. The last frame of the extended segment naturally connects to the first frame of @Video1: keep the central aisle of the greenhouse, the planting beds on both sides, the glass framework, the soft morning light, and the fixed wide-angle composition consistent with the first frame of @Video1; when it ends, the shading blinds are fully raised, the aisle has no people, and the leaves are still swaying slightly.
Declare each asset's role one by one, and state which assets are used in the forward extension segment and which are used only after the original video begins.
**Template**:
@Image1 defines <Person A>'s facial features. @Image2 defines <Person A>'s clothing. @Image3 defines the structure and material of <key prop>. @Video1 is the original video to be extended forward. Extend @Video1 forward. Before the original video begins, <Person A completes the preceding action or event>. The last frame of the extended segment naturally connects to the first frame of @Video1: <Person A's pose and orientation>, <the key prop's position and state>, <the positions of the other people>; keep <background and spatial relationships>, <camera position and composition>, <lighting>, <sound state>, and <motion trend> consistent with the first frame of @Video1. Throughout the extension, keep <character identity and clothing>, <key props>, <background layout>, <camera axis>, and <original sound environment> continuous. The same subject always remains a single continuous object, never duplicated or split; the number of character features or object parts stays stable. <Assets that only appear after the original video begins> do not appear early in the forward extension segment.**Example**:
@Image1 defines the <curator>'s facial features. @Image2 defines the <curator>'s dark-blue work jacket. @Image3 defines the structure and material of the <wooden display box>. @Image4 defines the gray work uniforms of the two <exhibition setup assistants>. @Image5 defines the space and lighting of the <exhibition preparation room>. @Video1 is the original video to be extended forward. Extend @Video1 forward. Before the original video begins, the <curator> walks to the worktable, picks up the closed <wooden display box>, and opens its lid. The last frame of the extended segment naturally connects to the first frame of @Video1: the <curator> stands in the center of the frame, holding the opened <wooden display box> with both hands; the two <exhibition setup assistants> stand to the left and right behind him respectively. Keep the vertical frontal medium shot, worktable position, preparation-room background, and left-side morning light consistent with the first frame of @Video1.
**Note**: The boundary frames should aim for a visually natural connection, and should not be understood as being pixel-by-pixel identical.
The volume of the extension result may vary slightly from the original video; when videos generated by the same-generation model are used for further extension, the picture and sound usually connect more naturally. During acceptance, both the picture and sound before and after the boundary, as well as the complete extended segment, need to be checked.
---
In multimodal reference mode, you can directly write @Image1 as the first frame and @Image2 as the last frame in the first sentence of the prompt, without switching to a separate first/last-frame mode. The system locks the output ratio to the first frame's aspect ratio, and the duration is set on the generation page or interface; the first frame and last frame should use the same aspect ratio, and the last frame may be stretched if the ratios differ.
**Basic template**:
@Image1 serves as the first frame, defining the composition, subject position, pose, prop state, scene, and camera direction at the start of the video.
@Image2 serves as the last frame, defining the composition, subject position, pose, prop state, scene, and camera direction at the end of the video.
@Image3 is used for the <appearance, clothing, structure, or material> of <Subject A>, without changing the first-frame composition defined by @Image1 or the last-frame composition defined by @Image2.
@Image4 is used for the <specified attributes> of <Subject B, prop, or scene>, without changing the first-frame composition defined by @Image1 or the last-frame composition defined by @Image2.
<Describe a continuous action or event>. The picture naturally starts from the first frame defined by @Image1, passes through continuous actions, and arrives at the last frame defined by @Image2. Between the first and last frames, keep <character identity, prop structure and ownership, scene layout, and camera direction> continuous.**Example**:
@Image1 serves as the first frame, defining the perfume workshop's composition, character position, pose, tabletop prop state, and camera direction at the start of the video. @Image2 serves as the last frame, defining the perfume workshop's composition, character position, pose, tabletop prop state, and camera direction at the end of the video. @Image3 is used for the <perfumer>'s face, hairstyle, and dark green apron, without changing the first-frame composition defined by @Image1 or the last-frame composition defined by @Image2. @Image4 is used for the <glass perfume bottle>'s shape, material, and label position, without changing the first-frame composition defined by @Image1 or the last-frame composition defined by @Image2. The <perfumer> starts from the first-frame pose, picks up the dropper and the <glass perfume bottle>, drips the amber perfume essence into the bottle, gently swirls to mix, caps the bottle, places the finished product in the center of the tabletop, and finally naturally arrives at the last frame defined by @Image2. Between the first and last frames, keep the <perfumer>'s identity and clothing, the number and structure of the perfume bottles, the wooden table layout, the warm side lighting, and the camera direction continuous.
**Note**: Each anchor image needs to be explained individually; do not combine them into "images 1 and 2 serve as the first and last frames." The first frame and last frame should use the same aspect ratio. The other reference images only provide the specified attributes; they do not replace the first/last-frame composition.
When multiple independent images each define a different stage of the process, write "use @Image1 through @ImageN in order as keyframes" as the first sentence of the prompt, then state the key state corresponding to each image one by one.
**Basic template**:
Use @Image1 through @ImageN in order as keyframes.
@Image1 serves as the first frame, defining <the composition, subject position, pose, prop state, and camera direction at the start>.
@Image2 defines the second keyframe: <the visible state when the first stage ends>.
@Image3 defines the third keyframe: <the visible state when the second stage ends>.
@ImageN serves as the last frame, defining <the composition, subject position, pose, prop state, and camera direction at the end>.
The picture passes in sequence through the states defined by @Image1, @Image2, @Image3, up to @ImageN, with continuous actions providing natural transitions between stages. Throughout the process, keep <subject identity, prop structure and ownership, scene layout, lighting, and camera axis> continuous.**Example**:
Use @Image1 through @Image4 in order as keyframes. @Image1 serves as the first frame, defining the orange paper airplane resting on the left side of the classroom wooden desk, with its nose pointing toward the right of the frame, in a fixed medium shot. @Image2 defines the second keyframe: the same orange paper airplane is lifted from the desktop by a hand, with the nose direction unchanged. @Image3 defines the third keyframe: the same orange paper airplane sweeps past the window, with the curtain swaying slightly to the right. @Image4 serves as the last frame, defining the same orange paper airplane landing on the middle shelf of the bookcase on the right, with its nose still pointing toward the right of the frame. The picture passes in sequence through the states defined by @Image1, @Image2, @Image3, and @Image4, keeping the flight direction and speed continuous between stages. Throughout the process, keep the paper airplane's orange material, size, and creases, as well as the classroom layout, afternoon side lighting, and camera axis continuous.
Grid storyboards are used to provide the overall story, shot order, and rough composition; they are not suitable for requiring strict reproduction of every cell's details. It is recommended to keep them within 15 cells, using simple line drawings or clean schematic images, and to reduce text annotations.
**Basic template**:
@Image1 provides the <shot order and rough composition>, read in the <left-to-right, top-to-bottom> order; the image's <line-drawing style, text annotations, or placeholder people> are not used.
@Image2 defines the <appearance and clothing> of <Subject A>.
@Image3 defines the <structure, material, or lighting> of <key prop or scene>.
Shot 1: <shot size, subject action, scene state>.
Shot 2: <shot size, subject action, camera movement or transition method>.
...
Shot N: <ending action and ending frame state>.
The final picture uses <visual style>, and the sound includes <dialogue, ambient sound, action sound effects, or music>.**Example**:
@Image1 provides the shot order and rough composition of a four-panel pottery-making storyboard, read in the left-to-right, top-to-bottom order; the image's line-drawing style and text annotations are not used. @Image2 defines the <ceramic artist>'s face, short hair, and dark gray apron. @Image3 defines the <blue-glazed cup>'s body proportions, glaze color, and curved handle.
Shot 1: a wide shot presents a quiet pottery studio, with the <ceramic artist> sitting in front of the pottery wheel.
Shot 2: a medium side shot of the <ceramic artist>'s hands steadying the spinning wet clay body as the cup takes shape.
Shot 3: a close-up of fingers trimming the joint between the cup rim and the handle, with clay slurry slowly sliding along the fingertips.
Shot 4: a medium close-up presents the fired <blue-glazed cup> being placed onto the wooden shelf, with the <ceramic artist> withdrawing both hands.
The final picture uses a realistic documentary texture, keeping the sound of the pottery wheel spinning, the friction of wet clay, and the studio ambient sound.
White-box model references are divided into coarse-grained and fine-grained white-box models.
| Type | Suitable for | Asset requirements | Prompt focus |
|------|----------|----------|------------|
| **Coarse-grained white-box model** | Rehearsing actions, movement lines, positions, camera movement, or cuts with simple geometric shapes | Clear geometric relationships and a complete action sequence; character, prop, and scene images can be overlaid | Map each white-box subject one by one, and state which temporal and spatial information to inherit |
| **Fine-grained white-box model** | A complete model already exists and characters, materials, colors, scenes, or styles need to be replaced | Complete model structure and a clean image, avoiding trajectory lines, coordinate lines, and camera frustums | Keep the structure, actions, and camera, and clearly state which attributes need to be re-rendered |
Coarse-grained white-box models are suitable for locking down action trajectories, motion directions, character positions, entrances and exits, camera paths, cut positions, lighting changes, and sound pacing.
**White-box information → content that needs to be stated**:
| Information type | Content to state |
|----------|----------|
| **Movement line** | Action trajectory, motion direction, character positions, and entrance/exit order |
| **Camera movement** | Camera position, camera path, motion direction, and speed changes |
| **Lighting** | Light direction, brightness changes, and the timing of the changes |
| **Cut** | The cut position, and the subject and composition before and after the cut |
| **Sound** | Whether dialogue, music, ambient sound, or action sound effects are inherited |
**Basic template**:
@Video1 is a coarse-grained white-box model reference, providing only <action path, character positions, camera position, camera movement, cuts, lighting changes, sound pacing, or spatial relationships>, not its white-box appearance, materials, and scene.
The <white-box Subject A> in @Video1 corresponds to <Subject A>.
The <white-box Subject B or geometric prop> in @Video1 corresponds to <Subject B or key prop>.
@Image1 defines the <appearance, clothing, or structure> of <Subject A>.
@Image2 defines the <specified attributes> of <Subject B, key prop, or scene>.
<Subject> completes <main action or event> in <scene>. Keep the <action path, positions, camera movement, cuts, lighting, or sound pacing> in @Video1. The final picture uses <people, scene, materials, and visual style>, and the sound includes <dialogue, ambient sound, or action sound effects>.**Example**:
@Video1 is a coarse-grained white-box model reference, providing only the character's walking path, the direction of the mobile display cart's movement, a fixed camera position, one push-in, and two cuts, not its gray geometric appearance and empty scene. The tall cylinder in @Video1 corresponds to the <guide>. The rectangular box in @Video1 corresponds to the <mobile display cart>. @Image1 defines the <guide>'s face, blue uniform, and name badge. @Image2 defines the <mobile display cart>'s white metal frame and transparent display cover. @Image3 defines the technology exhibition hall's curved walls, gray floor, and overhead linear lights. The <guide> pushes the <mobile display cart> forward along the curved wall, stops in front of the central display platform, and opens the transparent display cover. Keep the walking path, subject positions, push-in direction, and cut positions in @Video1. The picture uses a bright, realistic exhibition-hall documentary style, keeping footsteps, the sound of wheels, and the hall ambient sound.
A fine-grained white-box model already has complete character, prop, or scene structure, and is suitable for replacing materials, colors, character images, scenes, and the overall visual style. The white-box image should be as clean as possible, without production markers such as trajectory lines, coordinate axes, controllers, or camera frustums.
**Basic template**:
@Video1 is a fine-grained white-box model reference, keeping <subject structure, actions, spatial layout, camera position, camera movement, and cuts>, not the original gray-model materials and blank background.
@Image1 defines the <character image, material, color, or surface details> of <subject>.
@Image2 defines the <space, materials, lighting, or visual style> of <scene>.
Re-render the <subject> in @Video1 into <final subject> and re-render the scene into <final scene>. Keep the <structure, actions, camera, and spatial relationships> in @Video1; the picture presents <materials, colors, and style>, and the sound includes <ambient sound, sound effects, or music>.**Example**:
@Video1 is a fine-grained white-box model reference, keeping the ring device's complete structure, the three-layer ring rotation relationship, the display platform position, the orbiting camera movement, and the cuts, not the original gray-model materials and blank background. @Image1 defines the brushed brass material of the device's outer ring. @Image2 defines the translucent blue glass material of the inner blades. @Image3 defines the contemporary art gallery's white curved walls, dark gray floor, and overhead soft lighting. Re-render the ring device in @Video1 into a dynamic sculpture of brass and blue glass, and re-render the scene into a contemporary art gallery. Keep the structure, rotation pacing, orbiting camera movement, and cuts in @Video1, keeping the device's low rotation sound and the quiet indoor ambient sound.
---
One-click video creation is suitable for organizing multiple images, or images plus a style reference video, into a complete video with unified pacing and packaging style. The prompt needs to state each asset's role, the image order, on-screen dynamics, editing pacing, visual packaging, and sound; don't just write "make these assets into a video."
**ORDER: asset roles → image order → dynamics amplitude → editing style → visual packaging → sound**
[Asset Roles]@Image1 is used for <a person, product, scene, or opening shot>. @Image2 is used for <a person, product, scene, or process shot>. @Image3 is used for <a person, product, scene, or ending shot>. @Video1 is used only for <editing pacing, transitions, subtitle packaging, or music style>, not its character identities and scenes (optional).
[Arrangement]The images appear in <upload order, specified order, or freely arranged by theme>. <State the relationships between people, products, places, and events that need to be kept>.
[On-screen Dynamics]Each image uses <subtle Live motion effects, parallax, push-pull, lateral movement, or localized motion>. Keep <subject appearance, product structure, text, or background relationships> stable.
[Finished Video Style]Use <editing pacing, transition method, subtitle or graphic packaging, color style>.
[Sound]Includes <dialogue, ambient sound, sound effects, or music>.**Asset Roles**
- @Image1 is used for the night market entrance and the opening environment.
- @Image2 is used for the <traveler>'s shot walking along the street.
- @Image3 is used for the lantern stall and handcraft details.
- @Image4 is used for three friends dining together around a table.
- @Image5 is used for the riverside night scene and reflections.
- @Image6 is used for the ending shot of the three people taking a group photo by the bridge.
- @Video1 is used only for the brisk editing pacing, hand-drawn stickers, and transition method, not its character identities and places.
**Arrangement**
The pictures appear in the order of @Image1 through @Image6, forming a complete process of "arriving at the night market, strolling along the street, dining, walking, and taking a group photo." Keep the three friends' appearance and clothing stable, without mixing them together.
**On-screen Dynamics**
The environment images use a slow push-in and subtle parallax; the people images only add natural blinking, head-turning, raising a glass, and clothing swaying in the wind. Keep the stall structure, table position, and bridge railing stable.
**Finished Video Style**
Use a bright travel-short-film pacing, connecting scenes with natural occlusion and similar colors, with hand-drawn stickers appearing only at the edges of the frame.
**Sound**
Keep the night market crowd sounds, the light clinking of utensils, and the riverside wind, accompanied by brisk but not overly prominent instrumental music.
**Note**: When the image order matters, the order of appearance should be written out for each image; when the model is allowed to arrange freely, this should also be stated explicitly as "may be arranged freely by theme." When multiple people or multiple products are involved, they still need to be named one by one and bound to assets.
---
A seamless video transition generates continuous transitional content between two videos. The prompt needs to first state which video is the pre-transition video and which is the post-transition video, then write out the trigger action, camera movement, how the picture changes, the final state reached, and the sound connection.
**ORDER: pre-transition video → post-transition video → trigger action → camera movement → picture transformation → reached state → sound**
| Transition method | Writing focus |
|----------|----------|
| **Dive or backtrack** | State the camera direction, speed changes, and when to enter the next scene |
| **Character rotation** | State the character pose, rotation direction, and how the clothing or background changes continuously |
| **Foreground occlusion** | State when the occluding object covers the full frame, and the composition that appears after the occlusion |
| **Object transformation** | State the shapes, materials, and transformation process of the objects before and after |
| **Push-pull or focus change** | State the camera movement, focus target, and the continuous relationship between the spaces before and after |
@Video1 is the pre-transition segment, using its <trailing subject, action, composition, camera direction, and sound>.
@Video2 is the post-transition segment, using its <opening subject, composition, camera direction, and sound>.
Keep the <character identities, product structures, scenes, and main actions> in the original segments of @Video1 and @Video2 stable.
At the tail of @Video1, the <subject or foreground object> triggers the transition through <action>. The camera <motion direction and speed changes>, and the <shape, material, lighting, or space> in the picture gradually changes into the <corresponding element> at the opening of @Video2. When the transition ends, it naturally arrives at the opening composition of @Video2, keeping <subject position, camera direction, and motion trend> continuous.
The sound transitions smoothly from <pre-segment sound> to <post-segment sound>.@Video1 is the pre-transition segment, using a rainy night street, a red umbrella, a slow push-in camera, and the sound of rain. @Video2 is the post-transition segment, using a circular skylight in the exhibition hall, an upward-rising camera, and quiet indoor reverb. Keep the people, street, exhibition hall structure, and main actions in the two original videos stable. At the tail of @Video1, the red umbrella approaches the camera and gradually covers the entire frame, triggering the transition. The camera continues to push forward, the circular edge of the umbrella gradually becomes the metal ring of the exhibition hall skylight, and the red umbrella surface gradually transitions into the white daylight filtering through the skylight. When the transition ends, it naturally arrives at the low-angle composition at the opening of @Video2, with the camera smoothly shifting from a push-in to an upward rise. The rain sound gradually fades and smoothly transitions into the echo of footsteps inside the exhibition hall.
**Note**: The goal of a seamless transition is for the picture and sound to connect naturally. The prompt can require keeping the main content of the two original videos, but the generated transition should not be understood as a pixel-by-pixel-unchanged edit splice.
---
Writing only emotion words such as "tense, warm, oppressive" lets the model understand the overall direction, but the specific performance leaves more room for interpretation. If you want to control a character's acting stably, it is recommended to state the observable performances that can be directly seen or heard, such as the eyes, eyebrows, mouth corners, breathing, gaze, and hand movements.
There is no need to list every facial detail. For a single emotional turn, choosing 2 to 4 of the clearest performances is usually enough; only when the emotion involves multiple turns do you need to describe it in stages by triggering event.
The overall emotion shifts from <starting emotion> to <ending emotion>. After <triggering event> occurs, <subject> first shows <immediate observable reaction>. Then, <eyes, eyebrows, mouth corners, breathing, gaze, or hand movements> gradually undergo <change>. Finally, <subject> expresses <target emotion> through <restrained or explicit external performance>.When hearing or seeing <the first triggering event>, <the subject's first observable reaction>. After <the second triggering event> appears, <the changes in the subject's expression, gaze, or breathing>. After confirming <key information>, <the emotion the subject tries to restrain or conceal> gradually shows through <observable performances>. Finally, <the subject's final action, expression, or manner of speech>.Applause for the end of the performance comes from behind the stage. The young actress's fingers holding the program suddenly stop, her gaze slowly turns toward the direction of the curtain, and her shoulders remain tense. After confirming the curtain call, she first lets out a soft breath, her shoulders gradually relax, a restrained smile appears at the corners of her mouth, and her eyes slowly grow moist, but she never turns to leave.
---
Basic shot language and popular camera movements can be written directly into the prompt. When a term is uncommon, may have multiple interpretations, or needs precise control of the picture change, you should also state who the term acts on, how the picture changes, and the result you want to see.
| Category | Terms |
|------|------|
| **Shot size** | Extreme wide shot, wide shot, medium shot, close-up, extreme close-up |
| **Camera movement** | Push in, pull out, pan, tilt, dolly, tracking, orbit, dive, pull back, tilt up, handheld shake |
| **Camera position and viewpoint** | Low angle, top-down/high angle, first-person |
| Camera movement | Recommended writing |
|----------|----------|
| **One-shot / long take** | State the continuous order of subjects, spaces, and events the camera passes through |
| **Dolly zoom** | State the size the subject keeps, and the effect of the background space being pulled closer or pushed farther |
| **Aerial view** | State the overhead height, movement direction, and the range of environment to be shown |
| **FPV** | State the first-person flight or traversal path, speed, and turning |
| **Bullet time** | State the action being frozen or slowed down, and the direction of the camera orbit |
| **Handheld shot** | State the tracking target and the degree of shake, avoiding a mere "handheld feel" without a camera subject |
| **Rubber-hose speed / snapback speed change** | State at which point the action accelerates, decelerates, or snaps back, and the final resting state |
For terms that are overly niche, have inconsistent industry meanings, or may be unfamiliar to the model, keep the term name while also translating it into directly observable picture changes.
**FORMULA: professional term + acting subject + picture change + foreground/background relationship + direction or speed**
**Example**:
Rack focus: the focus smoothly shifts from the foreground leaves to the background person. The leaves gradually blur, and the background person's face transitions from blurry to sharp.
---
---
---
The examples in this guide are used only to illustrate prompt writing. Actual generation results may be affected by the input assets, task complexity, and generation parameters.
The commercial guide route contains the Sunday blog experience under a clearer public URL for users looking for practical AI video business workflows.
Dynamic guide articles load in the client app, while static HTML gives crawlers a stable index page.
The route frames commercial video creation around usable prompts, platform formats, and production outcomes.
Creators can study how reference images, product shots, and delivery formats fit together before generating.
Static HTML makes the guide hub readable even when the full article list is loaded client-side.
It covers ad prompts, product demos, monetization workflows, creator offers, and repeatable media pipelines that can be reused in Seedance.
Guide pages often target search-driven traffic, so route summaries and business-use context should be visible directly in source HTML.