HOME>Blog>How to Use Seedance 2.5: A Prompt Framework for References, Continuity, Audio, and Text
Guide15 min read

How to Use Seedance 2.5: A Prompt Framework for References, Continuity, Audio, and Text

Seedance 2.5 is more than a longer video generator. Its official guides show a practical way to define reference roles, stage a 30-second story, preserve continuity, and separate dialogue, sound effects, music, and subtitles.

Kylon TeamProduct

Official Seedance 2.5 guides

These two ByteDance guides are the primary references for the prompt structures and examples in this article.

The short version

Seedance 2.5 works best when a prompt is treated as a production brief, not a single descriptive sentence. The official User Guide and Prompt Guide make that clear: identify what each reference contributes, divide a longer story into observable stages, state what must remain consistent, and specify audio and on-screen text with separate syntax.

That is why Seedance 2.5 is attracting attention alongside MiniMax H3. Both launches point toward multimodal video systems that take more of the production brief as input. Seedance's documentation is especially useful because it explains how to structure that brief.

What Seedance 2.5 adds

ByteDance launched Seedance 2.5 on July 31. The company says the model can generate audio-video clips of up to 30 seconds in one pass, supports multi-round extension, accepts image, video, and audio references, and offers timestamp-level editing controls. It is rolling out through Jimeng AI and Doubao Pro, with API access planned through BytePlus ModelArk.

The product announcement matters. The guides matter more for teams trying to get repeatable output.

Read ByteDance's Seedance 2.5 announcement and the official User Guide and Prompt Guide.

The core formula: start with the production brief

Seedance's Prompt Guide proposes a flexible structure:

Subject + action or event + scene and environment + visual style + camera movement or cut + audio

Not every prompt needs every field. But separating them prevents a familiar failure mode: a single dense paragraph where the important requirements compete with one another.

Prompt layerWhat to defineWhy it matters
Subject and eventWho or what is acting, and what changesEstablishes the narrative center
SceneLocation, time, spatial relationship, background stateAnchors the environment
Visual treatmentLighting, materials, color, image texture, moodSets the visual language
CameraShot size, angle, movement, cut, focusMakes direction explicit
AudioDialogue, voice, ambience, sound effects, musicKeeps sound intentional rather than incidental

For a product team, this turns a request such as “make a teaser” into a usable creative brief: product, action, environment, camera plan, and audio plan each have a place.

1. Give every reference a role

The strongest rule in the Prompt Guide is simple: do not upload references and expect the model to infer their roles.

Instead, map each one explicitly:

@Image 1 defines the product's front view, material, and brand colors.
Do not use its background.

@Image 2 defines the studio layout and warm afternoon lighting.
Do not use the people in the image.

@Video 1 defines the camera pacing and the hand movement.
Do not use the person, clothing, or setting from the video.

@Audio 1 defines the narrator's calm voice and room ambience.

This syntax is valuable because it separates what to inherit from what to exclude. A product reference can define shape and color without importing an unwanted background. A motion reference can define pacing without carrying over the source person's identity.

The guide also recommends describing multi-angle materials as one object. For example, name each view of the same product and state that the output contains only one instance throughout. That reduces ambiguity when the model must maintain an object across cuts.

2. Write a 30-second story as stages and end states

Longer generation changes the prompt problem. A 30-second sequence often contains more than one event, so the guide recommends consecutive stages. Each stage should contain one primary state change and an end state that is visible on screen.

[Generation Goal]
Generate a 30-second product launch film. The central subject is one olive-green desk lamp.

[Stage 1]
Initial state: the lamp is folded on a pale wood desk in a quiet studio.
Primary event: a hand unfolds the lamp and turns it toward the desk surface.
End state: the lamp stands upright, facing left, with the light still off.

[Stage 2]
Continue from the previous stage: keep the same lamp, desk, left-facing direction, and camera side.
Primary event: the hand turns on the lamp and places a book beneath it.
End state: the warm pool of light falls on the open book.

[Stage 3]
Primary event: the camera pulls back to reveal the complete workspace.
End state: one lamp remains on the desk, the book stays open, and the room is quiet.

[Maintain Consistency]
Keep the lamp's identity, color, hinge structure, desk position, light direction, and room ambience consistent.

The practical benefit is operational: reviewers can identify which stage needs adjustment instead of discarding an entire prompt. It also gives the model observable checkpoints for continuity.

3. Use precise timing only for critical moments

The official guidance recommends stages for normal narratives. Use timestamps when a handoff, entrance, exit, cut, or beat has to land at a particular moment.

0–5 seconds: The folded lamp sits alone on the desk. A hand enters and unfolds it.
End state: the hand has left the frame and the lamp is upright.

5–12 seconds: Turn on the lamp. The camera tracks across the desk toward the book.
End state: warm light reaches the open pages.

12–20 seconds: Pull back to a medium-wide view of the workspace.
End state: the lamp, book, and desk remain in their established positions.

This is not an instruction to timestamp every movement. It is a way to protect the beats that affect continuity or editorial timing.

4. Keep continuity in its own instruction block

A good continuity instruction is concrete. “Keep it consistent” is not enough. The guide calls out the elements worth naming: character identity, character count, clothing, prop ownership, spatial direction, and audio relationships.

For branded work, add product geometry, logo treatment, and approved copy rules. If a wordmark or interface label must be accurate, use a real asset or add it in post-production rather than relying on generative rendering.

The same discipline applies to extension. Before describing the next event, establish what the last frame already contains. Then describe only the new action. Otherwise, the extension prompt can accidentally reset the scene.

5. Treat audio and text as separate layers

Seedance supports natural-language audio direction, but the Prompt Guide documents syntax for cases where separation helps:

  • () for music
  • <> for sound effects
  • {} for dialogue

For example:

(Soft rhythmic piano plays under the scene.)
<The lamp switch clicks once.>
Dialogue language: Japanese. The presenter says in a calm, natural Japanese voice: {今日は、作業に合う明かりを選びます。}

If a language or regional variety matters, state it before the dialogue. This is particularly useful for multilingual creative work and for production teams reviewing spoken copy.

Use subtitles and on-screen text as an explicit layer as well. The User Guide notes that Seedance 2.5 improves control over unwanted subtitles and background music, but teams should still define which text belongs in the generated scene and which text will be composed in editing.

A reusable Seedance 2.5 prompt template

[References]
@Image 1 defines <subject or product>.
@Image 2 defines <scene and lighting>. Do not use <unwanted element>.
@Video 1 defines <motion, camera movement, or pacing>. Do not use <unwanted attributes>.
@Audio 1 defines <voice, dialogue style, ambience, or music>.

[Generation Goal]
Generate a <duration and video type>. The central subject is <subject>. The primary event is <story summary>.

[Stage 1]
Initial state: <visible opening state>.
Primary event: <one action>.
End state: <visible result>.

[Stage 2]
Continue from the previous stage: <facts that must remain unchanged>.
Primary event: <one action>.
End state: <visible result>.

[Maintain Consistency]
Keep <identity, object geometry, wardrobe, positions, direction, and audio relationships> consistent.

[Audio]
(<music>)
<sound effects>
Dialogue language: <language and delivery>. <speaker> says: {<line>}

What public demos are showing, and what they do not prove

The release materials establish what ByteDance says Seedance 2.5 can do. Public creator posts are useful for a different reason: they show which parts of the workflow people are actually trying first. They are demonstrations, not controlled benchmarks. A successful clip does not show retry count, cost per approved result, or whether the same result can be reproduced on another prompt.

English-language demo posts to inspect

The posts below are selected for a combination of engagement and a clearly stated test setup or prompt. Engagement figures change, so treat these as discovery links rather than performance claims.

Rewind and timed-camera prompt. TechHalla shared a 1950s diner sequence with a long-form prompt that specifies visual treatment, camera behavior, and an effect sequence. It is a useful example of why camera and temporal instructions should be written separately from the subject description.

Reference-led travel-vlog test. BubbleBrain published a realistic vlog case with a prompt that explicitly asks to retain the reference subject's facial identity, hairstyle, and features. It is a practical demonstration to inspect when evaluating identity instructions across a moving, handheld sequence.

Minimal-input product test. Ori Silver shared an unboxing test described as using an image of a box and an image of shoes. This is a useful counterpoint to a 50-reference workflow: start with only the recurring elements that really need to persist, then add references only when the brief requires them.

Japanese-language demo posts to inspect

Japanese creator posts expose a different set of practical questions: Japanese dialogue, the cost of a 30-second render, multi-shot continuity, and what still fails in real use.

Japanese speech test. チビクロ🧩AI錬金術士 tested a food-report style video and specifically discussed Japanese speech. This is a good clip to review alongside the guide's instruction to declare a dialogue language and delivery style before the spoken line.

A 30-second one-take with prompt and cost context. タナベ | AI動画 × マーケティング shared a 30-second action sequence, including the prompt and an estimated generation cost. It is useful for understanding that prompt length, iteration budget, and quality review need to be part of production planning.

Audio-reference music-video experiment. ヤノ shared a longer music-video experiment that used audio as a reference and also documented continuity and motion issues. This is the kind of honest field note teams should collect internally rather than judging a model solely from its best clip.

A practical demo-review checklist

When reviewing any Seedance 2.5 demo, ask the same questions before treating it as a production signal:

Review questionWhat to inspect
Reference controlDoes the creator state which material governs identity, product, motion, environment, or audio?
ContinuityAre identity, product geometry, screen direction, props, and audio relationships stable from start to finish?
Temporal controlCan you see a planned handoff, entrance, exit, or cut land near its intended beat?
AudioIs the dialogue language correct? Do music, effects, and voice fit the action?
EvidenceDoes the post reveal the prompt, assets, settings, retries, duration, and any editing afterward?
EconomicsIs there enough information to estimate cost per approved output rather than cost per generation?

The official guides are strongest as a grammar for production briefs. Creator demos are strongest as hypotheses to test against your own assets and review criteria.

A detailed workflow from brief to approved cut

A repeatable workflow has six parts.

1. Build a reference pack before writing the final prompt

Prepare a separate asset for every recurring element that must remain recognizable: one product, one lead character, a location, a key prop, a motion example, and a voice or music reference when applicable. The official Prompt Guide permits up to 30 images, 10 videos, and 10 audio clips, but its recommended ranges are lower. More material is not automatically better. Every asset should have a stated role and a stated exclusion.

For a physical product, use separate front, side, rear, and material references when those views matter. State that they all represent one product. For an actor or presenter, separate identity, wardrobe, and voice references rather than asking one ambiguous source image to carry every requirement.

2. Write a scene ledger

Before generating, list the starting state, one primary event, and a visible end state for each scene. If the video has three beats, write three entries. This makes prompt changes reviewable: the team can revise Stage 2 without weakening the established start of Stage 1.

3. Decide which requirements belong in generation and which belong in finishing

Use generation for movement, setting, mood, natural sound, and story progression. Use approved source assets and finishing tools for exact brand wordmarks, interface text, legal copy, product labels, and calls to action. This boundary keeps the generated portion expressive while keeping brand-critical elements governed by approved materials.

4. Generate a draft and review it against the ledger

Do not review the clip with a general question such as “does it look good?” Review it against the staged brief: correct product? correct person? correct directional continuity? correct audio? correct closing frame? Record which reference was used and which instruction changed in each iteration.

5. Use targeted edits when the rest of the take is usable

ByteDance's announcement describes timestamp-level editing, green-screen editing, camera perspective editing, and reference-based editing. The production decision is not whether the feature exists, but whether a local revision preserves everything that is already approved. Test edits on a representative clip before relying on them for a time-sensitive campaign.

6. Maintain a reusable prompt and evidence library

Save the prompt, the mapped source assets, the generated version, review notes, final edit, and usage rights in one place. This lets the next campaign begin with an approved reference pack instead of rediscovering the same constraints through ad hoc prompting.

Where MiniMax H3 fits into the same trend

MiniMax released H3 on the same day as Seedance 2.5. MiniMax describes H3 as a general-purpose multimodal generation model that understands text, image, video, and audio together and produces up to 15 seconds of 2K video with native stereo sound. It also says it plans to release the model weights under its community license, subject to applicable requirements.

The two launches are not interchangeable. Seedance 2.5's public documentation emphasizes a prompt workflow for multi-reference storytelling, staged sequences, extension, and editing. H3 emphasizes a unified multimodal model and open-weight direction. For a team selecting a workflow, the useful question is not which demo looks best. It is whether the system can translate a creative brief into repeatable production instructions.

Read MiniMax's H3 announcement.

What teams should do next

  1. Create a one-page prompt brief template with sections for references, exclusions, stages, continuity, camera, and audio.
  2. Store approved product imagery, voice direction, and copy rules in a shared workspace so every new video begins with the same source material.
  3. Review outputs stage by stage. Log what changed and which reference or instruction caused the change.
  4. Keep generated footage and final composited text separate when brand accuracy matters.

Sources, evidence boundaries, and further reading

Primary sources

  1. ByteDance Seed Team, “One-take Creation, Flexible Referencing: Introducing Seedance 2.5”, July 31, 2026. Release scope, rollout, 30-second generation, multi-round extension, references, and editing claims.
  2. ByteDance / Dreamina, Seedance 2.5 User Guide, updated July 31, 2026. Parameters, extensions, language support, and editing guidance.
  3. ByteDance / Dreamina, Seedance 2.5 Prompt Guide, updated July 31, 2026. Reference-role mappings, stage-and-end-state structure, continuity guidance, timestamps, and audio and dialogue syntax.
  4. MiniMax, “MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities”, July 31, 2026. H3's stated multimodal inputs, 2K 15-second output, native stereo sound, and open-weight direction.

Independent context

  • VidMuse’s Seedance 2.5 workflow guide is useful for a production-planning view, but platform availability and pricing should always be checked against the first-party product surface.
  • Seedance.tv’s evidence review usefully distinguishes official claims, curated official demos, and independently reproduced results. Its release-access snapshot predates the July 31 launch, so it should not be used as a current availability source.

How this article evaluates evidence

We separate official claims, creator demonstrations, and repeatable production evidence. Official materials establish what the provider says it supports. Creator posts reveal real prompt approaches and failure modes, but they do not establish average reliability. For campaign decisions, run matched tests with your own assets, record retries and approval rates, and evaluate cost per approved output.

Bottom line

Seedance 2.5 is worth tracking not only because of its 30-second generation window. Its official guides describe a practical creative language: assign roles to references, express a story as stages with end states, name continuity constraints, and keep music, sound effects, dialogue, and text distinct.

That structure turns a video prompt from an improvised request into an asset a team can review, reuse, and improve.

Hire the AI agent team that runs your entire business.