The closest thing to a useful winner is a conditional one. Start with Seedance 2.5 when your expensive failures come from juggling many references, preserving a longer story, or regenerating an entire clip for one local change. Start with Wan 3.0 when native 1080p is a delivery requirement, the source is a document or web page, or finance needs a per-second ceiling before a job runs.
That is the decision as of August 26, 2026. Both models advertise native audio-video creation, multimodal reference, editing, extension, and up to 30 seconds in one generation. A montage of selected outputs cannot tell you which one will preserve your particular face, product, action, language, or brand system. The official API contracts can tell you which tests are worth paying for.
Four contract differences change the shortlist
| Production question | Seedance 2.5 on BytePlus | Wan 3.0 on Alibaba Cloud Model Studio |
|---|---|---|
| What route is documented? | dreamina-seedance-2-5-260628 | wan3.0-video |
| What can one request reference? | Up to 30 images, 10 videos, and 10 audio clips | Up to 10 images, 5 videos, and 5 audio clips |
| What else can it ingest? | Text, images, video, and audio | Text, images, video, audio, documents, and web pages |
| What is the native API output ceiling? | 480p or 720p; 4–30 seconds | 480p, 720p, or 1080p; up to 30 seconds |
| How is the route billed? | Returned video tokens, with different rates when input includes video | Generated seconds, priced by output resolution |

ByteDance's Seedance 2.5 launch note, the current BytePlus task API, and Alibaba Cloud's Wan 3.0 release support those facts. They are provider specifications and demonstrations, not a neutral output benchmark.
The resolution row is especially easy to blur. BytePlus currently documents Seedance 2.5 at 480p and 720p and explicitly excludes it from 1080p support in this route. Wan 3.0 lists 1080p. A creator app may upscale or export a larger file, but a 4K container is not proof of native 4K generation. If your downstream delivery contract says “native 1080p,” this distinction can end the first round before a subjective review begins.
Seedance 2.5 is the reference-heavy route
Thirty images, ten video clips, and ten audio clips create a much larger control envelope than most briefs need. The advantage appears when the dependencies are real: a performer and wardrobe, several products, a location, a blocking reference, camera language, voice, music, and an existing edit rhythm may all need to survive the same 20- or 30-second sequence.
Seedance 2.5 also supports timestamp-directed edits. A team can target a section of a clip instead of throwing away the whole generation because one beat, line, or movement failed. ByteDance also names reference editing, green-screen work, perspective changes, and clay-render control. For a production team, the hypothesis is not merely “the raw clip will look better.” It is “the model may reach an approved revision with fewer full reruns.”
That remains a hypothesis until tested. ByteDance's own release acknowledges room to improve the physical plausibility of complex movement and the stability of multi-subject interactions. Hands, fights, occlusion, prop exchange, and reappearing products still deserve explicit acceptance checks.
If the decision is actually between Seedance tiers rather than vendors, use the 2026 Seedance model guide to separate 2.5, 2.0, Fast, and Mini before paying for a cross-vendor test.
Wan 3.0 turns source material into a delivery route
Wan 3.0 takes up to ten images, five video clips, and five audio clips. It adds a different kind of leverage: a document or web page can be part of the request. Alibaba Cloud documents one file or link, up to 100 MB and 50 pages, with common office and text formats supported.
That matters for a product launch deck, training manual, destination page, or report that would otherwise require manual extraction before storyboarding. Wan 3.0 also carries smart duration, video extension, native audio-video, and editing of visuals, plot, and dialogue. Its 1080p tier makes it the cleaner first integration when the final pipeline cannot accept a 720p native master.
Do not convert the history of the Wan brand into an unstated deployment promise. Older Wan releases have published weights, but the current first-party evidence here establishes the Model Studio cloud route. A local Wan 3.0 deployment, ComfyUI workflow, or weight license needs its own current official artifact before it belongs in an architecture plan.
Wan has a clip price; Seedance has a usage formula
Wan 3.0 list pricing is $0.05 per second at 480p, $0.10 at 720p, and $0.20 at 1080p. A 30-second generation therefore lists at $1.50, $3.00, or $6.00. The release page shows a 30% Standard promotion through September 24, 2026 at 00:00 UTC+8, reducing those 30-second examples to $1.05, $2.10, and $4.20. The console price controls, and Prime, tax, currency conversion, and retries sit outside those examples.
BytePlus publishes Seedance 2.5 at $6.40 per million tokens when the input includes video and $10.70 per million tokens without video input. The task response provides usage.total_tokens, so the cost calculation is:
task cost = usage.total_tokens / 1,000,000 × applicable token rate
The two units are not interchangeable. Comparing Wan's dollars per second with Seedance's dollars per million tokens is like comparing a hotel room with a metered taxi before knowing the trip. Run the task, capture usage and all retries, then normalize by accepted output.
A fair first round uses the common denominator
Use two briefs that represent costly failures, not easy beauty shots. One could be a two-person dialogue with a prop exchange; the other a product demonstration with hand interaction, a camera move, and an exact closing pack shot.

For the matched round, stay inside both contracts: no more than eight images, two reference videos, and one audio clip; 10 seconds; 720p; audio enabled. Give each model the same creative goal, assets, timing requirements, and acceptance criteria. The prompt wording does not need to be identical because the APIs have different control languages. Run three attempts per brief and hide the model name during review where practical.
Record evidence that can change a production decision:
- Did identity, product geometry, and wardrobe survive occlusion and cuts?
- Did required actions occur in order without intersections, drift, or missing beats?
- Were speech, mouth movement, ambience, and timing usable without replacement?
- How many full reruns and local revisions were needed for one accepted result?
- What were queue time, generation time, returned tokens or seconds, and total bill?
Then run separate advantage tests. Give Seedance the larger reference pack and a timestamp revision. Give Wan a source document and a 1080p delivery. Do not add those results to the matched score as though the inputs were equivalent; they answer whether each model's unique capability removes work from your pipeline.
The durable cost metric is:
effective cost per accepted clip = all generation and repair cost / accepted clips
Add human repair minutes beside it. A higher-priced request can be cheaper to deliver if it passes in one attempt. A lower-priced route can still win for tolerant background shots that are safe to regenerate automatically.
Make the default reversible
Choose Seedance 2.5 as the default when reference capacity, long-form continuity, and targeted revision reduce the most expensive part of your process. Choose Wan 3.0 when 1080p, document/web ingestion, or per-second budget predictability is the hard gate. Keep the other model behind a written trigger instead of deleting it from the stack.
If neither set of conditions dominates, the matched round is the answer: your faces, products, actions, language, and sound decide the primary route. For a wider shortlist after this binary decision, compare the same briefs against the current eight-model video guide, not an overall leaderboard assembled from unrelated clips.



