The most important AI video fact in August 2026 is not a quality score. It is a date: OpenAI’s Sora 2 video API is scheduled to shut down permanently on September 24, 2026. A famous model can still be the wrong foundation for a new product.
At the same time, the current field has expanded in directions that a single “best generator” label cannot capture. Seedance 2.5 can make longer audio-video clips from mixed references. MiniMax H3 combines open weights with an official 2K workflow. Wan 2.7 offers region-scoped cloud routes. Kling 3.0 brings native multilingual audio into an all-in-one system. Veo 3.1 has quality, fast, and lower-cost tiers. Runway and Luma serve very different production workflows.
The useful question is therefore: which model removes the hardest constraint in your shot, without creating a new access, cost, or lifecycle problem?
The short list, with the boundary that matters
| Route | What makes it a current candidate | Boundary to verify before committing |
|---|---|---|
| Seedance 2.5 | Up to 30-second audio-video generation, multi-round extension, image/video/audio references, editing | Exact API surface, region, quota, and price |
| MiniMax H3 | 4–15 seconds, stereo audio, rich multimodal references, open H3-Base weights | Open weights produce 768p-short-side base output; official regeneration provides 2K |
| Wan 2.7 | 2–15 seconds, 720p/1080p, 30fps, audio, explicit international deployment routes | Global, International, US, and mainland China scopes are not interchangeable |
| Kling 3.0 / Omni | Up to 15 seconds, native multilingual audio, unified generation/reference/editing | Product and API details vary by surface |
| Veo 3.1 / Fast / Lite | 4/6/8-second audio-video, landscape/portrait, tiered quality/latency/cost | Gemini, Flow, and Vertex AI do not expose identical controls |
| Runway Gen-4.5 | Text- or image-to-video inside a mature creator workflow | 2–10 seconds, 720p, 12 credits per second on the documented web route |
| Luma Ray3.2 | Up to 16 keyframes, video transformation, 20-second 1080p workflow, HDR/EXR | It is a source-video/keyframe control tool, not a general text-to-video replacement |
| Sora 2 | Synced audio, 4/8/12 seconds, familiar OpenAI API | Deprecated; official shutdown scheduled for September 24, 2026 |
These are official specifications and lifecycle facts, not a universal quality benchmark. Vendor leaderboards and “best-in-class” labels can be useful release evidence, but they do not tell you which model will preserve your product, character, typography, or edit requirements.

Three changes make the old rankings obsolete
A shot can now start with more than a prompt
Text-to-video is only one input mode. A production shot may depend on a first frame, a final frame, several character and product images, a motion reference, an existing clip, or a voice and soundtrack. The model that looks strongest from text alone may be the weaker option once your real reference package is included.
ByteDance’s Seedance 2.5 release emphasizes mixed references, editing, extensions, and clips up to 30 seconds. MiniMax’s H3 documentation allows as many as nine images, three video clips, and three audio clips in its reference mode, with no more than 12 files total. Kling’s 3.0 launch describes text, image, audio, and video within one multimodal system.
Those capabilities change the test. Do not flatten every model to the same text prompt. Preserve the same shot goal and acceptance criteria, but let each candidate use the controls it was built around.
Audio is part of the generation decision
Native audio is no longer a rare extra. H3 generates stereo audio and lists stable dialogue support for 11 languages, including English, Spanish, Russian, Japanese, Korean, and Chinese. Kling 3.0 advertises multilingual dialogue across languages, dialects, and accents. Wan 2.7 and Veo 3.1 also offer audio-video generation on documented routes.
“Supports audio” still needs unpacking. It may mean generated dialogue, ambience and effects; audio used as a reference; or a separate application step. Test lip sync, voice continuity, intelligibility, environmental timing, and whether the audio survives your editing workflow. A convincing silent demo says nothing about those requirements.
Model access is now a production feature
A model name does not define one consistent product. The creator app may add editing, stock assets, export options, or safety controls. An official API may expose a narrower parameter set. An aggregator may use different model IDs, billing units, regions, and version timing. Open weights can have a different resolution or license boundary than the vendor’s hosted workflow.
H3 is the clearest example. The open H3-Base checkpoints generate at a 768-pixel short side. MiniMax’s hosted H3-Regenerate-2K step uses the base result plus the original context to generate a 2K result, but that regeneration module is not open source yet. “Open-source 2K H3” collapses two materially different surfaces.

Where each route earns a place in a pilot
Seedance 2.5: when continuity and mixed references dominate
Seedance 2.5 deserves an early pilot when the expensive problem is maintaining a long performance, complex camera choreography, or multiple reference assets. Its 30-second ceiling and multi-round extension can reduce the number of boundaries at which a character, object, or sound must be reconstructed.
Do not infer that every Seedance-branded endpoint exposes the same model. BytePlus documentation lists a specific dreamina-seedance-2-5-260628 route alongside 2.0, Fast, and Mini variants. Verify the route, input combination, region, and current price you will actually use. For a deeper model-family breakdown, see the 2026 Seedance model guide.
MiniMax H3: when openness, language, and references meet
H3 is unusual because it can enter both a hosted API shortlist and an open-deployment investigation. It generates 4–15 seconds at 24fps with 32kHz stereo audio. Its reference system is useful for character, motion, voice, and mixed-context work; its stable 11-language dialogue list makes it especially interesting for localized creative.
The open checkpoints and hosted 2K pipeline should be evaluated separately. Record local GPU requirements, base-resolution quality, license obligations, hosted regeneration cost, and the operational work required to reproduce an acceptable take.
Wan 2.7: when deployment scope must be explicit
Alibaba Cloud’s current video documentation lists Wan 2.7 for International deployment with 2–15-second audio-video output at 720p or 1080p and 30fps. It also documents separate Global, International, US, and mainland China scopes. That makes Wan worth testing when data and inference location are part of the architecture rather than an afterthought.
For one provider-specific reference point, Alibaba currently lists International Wan 2.7 reference-to-video at $0.10 per second for 720p and $0.15 per second for 1080p. Those numbers do not apply to every Wan modality, region, or fast variant.
Kling 3.0: when performance and native dialogue matter
Kling 3.0 and Omni belong on a shortlist for character performance, multilingual dialogue, and stories assembled from mixed media. Kuaishou’s official release states a maximum duration of 15 seconds and brings generation, reference-to-video, and in-video editing into one framework.
The provider’s quality claims are not a substitute for your test footage. Give Kling a scene with interacting people, a line of dialogue, an occlusion, and a camera move. That reveals more than a beauty shot.
Veo 3.1: when Google’s ecosystem is already an advantage
Google positions Veo 3.1 as the quality tier, Fast as the faster production tier, and Lite as the volume-oriented tier. The current lineup supports native audio, landscape and portrait output, and 4-, 6-, or 8-second clips. Google also offers improved 1080p and 4K enhancement through its current product surfaces.
On Vertex AI, the documented 3.1 routes support 720p/1080p, 24fps, image-to-video, and first/last-frame generation, with English listed as the API prompt language. That is not a promise that every Gemini or Flow feature maps one-to-one onto Vertex.
Runway Gen-4.5: when the surrounding workflow has value
Gen-4.5 creates 2–10-second, 720p clips from text or an image at 24 or 25fps. Runway’s documented web cost is 12 credits per second. Its strongest reason to enter a pilot may be the surrounding Runway workflow and team familiarity, not a spec-sheet win.
Runway also publishes useful limitations: causal order can invert, objects can disappear after occlusion, and actions show a “success bias.” Include a failed action and an occluded branded object in your trial if those risks matter.
Ray3.2: when regeneration is the wrong operation
Ray3.2 starts from footage or keyframes and returns a controlled transformation. Luma describes up to 16 keyframes, clips up to 20 seconds at 1080p, and professional HDR/EXR workflows. If you already have a plate, previs, or approved story beats, transforming that asset may preserve intent better than generating from scratch.
If all you have is a text prompt, Ray3.2 is not the equivalent starting point. Compare it with a generation route only when both can satisfy the same final deliverable.
Sora 2: a migration case, not a default
OpenAI’s Sora 2 model page lists synced audio and a standard 720p price of $0.10 per second. The video API accepts 4-, 8-, and 12-second durations. But the create-video reference is deprecated and schedules permanent shutdown for September 24, 2026; the consumer Sora product ended on April 26.
Existing users should export representative prompts, reference images, output settings, accepted results, and failure examples now. A third-party route that continues to use “Sora” in its name does not extend OpenAI’s official lifecycle.
A two-pass test that produces a defensible choice
Create 8–12 shot tasks from your actual backlog. Include a static product, close human performance, multiple interacting subjects, large motion, camera motion, brand or text detail, native audio, and one action that is supposed to fail. Keep the deliverable and acceptance criteria fixed; use each model’s native reference and control features.
The first pass is a functional gate:
- Did the required subject, action, framing, and timing appear?
- Did identity, product shape, and text survive motion and occlusion?
- Is dialogue, lip sync, ambience, and sound timing usable?
- Can the output enter your edit, color, and compositing pipeline?
- Is failure obvious enough to automate or repair safely?
Only then compare economics. Track every generation, accepted take, enhancement or download fee, editor intervention, and retry. Use:
cost per usable take = all generation and post-processing cost for the task / accepted takes
A $0.05-per-second route that needs four rerolls can cost more than a $0.15 route that passes once. The opposite can be true for low-risk background shots with automated checks.
Make the choice maintainable
For each production route, record a default model and version, the condition that upgrades or downgrades it, and an exit condition. Re-run representative shots when a model version, price, region, policy, or lifecycle changes.
That policy survives the next release better than a permanent ranking. In August 2026, the practical default is not one model for everything: it is a small, evidence-backed routing rule that matches the shot, the input package, and the cost of failure.



