Say who or what the shot is about, then add two or three visible details such as age, clothing, material or color. āA tired nurse in teal scrubsā gives the model far more to draw than āa personā. Keep one main subject per clip so the frame has a clear center of attention.
IAllegro
Text to Video AI Generator
Describe a shot in plain words and get a video clip back. This text to video generator lets you choose MiniMax H3, Wan 3.0 or Seedance 2.5, set the length, resolution and frame shape, and see the credit cost before you submit. Below the generator: how to write a prompt a video model can follow, how to match a model to your shot, and how to test ideas cheaply.
Current modelMiniMax H3
Andante
The three models at a glance
Every setting is made in the generator above. The ranges differ from model to model, so read them against what your shot needs. None of the three is the right pick for every clip.
- Input
- A written prompt of up to 5,000 characters. Text is the default; a first frame or reference images can be added when you would rather start from a picture.
- Models
- MiniMax H3, Wan 3.0 and Seedance 2.5, switched in the model picker on this page.
- Length
- MiniMax H3: 4 to 15 seconds. Wan 3.0: 2 to 30 seconds. Seedance 2.5: 4 to 15 seconds.
- Resolution
- MiniMax H3: 768P or 2K. Wan 3.0: 480P, 720P or 1080P. Seedance 2.5: 480p, 720p or 1080p.
- Aspect ratio
- 16:9, 9:16, 1:1, 4:3 and 3:4 on all three. 21:9 on MiniMax H3 and Seedance 2.5. An adaptive option on Wan 3.0 and Seedance 2.5.
- Sound
- MiniMax H3 always generates a stereo soundtrack together with the picture. On Wan 3.0 and Seedance 2.5 the audio can be switched off for a silent clip.
- Cost
- Paid in credits. The figure on the Generate button updates as you change model, duration and resolution, so you know the price before you submit.
- Failed jobs
- If a generation fails, its credits are refunded automatically.
- Your videos
- Saved in History under your account. You need to sign in to generate.
Adagio
Six things a text to video prompt should cover
A video model works only from what you write. Whatever you leave unsaid, it decides on its own, and that guess may not match the picture in your head. Covering these six parts, roughly in this order, gives a single shot everything it needs.
Scherzo
From first draft to final clip
Every generation costs credits, so catch problems on cheap drafts and pay for full length and resolution only once the prompt works.
Sign in and choose a model
After signing in, pick MiniMax H3, Wan 3.0 or Seedance 2.5 by what the shot needs: how long it runs, how sharp it must be, which frame shape it uses and whether you want generated sound. The ranges for each model are in the comparison table above.
Write one paragraph
Put the six parts into a few plain sentences, subject first. The field takes up to 5,000 characters, but a clip of a few seconds cannot show a page of detail, so spend your words on what the viewer will actually see and hear. Text is the default; images are optional.
Run a cheap draft
Choose the shortest duration and the lowest resolution the model allows, for example 480P on Wan 3.0 or Seedance 2.5, or 768P on MiniMax H3, and watch the figure on the Generate button drop. A draft is enough to judge composition, motion and mood.
Change one thing per retry
If the camera is wrong, rewrite only the camera sentence; if the light looks flat, change only the light. Editing everything at once hides which change made the difference. Every result stays in History, so you can go back and compare versions before deciding on the final wording.
Render the final version
When a draft looks right, raise the duration and resolution and generate again. The final is a new generation, not an enlarged copy of the draft, so expect small differences; a longer clip may also need a second beat in the action. If a job fails, the credits return automatically.
Adagio
Matching a model to the job
No model here is right for everything. Start from what the finished clip has to do, then check whose limits fit it.
Adagio
One prompt, taken apart
Below is an example written for this page, followed by the reason each piece is there. Borrow the structure rather than the words, and put your own subject in.
Coda
Text to video FAQ
Short answers about prompts, models, credits and results on this page.
01What does a text to video generator actually do?
It reads a written description and generates a short clip that tries to match it: the subject, the motion, the setting, the camera and, on models with audio, the sound. The clip is generated, not cut together from existing footage, which is why the wording of your prompt matters so much.
02Which of the three models should I use?
That depends on the clip, not on a ranking. Need more than 15 seconds? Wan 3.0 goes up to 30. Need 21:9? MiniMax H3 or Seedance 2.5. Want sound every time? MiniMax H3 always adds a stereo track. Want silence? Wan 3.0 or Seedance 2.5 with audio off. Need 2K? MiniMax H3.
03How long should my prompt be?
The field accepts up to 5,000 characters, but a single shot is usually well described in three to six sentences. Extra words do not add screen time; a clip of a few seconds can only show so much. If a prompt keeps growing, it often contains two scenes, which work better as two generations.
04How many credits does a video cost?
It depends on the model, the duration and the resolution you choose. The cost updates as you change those settings and is shown on the Generate button before you submit. Short, low-resolution drafts cost the least, which is what makes them useful as test runs.
05What happens if a generation fails?
The credits for that job are refunded automatically; there is nothing to claim. A request can fail when the model declines the content of the prompt or when a temporary error occurs on the model side. If the same prompt fails twice, rephrase the part most likely to be refused and try again.
06Why doesnāt my video look like what I wrote?
The usual causes are too many actions for the length of the clip, camera directions that contradict each other, vague words like ābeautifulā or āepicā with nothing concrete behind them, or two subjects fighting for attention. Cut back to one subject and one action, run a short draft, then add detail.
07Do the videos have sound?
With MiniMax H3, always: every clip comes with a generated stereo soundtrack. Wan 3.0 and Seedance 2.5 can generate audio as well, and both have a switch to turn it off when you want a silent clip for your own music or voice-over.
08Can I start from an image instead of text?
Yes. Text is the default here, but the same page also takes a first frame or reference images. If your project is mainly image-first, such as animating a product photo, the individual model pages linked above cover those workflows in more depth.
09Where are my videos saved?
Finished videos are saved in History in your account, so you can come back to them later and compare drafts with the final version. Generating requires signing in, because credits and History are both tied to your account.
10Is Sonata AI affiliated with MiniMax, Alibaba or ByteDance?
No. Sonata AI is an independent platform with no affiliation to MiniMax, Alibaba or ByteDance. MiniMax H3, Wan 3.0 and Seedance 2.5 are named only to identify the models you can use here; those names are trademarks of their respective owners.