Taxis, buses, street lamps, crosswalk stripes and people on the sidewalk act as measuring sticks. Viewers know roughly how tall a person or a car is, so a figure that dwarfs them reads as enormous within a second. That is why the scene keeps the traffic small and busy around the giant's feet instead of leaving the street empty.
IAllegro
Make a City Giant Video From One Photo
City Giant is a photo-to-video template: upload one picture of a person or a character, and the video model rebuilds them as a colossal figure in a downtown street, taller than the buildings around them. Cars and pedestrians crawl past their feet while the camera slowly pulls back and tilts up, until the whole giant stands against the skyline. The scene is ready-made; optional details let you adjust it.
Current modelMiniMax H3
Andante
The template at a glance
The scene and the camera move are set by the template. Everything else in this list can be changed in the generator before you submit.
- Input
- One photo of a person or a character. A single, clearly visible subject works better than a group.
- Scene
- Built in: the subject stands as a colossal giant in a downtown street, with small cars and pedestrians moving at street level.
- Camera
- A slow pull-back combined with an upward tilt that ends on the full figure against the skyline.
- Optional details
- Clothing, weather, time of day and what the giant is looking at.
- Aspect ratio
- Vertical 9:16 by default; switch to another ratio, such as 16:9, for widescreen.
- Duration
- 6 seconds by default, adjustable in the generator.
- Models
- MiniMax H3 (default), Wan 3.0 or Seedance 2.5.
- Sound
- City ambience generated together with the clip on models that support audio.
- Cost
- Paid in credits after sign-in. The exact cost is shown before you generate, and failed generations are refunded automatically.
Adagio
What makes a City Giant shot feel enormous
Nobody judges size in a vacuum; the eye compares. The template builds in the main comparisons, and knowing how they work helps you pick a photo and details that support the illusion instead of fighting it.
Scherzo
How to make the clip
Five short steps. The only one that needs real thought is choosing the photo.
Sign in
Sign in with Google or email. The template is paid with Sonata AI credits, so check your balance first; the Generate button shows the cost of your current settings before anything is charged.
Upload one photo
Pick a single picture in which one person or character is clearly visible, face and outfit included. A sharp, well-lit image without heavy filters gives the model the most to work with; the photo tips further down explain why.
Add details if you want
Leave the details field empty for the standard scene, or add a short note about clothing, weather, time of day or where the giant looks. One or two phrases are enough, because the scene prompt itself is already written.
Choose model, duration and ratio
MiniMax H3 is preselected with a 6-second vertical 9:16 clip. Change the model, the length or the aspect ratio if your post needs it; the credit cost updates as you change model, duration and resolution.
Generate and review
Submit and wait while the video renders. The finished clip appears in the preview and is stored in your History, so you can return to it later. If a generation fails, the credits go back to your balance automatically.
Adagio
Where a giant clip fits
The default vertical, six-second version is made for phone feeds, and a couple of setting changes adapt it to other places.
Adagio
Choosing the photo and the details
Two things shape the result more than any setting: the picture you upload and the few words you add. The scene prompt is already written, so neither needs to be elaborate.
Coda
City Giant: questions and answers
Photos, likeness, models, sound, credits and responsible use.
01Will the giant look exactly like the person in my photo?
Not exactly. The model interprets your photo and rebuilds the person at a very different scale in a new scene, so face, outfit and proportions come out as a likeness rather than a copy. Results vary from one generation to the next; a clear, well-lit photo helps, and a second attempt may land closer.
02Do I need to write a prompt?
No. The scene, the street traffic and the camera move are already described by the template. The details field is optional and meant for small changes: clothing, weather, time of day, or what the giant looks at. Leave it empty to get the standard version.
03Can I use a group photo or a drawing?
The template is built for one subject, either a person or a character. In a group photo the model may enlarge the wrong person or blend several people into one, so crop to a single figure first. Character art can work as long as the figure is clear and not crowded by other elements.
04Which model should I pick?
Start with MiniMax H3, the default. Wan 3.0 and Seedance 2.5 run the same scene, but each draws the city, the light and the movement in its own way, and the credit cost differs between them. If you like the idea but not the look, keep the photo and details and change only the model.
05Does the video have sound?
On models that support audio, the clip comes with city ambience generated together with the picture, so the street sounds like a street. Whether audio is included depends on the model you choose in the generator.
06Why does my giant look small or out of place?
Usually the photo or the details are working against the scale. Tight headshots, several people in frame, heavy filters and dark images give the model less to work with, and details that add other actions or characters distract from the reveal. A very short duration may also leave less time for the pull-back.
07Can I change the length or make it horizontal?
Yes. The template starts at 6 seconds in vertical 9:16, and both settings can be changed in the generator. Choose a wide ratio such as 16:9 for YouTube or presentations; a longer clip gives the slow reveal more time, and the credit cost updates to match.
08How are credits charged?
You need to be signed in, and each video is paid with credits. The cost depends on the model, duration and resolution, and it updates on the Generate button as you change them, so you see it before you submit. If a generation fails, the credits are refunded automatically.
09Where do I find my finished videos?
Every finished video is saved in your History on Sonata AI. That makes it easy to go back to earlier takes, compare the same photo on different models and keep the version you like best.
10Can I make a giant video of someone else?
Only with their permission. Upload photos of people who agreed to it, and do not use the result to mislead anyone, for example by passing it off as real footage, or to mock or harass the person shown. Clips of strangers made without consent are not what this template is for.
11Is Sonata AI connected to MiniMax, Alibaba or ByteDance?
No. Sonata AI is an independent platform and is not affiliated with MiniMax, Alibaba or ByteDance. MiniMax H3, Wan 3.0 and Seedance 2.5 are trademarks of their respective owners; the names appear here only to tell you which model renders your clip.