Many photo animations only nudge a still picture into gentle motion. This template places both people on a stage and generates them performing a rap together, as a short scene with sound. Think of it as a tiny music video starring the two people in your photos.
IAllegro
Hotel Lobby AI: Put Two People on Stage for a Rap Duet
Hotel Lobby is a photo-to-video template for two people. Upload a picture of a duo, or one photo of each person, pick a stage such as the marble lobby under its chandelier, and the video model creates a short vertical clip of the pair rapping together. The vocals, the beat and the lyrics are generated along with the picture, so there is no track to bring.
Current modelMiniMax H3
Andante
Template settings at a glance
What you upload, what you get back and what you can change before generating.
- Photos
- One photo with both people in it, or two photos with one person each
- Stages
- Six presets: orange-walled recording booth, marble hotel lobby with a chandelier, neon-lit music studio, rooftop at night, graffiti alley, house party with confetti
- Topic
- Optional short line that tells the rap what it is about
- Aspect ratio
- Vertical 9:16 by default
- Length
- About 10 seconds by default; the duration can be changed before generating
- Sound
- Vocals, beat and lyrics generated by the video model together with the picture
- Models
- MiniMax H3 (default, always with sound), Seedance 2.5, Wan 3.0 (audio on)
- Access
- Sign-in and credits required; failed generations are refunded automatically; clips are saved in History
Adagio
What the Hotel Lobby template does
Two people, one stage and a soundtrack written on the spot. Here is what the template handles and what stays in your hands.
Scherzo
Making a rap duet clip, step by step
Five steps from photos to a finished video. Most of the effort goes into picking good photos.
Add your photos
Upload one photo that shows both people, or two photos with one person in each. Zoom in on the faces first: if either is blurry, half hidden or tiny, swap in a better picture now rather than after spending credits.
Choose a stage
Pick one of the six presets: booth, hotel lobby, neon studio, night rooftop, graffiti alley or house party. If you cannot decide, go with the mood that fits the person you plan to send the clip to.
Write a topic, or skip it
Type a short line about what the rap should cover, for example "our first apartment" or "the team that never wins". An empty field is fine too; the model will choose a subject on its own.
Check model, duration and cost
MiniMax H3 is selected by default. Change the model, duration or resolution if you like and watch the credit cost on the button update. Sign in when asked, since each generation uses credits from your account.
Generate, listen, keep the best take
Start the generation and wait for the clip. Play it with the sound on, because the vocals are half of the result. Download the take you like; every finished clip also stays in History.
Adagio
From finished clip to finished post
Where a short rap duet fits, and the small edits that make it ready to share.
Adagio
Photo and topic tips for a better duet
Recognisable faces and focused lyrics both start before you click generate.
Coda
Hotel Lobby AI: common questions
Practical answers on photos, lyrics, sound quality, credits and fair use.
01Is the hotel lobby the only setting?
No. The template is named after one of its stages, but there are six: a small recording booth with warm orange walls, a marble hotel lobby with a chandelier, a neon-lit music studio, a rooftop at night, a graffiti alley and a house party with confetti.
02Do both people have to be in the same photo?
No. One photo showing both people works, and so do two separate photos, one per person. With separate photos, similar framing and lighting help the result look as if the two of you really stood on the same stage.
03Can I decide the exact lyrics?
Not word for word. Your topic guides what the rap is about, but the model writes and performs the lyrics itself. The words may differ from what you had in mind, and two runs with the same topic can sound quite different.
04How good is the audio?
It is made for fun, shareable clips, not as a studio recording. Voices, timing and pronunciation vary from run to run, and some lines may be hard to make out. If you need a specific track, replace the sound afterwards with the add-audio tool.
05Why does one of us not look like ourselves?
Usually a face is partly hidden, small in the frame, blurry or turned away, or the two photos are framed very differently. Sunglasses and masks are common culprits. Try clearer, front-facing photos, or run the same photos on another model.
06What does a clip cost?
You need to sign in, and each generation uses credits. The exact cost appears on the generate button and changes with the model, duration and resolution you choose. If a generation fails, the credits are refunded automatically.
07Can I use photos of other people, or imitate a famous artist?
Only upload photos of people who have agreed to it. Do not use the template to imitate real artists, to embarrass someone by putting words in their mouth, or to mock or harass anyone. Make something both people in the clip would be happy to watch.
08Which model should I pick?
Start with MiniMax H3, the default, which always generates sound. Wan 3.0 runs with audio on here, and Seedance 2.5 is the third option. If one model gives a stiff performance or a weak likeness, the same photos on another model may turn out differently.
09Is Sonata AI affiliated with MiniMax, Alibaba or ByteDance?
No. Sonata AI is an independent platform and is not affiliated with MiniMax, Alibaba or ByteDance. MiniMax H3, Wan 3.0 and Seedance 2.5 are trademarks of their respective owners; Sonata AI offers these models inside its own templates.