IAllegro

Hotel Lobby AI: Put Two People on Stage for a Rap Duet

Hotel Lobby is a photo-to-video template for two people. Upload a picture of a duo, or one photo of each person, pick a stage such as the marble lobby under its chandelier, and the video model creates a short vertical clip of the pair rapping together. The vocals, the beat and the lyrics are generated along with the picture, so there is no track to bring.

Current modelMiniMax H3
Photos0/2

One photo of both of you, or two photos with one person each. Faces clearly visible, no sunglasses or masks. Only use photos of people who have agreed to it.JPG, PNG or WebP, up to 20 MB.

0 / 120
Advanced settingsMiniMax H3 Ā· 768P Ā· 9:16 Ā· 10 s
Model
10 s
Resolution
Aspect ratio
Sound

MiniMax H3 always generates the video with sound.

The price updates as you change the settings. If a generation fails, the credits are refunded automatically.

Preview
No video available. Please generate a video first!
Andante

Template settings at a glance

What you upload, what you get back and what you can change before generating.

Photos
One photo with both people in it, or two photos with one person each
Stages
Six presets: orange-walled recording booth, marble hotel lobby with a chandelier, neon-lit music studio, rooftop at night, graffiti alley, house party with confetti
Topic
Optional short line that tells the rap what it is about
Aspect ratio
Vertical 9:16 by default
Length
About 10 seconds by default; the duration can be changed before generating
Sound
Vocals, beat and lyrics generated by the video model together with the picture
Models
MiniMax H3 (default, always with sound), Seedance 2.5, Wan 3.0 (audio on)
Access
Sign-in and credits required; failed generations are refunded automatically; clips are saved in History
Adagio

What the Hotel Lobby template does

Two people, one stage and a soundtrack written on the spot. Here is what the template handles and what stays in your hands.

Scherzo

Making a rap duet clip, step by step

Five steps from photos to a finished video. Most of the effort goes into picking good photos.

  1. Add your photos

    Upload one photo that shows both people, or two photos with one person in each. Zoom in on the faces first: if either is blurry, half hidden or tiny, swap in a better picture now rather than after spending credits.

  2. Choose a stage

    Pick one of the six presets: booth, hotel lobby, neon studio, night rooftop, graffiti alley or house party. If you cannot decide, go with the mood that fits the person you plan to send the clip to.

  3. Write a topic, or skip it

    Type a short line about what the rap should cover, for example "our first apartment" or "the team that never wins". An empty field is fine too; the model will choose a subject on its own.

  4. Check model, duration and cost

    MiniMax H3 is selected by default. Change the model, duration or resolution if you like and watch the credit cost on the button update. Sign in when asked, since each generation uses credits from your account.

  5. Generate, listen, keep the best take

    Start the generation and wait for the clip. Play it with the sound on, because the vocals are half of the result. Download the take you like; every finished clip also stays in History.

Adagio

From finished clip to finished post

Where a short rap duet fits, and the small edits that make it ready to share.

Adagio

Photo and topic tips for a better duet

Recognisable faces and focused lyrics both start before you click generate.

Coda

Hotel Lobby AI: common questions

Practical answers on photos, lyrics, sound quality, credits and fair use.

01

Is the hotel lobby the only setting?

No. The template is named after one of its stages, but there are six: a small recording booth with warm orange walls, a marble hotel lobby with a chandelier, a neon-lit music studio, a rooftop at night, a graffiti alley and a house party with confetti.

02

Do both people have to be in the same photo?

No. One photo showing both people works, and so do two separate photos, one per person. With separate photos, similar framing and lighting help the result look as if the two of you really stood on the same stage.

03

Can I decide the exact lyrics?

Not word for word. Your topic guides what the rap is about, but the model writes and performs the lyrics itself. The words may differ from what you had in mind, and two runs with the same topic can sound quite different.

04

How good is the audio?

It is made for fun, shareable clips, not as a studio recording. Voices, timing and pronunciation vary from run to run, and some lines may be hard to make out. If you need a specific track, replace the sound afterwards with the add-audio tool.

05

Why does one of us not look like ourselves?

Usually a face is partly hidden, small in the frame, blurry or turned away, or the two photos are framed very differently. Sunglasses and masks are common culprits. Try clearer, front-facing photos, or run the same photos on another model.

06

What does a clip cost?

You need to sign in, and each generation uses credits. The exact cost appears on the generate button and changes with the model, duration and resolution you choose. If a generation fails, the credits are refunded automatically.

07

Can I use photos of other people, or imitate a famous artist?

Only upload photos of people who have agreed to it. Do not use the template to imitate real artists, to embarrass someone by putting words in their mouth, or to mock or harass anyone. Make something both people in the clip would be happy to watch.

08

Which model should I pick?

Start with MiniMax H3, the default, which always generates sound. Wan 3.0 runs with audio on here, and Seedance 2.5 is the third option. If one model gives a stiff performance or a weak likeness, the same photos on another model may turn out differently.

09

Is Sonata AI affiliated with MiniMax, Alibaba or ByteDance?

No. Sonata AI is an independent platform and is not affiliated with MiniMax, Alibaba or ByteDance. MiniMax H3, Wan 3.0 and Seedance 2.5 are trademarks of their respective owners; Sonata AI offers these models inside its own templates.