How to make the Hotel Lobby AI video

There are three ways to make it. They differ in cost and effort, and above all in how closely the result follows the original moves. Start with the photos, then pick a method.

Updated September 25, 2026

Before you start: the two photos that work

Every method starts from photos, and the photos decide most of the result. Get these right once and you can reuse them for all three methods.

  • One person per photo. Crop group shots so only the person you want is left in the frame.
  • Facing the camera. A straight-on face with both eyes visible swaps far better than a profile or a tilted head.
  • Waist-up or full body. The outfit carries over into the video, so include it.
  • Good, even light. Skip heavy filters, sunglasses and hands over the face.
  • A clear original. Use the photo itself, not a blurry screenshot of it.

Decide now who stands on the left and who stands on the right. In the original performance the two sides have different moves, so the choice changes who does what.

Method 1 – One-click generator (same moves, original audio)

This is the only method that keeps the original choreography. The generator uses the real performance as its base and replaces the two performers with the people in your photos, so every point, turn and beat lands where it should.

  1. 1Upload two photos

    Open the Hotel Lobby AI generator and add one photo to the left slot and one to the right slot. One person per photo.

    Generator screen with two empty photo slots, left and right
  2. 2Pick who stands left

    The left photo replaces the performer on the left. Use the swap button if you want them the other way round, then choose 720p or 1080p. The price is shown before you generate.

    Both photo slots filled, with the swap button and resolution choice
  3. 3Generate and download

    Press Generate. The clip is usually ready in 3 to 6 minutes and is saved to your creations, so you can leave the page. Download the vertical MP4 and post it.

    Finished vertical video with a download button
Time
Usually 3-6 minutes
Output
Vertical 9:16 MP4, 720p or 1080p
Cost
Paid per video, price shown up front
Watermark
None

Method 2 – CapCut template (free, but it's a face swap)

CapCut has user-made Hotel Lobby templates. They are free and quick, but a template is a fixed video with your photo pasted over the faces, so the bodies, the moves and the timing belong to whoever made the template.

  1. Open CapCut (phone app or desktop) and go to Templates.
  2. Search for hotel lobby and open a template whose preview shows the orange booth.
  3. Tap Use template and add your photos in the order the template asks for.
  4. Preview the result, then export and share.

The catch

  • The moves don't match the original. You get the template creator's clip with new faces on top.
  • Free exports can carry a CapCut watermark or end screen.
  • Length, cuts and framing are fixed by the template, so you can't extend or re-time it.
  • Templates come and go. If one disappears, search again.

Method 3 – Write your own prompt (Kling, Veo, Sora)

If you would rather build it yourself, any current video model can get close to the look. Kling, Veo and Sora all take a text prompt plus an optional start image. Starting from a photo of your two people (image-to-video) keeps their faces far closer than text alone.

Start with this prompt and adjust the people and outfits:

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. Two people stand side by side, facing the camera. They take turns performing a rap verse straight to the lens: pointing at the camera, gesturing with their hands, stepping and turning on the beat, while the other one nods along. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

Variants for couples, pets, kids and solo videos, plus a negative prompt and settings, are on the Hotel Lobby AI prompts page.

Expect randomness. The same prompt gives different moves every time, and no model reproduces the original choreography from text alone. Budget several attempts and keep the best one.

Common problems

The face doesn't look like the person
Use a sharper, front-facing photo with even light, and crop so the face is larger in the frame. Remove sunglasses and heavy filters.
The wrong person is on the wrong side
In the generator, swap the left and right photos. In a prompt, describe each person by position, for example "on the left, a woman in a red jacket".
The background isn't orange
This happens with prompts and some templates. Put "seamless matte-orange studio booth" early in the prompt and add "background color other than orange" to the negative prompt.
The lip-sync or audio is out of sync
Templates and prompt videos are not timed to the song. Line up the start in your editor, or add the sound in TikTok and nudge it. The generator keeps the original timing.

Which method should you pick?

CompareOne-click generatorCapCut templateYour own prompt
CostPaid per videoFreeFree tiers, then paid credits
Time3-6 minutesAbout 10 minutesVaries, several attempts
MovesSame as the originalThe template creator'sRandom, rarely close
WatermarkNonePossible on free exportsDepends on the tool
AudioOriginal trackThe template's soundAdd it yourself

Short version: if you want it to look like the real thing, use the generator. If you want free and fast and don't mind different moves, use a CapCut template. If you enjoy tinkering, write a prompt.

How-to FAQ

How long does it take to make a Hotel Lobby AI video?

With the one-click generator, usually 3 to 6 minutes. A CapCut template takes about ten minutes including export. Writing your own prompt takes longer, because you will usually need several attempts.

Can I make the Hotel Lobby AI video for free?

Yes, with a CapCut template or the free tier of a video model, but neither copies the original moves. The one-click generator is paid per video.

Why doesn't my video match the original moves?

CapCut templates and text prompts don't use the original performance, so the moves belong to the template creator or are invented by the model. Only a tool that uses the original performance as its base keeps the choreography.

What photos should I use?

One person per photo, facing the camera, in good light, waist-up or full body so the outfit carries over. Crop group photos first.

Ready to put your people in the booth?

Upload two photos and get a 1080p vertical video in a few minutes.

Make yours