Hotel Lobby AI prompts

Copy-paste prompts for the Hotel Lobby look: a matte-orange booth, one hanging mic, two performers, vertical 9:16. They work in Kling, Veo, Sora and most other video models.

Updated September 25, 2026

The base prompt

This prompt describes the whole scene: the booth, the hanging mic, two performers and a fixed vertical frame. Paste it as is, or replace "Two people" with a short description of your own pair.

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. Two people stand side by side, facing the camera. They take turns performing a rap verse straight to the lens: pointing at the camera, gesturing with their hands, stepping and turning on the beat, while the other one nods along. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

For faces that look like real people, use image-to-video: give the model a photo of your two people standing side by side, and let the prompt handle the room and the movement.

Variants

Each variant is a complete prompt, ready to copy. They keep the same booth and camera and change who is performing and how.

Two friends

The classic version. Swap in your own outfits.

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. Two friends in casual streetwear, one in a hoodie and one in a denim jacket, stand side by side, facing the camera. They take turns performing a rap verse straight to the lens: pointing at the camera, gesturing with their hands, stepping and turning on the beat, while the other one nods along. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

Couple

Adds glances between the two, which reads well for partners.

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. A couple stands side by side, facing the camera, glancing at each other between verses. They take turns performing to the lens, pointing at the camera and gesturing on the beat, and finish leaning in toward the microphone together. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

Two pets

The version that started the trend. Works best with a photo of the animals as the start image.

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. Two cats sit upright on the orange floor side by side, facing the camera like performers. They bob their heads to the beat, take turns lifting a paw toward the lens and look up at the hanging microphone. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

Solo twice

One person in both spots. Describe the outfit so both copies match.

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. The same person appears twice, side by side, with an identical face, hairstyle and outfit, facing the camera. The two copies trade lines with each other: one points at the camera while the other reacts, then they switch, turning on the beat. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

Kids

Bigger, playful gestures suit younger performers.

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. Two kids in bright casual clothes stand side by side, facing the camera. They take turns performing to the lens with big, playful gestures, pointing at the camera, jumping on the beat and laughing between lines. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

Coworkers

Office clothes against the orange booth make the joke land.

Vertical 9:16 video. A seamless matte-orange studio booth: curved backdrop and floor in the same flat orange, soft even light, no visible edges or corners. A single black microphone hangs from the ceiling on a thin cable, centered between the two performers. Two coworkers in office clothes, one in a button-up shirt and one in a blazer, stand side by side, facing the camera. They take turns performing a rap verse straight to the lens: pointing at the camera, gesturing with their hands, stepping and turning on the beat, while the other one nods along. Locked-off camera at chest height, full bodies in frame, no zoom, no cuts. No captions, no subtitles, no text, no logos.

Negative prompt & settings

If your tool has a negative prompt field, paste this in to keep extra people, text and camera moves out of the shot.

extra people, crowd, audience, text, captions, subtitles, logo, watermark, camera shake, zoom, cuts, fast camera moves, blurry faces, distorted hands, extra fingers, background color other than orange, second microphone

Aspect ratio
9:16 (vertical)
Length
10-15 seconds
Quality
Standard (std) for drafts, Pro for the final render
Start image
A photo of your two people side by side, if the model supports image-to-video
Camera
Static or fixed, if the tool has camera control

Why prompts alone don't nail the moves

A prompt describes a scene; it can't hand over choreography. The model invents new movement on every run, so you will get two people in an orange booth, but never the exact points, turns and timing of the original performance.

Motion-control and video-editing models get around this by using the real performance as a reference and only replacing the people in it. That is what the one-click generator does: two photos in, original moves out. The how-to guide compares all three methods side by side.

Prompt FAQ

Which AI video model works best with these prompts?

Any current model that takes a text prompt will render the scene, including Kling, Veo and Sora. Starting from an image of your two people (image-to-video) keeps their faces far closer than text alone.

Can I use these prompts for free?

Yes. Copy them and adapt them however you like. Whether generating the video is free depends on the tool and plan you use.

Why does my result never match the original performance?

Text prompts don't carry choreography, so the model invents its own moves on every run. To keep the original moves you need a tool that uses the performance as a reference, like the one-click generator.

Ready to put your people in the booth?

Upload two photos and get a 1080p vertical video in a few minutes.

Make yours