← Blog

AI videoSeptember 28, 2026

Why Vaaya’s Hotel Lobby recipe is built for the Migos AI trend

See the supplied Vaaya Hotel Lobby video beside an H3 Max sample, compare photographic FLUX and Nano Banana Pro images, and learn how Vaaya builds prompts for each character.

The appeal of the Migos AI trend is immediate: put yourself and a friend into the Hotel Lobby performance, keep the energy, and make the clip unmistakably yours. The hard part is keeping it yours once the performers start moving.

A face has to survive a turn. Two identities have to stay separate. Your outfit should carry into the performance. The camera, gestures and soundtrack have to work together.

That is the standard behind Vaaya’s Colors / Hotel Lobby recipe. We build for a best-in-class Hotel Lobby experience by treating the whole performance as the product: a model selected for this edit, a prompt built around each character’s uploads, and a finished video with the original audio restored. You bring the people. The recipe handles how they fit into the performance.

The comparison: the same people, different creative treatments

The photos below provide the character references for our new image samples and H3 Max video. The leather-jacket performer is assigned to the viewer’s left; the blue-blazer performer is assigned to the right, matching the supplied Vaaya example.

Left performer reference Right performer reference
Supplied reference photo of the gray-haired adult man wearing glasses and a black leather jacket Supplied reference photo of the short-haired adult woman wearing glasses and a dark blazer

These are AI-generated portrayals, not recordings of the pictured people performing. The Vaaya video is an existing sample supplied for this article. We generated the H3 Max clip and the image variations separately.

Stage Comparison treatment Vaaya demonstration
Character images FLUX.2 [klein] 9B, prompted for a photographic portrait with light editorial polish Nano Banana Pro, prompted for natural lighting, skin texture and realistic clothing
Video New MiniMax H3 Max generation through fal, using the two uploaded photos directly Supplied Vaaya / ReAPI example, embedded without a new ReAPI generation
Performer placement Leather jacket on the left; blue blazer on the right The same left/right arrangement in the supplied video
Duration 7.29 seconds returned from a seven-second request 25.5 seconds in the supplied file

The image prompts use different photographic treatments. FLUX asks for light editorial polish; Nano Banana Pro emphasizes natural lighting and surface texture. This shows the effect of model choice plus creative direction. The samples are not attributed to competing platforms, and the differing video inputs and durations make this an illustrative comparison rather than a controlled benchmark.

Start with the look you want to keep

FLUX: photographic with a polished finish

The FLUX brief asks for believable proportions, soft studio lighting, natural hair and clothing detail, with a lightly retouched finish. The result should read as a photograph, with a subtle AI-polished look.

Nano Banana Pro: realism in the details

The Nano Banana brief asks for soft studio lighting, believable shadows, age-appropriate skin texture, individual hair strands and visible leather or fabric grain. For a performance that should feel photographic, those are the details we want to carry into motion.

FLUX — photographic, lightly polished Nano Banana Pro — realism-directed
FLUX photographic portrait of the left performer in a black leather jacket against an orange backdrop Nano Banana Pro realism-directed portrait of the left performer in a black leather jacket against an orange backdrop
FLUX photographic portrait of the right performer in a blue blazer against an orange backdrop Nano Banana Pro realism-directed portrait of the right performer in a blue blazer against an orange backdrop

These stills were made from the two supplied photos. They were not extracted from the supplied video, and neither set was used as input to the new H3 Max clip. H3 Max received the original photos directly.

Nano Banana Pro illustrates our preferred photographic treatment here; Seedance 2.5 is the performance editor used by the recipe. In the live upload workflow, Vaaya validates and prepares your photos for Seedance without regenerating your face in an image model.

Watch both performance samples

MiniMax H3 Max — generated from the supplied photos

Silent H3 Max AI performance sample: the leather-jacket performer gestures on the left while the blue-blazer performer points and turns on the right in the orange studio. · Download MP4

The video stream is preserved without re-encoding. We removed its output audio for a silent visual comparison with the supplied example.

The new H3 Max request uses the original photos, explicit performer assignments and the recipe’s seven-second source performance as its motion reference. We asked for photorealistic faces, natural lighting and realistic clothing. The FLUX portrait prompt was not used for this video.

Vaaya / ReAPI — supplied reference video

Silent supplied AI performance sample: the leather-jacket performer stays on the left while the blue-blazer performer gestures, turns and performs on the right beneath a hanging microphone. · Download MP4

This is the supplied 25.5-second file, preserved without replacing its visuals or soundtrack. It has no audible soundtrack. Its length and export are properties of this supplied example, not the recipe’s seven-second Standard or 29-second Extended options. We did not generate a new ReAPI video or independently record the settings used for this existing sample.

Both clips keep the leather jacket on the left and the blue blazer on the right, with the orange set and hanging microphone intact. In the H3 Max clip, the right performer points forward, then turns toward the side. The longer supplied example also shows her turn her back to the camera and return. Those changes of angle are useful places to inspect facial consistency, hair and clothing.

The supplied example demonstrates a longer sequence, while H3 Max covers approximately seven seconds. We do not use that difference alone to declare a quality winner.

What to look for in both clips

Detail What a strong result should preserve
Character identity A recognizable face through head turns and changes of expression
Performer placement Each person remains on their assigned side, with two distinct identities
Wardrobe The intended outfit’s recognizable colors, shape and details
Performance Gestures, framing and timing that follow the source clip
Faces and hands Coherent features during movement, including near the microphone
Finished playback The expected clip length and the intended soundtrack

Follow each person through the full clip. A good opening frame is only the beginning; the result has to hold together as they perform.

A prompt built for each character upload

“Put us in the video” leaves important decisions unstated. Which person goes where? Do two photos show one person or two? Should an outfit come from the source performance or the upload?

Vaaya turns those choices into explicit instructions. You assign the viewer’s left, right, or both performers. Each assigned person gets a separate set of one to five photos. The recipe then maps that set to the matching performer in the prompt.

With two people, it explicitly instructs the model to keep their identities separate, preserve their sides, and avoid putting one face on both performers. With one person, it asks the model to preserve the unassigned performer. Your creative notes are appended within those boundaries.

That means the prompt changes with the number of people, their assignments, their reference-photo counts and your notes. Each upload becomes part of an explicit character-to-performer mapping.

For example, you can ask: “Put me on the left and my friend on the right. Keep our outfits.” After you assign the photos, the recipe builds instructions around those choices:

Your choice What the recipe tells the model
Your photos assigned to the left Replace the viewer’s left performer with the person in those references
Your friend’s photos assigned to the right Use that separate identity and wardrobe for the right performer
Both performers selected Keep two distinct people; do not swap sides or merge faces
“Keep our outfits” Apply the wardrobe notes to both assigned people while preserving the source scene

A clear face photo gives the model an identity reference; an outfit photo gives it wardrobe detail. Keeping each person’s photos together helps the recipe express exactly whose appearance belongs in each position.

A model selected for this performance

The Hotel Lobby recipe uses Seedance 2.5 through ReAPI in video-edit mode. It starts with the saved performance and instructs the model to preserve its motion, timing, camera and scene while replacing the assigned performers.

This is why the recipe matters. You are asking for a specific performance with your characters in it. The source clip carries the choreography; the reference photos carry identity and wardrobe; the prompt connects them.

The recipe pairs a selected video model with character-specific reference mapping and prompt construction. That is how each upload gets tailored instructions while every run follows the same performance-editing workflow.

The soundtrack is part of the result

The production workflow generates the visual edit, then restores the original saved clip’s audio with FFmpeg. It checks the duration and verifies that the final export contains an audio stream before delivering the MP4.

The result uses the soundtrack people recognize, with the saved performance as its timing reference. Lip synchronization still depends on the generated visual edit; restoring the audio ensures the finished file contains the original track.

Standard uses a seven-second saved clip. Extended uses 29 seconds from the same performance. You choose the length and the people; the recipe handles the photo preparation, performer instructions, video edit and final export.

How these samples were made

The supplied photos and video were added on September 27, 2026. We made a new image-edit request for each person with each image model: FLUX.2 [klein] 9B via its LoRA-capable edit endpoint, with no custom LoRA, and Nano Banana Pro via its edit endpoint. FLUX was requested at 1536 × 2048; Nano Banana Pro at 4K. Presentation images are resized WebP files, without retouching.

The revised FLUX prompt requests a photographic portrait with light editorial retouching. The Nano Banana prompt emphasizes photographic lighting and realistic surface texture. They share character references and intended outfits, but they are not identical-prompt model tests.

H3 Max received the original photo uploads directly, plus the saved seven-second Hotel Lobby source clip. It was requested at 1080P with prompt expansion disabled. Its video stream was copied without re-encoding and its audio removed for this silent comparison. The supplied Vaaya MP4 was copied unchanged; its generation inputs and settings were not independently verified in this session.

Read the prompts, model identifiers and asset provenance. The H3 Max reference endpoint supports image and video references; see fal’s API documentation.

Make the Hotel Lobby trend yours

For us, best-in-class means caring about all the details that make this particular trend work: recognizable characters, deliberate performer placement, the original performance and a complete, downloadable video.

Choose one or two people, add their photos, and tell Vaaya who belongs on each side. Use your own photos or photos of adults who have given permission.

Create your Hotel Lobby video with Vaaya.

Questions

What is the Migos AI trend?

The Migos AI or Hotel Lobby AI trend puts new characters into the familiar two-performer rap performance. Vaaya’s Colors recipe uses a saved performance clip and your character photos to make that transformation.

Can I replace both performers?

Yes. Assign one person to the viewer’s left and another to the viewer’s right, with a separate photo set for each. You can also replace just one performer.

Does every uploaded photo go through Nano Banana Pro?

No. Nano Banana Pro created the realistic demonstration portraits in this article. The production recipe prepares uploaded photos and sends them to Seedance 2.5 through ReAPI with a prompt built from the performer assignments.

How long is the Hotel Lobby video?

Standard is seven seconds and Extended is 29 seconds. Both use the saved performance and restore its original audio.

Try Vaaya with your agent.

Connect your agent, choose a service, and try your first call with Vaaya.

npx @vaaya/mcp install