← Blog

videoSeptember 16, 2026

How to turn any photo into a video with your AI agent

Upload a photo, describe the motion and let your agent run an image-to-video model. See the original image, actual generated clip, exact prompt and recorded cost.

To turn a photo into a video, give your AI agent the image, describe what should move and set a clip length and budget. The agent sends the photo to an image-to-video model, waits for the render and returns a playable video. With Vaaya, it can discover and call the model through the same tool connection it uses for other paid work. Start with a short, simple movement: clouds drifting, water rippling or a slow camera move. Supported files and output quality vary by model; “any photo” is not a guarantee that every detail will survive unchanged.

A small mechanical assistant watches a mountain-lake photograph become a sequence of frames with moving clouds and water.

Original generated illustration. The working example below uses a real, separately credited photograph.

See the original photo and the actual generated video

We used this photograph of Lake McDonald in Glacier National Park. The U.S. Fish & Wildlife Service credits Ryan Hagerty/USFWS and marks the image public domain. We downloaded its original 1,200 × 900 JPEG and supplied those unchanged bytes to the model. Photo source and credit.

Original photograph of Lake McDonald: dark mountains below a cloudy sky, with a lake and wooded shoreline in the foreground.

Before: the original photograph, Ryan Hagerty/USFWS. Download the source JPEG.

We then asked Seedance 2.0 Fast, through Vaaya, for a five-second silent animation using its 720p setting and the same 4:3 framing. This is the resulting file:

AI-generated animation of the Lake McDonald photo: clouds drift above the mountains and small ripples move across the lake. Silent clip. · Download MP4

After: AI-generated motion from the photograph. This is not footage recorded at the lake, and the source agency did not produce or endorse the animation.

The downloaded MP4 is 5.04 seconds, 1,112 × 834 pixels and 24 frames per second, with no audio track. The clouds visibly change and the lake surface moves. The model also redraws fine mountain and cloud detail; it does not preserve the original photograph pixel for pixel. Those are observations from this result, not a promise about every run.

The generation request, recorded result and artifact provenance document this run. The video above is hosted with this article, rather than depending on a temporary provider link.

What prompt did we use?

Here is the exact motion prompt from the request:

Animate this landscape photograph into a natural five-second shot. Keep the camera completely locked and preserve the original framing, mountain shapes, shoreline, trees, colors and overcast lighting. Add gentle small ripples moving across the lake and slow, subtle cloud drift. Keep the mountains and shoreline still. No camera movement, no zoom, no cuts, no new people, boats, birds, buildings or text. Natural restrained motion throughout.

The photo supplies the scene. The prompt supplies the change. “Make this cinematic” leaves the model many decisions to invent; naming the movement gives you something concrete to evaluate.

For your own image, try this instruction:

Turn this photo into one short video. Move [the specific subject] gently, keep [the important details] stable and use [a fixed camera or one slow camera move]. Keep the original framing. Tell me the model, duration and quoted cost before running it, and stay within my approved budget. Return the finished video and save the source image and prompt.

What does the agent actually do?

First, inspect the image and choose a compatible model. Our example uses the Vaaya model key seedance-2-0--fast--image-to-video. Its provider documents JPEG, PNG and WebP inputs up to 30 MB, with 480p or 720p output. These are this model's limits, not a universal rule for every video tool. Seedance 2.0 Fast input schema.

Next, make the image reachable. An attachment in your chat or a filename on your laptop is not automatically a URL the video provider can read. For this run, the agent requested an upload slot through fal/upload, uploaded the original JPEG bytes and passed the returned file URL to the model.

Submit one generation with explicit settings. We requested duration: "5", resolution: "720p", aspect_ratio: "4:3" and generate_audio: false. The exact request is linked above. Duration and image-field names differ across models, so the agent should use the current schema instead of copying these fields into an unrelated endpoint.

Check that job until it finishes. Vaaya returns a job ID for the asynchronous render. Use result with that ID to retrieve status and output. Calling use again creates another generation; it is not a refresh button. The Vaaya tool reference explains this workflow.

Download and review the result. Save the MP4, watch the entire clip and compare it with the still. Check the beginning and end for altered objects, jumping edges or unwanted camera movement. A successful render means a file was produced, not that every instruction was followed perfectly.

What should you ask to move?

Keep the first attempt to one main action. These are starting prompts, not guaranteed outcomes:

Photo Useful first motion What to inspect
Landscape Gentle clouds or water movement; fixed camera Mountains, buildings and the horizon stay coherent.
Food or drink A little steam rising The cup, plate and food keep their shape.
Product One slow camera move Labels, logos and product geometry remain usable.
Person Subtle breathing or a small head movement Face, hands and clothing do not distort.

A strong camera move asks the model to invent parts of the scene the photograph never captured. If preserving the source matters more than dramatic movement, begin with a fixed camera and modest subject motion. Use photos you have permission to upload and animate, especially when they show other people.

How much did this example cost?

This example cost $1.36: 1¢ to stage the photo and $1.35 for one successful video generation. The generation call had a $1.50 ceiling. We made one render, then retrieved that job's result; we did not buy repeated attempts to produce the clip shown here. See the upload record and generation result. Recorded September 16, 2026, Pacific time; artifact timestamps use UTC.

Treat that as the observed price of this run, not a promise for every model, duration or future request. Ask the agent to check the current quote and pass a per-call ceiling. A fresh generation is another purchase, even if you are only changing one word in the prompt. Vaaya's spending controls explain the available limits.

If the account cannot fund the next attempt, keep the image, prompt and any completed video. The recovery steps are in what happens when your AI agent runs out of money mid-task.

Can I do this by texting Instinct?

With Instinct on WhatsApp or iMessage, the instruction can be conversational: “Turn this lake photo into a five-second silent video. Move only the clouds and water, and keep the camera still.” Where your Instinct setup has access to Vaaya, the assistant can use that connection to find and run the video tool within your approved budget. This is an example instruction, not a captured Instinct conversation; the actual artifact on this page was generated through Vaaya's tool call. The useful handoff is the finished video together with its source, prompt and cost.

What if the result looks wrong?

Change one thing at a time. If the camera drifts, ask for a locked shot. If a face or product warps, reduce the requested movement or choose a clearer source. If the composition changes, check the requested aspect ratio. Instructions can improve the next attempt, but cannot guarantee exact preservation.

If the job fails before producing a video, read the returned error first: an inaccessible image or unsupported setting needs a different fix from an unsatisfactory render. If it is still running, keep checking the same job.

For a first try, choose a photo with one clear subject, ask for a short movement and judge the actual file. Save the version that works along with the request that produced it.

Questions

How do I turn a photo into a video with an AI agent?

Give the agent a supported image, describe the movement and specify the clip length, framing and budget. The agent uploads the image, runs an image-to-video model, checks the existing job until it finishes and returns the downloaded video.

Can any photo become an AI video?

Many photos can, but each model has file limits and the result is not guaranteed. Clear subjects and modest motion are good starting points. Faces, hands, lettering and fine product details can change during generation.

Does the video show what really happened after the photo was taken?

No. Image-to-video models invent motion from a still image and a prompt. The generated clip is an animation, not recovered footage or evidence of a real event.

Does a photo-to-video clip need sound?

No. This example requests a silent clip. On the Seedance 2.0 Fast model used here, audio can be disabled with generate_audio: false. Other models may use different parameters.

Should I submit the same request again while the video is rendering?

No. Keep the returned job ID and check it with result. Submitting another generation starts another job and can create another charge. Download the completed MP4 rather than relying on the provider URL as permanent storage.

Try Vaaya with your agent.

Connect your agent, choose a service, and try your first call with Vaaya.

npx @vaaya/mcp install