Vixlens
FeaturesModelsPricingPresetsGallery
Log inStart for Free
Vixlens
FeaturesPricingModelsGuidesPresetsGalleryAboutContactTermsPrivacy

© 2026 Vixlens. All rights reserved.

  1. Home
  2. Guides
  3. How to turn a photo into an AI video

8 September 2026 · 7 min read

How to turn a photo into an AI video

Image-to-video is the most reliable way to get an AI video of something specific — your product, your face, your artwork — because the model starts from a picture instead of guessing what you meant. Here is how to do it well.

Text-to-video is impressive and unpredictable. You describe a scene, the model invents everything in it, and the thing you get back is a plausible interpretation rather than the shot you had in mind. That is fine for a mood piece and useless when the video has to contain your actual product.

Image-to-video removes most of that uncertainty. You give the model a still — a photograph, a render, a piece of artwork — and the prompt stops describing what exists and starts describing what happens. The subject is settled before generation begins.

Pick a model that actually reads the image

Not every model that accepts an upload uses it. Some providers list an image field in their schema, take the file, and then generate as though you had sent nothing — the request succeeds, you get billed, and your product is not in the video. On Vixlens the model pages state this explicitly: the reference-image row on each one says Required, Optional or Not used.

Three broad groups are worth knowing about:

  • Image-required models. HappyHorse 1.1 and Wan 2.7 have no text-only mode at all. The still you upload is the subject, and the prompt only directs motion. These are the safest choice when fidelity to the source photo matters more than anything else.
  • Image-optional models with strong reference handling. Seedance 2.5 holds a subject's identity across a long take from a single reference; Seedance 2.0 takes up to four subject references at once, which is what you want when a clip has to feature more than one recurring person or object.
  • First-and-last-frame models. Vidu Q3 Pro and P-Video accept both an opening and a closing image. That turns the generation into an in-between: you supply where the shot starts and where it ends, and the model fills the motion.

Start from a better still

Everything about the output is downstream of the input. A model cannot invent detail that was never in the source, and it will happily invent detail that was ambiguous in it.

  1. Use the highest resolution you have. A phone photo straight from the camera roll beats the same photo after a messaging app has recompressed it.
  2. Crop to the aspect ratio you want out. Models that take a reference image usually ignore the aspect-ratio parameter and follow the image instead. Crop to 9:16 before uploading if the result is going on TikTok.
  3. Keep the subject unobstructed. A hand across a product label, a face half out of frame, motion blur on the thing that matters — all of these give the model licence to redraw the ambiguous part.
  4. Prefer even, directional light. Hard mixed lighting is the most common cause of features drifting between frames.

Write a motion prompt, not a scene prompt

This is where most first attempts go wrong. People upload a photograph of a bottle and then write the prompt they would have written for text-to-video: a matte black perfume bottle on a mirrored plinth, studio lighting, luxury advertisement. The model now has two descriptions of the subject — the image and the sentence — and where they disagree, it compromises. The bottle changes shape.

Describe only what changes. Assume the model can already see everything else:

  • Slow push in on the subject, background falling out of focus, everything else held still.
  • The subject smiles slowly and turns their head toward the camera, hair moving slightly.
  • The camera cranes up and back to reveal the wider room.

Camera vocabulary is worth learning because it is unusually well understood by these models: push in, pull back, pan left, tilt down, orbit, crane up, handheld follow, locked off. So are motion qualifiers — slow, gentle, sudden, continuous — which do more work than any adjective about the subject.

Use the negative prompt to hold things still

Both Alibaba image-to-video models take a negative prompt, and on an animation job its most useful job is suppression rather than aesthetics. Background people moving, camera shake, text appearing, hands changing shape is a more effective negative prompt than low quality, blurry.

Keep clips short, then join them

Drift accumulates. A face that is perfect for two seconds may be subtly wrong by eight, because each frame is conditioned on the last. If a shot only needs three seconds, ask for three seconds — it is cheaper, faster, and more likely to hold.

For anything longer, generate several short clips from the same reference image and cut them together rather than asking for one long take. The exception is Seedance 2.5, which is built for long single takes and is the model to reach for when a shot genuinely has to run without a cut.

A workable process

  1. Pick the still and crop it to your delivery ratio.
  2. Draft on a cheap model first — Seedance 2.0 Mini or Vidu Q3 Turbo — to find out whether the motion idea works at all.
  3. Rewrite the prompt so it describes only motion and camera.
  4. Re-run the same prompt on the model you actually want to deliver from.
  5. Generate two or three takes; motion is stochastic and the second take is often the good one.
Every generation costs credits based on the model, resolution and duration, and the exact cost is shown before you press generate. New accounts get 50 credits to experiment with — see pricing for what the paid plans include.

Common failure modes

  • The subject morphs. Usually too long a duration, or a prompt that re-describes the subject. Shorten the clip and strip the subject description out of the prompt.
  • Nothing moves. The prompt described a state rather than an action. A calm lake at dawn gives a still image with grain; mist drifting across a calm lake at dawn, reeds moving in the foreground gives a video.
  • The wrong thing moves. Add the unwanted motion to the negative prompt, or lock the camera if the model supports it — Seedance 2.0 has a genuine fixed-camera mode.
  • The upload was ignored. Check the model's reference-image row. Some models accept a file and use nothing from it.

Models mentioned here

Alibaba

HappyHorse 1.1 — Live

Alibaba's image-to-video model, tuned for faces and close-ups

Alibaba

Wan 2.7

Alibaba Wan 2.7 image-to-video, from a single still

ByteDance

Seedance 2.5

Long-form AI video with reference control and native audio

Vidu

Vidu Q3 Pro

Vidu Q3 at up to 1080p with synced audio and end-frame control

Put this into practice

50 free credits on sign-up, no credit card required.

Start generatingSee pricing

Keep reading

  • Text to video: the complete guide
  • How to create ad creative with AI
  • AI video generators without a watermark: what to look for in 2026