To turn a photo into an AI video in 2026: (1) pick a model that supports image-to-video — Luma Ray 3, Runway Gen-4.5, Veo 3.1, or Kling 3.0. (2) Set your first frame as the anchor. (3) Optionally lock the last frame for predictable endings. (4) Describe the motion in plain language. (5) Generate at 5-10 seconds; longer needs stitching. Costs $0.50-$5 per clip.
Image-to-video is the most underrated workflow in AI in 2026. Text-to-video gets the headlines, but starting from a real photo gives you something text-only generation cannot — a fixed, controllable starting point that already looks exactly right. Your composition, lighting, and subject identity are locked before the AI does anything. The model only has to invent motion, not the entire frame.
This guide walks through the actual workflow: which model to pick, how first-frame and last-frame anchoring work, how to describe motion so the model honors it, the mistakes that produce garbage clips, and pricing across the five 2026 leaders. It is updated for May 2026 — verify the latest figures on each provider's page before committing budget.
TL;DR — the 5-step image-to-video workflow
- 1. Pick a model: Luma Ray 3 for portraits and best preservation, Runway Gen-4.5 for region-specific motion, Veo 3.1 for audio, Kling 3.0 for cheap clips, Pika 2.5 for effects.
- 2. Set the first frame: upload your photo at 1024px or higher and lock it as frame 0.
- 3. Lock the last frame (optional): upload a second image on Luma or Kling to interpolate predictable endings.
- 4. Describe the motion: use specific verbs ("zoom out slowly", "pan left", "character turns head right"). Avoid "move".
- 5. Generate at 5-10 seconds: longer means stitching multiple clips. Budget $0.50-$5 per finished clip.
Which AI tools support image-to-video in 2026?
Five tools dominate image-to-video in 2026, and each has a different sweet spot. Picking the right one matters more than tuning the prompt — a bad-fit model produces unfixable output no matter how well you describe the motion.
Luma Ray 3 is the standout for image-led storytelling. Its first-frame preservation is the strongest in the category — faces, products, and landscapes hold their identity tighter than on any competitor. Luma is also one of the only two models that exposes last-frame anchoring (more on that in section 3). The trade-off: shorter native clip lengths (5 and 10 seconds) and no audio. For portraits, product shots, and any clip where the source photo must look like itself, Luma Ray 3 is the default pick.
Runway Gen-4.5 is the editing suite. Its motion brush lets you paint which regions of the frame should move and which should stay locked — invaluable when you want hair to blow but the background to stay still, or eyes to move while the rest of the face is frozen. Runway is also the most production-ready end-to-end tool with timeline editing, color grading, and audio mix all in one app. Best for narrative work that needs region-specific motion.
Google Veo 3.1 can do image-to-video on the Pro tier and ships with native synchronized audio — dialogue, ambient sound, foley. If you need the clip to come out with sound already baked in, Veo is the only tool in the category that does it natively in 2026. Per-second pricing is higher than Luma and Kling, but the audio integration saves a full post-production pass.
Kling 3.0 is the cheap option. Quality has grown substantially through 2025-2026 releases and it now supports both first-frame and last-frame anchoring, which makes it the best low-cost choice for predictable clips. Quality is below Luma and Runway on demanding shots but acceptable for social media and rough cuts.
Pika 2.5 sits in a category of its own. Its image-to-video is solid but its real differentiator is "Pikaffects" — preset trendy effects (squish, melt, inflate, explode) that turn a still photo into a viral 3-second clip with one click. If the goal is social engagement rather than cinematic motion, Pika is faster than any of the others.
For a side-by-side breakdown of three of these on the same prompts, our Veo vs Kling vs Runway comparison walks through identical test shots. For the broader 2026 video market overview, the AI video generators compared 2026 pillar covers every tool worth using.
The first-frame trick — anchoring your photo as the start
The single most important concept in image-to-video is that your photo becomes literal frame 0 of the output. The model does not interpret your image as "inspiration" or "reference" — it uses the exact pixels as the starting state of the video and generates motion outward from there.
This sounds obvious but it has three practical implications most beginners miss.
First, resolution matters more than you think. A 1080px or 1024px source image is the floor. Below that, the model upscales internally before generating motion, which introduces blur and identity drift in the first 12 frames. Above 2048px, returns diminish but you get sharper edges on fast-moving subjects. The sweet spot is 1536-2048px on the longest edge.
Second, composition lock-in is your superpower. Because frame 0 is fixed, you keep complete control of the opening composition — your photographer's eye on framing, lighting, and color is preserved. Compare that to text-to-video, where you describe the scene and pray the model interprets "rule of thirds" the way you meant it. Image-to-video skips that lottery entirely.
Third, the photo's flaws come along for the ride. If the source has a dust spot, an awkward crop, or a distracting background element, those are now baked into the video. Clean up the source in your image editor before generating — it is far easier to spot-heal a still than to ask the model to fix a recurring artifact across 120 frames.
The first-frame trick is the reason image-to-video usually beats text-to-video for any clip that needs a specific look. You are not trying to coax the model into generating a specific scene — you are showing it the scene and asking only for motion.
Anchoring start AND end frames — predictable endings
The most underused capability in image-to-video is last-frame anchoring. Luma Ray 3 and Kling 3.0 both support it. Upload a first frame and a last frame, and the model interpolates the motion in between automatically.
This eliminates the dice-roll quality of free-form image-to-video. Instead of describing motion and hoping the ending lands somewhere acceptable, you specify both endpoints and let the model figure out the path between them. Three real-world use cases where it shines:
- Product transformation reels: first frame is the product closed, last frame is the product open. The AI generates the opening motion.
- Before/after reveals: first frame is a retouched portrait, last frame is the final composite. The model generates a smooth transition for social.
- Character pose shifts: first frame is the subject looking left, last frame is the subject looking right. The model interpolates the head turn naturally.
For multi-clip narratives, the trick scales further. Use the last frame of clip A as the first frame of clip B. The visual identity carries across the cut, and you can chain three or four 5-second clips into a coherent 15-20 second sequence. Read our broader how to turn photos into AI videos guide for the full multi-clip stitching workflow.
Runway Gen-4.5 and Veo 3.1 do not currently expose last-frame anchoring directly. Runway compensates with motion brush region control; Veo compensates with stronger temporal consistency on long single clips. But for predictable endings, Luma and Kling are the only two correct answers in 2026.
How to describe motion in your prompt
The biggest beginner mistake is treating motion prompts like image prompts. They are not the same thing. An image prompt describes a static scene; a motion prompt describes change over time. Three rules separate clips that work from clips that drift.
Rule 1: specific verbs, not vague ones. "Zoom out slowly" is a specific verb with a clear direction and pace. "Move" is not a verb that means anything to the model. The most reliable motion verbs in 2026 are: zoom in/out, pan left/right, tilt up/down, push in, pull back, dolly forward, dolly back, orbit, tracking shot, rack focus, hair blows in wind, fabric flutters, head turns, eyes blink, smoke rises, water flows.
Rule 2: one or two motions max, not five. A prompt like "the woman turns her head, the wind picks up, the leaves fall, the camera zooms in, the lighting shifts to sunset" is asking for five different motion vectors in five seconds. The model will compromise on all of them. Pick one primary motion (the head turn) and one secondary motion if appropriate (subtle hair movement). That is the entire prompt budget.
Rule 3: do not violate physics. "The car drives away then comes back" implies a reversal that no current image-to-video model handles cleanly within 5 seconds. "The character runs across the frame and disappears" requires the subject to fully exit, which most models fight against because they want to preserve the subject identity. Describe motion that obeys conservation laws and is consistent with one continuous take.
Common mistakes (and how to avoid them)
After watching thousands of failed image-to-video generations, the failure modes cluster into four recurring patterns.
Mistake 1: too many motion verbs at once. Five things cannot all happen smoothly in five seconds. The fix is brutal editing — cut your prompt to one primary motion and one optional ambient motion. If you want multiple distinct motions, generate multiple clips and cut between them.
Mistake 2: low-resolution source images. Anything under 1024px on the longest edge will produce a soft, drifting opening. The model has to upscale internally and that upscale fights the motion generation. Fix it by upscaling your source first — Crystal, Magnific, or even Photoshop's Preserve Details 2.0 are all fine for this. Aim for 1536-2048px.
Mistake 3: describing physically impossible motion. Liquid flowing uphill, characters teleporting between locations, objects passing through walls — these break the model's temporal consistency. The output will either ignore the impossible motion or produce uncanny morphing. Stay within continuous, single-take physics.
Mistake 4: expecting identity to hold over 10+ seconds. Most current models will drift by 6-7 seconds even with anchoring. A face that started as recognizably yours will subtly shift toward a generic average. The fix is to cap clips at 5-6 seconds and stitch. For multi-clip work where the character must look identical across cuts, read our AI character consistency deep dive.
Pricing comparison across image-to-video tools
The economics of image-to-video moved fast through 2025-2026. Five-second clips that cost $5 in 2024 now run $0.50-$1.50 on the best model. Here is what the May 2026 pricing actually looks like per 5-second clip via official APIs.
| Tool | First-frame | Last-frame | Audio | Max clip | Cost per 5-sec |
|---|---|---|---|---|---|
| Luma Ray 3 | Yes (best preservation) | Yes | No | 10 sec | $0.50-$1.50 |
| Runway Gen-4.5 | Yes | No | Add-on | 10 sec | $0.25-$0.50 |
| Veo 3.1 (Pro) | Yes | No | Yes (native) | 8 sec | $2.00-$3.20 |
| Kling 3.0 | Yes | Yes | No | 10 sec | $0.50-$1.50 |
| Pika 2.5 | Yes | No | Add-on | 10 sec | $0.25-$0.75 |
Free tiers exist on all five but add visible watermarks, queue delays, and lower-resolution output. For client work, the paid API or pro-tier subscription is non-negotiable.
Worth knowing: Rangy includes Luma Ray 3 (currently in beta) for image-to-video, using your own Replicate API key. If you also do image work in Rangy, the source image is one click away — no exporting and re-uploading. Costs about $0.50-$1.50 per 5-second clip in API spend, with no Rangy subscription on top. The same workflow through Luma's direct subscription runs $30-$95 per month for comparable usage.
Which is best for portraits? For landscapes? For products?
The right model depends as much on subject as on feature checklist. After hundreds of test clips across the three most common use cases, the honest recommendations:
For portraits — Luma Ray 3 is the default. Its identity preservation is the tightest in the category, which is the whole game for face-driven content. For portraits that need region-specific motion (eyes move but the rest of the face is still), Runway Gen-4.5's motion brush is the secondary pick. Avoid Pika for portraits; its style biases toward effects-driven warping that distorts facial features.
For landscapes — Veo 3.1 wins when audio matters (wind, water, ambient nature sound). Luma Ray 3 wins on pure visual fidelity for natural environments. Kling 3.0 is the cost-effective third option and produces acceptable landscape motion at a fraction of Veo's price.
For products — Luma Ray 3 with last-frame anchoring is the killer combo. Lock the closed-product frame as the start and the open-product frame as the end; the model interpolates a clean transformation. For products with text or labels that must stay legible, Runway Gen-4.5 holds typography slightly better than Luma during motion.
The verdict — your 5-step image-to-video workflow
The default workflow for a 5-second cinematic portrait clip
Step 1: open your source portrait, confirm it is at least 1536px on the longest edge. Step 2: pick Luma Ray 3 for the strongest identity preservation. Step 3: set the portrait as the first frame. Step 4: write a single specific motion verb — "subtle wind moves the hair, eyes blink once". Step 5: generate at 5 seconds, expect $0.50-$1.50 in API cost, and you have a publishable clip in roughly 60 seconds of wall time. For anything else — landscapes, products, multi-clip narratives — substitute the model per the section above and the rest of the workflow is identical.
Frequently asked questions
Can I animate any photo with AI?
Almost any photo at 1024px resolution or higher will work. Source images under 1024px tend to produce blurry frames and identity drift. Faces, products, and landscapes all animate well. Heavily stylized illustrations sometimes lose their style during motion — for those, Pika and Runway preserve style better than Luma.
How long can image-to-video clips be?
Most models cap a single generation at 5-10 seconds in 2026. Luma Ray 3 supports 5 and 10 second clips natively, Runway Gen-4.5 supports up to 10 seconds, Kling supports 5 and 10 seconds. For longer outputs, use the last frame of the previous clip as the first frame of the next clip and stitch — up to roughly 60 seconds is practical before identity drift compounds.
Which is best for portraits?
Luma Ray 3 is the strongest for portraits because its first-frame preservation holds the face identity better than competitors. For portraits with significant motion (turning head, walking), Runway Gen-4.5 with motion brush gives more precise control over which regions move and which stay still.
How do I keep the character looking the same?
Three rules. (1) Keep clips under 6 seconds — identity drift compounds with length. (2) Anchor both the first and last frame when the model supports it (Luma, Kling). (3) Avoid motion verbs that imply major pose changes — instead of "character turns around", describe "character looks over their shoulder". For multi-clip narratives, read our guide to AI character consistency.
Can I lock the last frame?
Yes, on Luma Ray 3 and Kling 3.0. Upload a first frame and a last frame, and the model interpolates the motion in between. This is the single most underused trick in image-to-video — it eliminates the dice-roll quality of free-form generation and gives you predictable endings. Runway and Veo do not currently expose last-frame anchoring.
How much does image-to-video cost?
Luma Ray 3 costs roughly $0.50-$1.50 per 5-second clip via API. Runway Gen-4.5 runs $0.05-$0.10 per second. Veo 3.1 Standard is about $0.40 per second on Google's API. Kling 3.0 is $0.10-$0.30 per second on paid plans. Pika 2.5 ranges from $0.05-$0.15 per second. Free tiers exist on all five but usually add visible watermarks and queue delays.
Why does my AI video drift after 5 seconds?
Identity drift is the model gradually forgetting the original subject. It happens because each frame is generated conditioned on the previous frame, not the source image — error accumulates. Fix it by capping clips at 5-6 seconds, using last-frame anchoring when available, and using motion verbs that imply small movements (head tilt, slow zoom) rather than large ones (run, jump, walk away).
For text-only video generation (no source photo), see our companion guide on how to generate AI videos from text.
Image-to-video with one API key
Rangy includes Luma Ray 3 (beta) for image-to-video using your own Replicate API key. Plus the full 2026 AI image stack.
Download Rangy free