Keeping a character consistent across multiple AI images in 2026 follows a hierarchy: seed control gives mild consistency, reference images give strong consistency, and LoRA training gives near-perfect consistency. Nano Banana Pro and Flux 2 Pro lead for multi-reference (up to 10 input images). Midjourney's --oref helps but is less reliable. For commercial work, train a LoRA on Replicate.
The right choice depends on how many images you need and how strict the match has to be. Three lifestyle scenes for a single product? Reference images. A brand mascot across a whole campaign? LoRA. This guide walks all three methods in order so you do not waste hours fighting drift that a better method would have prevented.
TL;DR — the consistency hierarchy
- Seed control — mild consistency. Same prompt + same seed = similar (not identical) output. Free, instant, no setup.
- Reference images — strong consistency. Upload 1–10 images of the character. Nano Banana Pro and Flux 2 Pro lead.
- LoRA training — near-perfect consistency. Train once on 20–50 images, reuse forever. $2–$10 per training run on Replicate.
- Best for 1–3 images: seed + strong prompt vocabulary.
- Best for 4–15 images: reference images via Nano Banana Pro or Flux 2 Pro.
- Best for 15+ images or ongoing work: custom LoRA.
- Multi-subject scenes: Nano Banana Pro (up to 5 subjects) or Flux 2 Pro (up to 10 references with role assignment).
Why does AI change the character every time?
Image models are probabilistic. Every generation starts from noise and walks toward an image that satisfies your prompt. Same prompt twice = different noise, different path, different face. That randomness is what lets the model produce variety — and it is also the entire reason consistency is hard.
Three forces pull the model away from a stable character. Prompt ambiguity — "a man with brown hair" describes millions of men. Stochastic sampling — the random seed is different every run. Latent drift — when you chain edits, each step inherits small interpretations from the previous step and they compound.
Every method in this guide attacks one or more of those forces. Seed control freezes the sampling. Reference images replace ambiguity with pixels. LoRA training rewires the model's understanding of the character so even a vague prompt returns the same person.
The consistency hierarchy — seed, reference, LoRA
Treat these three methods as a ladder. Start at the bottom and move up only when the rung below fails for your job. Most people skip straight to reference images because they are the practical winner for everyday work — but understanding all three stops you over-investing in LoRA training for a job that needed thirty minutes of prompt discipline.
| Method | Setup time | Cost | Consistency level | Best for |
|---|---|---|---|---|
| Seed control | None | Free | Mild — same vibe, not same face | 1–3 images, exploration, same composition with small prompt tweaks |
| Reference images | 5 minutes | Standard per-image API cost | Strong — recognisable identity across 4–15 images | Campaign shoots, single product across multiple scenes, lookbook variants |
| LoRA training | 30–90 minutes setup + training time | $2–$10 per training run + standard inference | Near-perfect — same face indefinitely, any scene | Brand mascots, recurring characters, agency work, ongoing series |
| Combination (reference + LoRA) | Same as LoRA | Same as LoRA | Maximum — locks identity and pose | Highest-stakes commercial work where one drift kills the deliverable |
Method 1: Seed control (the easy starting point)
The seed is the random number that initialises the diffusion process. Most image APIs let you set it explicitly. Same prompt + same seed = same image. Same seed + one changed word = a similar image with that one change applied.
This is the cheapest form of consistency and it is right for one specific job: small variations of the same composition. Set the seed, generate a base image you like, swap a single attribute ("blue shirt" to "red shirt", "smiling" to "looking away") with the seed locked. Result: the same person in the same scene with that one change.
Where it fails: any meaningfully different scene. Change "in a coffee shop" to "on a mountain at sunset" and even with the seed locked, the face drifts because the new prompt activates different parts of latent space. Treat seed control as a tool for micro-variations, not full scene changes. When a seed produces a face you like, write it down — a seed plus a prompt is the most portable representation of a character you can carry between tools.
Method 2: Reference images (the practical winner)
Reference images replace prompt ambiguity with actual pixels. Instead of describing "a man with brown hair and green eyes wearing a navy jacket," you upload a photo of that man and let the model copy the identity-defining details directly. This is what most working creators use, and the 2026 landscape has made it dramatically more reliable than a year ago.
Nano Banana Pro leads multi-subject reference work. It accepts up to five reference subjects per scene and keeps each one identifiable across multiple frames in the same generation. "The same woman from reference 1 next to the same dog from reference 2, in a park" actually returns that woman with that dog. For lookbooks and brand campaigns, the cleanest workflow.
Flux 2 Pro goes further on count and control. It accepts up to ten reference images per generation with explicit roles you assign in the prompt — "color from image 1, lighting from image 2, composition from image 3, character from image 4." That role assignment is the feature you cannot get anywhere else in 2026, and the only tool that maps cleanly to layered moodboard direction.
GPT Image 2 uses chat memory. Describe the character once at the start of a conversation, and the model holds that description across every subsequent generation in the session. Works well for short sessions of three to seven images. On longer sessions it drifts. Right for a focused half-hour of variants, not a full-day campaign.
Midjourney's --oref (omni-reference) takes one reference image URL appended to your prompt. Simplest syntax of any reference workflow. Works well for single-subject portraits with a clear front-facing reference; less reliable when the scene includes multiple characters or when the target pose differs significantly from the reference angle.
Vocabulary discipline matters even more with reference workflows. Describe the character with identical words across every prompt — same hair colour, eye colour, body type, outfit. Our guide on how to write better AI prompts covers the patterns that make this reliable.
Method 3: Training a LoRA (when you need it)
LoRA (Low-Rank Adaptation) is a small set of additional weights you train on top of a base model (Flux, Stable Diffusion, SDXL) that teaches it a specific concept — a face, a style, an object. Once trained, the LoRA is a portable file you load alongside the base model at inference. After training, prompt "her smiling on a beach," "her hiking a mountain," "her at a wedding" and get the same person every time, no upload step. No upper bound on consistent images, and the LoRA is reusable for the life of the base model.
Training on Replicate: gather 20–50 source images of the subject, varied across angles, lighting, expressions, and distances so the model learns the invariant identity rather than overfitting. Square crops at 1024px are the convention. Replicate's LoRA training endpoint takes a zip plus a few hyperparameters (steps, learning rate, trigger token). Training runs 20–60 minutes and costs $2–$10 depending on base model and config.
Civitai hosts thousands of community-trained LoRAs free. Useful for personal exploration, but commercial use depends entirely on each LoRA's license, set by the uploader and varying enormously. Always check the license tab before using one in client work, and bias toward LoRAs you trained yourself when commercial stakes are real.
When to train: more than fifteen images of the same character, recurring use across projects, or stakes high enough that training cost is a rounding error.
Which tools have the best consistency in 2026?
Ranked by how reliably each tool holds character identity across multiple generations, with the right method applied:
- Custom LoRA (Replicate, on Flux 2 base) — near-perfect identity preservation indefinitely. The ceiling.
- Nano Banana Pro multi-reference — strongest commercially available reference workflow without training. Up to 5 subjects per scene, recognisable across frames.
- Flux 2 Pro 10-reference — most controllable for layered direction. Explicit role assignment per reference image.
- Flux Kontext Pro — best for precise reference-driven edits where you want to keep almost everything identical and change one specific element.
- GPT Image 2 chat memory — best for short conversational sessions. Drifts on longer ones.
- Midjourney
--oref— fine for single-subject portraits. Less reliable for multi-subject or large pose changes. - Seed control on any model — minimum-effort method for micro-variations only.
Running this workflow needs Nano Banana Pro for multi-subject reference, Flux Kontext for precise edits, and Flux 2 Pro for 10-image reference stacking. Rangy gives you all three through bring-your-own-API-keys (Replicate + Kie), no Rangy subscription. LoRA training also runs through Replicate using the same key.
How to keep your product (not character) looking the same
Product consistency is the same problem with two practical differences that make it easier. Products do not express emotion, and they do not vary their pose. A single clean reference shot is usually enough.
Flux Kontext Pro is the cleanest tool for reference-driven product edits. Feed it a studio shot of the product plus a prompt describing a new scene ("the same bottle on a marble countertop with morning light") and it preserves the product geometry, label artwork, and material finish while changing everything around it. Nano Banana Pro multi-subject also handles products well — particularly when the scene needs both a consistent product and a consistent model. Upload references of both, prompt the scene, and the output preserves both.
For a full e-commerce pipeline, our guides on the best AI tools for product photography and AI for e-commerce product photos walk through catalogue generation and on-model shots. For creator-focused workflows where characters recur across thumbnails and scenes, see AI for content creators.
Common mistakes and how to avoid them
Four mistakes account for almost every "my character keeps drifting" complaint:
- Inconsistent prompt vocabulary across images. If image 1 says "young woman with chestnut hair," image 2 says "woman with brown hair," and image 3 says "her," the model parses three different descriptions. Pick one exact wording for every identity-defining attribute and copy it verbatim into every prompt.
- Reference image at the wrong angle. A three-quarter portrait reference will not produce a clean straight-on shot. The reference angle leaks into the output. For multiple angles, supply multiple references (Flux 2 Pro) or train a LoRA.
- Reference too dark, blurry, or low resolution. Below 1024px the model cannot resolve enough detail to lock identity. Below adequate exposure, the model interpolates shadows and the interpolation is what drifts. Clean, evenly-lit, sharp 1024px+ is the minimum.
- Expecting consistency past 10–15 images without a LoRA. Reference workflows fray around image 10–15 — small drift in jawline, eye spacing, proportions. Past that count, switch to a LoRA.
One more worth naming: do not confuse style references with identity references. "In the style of" is a different problem with a different method — our guide on how to extract style prompts from any image covers separating the two.
The verdict — your character consistency workflow
The most common job: one product across ten lifestyle scenes
Source one high-quality reference image at 1024px+, well-lit, front-facing. Open Nano Banana Pro or Flux Kontext Pro. Upload the reference. Write your scene prompt with consistent vocabulary. Generate the first three and review them side by side. If identity holds, run the rest. If you see drift past scene seven, switch to Flux 2 Pro and add the first clean generation as an additional reference. If drift persists past scene ten, you needed a LoRA — pause, train one on Replicate, and finish the batch with identity locked. Reference-only cost: a few dollars in API fees. LoRA escalation: $10–$20 plus a few hours, and the LoRA is reusable forever.
Frequently asked questions
How does Midjourney's --oref work?
Midjourney's --oref (omni-reference) parameter takes one reference image URL appended to your prompt. It is reliable for one-subject scenes with a clear front-facing reference, but less reliable for multi-character scenes or when the reference angle differs from the target pose. For multi-subject work, Nano Banana Pro or Flux 2 Pro is more dependable.
Do I need to train a LoRA?
Only if you need more than 10–15 consistent images of the same character. For short campaigns or one-off projects, reference-image workflows with Nano Banana Pro or Flux 2 Pro are faster and cheaper. LoRA training makes sense when the character recurs across many projects or when you want unlimited variants without re-uploading references.
Can I keep my product looking the same in every shot?
Yes — often more reliable than character consistency because products do not express emotion or pose. Flux Kontext Pro is the cleanest tool: feed a clean product photo plus a new background prompt, and it preserves the product geometry while changing the scene. Nano Banana Pro multi-subject also handles products well.
What's the easiest way to keep a character consistent?
Upload one clean reference image to Nano Banana Pro or Flux 2 Pro with a prompt describing the new scene. Fastest setup for 1–10 images, no training. Reference should be 1024px+, well-lit, shot from the angle you most want preserved. Use identical vocabulary across all prompts.
How many reference images can I use?
Depends on the model. Nano Banana Pro accepts up to 5 reference subjects per frame, each identifiable across multiple frames. Flux 2 Pro accepts up to 10 with explicit role assignment ("color from image 1, lighting from image 2"). Midjourney's --oref takes one. GPT Image 2 holds character descriptions in chat memory across a session.
Why does my character drift after a few images?
Drift compounds when each new image is influenced more by the previous output than your original reference. Four causes: inconsistent prompt vocabulary, reference at the wrong angle, reference too dark/blurry/low-res, or pushing past 10–15 generations without a LoRA. Re-anchor to your original reference each round instead of chaining outputs.
How much does training a LoRA cost?
Roughly $2–$10 per training run on Replicate, depending on base model and duration. You need 20–50 source images — varied angles, lighting, expressions. Output is a reusable model file at standard inference pricing (a few cents per image). Civitai hosts thousands of community-trained LoRAs free, but commercial-use rights depend on each model's license — always verify.
Character consistency in one app
Rangy gives you Nano Banana Pro, Flux Kontext, and Flux 2 Pro through your own API keys. Plus LoRA training via Replicate using the same key. No Rangy subscription.
Download Rangy free