AI Comic Book: Panels, Faces and Lettering

What actually comes back when you ask a model for a six-panel page, why the character drifts between panels, and why every professional comic workflow letters the art last.

In this article
  1. Can AI make a comic page?
  2. Speech bubbles and lettering
  3. The actual problem
  4. Whose story is it?
  5. When to generate panels
  6. Holding a character
  7. Is it big enough to print?
  8. What a comic costs
  9. What AI still cannot do
  10. Frequently asked questions
  11. The bottom line
Short answer

Models can now draw a full comic page, keep the character consistent across it, and letter your exact dialogue into the right panels. All of that is solved. What is not solved is control: nothing on the page can be edited afterwards, and the art style will not match the next page.

Three comic panels of the same detective character, the second and third generated from the first as a reference
Panel one came from a prompt. Panels two and three were generated with panel one attached as a reference image. Total cost: twenty-three cents.

Almost everything written about making comics with AI is now out of date, and this article started out repeating it. The plan was to show that asking for a whole page fails, that speech bubbles come back as gibberish, and that you have to build everything panel by panel.

Then the tests were run. All three turned out to be wrong.

So this is a guide to where the difficulty actually sits in 2026, with the failures and successes shown as they came back, unretouched.

Want to try the panel workflow? Rangy reference-chains a character across a sequence on your own API key, from about three cents a panel.

Try it free

Can AI actually make a comic page?

Yes, and considerably better than the current advice suggests. A single request for a six-panel page returned a correct three-by-two grid, the same character in every panel, the story beats in the right order, and readable dialogue in every bubble. It cost five cents and took under two minutes.

A six-panel comic page generated from a single prompt, with consistent character and readable dialogue
One prompt, one generation, five cents, unretouched. The brief gave six story beats and asked for dialogue; everything else here is the model's.

Look closely and it is doing more than holding together. The door reads "317 PRIVATE", the file on the desk reads "PROJECT NIGHTFALL", and the panel order carries the sequence properly — arriving, finding the lock, picking it, entering, discovering the file, realising she is not alone.

If you have read that AI cannot do comic pages, that was true and is not any more. Which moves the interesting question somewhere else entirely.

Can AI do speech bubbles and lettering?

Yes. A panel requested with three specific text elements returned all three exactly as written, apostrophes included, in correctly shaped bubbles with tails pointing at the right speakers. It also correctly lettered a door sign, a file label and a desk nameplate that were never asked for.

A comic panel with two speech bubbles and a caption box, all text rendered correctly
Every string came back exactly as specified. The door, the file and the nameplate were the model's own additions and are also legible.

This is the claim that has aged worst. Garbled comic lettering was universal two years ago and is the reason every guide still tells you to letter separately. On short uppercase dialogue, current models handle it.

The advice to letter separately is still right — but for a completely different reason than the one usually given, and that reason is the next section.

So what is the actual problem?

You cannot change anything. Every line of dialogue, every bubble position and every panel is baked into one flat image. Fixing a single typo means regenerating the entire page, which returns six different panels — including the five that were fine.

This is the same constraint that governs book cover typography, except a comic page has ten or fifteen text elements instead of two, and a comic has twenty-two pages instead of one cover.

Work out what that means in practice:

  • An editor's note becomes a re-roll. "Cut the second line" is a five-second change in a lettering file and a full regeneration here, with no guarantee the new art is as good.
  • Translation is impossible. A lettered file can be re-lettered in another language. A generated page has to be redrawn.
  • Your typeface is an approximation. The model produces something comic-like, and it will differ subtly between pages.
  • Bubble placement is not yours. Where a bubble sits controls reading order, and reading order is craft rather than decoration.

The rule that follows. Generate the art with deliberate empty space where bubbles will go, and letter it afterwards in a layout tool. Not because the model gets the letters wrong — it does not any more — but because text you can edit is worth far more than text you cannot.

Can you get your exact script into the panels?

Yes, and this was the third assumption that failed. A page requested with six specific lines of dialogue, one per panel in a stated order, returned all six verbatim in the correct panels — and obeyed an instruction to add no other text anywhere on the page.

A six-panel comic page where all six specified lines of dialogue appear verbatim in the correct panels
Six exact lines, six panels, correct order, no stray signage. Also a completely different art style from the page above — which is the next problem.

That is worth sitting with, because it removes the last technical objection people usually raise. The model will draw your page, letter your words, and put them where you said.

But compare this page with the one earlier in the article. Same subject, same character description, same noir brief, two generations apart — and one is full-colour with amber street lamps while the other is near-monochrome sepia. Within a page, consistency is excellent. Between pages, the style is a coin toss.

For a twenty-two page comic that is the difference between a book and a collection of unrelated illustrations. It is also the strongest practical argument for building pages out of individually referenced panels, because a reference image holds a look across generations in a way a description does not.

When should you generate panels individually?

Whenever the page has to be right rather than merely good, and whenever it has to match the page before it. Individual panels give you six times the pixels, per-panel iteration without losing the others, and a style that holds across a whole book — at the cost of assembling the page yourself.

The measurement is stark. The one-shot page came back at 2,048 × 3,072 pixels, so each of its six panels carries roughly 1.05 megapixels. A single panel generated on its own came back at 3,072 × 2,048 — 6.29 megapixels, six times as much. That headroom is what lets you crop in, enlarge a panel for a cover, or fix one shot without touching the rest.

The panel-by-panel workflow, which is what the three panels at the top of this article are:

  • Generate one panel you are happy with and treat it as the reference for everything after it. Spend real effort here; it defines the whole look.
  • Attach it as a reference for every subsequent panel and describe only what changes — the new action, the new angle, the new setting.
  • State what must not change. The face, the hair, the coat, the palette, the linework. Reference-guided models drift toward a generic version without it.
  • Assemble and letter in a layout tool. Panels, gutters, bubbles and captions as real editable objects.

How do you hold a character across a whole book?

By anchoring every panel to one approved image rather than to a description. Words drift; a reference image does not. The three panels above share a face, a haircut, a tan coat, a red scarf and a palette because panels two and three were generated from panel one, not from a re-typed description of it.

Over twenty-two pages this matters more than any other decision. A character described in words will be a slightly different person by page eight, and readers notice long before they can say why.

  • Keep a character sheet. One approved image per recurring character, plus a front and three-quarter view if you can get them consistent. That is your source of truth.
  • Re-anchor, never chain. Always reference the original approved image, not the most recent panel. Chaining panel to panel compounds drift.
  • Lock costume and palette explicitly in every prompt, because those drift faster than faces do.
  • Accept a re-roll rate. Roughly one panel in three needs regenerating. Budget for it rather than fighting it.

The broader techniques, including when a trained model beats reference images, are in the character consistency guide.

Is the output big enough to print?

For a standard comic page, yes, almost exactly. A US comic trim of 6.625 by 10.25 inches at 300 DPI needs 1,987 by 3,075 pixels, and the one-shot page came back at 2,048 by 3,072 — 103% of the width and 100% of the height required.

That is a genuinely close fit, and it means a generated page is print-ready as-is for a standard trim. It also means there is no margin: no room to crop, no room to bleed, no room to enlarge.

Individual panels are the opposite. At 6.29 megapixels each, a single panel has enough resolution to run full-page, become a cover, or be cropped several ways — which is another argument for generating them separately when the work matters.

Check the file, do not trust the setting. Requested resolutions are not guarantees; the same 2K request in different generations returns different pixel dimensions. Read the actual size of every file before laying out a page for print.

What does an AI comic cost?

Almost nothing per image, which is not the same as almost nothing per book. The three panels at the top of this article cost twenty-three cents together. A one-shot page costs five. A twenty-two page comic built panel by panel, with a realistic re-roll rate, lands in the region of twenty to forty dollars.

Approach What you get Cost
One-shot page A complete lettered page, uneditable, ~1 MP a panel $0.05
Panel from a prompt Establishes the look; 6.29 MP $0.05
Panel from a reference Holds the character; needs a stronger model $0.09
A six-panel page, built properly Editable, high resolution, your script ~$0.50 plus re-rolls

Rates from Rangy's live pricing tables at the cheapest configured provider, checked August 2026. Reference-guided panels here used Nano Banana Pro; providers change rates.

The cost is not the interesting number. Twenty-two pages of comic art has never been a materials problem — it has been a time and skill problem, and what changed is that the drawing stopped being the bottleneck.

What can AI still not do for comics?

Tell a story. It draws well, letters well and holds a character, but pacing, panel rhythm, where to cut, what to withhold and how a page turn lands are the actual craft of comics — and none of it is a rendering problem.

Stated plainly, because this guide is published by a company that sells image generation:

  • Sequential storytelling. The gap between panels is where comics happen. A model fills panels; it does not decide what to leave out.
  • A consistent style across pages. Demonstrated above — two pages from near-identical briefs came back in different palettes. Within a page it holds; between pages it does not, without reference images.
  • Consistency at book length. One page is easy now. Ninety panels across twenty-two pages, with several recurring characters and recurring locations, is still hard work.
  • Editability. Every note from an editor, translator or letterer costs a full regeneration unless you kept the text out of the art.
  • A voice. Generated pages tend toward a competent house style. Distinctive linework is why people follow artists.

The version of this that works is not the model replacing the artist. It is the model doing the rendering while a person does the writing, the pacing and the lettering — which is roughly how comics have always been made, with the labour redistributed.

Frequently asked questions

Can AI generate a whole comic page in one go?

Yes. A single request for a six-panel page returned a correct grid, one consistent character across all six panels, the story beats in order and readable dialogue in every bubble, for five cents. The limitation is not quality but control: the page is one flat image, so changing anything means regenerating all six panels.

Can AI write text in speech bubbles correctly?

On short uppercase dialogue, yes. A test panel requested with three specific strings returned all three exactly, apostrophes included, in properly shaped bubbles with tails pointing at the right speakers. This is a genuine change from a couple of years ago, and most guides still saying otherwise are out of date.

Why should I letter separately if the model gets it right?

Because generated text cannot be edited. An editor's cut, a translation, a repositioned bubble or a typo all require regenerating the whole image and losing the art you approved. Text set in a layout tool can be changed in seconds. Generate the art with clear space where bubbles go, then letter it.

How do I keep the same character across every panel?

Approve one image of the character and attach it as a reference for every subsequent panel, describing only what changes. Always re-anchor to that original image rather than to the most recent panel, since chaining panel to panel compounds drift. Lock costume and palette explicitly, because those drift faster than faces.

Is a generated comic page big enough to print?

For a standard US comic trim of 6.625 by 10.25 inches at 300 DPI you need 1,987 by 3,075 pixels, and a 2K page came back at 2,048 by 3,072 — just clearing it with no margin for bleed or cropping. Individual panels generated separately carry about six times the pixels of a panel inside a one-shot page.

Can I sell a comic made with AI art?

Generally yes, subject to each model provider's terms, though several platforms now require you to declare AI-generated content and those rules have changed repeatedly. Copyright protection for purely AI-generated images is also unsettled and varies by jurisdiction, so take advice before building a business on exclusivity. Check the current policy wherever you publish.

Can I control the exact dialogue in each panel?

Yes. A page requested with six specific lines, one per panel in a stated order, returned all six verbatim in the correct panels and obeyed an instruction to add no other text. The remaining problem is not getting the words in, it is that once they are in you cannot change them without regenerating the page.

Which model is best for comic panels?

Different ones for different jobs. GPT Image 2 handles pages, layout and in-image lettering best, which makes it the choice for the establishing panel and for anything with text. Nano Banana Pro is stronger at holding a supplied character across new scenes, which makes it the choice for every panel after the first.

The bottom line

The rendering problem is solved and the control problem is not. Models will draw your page, keep your character and letter your dialogue — but hand you a flat image you cannot change. Generate the art, keep the words out of it, and assemble the page yourself.

The honest summary of the tests in this article is that all three of the things it set out to warn about are no longer true. Whole pages work. Lettering works. Exact scripts work. It is worth checking claims like these yourself rather than inheriting them, because this field invalidates its own advice every few months.

What has not moved is the part that was never about rendering. Deciding what happens in the gap between two panels is still the job, and no model has been asked to do it.

"This software has increased my workflow speed tenfold, and the output quality it has delivered in my work has been exceptional."

T @Taste.budtales · YouTube comment

Reference-chain a whole sequence

Rangy runs GPT Image 2 and Nano Banana Pro side by side on your own API key, so you can establish a panel with one and hold the character across the rest with the other — with every panel saved to a project folder on your disk.

Download Rangy Free Or read the character consistency guide →

Mac & Windows · Free plan, no credit card · 5 generations a day

How this article was made

Every image here is a real generation shown unretouched. The six-panel page was one GPT Image 2 request at 2K for $0.05; the brief supplied six story beats and asked for dialogue, and the model wrote the lines. The lettering panel was a separate GPT Image 2 request specifying three exact strings, all of which came back correct. A third test asked for a six-panel page carrying six exact lines of dialogue in a stated order, and all six came back verbatim in the correct panels with no stray signage. The three-panel sequence is one GPT Image 2 panel at $0.05 plus two Nano Banana Pro panels at $0.09 each, generated with the first attached as a reference image. Pixel figures are read from the files: the page at 2,048 × 3,072 giving roughly 1.05 MP a panel, against 6.29 MP for a panel generated alone. The print requirement is arithmetic — 6.625 by 10.25 inches multiplied by 300. Model rates come from Rangy's live pricing tables, checked on 14 August 2026, and will drift as providers change them. No platform's AI-disclosure policy is quoted, because they differ and change.

This guide is published by Rangy, which makes one of the tools it describes, so it is not a neutral source. The case against the approach is in what you cannot edit and what AI still cannot do, and both are the honest ones.