Claude Code cannot generate images on its own. It can once you connect an MCP server that exposes an image tool. Setup takes about two minutes, after which you ask in plain language and the assistant picks the model, runs the generation and writes the file to disk — on your own API key, at two to nine cents an image.
Claude Code cannot draw. Neither can Codex, or Gemini CLI, or any of the coding assistants people now spend their day inside. They are text models.
Ask one for a hero image for the landing page it just built and it will write you a very polite paragraph explaining that it cannot do that, then offer to add a placeholder div.
That gap is annoying in a specific way, because the assistant already knows exactly what image you need. It has the brand colours in context. It knows the component is 16:9. It read your copy. All the information required to write a great image prompt is sitting right there, and then the workflow breaks and you go open a browser tab.
MCP closes that gap. With one connection configured, your assistant gains a real image generator as a tool: it writes the prompt, picks the model, sets the aspect ratio, runs the generation and hands you a file path. This guide covers how that works, the three ways to set it up, what it costs, and when it is the wrong approach.
Want to wire it up now? Rangy 2.7 ships an MCP server — one toggle, one snippet, works with Claude Code, Codex, Gemini CLI and Cursor.
Try it freeCan Claude Code generate images?
Not by itself. Claude Code is a text and code model with no image generation built in. It gains the ability the moment you connect an MCP server that exposes an image tool — then you ask in plain language, and the assistant calls the tool, chooses the model and settings, and saves the finished file.
This is worth stating plainly because the question gets answered badly all over the web. There is no hidden flag, no plugin, no "image mode". The capability comes entirely from tools you attach.
The same applies to Codex CLI, Gemini CLI, Cursor and Windsurf. They are all MCP clients, so one server configuration serves all of them.
Why generate images from your terminal at all?
Because the assistant already holds the context. It knows your palette, your copy and your layout, so it can write a better prompt than you would retyping the request into a browser tab. The value is not saved clicks — it is not having to be the integration layer between two tools.
When you generate an image in a separate tab, you become that integration layer. You retype what you want, translate the assistant's understanding into a prompt, download the file, rename it, move it into the project, update the reference. Every step is somewhere for an error or a compromise to creep in, and by the third round trip most people quietly settle for a worse image because the loop is tedious.
With the generator as a tool, the loop collapses into conversation:
- "Generate a hero image for this landing page that matches the palette we're using." The assistant already knows the palette.
- "Make twelve item icons for the inventory system, same style, transparent background." It already knows what the items are.
- "That one's too dark. Redo it with the light coming from the left." It still has the original prompt.
The iteration case is where it really shows. Refining an image normally means editing a prompt string by hand and losing track of which version produced which file. In a chat loop the assistant keeps the history, and "warmer, and move the subject left" is a complete instruction.
How does MCP image generation actually work?
MCP is an open standard for letting an assistant call external tools through a uniform interface. Your assistant calls a tool such as generate_image, a server receives it, a generator executes the job against a real model, and the model provider bills your account directly. Configure it once and every MCP client can use it.
The chain has four links:
- Your assistant decides an image is needed and calls a tool such as
generate_imagewith a prompt, a model and an aspect ratio. - The MCP server receives that call. If it runs locally, it is a small process on your own machine.
- The generator executes the job against a real image model.
- The model provider does the inference and bills your account directly.
If link three is a cloud service, your prompts and outputs pass through someone else's infrastructure and you pay their margin. If it is an app already on your machine, generation runs on your own API keys, files land in your own folders, and nothing leaves except the API call the provider needs anyway.
What are the three ways to add images to Claude Code?
A hosted MCP image service, a server you write yourself, or a desktop app that exposes its own MCP server. Hosted is fastest to start and charges a margin. Rolling your own is a reasonable afternoon until you hit retries, queueing and per-model parameters. A desktop app already solved those.
Option 1: A hosted MCP image service
Several providers now offer a hosted MCP endpoint. Nothing to install, works from any machine, usually a free tier. The trade-offs are the usual hosted ones: a monthly fee or credit system on top of the underlying model cost, your prompts and images sitting in their storage, and whatever model selection they have chosen to support.
Option 2: Roll your own MCP server
A minimal image MCP server is maybe two hundred lines: accept a prompt, call a provider API, poll until done, write the file, return the path. If you only ever want one model with one set of options, this is a genuinely reasonable afternoon project and there are open-source starting points on GitHub.
Where it gets tiresome is everything after the happy path — per-model parameter differences, provider rate limits, retry and backoff, job queueing, cost accounting, and the fact that every model expects its aspect ratios and resolutions in a slightly different shape. That is the part that turns an afternoon into a side project.
Option 3: A desktop app that exposes its own MCP server
This is the approach Rangy takes in version 2.7. Rangy is a desktop AI image and video app for Mac and Windows; the MCP server is a toggle in its settings. Flip it on and your assistant gets the app's full toolset — every image model, every video model, the upscalers, the thumbnail maker — running on your own Replicate, Kie.ai or Google keys.
The practical advantage over rolling your own is that all the boring infrastructure already exists because the desktop app needed it: the job queue, the retry logic, the per-model parameter normalisation, the cost estimates. And everything generated through chat lands in the same local library as everything you make by hand, so there is one gallery rather than two.
| Hosted service | Roll your own | Desktop app | |
|---|---|---|---|
| Setup | Seconds | An afternoon | ~2 minutes |
| Cost on top of model | Their margin | None | None |
| Where files land | Their storage | Your disk | Your disk |
| Model choice | Their selection | Whatever you wire | Full lineup |
| Queue & retries | Included | You build it | Included |
| Works offline from the machine | No | App must run | App must run |
How do you set it up?
Turn on the MCP toggle in the app's settings, copy the snippet for your client, and restart it. In Claude Code that is one terminal command; in Codex and Gemini CLI it is a few lines in a config file. Then type /mcp and confirm the server is listed.
Step 1: Turn on the MCP server
Open Rangy, go to Settings, and find the Connect AI assistants (MCP) card. Flip the toggle on. It requires a Pro license, and Rangy has to be running for the connection to work — the assistant is talking to the app, so the app needs to be awake.
Step 2: Add it to your client
The same settings card has a tab per client, each with a copy button, so you never hand-write JSON:
| Client | What you do |
|---|---|
| Claude Desktop | Click Add to Claude Desktop, then fully quit and reopen the app |
| Claude Code | Copy one terminal command and run it once |
| Codex CLI | Paste the snippet into ~/.codex/config.toml, restart Codex |
| Gemini CLI | Merge the snippet into ~/.gemini/settings.json, restart |
| Cursor / Windsurf | Save the snippet to that client's MCP config file |
Step 3: Confirm it connected
Start a new Claude Code session and type /mcp. You should see rangy listed with its tools. In Claude Desktop, the tools appear under the sliders icon in the chat input.
Under the hood. Any MCP-compatible client works, because it is a standard stdio server. All a client needs is three values: the command, the path to the forwarder script, and the environment variable ELECTRON_RUN_AS_NODE=1. The settings panel generates all three for your install, so you never need to know the paths.
What can the tools do?
Beyond generate_image you get list_models, estimate_cost, generate_video, upscale_image, make_youtube_thumbnail, extract_assets, manage_folders and get_job_status. The pairing that matters most is list_models plus estimate_cost — together they let the assistant choose on real capability and real price rather than on half-remembered model names.
list_models— every image, video and upscaling model with prices, settings and which providers you have keys for.estimate_cost— the exact USD total for a planned batch, before anything runs.generate_image— the main event. Model, prompt, resolution, aspect ratio, reference images.generate_video— the same for video models, with duration and frame control.upscale_image— enlarge an existing file, up to very large output.make_youtube_thumbnail— the dedicated 16:9 thumbnail pipeline with title and logo placement.extract_assets— cut every object out of a generated sheet into separate transparent PNGs.manage_foldersandget_history— organise the output and look back at what was made.
Here is what a real batch looks like from the client side — the actual call sequence, not a mock-up:
> make 20 illustrations for the blog posts, 2K, dark theme
# 1. the assistant checks what is available and what it costs
list_models → gptimage2: kie $0.05 @2K · providers: kie, replicate
estimate_cost → 20 x gptimage2 @2K = $1.00 total
# 2. it asks you to approve the spend, then submits them all at once
generate_image x20 (wait: false)
→ job_4 running
→ job_5 running
→ job_9 queued, position 1
...
# 3. it polls the whole batch in one call
get_job_status [job_4 ... job_23]
→ done: 20, error: 0, spent_usd: 1.00
# every file is written to your local library, not a cloud bucket
Two details in that flow matter. The assistant calls estimate_cost before spending anything, so you approve a real number rather than a guess. And it submits all twenty with wait: false, then polls the batch in a single call — which is why twenty images take about as long as five.
How do you generate a batch?
Ask for the whole set in one instruction and let the server queue it. A good MCP image server runs a few jobs in parallel, queues the rest, staggers launches so providers do not rate-limit you, and retries throttled jobs automatically. Rangy runs five at once and queues up to sixty.
One image at a time is convenient. Thirty at a time is a different category of thing, and it is the reason this workflow is worth setting up.
Consider a game asset pack. You need forty props in a consistent style, each on its own transparent background. By hand that is forty prompt-edit-generate-download-crop cycles, an afternoon at minimum, and the style drifts by the end because you keep tweaking wording.
Through MCP it is one instruction. The assistant writes forty prompts sharing a style preamble, submits them all, and the queue handles the rest.
There is a neat trick for the transparent-background part. Rather than generating forty separate images, generate a handful of sheets with several props spaced out on one flat unused colour, then run extract_assets to cut each prop into its own trimmed transparent PNG. Fewer generations, lower cost, and the props come out visually consistent because they were rendered under the same lighting in the same pass.
Other batches that suit this shape: a blog's full image set in one go, storyboard frames, product shots across every aspect ratio you publish in, or the same hero concept in twenty variations.
Related Claude Code workflow Claude Code + Remotion = Free Motion Graphics (No Subscription) A different application of the same idea — one prompt scaffolds a working motion-graphics project, then you describe the animation in plain English. Not an MCP tutorial, but the closest thing on the channel to this workflow. Watch on YouTube →What does it cost?
Two to nine cents per image at 2K, paid directly to the provider with no markup. A forty-image asset pack on a mid-tier model costs about $2.40. Subscription image tools charge $20 to $30 a month whether you generate four hundred images or none.
| Model | Best for | Cost per image |
|---|---|---|
| Grok Imagine | Fast exploration, many options | ~$0.02 |
| Flux Klein | Quick text-to-image | ~$0.02 |
| Seedream 4.5 | High-detail generation | ~$0.03 |
| GPT Image 2 (2K) | Anything with text in it | ~$0.05 |
| Nano Banana 2 (2K) | Editing and composition | ~$0.06 |
| Nano Banana Pro (2K) | Maximum precision | ~$0.09 |
Kie.ai rates as of August 2026, checked against Rangy's live pricing tables. The same models through Replicate cost more, which is why per-model provider routing saves real money over a year.
The comparison worth making is against subscriptions. If you generate in bursts — which most people do — per-use pricing wins by a wide margin. That argument in full is in the no-subscription guide.
When is MCP image generation the wrong approach?
Skip it if you rarely need images, if you want to art-direct closely with reference images and side-by-side comparison, if you cannot leave a desktop app running, or if your work forbids sending client material to third-party APIs. A GUI beats a chat loop whenever you need to see options together.
I build one of these, so weigh this accordingly — but there are four cases where I would not use the MCP route.
- You need to art-direct visually. Comparing four variants side by side, zooming into detail, dragging reference images around — that is what a GUI is for. Chat is a poor medium for judging images.
- You generate images rarely. If images are an occasional need rather than part of a build loop, opening the app directly is simpler than maintaining a server connection.
- The app cannot stay running. A local server means the desktop app must be awake. On a headless box, in CI, or on a locked-down work machine, a hosted service is the realistic option.
- Client work with data restrictions. Generation still calls a third-party provider API. If your contract forbids sending client material to external services, MCP does not change that — the prompt and any reference images still leave the machine.
The honest framing: MCP is excellent when images are a byproduct of building something else. It is worse than a normal interface when the images are the work.
What goes wrong at first?
Four things trip up most people: the app is not running, the client was not restarted after configuring, no model was specified so the assistant picked a default, or a large batch was submitted without a cost estimate first. All four are quick to fix once you know them.
- The app has to be running. With a local server, the assistant talks to the app. If Rangy is closed, the tools return a "not running" error rather than silently failing — but it is still the most common first stumble.
- Restart the client after configuring. MCP servers are discovered at startup. Claude Desktop in particular needs a full quit, not just closing the window.
- Say the model name if you care. Left alone the assistant picks a sensible default, but if you want text rendered properly, ask for GPT Image 2 explicitly. The tools honour model overrides.
- Ask for a cost estimate before big batches. Thirty video clips is real money in a way thirty images is not.
- Watch aspect ratio support. Not every model accepts every ratio; some snap to the nearest supported one. If exact dimensions matter, check first.
- Generation is slow by tool standards. An image takes 30 to 90 seconds and video can take minutes. Submit batches as background jobs and poll rather than waiting on each one.
One toggle, one snippet. Rangy 2.7's MCP server works with Claude Code, Claude Desktop, Codex, Gemini CLI, Cursor and Windsurf.
Download RangyFrequently asked questions
Can Claude Code generate images?
Not on its own — it is a text and code model with no built-in image generation. It can generate images once you connect an MCP server that exposes an image tool. After that you ask in plain language, and the assistant calls the tool, picks the model and settings, and writes the finished file to disk.
What is MCP and why does it matter for images?
The Model Context Protocol is an open standard for letting AI assistants call external tools through a uniform interface. For images it matters because you configure the connection once and every MCP-compatible client can use it — Claude Code, Claude Desktop, Codex CLI, Gemini CLI, Cursor, Windsurf — with no per-app integration work.
Do I need an API key to generate images through MCP?
Yes, and that is the point. A local MCP image server runs on your own provider keys, typically Replicate or Kie.ai, so you pay provider rates of roughly two to nine cents per image instead of a monthly fee to a middleman. You top up credit directly and see exactly what each generation costs.
Is a local MCP image server better than a cloud one?
For most creative work, yes. Local keeps prompts and files on your machine, uses your keys with no markup, and imposes no limits beyond the provider's. Cloud services are easier to start with since there is nothing to install, but you pay their margin, your files live in their storage, and you are limited to the models they chose to support.
Can the assistant generate several images at once?
Yes, and it is where the workflow pays off. A good server accepts many jobs at once, runs a few in parallel, queues the rest and staggers launches so providers do not throttle you. Rangy runs five concurrently and queues up to sixty, so thirty variations is one instruction rather than thirty.
Does MCP image generation work with Codex and Gemini CLI?
Yes. MCP is a client-agnostic standard, so the same server works with Codex CLI, Gemini CLI, Cursor, Windsurf and VS Code with Copilot. Each needs the same three values — a command, an argument, and one environment variable — pasted into its own config file, and a restart afterwards.
Does the desktop app need to stay open?
Yes, for a local server. The assistant is talking to the running app over localhost, so if the app is closed the tools return a clear "not running" error. If you need generation from a headless machine or CI, a hosted MCP service is the right choice instead.
The bottom line
The reason to spend two minutes on setup is not that generating images gets faster. It is that generating images stops being a separate task — the assistant that wrote your copy can finish the job instead of handing you a placeholder.
The same applies to icon sets, storyboards, thumbnails and asset packs — anywhere the hard part was never the image, it was the twenty context switches on the way to it.
Turn on the toggle, paste one snippet, and ask for something. The first time an assistant answers "here's the hero image, saved to your library" instead of "I can't generate images", the workflow stops feeling like a demo.
"AMAZING! There is nothing like it out there. The functionality is superb. The one feature I love is local storage — I save what I need for the project and store them with the project for future work. Very happy with this developer!"
Connect Rangy to Claude Code
Rangy 2.7 ships a built-in MCP server — 25+ image and video models available to Claude Code, Claude Desktop, Codex and Gemini CLI, running locally on your own API keys.
Download Rangy Free Or read the setup docs first →Mac & Windows · Free plan to try the app · MCP needs a Pro licence
Every illustration in this article was generated through the MCP server it describes — submitted from a chat client as a single batch of twenty, at $1.00 total, and written straight into the local library. The call sequence shown above is that batch, not a mock-up. Prices come from Rangy's live pricing tables, checked on 14 August 2026, and will drift as providers change rates.
I develop Rangy, so this is not a neutral comparison. Where a hosted service or a normal interface is the better answer, I have said so in when MCP image generation is the wrong approach.