Server data from the Official MCP Registry
Generate images and video from text prompts across Gemini, OpenAI and Kie AI.
About
Generate images and video from text prompts across Gemini, OpenAI and Kie AI.
Security Report
mediagen is a well-structured MCP server for image and video generation with proper authentication, secure credential handling, and reasonable permissions aligned with its purpose. No critical security issues found. Minor quality observations around error handling and input validation are present but do not materially impact security. Supply chain analysis found 6 known vulnerabilities in dependencies (1 critical, 4 high severity). Package verification found 1 issue.
4 files analyzed · 10 issues found
Security scores are indicators to help you make informed decisions, not guarantees. Always review permissions before connecting any MCP server.
Permissions Required
This plugin requests these system permissions. Most are normal for its category.
How to Install
Add this to your MCP configuration file:
{
"mcpServers": {
"io-github-cripacx-mediagen": {
"args": [
"-y",
"mediagen"
],
"command": "npx"
}
}
}Documentation
View on GitHubFrom the project's GitHub README.
mediagen — an image and video generation skill for coding agents
Give your agent image and video generation across Google Gemini, OpenAI and Kie AI — one skill, one install, no API to learn.
Say roughly what you want and let it decide the rest:
You: Make a hero image for the landing page, wide, something moody and industrial.
Agent: npx -y mediagen image "A disused loading dock at dusk, wet concrete…"
--aspect-ratio 21:9 --json
→ ./output/image-20260823T094107Z.png
It wrote the prompt out in full, knew 21:9 rules OpenAI out, and read the path
back from --json.
Or be exact, and it stops deciding:
You: Same thing with gemini-3-pro-image, 21:9, 4K, marked as AI-generated,
saved as hero.png.
Agent: npx -y mediagen image "A disused loading dock at dusk, wet concrete…"
--model gemini-3-pro-image --aspect-ratio 21:9 --size 4K
--output-name hero.png --json
→ ./output/hero.jpg
[looks at the image]
npx -y mediagen mark ./output/hero.jpg --visible-label
--label-position top-right --json
→ ./output/hero.labelled.jpg
It marked in a second pass, after seeing where the subject ended up — and put the label on the flat roofline rather than over the loading bay.
Anything the CLI takes can be asked for in words — provider, model, aspect ratio, size, output name and directory, quality preset, video duration, an input image to edit, and whether to mark the result. Name any of them and the agent uses it as given; leave it out and it chooses, or falls back to what you configured.
Install
npx -y skills add Cripacx/mediagen --skill mediagen
That is the whole installation. Nothing else to set up — the skill runs the
CLI through npx, which fetches it on first use and caches it afterwards.
Then give it a key:
npx -y mediagen init
An interactive wizard: pick providers, enter each key without it being echoed, verify each against the live API, choose a model per provider, and set a default. One key is enough to start.
Now just ask your agent for an image.
What your agent can do with it
- Generate images and video from a description, or edit media you already have
- Choose the right provider for the task — the skill knows Gemini handles 21:9 and video, that OpenAI cannot do 16:9 at all, and that Kie aggregates around thirty third-party models
- Take your instructions literally when you give them. Name a model, a
ratio, a size or a filename and it is used as stated.
mediagen modelsis there when you want to see what is available before choosing - Write the prompt properly. The skill carries prompt-writing guidance, so "a hero image, moody and industrial" becomes a full description of subject, composition, light, camera and materials before anything is sent
- Mark AI-generated output when it matters — the skill suggests it for anything photorealistic or destined for publication
- Recover from failures on its own, because every error carries a machine-readable code and a next action
And what it will not do: ask you to paste an API key into the chat. On a
configuration error the skill tells you to run mediagen init in your own
terminal.
Prerequisites
- At least one API key:
- Google Gemini (get one here) — images and video, widest range of shapes
- OpenAI (get one here) — images
- Kie AI (get one here) — ~30 third-party models
- Node.js 20.11 or later
Configuration
mediagen init covers first-time setup. To change something later:
npx -y mediagen config edit
A menu of every setting with its current value and where that value came from —
provider, model, key, output directory, quality preset, and what mediagen mark does by default. Each change is written as you make it.
For CI and scripts, where there is no terminal, config set takes the same
settings as arguments:
echo "$GEMINI_API_KEY" | npx -y mediagen config set gemini --stdin
npx -y mediagen config set gemini-model gemini-3-pro-image
Which provider gets used
Set the order you prefer, most preferred first:
npx -y mediagen config set provider-priority gemini,kie,openai
It is a preference, not a whitelist. A request with no --provider goes to the
first provider in the order that has a key and can do the job — so a
missing OpenAI key never stops something Gemini can do, and asking for video
skips straight past the providers that do not make any. Naming --provider
overrides all of it.
mediagen models shows the resulting order, which providers are usable, and
what a request would get:
1. Google Gemini (gemini) [preferred, key from config file]
Would use: gemini-3.1-flash-image (provider default)
2. Kie AI (kie) [key from config file]
3. OpenAI (openai) [preferred, no key]
No key configured, so requests to it fail. Fix: mediagen config set openai
A request with no --provider uses gemini/gemini-3.1-flash-image.
With --json it adds wouldUse, usableProviders and providerPriority.
That is what the skill has an agent read before generating, so it never picks
a model from a provider you have no key for.
To check what is configured and whether the keys still work:
npx -y mediagen doctor
doctor reports, per provider, whether a key is configured, which layer it came
from, and whether the provider accepts it — keeping not configured,
rejected, unreachable and no cheap way to check distinct, because they
call for four different fixes.
[!WARNING] There is deliberately no flag that takes an API key as an argument. Arguments land in shell history and in the process list, where they outlive the command that used them.
Settings resolve from the environment first, then .env in the working
directory, then the config file:
| Variable | Purpose |
|---|---|
GEMINI_API_KEY, OPENAI_API_KEY, KIE_API_KEY | credentials; at least one |
MEDIAGEN_PROVIDER_PRIORITY | providers in preference order |
GEMINI_MODEL, OPENAI_MODEL, KIE_MODEL | default model per provider |
MEDIAGEN_OUTPUT_DIR | where media is saved |
MEDIAGEN_QUALITY | fast, balanced, quality |
MEDIAGEN_MARK | mark output by default |
MEDIAGEN_VISIBLE_LABEL | add a visible label by default |
[!TIP] A stale environment variable shadowing the key you just configured is the most expensive failure this kind of tool has.
mediagen config listmarks every shadowed value, so you can see it rather than guess.
Providers
| Provider | Images | Video | Editing | Key verification |
|---|---|---|---|---|
| Google Gemini | Nano Banana family, up to 4K | yes | yes | live probe |
| OpenAI | gpt-image family, DALL·E | — | yes | live probe |
| Kie AI | ~30 models: Flux, Imagen, Grok | — | most | no cheap probe; reported as such |
[!NOTE] OpenAI takes pixel dimensions rather than aspect ratios and genuinely cannot produce 16:9 — its widest image is 1536×1024, which is 3:2. Asking for 16:9 is refused by name rather than quietly served as something else. The skill knows this and routes wide shapes to Gemini.
A model absent from the listings is still sent to the provider, so a newly released model works before mediagen knows about it. Kie's catalogue is generated from Kie's own documentation rather than maintained by hand.
Content marking
The EU AI Act splits disclosure into two duties, so mediagen mark has two
independent switches:
| Flag | Duty | What it does |
|---|---|---|
| machine-readable | make it findable by tools | writes IPTC/XMP DigitalSourceType, on by default |
--visible-label | disclose it to people | composites the EU's official AI-content label |
Generating never marks. Marking is always a second command, run on the file afterwards:
npx -y mediagen image "a wide banner" --aspect-ratio 21:9 --size 2K --json
npx -y mediagen mark ./output/image-….jpg --visible-label --label-position top-left
That is not ceremony. A visible label has to go where the subject is not, and only the finished image can say where that is. The machine-readable marker is not free either: adding metadata to a JPEG or WebP means decoding and re-encoding it, so marking costs a second lossy pass — worth paying deliberately, not as a side effect of asking for an image. mediagen re-encodes at high quality to keep that cost small, but it cannot make it zero.
To have mediagen mark draw the visible label without being asked each time:
npx -y mediagen config edit # "AI marking by default"
A configured default is still overridable per run with --no-mark or
--no-visible-label.
The visible label is the European Commission's own icon, published with the
Code of Practice on Transparency of AI-generated Content and free to use
without attribution. Two of its three variants are used: AI GENERATED by
default, AI MODIFIED with --modified, for media a person made and a model
altered. The light or dark version is picked from what is actually under the
corner it lands in, because a label nobody can read is not a disclosure.
It sits bottom-right by default. --label-position moves it to another corner,
or auto puts it wherever the image has the least detail.
[!IMPORTANT] A visible label is never written over its source.
mediagen mark photo.png --visible-labelproducesphoto.labelled.pngand leaves the pixels ofphoto.pngexactly as they were, while writing the machine-readable marker into it — so whichever of the two you publish carries the disclosure. A label placed badly can be redone from an untouched original and from nothing else.--in-placeoverwrites if you really mean to.
Using it without an agent
The skill is a wrapper around a CLI, and the CLI stands on its own.
npx -y mediagen image "a wide banner" --aspect-ratio 21:9 --size 2K
npx -y mediagen image "make the sky stormy" --input ./photo.jpg
npx -y mediagen mark ./output/image-….jpg --visible-label
npx -y mediagen video "a marble rolling down a wooden track" --duration 6
npx -y mediagen models
Install it globally if you use it often enough to want the shorter command:
npm install -g mediagen
Options
| Option | |
|---|---|
--provider <name> | gemini, openai, kie |
--model <id> | see mediagen models |
--input <path> | source media to edit or transform |
--aspect-ratio <ratio> | 1:1, 16:9, 9:16, … |
--size <size> | 1K, 2K, 4K |
--duration <seconds> | video only |
--output-name <name> | the extension may select the format |
--output-dir <dir> | where to save |
--quality <preset> | fast, balanced, quality |
--json | exactly one JSON object on stdout |
--verbose --quiet | diagnostics on stderr |
[!TIP] Without the skill, prompt writing is on you. mediagen sends the prompt exactly as written — it does not expand or rewrite it. Decide subject, composition, light, camera or medium, materials and atmosphere, and say what you want rather than what you do not.
Scripting
With --json, stdout carries exactly one JSON object and nothing else.
Without it the saved path is the last line, and a failure writes nothing to
stdout at all — so reading the last line can never hand you an error message
where you expected a path.
{
"success": true,
"filePath": "./output/image-20260823T094107Z.png",
"kind": "image",
"provider": "gemini",
"model": "gemini-3.1-flash-image",
"mimeType": "image/png"
}
| Exit code | Meaning |
|---|---|
0 | success |
2 | invalid input or usage |
3 | configuration or credentials |
4 | generation, network, or file I/O |
A failure carries an errorCode — one of VALIDATION_ERROR, CONFIG_ERROR,
API_ERROR, NETWORK_ERROR, FILE_ERROR, CONTENT_BLOCKED or TIMEOUT —
and a hint naming a concrete next action.
Using it as an MCP server
For hosts that speak MCP rather than running skills. Same package, started with
mediagen mcp, exposing generate_media, list_models and
check_configuration.
claude mcp add mediagen --env GEMINI_API_KEY=your-api-key-here -- npx -y mediagen mcp
The -- separates Claude's own flags from the command that starts the server.
Add --scope project or --scope user to change where the entry is written.
Settings → Developer → Edit Config, or edit directly:
- macOS:
~/Library/Application Support/Claude/claude_desktop_config.json - Windows:
%APPDATA%\Claude\claude_desktop_config.json
{
"mcpServers": {
"mediagen": {
"command": "npx",
"args": ["-y", "mediagen", "mcp"],
"env": {
"GEMINI_API_KEY": "your-api-key-here"
}
}
}
}
Restart Claude Desktop afterwards.
code --add-mcp "{\"name\":\"mediagen\",\"command\":\"npx\",\"args\":[\"-y\",\"mediagen\",\"mcp\"]}"
Or create .vscode/mcp.json in your workspace — note that VS Code uses
servers, not mcpServers:
{
"servers": {
"mediagen": {
"command": "npx",
"args": ["-y", "mediagen", "mcp"]
}
}
}
Cursor, Windsurf, Zed and most others take the same shape as Claude Desktop, in their own config file:
| Field | Value |
|---|---|
command | npx |
args | ["-y", "mediagen", "mcp"] |
| transport | stdio |
env can be omitted from any of these once mediagen init has run — the
server reads the same configuration the CLI does.
How it works
One pipeline, three frontends. The skill drives the CLI; the MCP server is a second adapter over the same core. Neither contains behaviour of its own, so a capability cannot exist in one and be missing from the other.
agent skill · CLI · MCP server
↓
request → model resolution → capability check → provider client
↓
save to disk → result
mediagen mark → AI content marking
Each provider is one self-contained directory declaring what it supports.
Adding one touches a single line outside its own folder. Provider manifests
carry no vendor SDK imports, so doctor and config never pay to load one.
src/
├── types/ leaf types everything shares
├── core/ pipeline, errors, capability checks, file handling
├── config/ the three configuration layers, key verification
├── providers/ one directory per provider, plus the registry
│ ├── gemini/
│ ├── openai/
│ ├── kie/
│ └── shared/ polling for asynchronous providers
├── cli/ the command tree; output.ts owns stdout
├── mcp/ the MCP server
└── marking/ AI content marking
skills/mediagen/ the agent skill
scripts/ catalogue generation, version syncing
Contributing
See CONTRIBUTING.md for development setup, how to add a provider, and how releases are cut.
Reviews
No reviews yet
Be the first to review this server!
More Developer Tools MCP Servers
Fetch
Freeby Modelcontextprotocol · Developer Tools
Web content fetching and conversion for efficient LLM usage
Git
Freeby Modelcontextprotocol · Developer Tools
Read, search, and manipulate Git repositories programmatically
Toleno
Freeby Toleno · Developer Tools
Toleno Network MCP Server — Manage your Toleno mining account with Claude AI using natural language.
mcp-creator-python
Freeby mcp-marketplace · Developer Tools
Create, build, and publish Python MCP servers to PyPI — conversationally.
MCP Marketplace
Freeby mcp-marketplace · Developer Tools
Search and install MCP servers from inside your AI client.
MarkItDown
Freeby Microsoft · Content & Media
Convert files (PDF, Word, Excel, images, audio) to Markdown for LLM consumption
