media_gen.py — Gemini + OpenRouter media generation
Companion to video_gen.py (which runs video on OpenRouter, budget-capped).media_gen.py generates speech, images, and music. Speech runs on the
Gemini API free tier (GEMINI_API_KEY) — free, no spend guard needed.
Images and music default to OpenRouter (OPEN_ROUTER_API_KEY) — the same
$5-capped key video_gen.py uses, real dollars per call, guarded by the same
estimate → check real balance → refuse pattern. Both credentials come fromprototypes/.env.
Verified against the live Gemini and OpenRouter APIs on 2026-08-11. Results
below are real — see trials/speech/, trials/images/, trials/music/,trials/gen-log.md (every Gemini call), and trials/spend-log.md (every
paid OpenRouter call, same table video_gen.py writes to).
Bottom line
| Modality | Default backend | Status | Real cost |
|---|---|---|---|
| Speech (Arabic TTS) | Gemini (free) | WORKS | $0 |
| Image | OpenRouter | WORKS | ~$0.04/image (gemini-2.5-flash-image) |
| Music (Lyria) | OpenRouter | WORKS | $0.04/30s clip, $0.08/song |
| Image, Music | Gemini (--backend gemini) | BLOCKED | free tier limit: 0 — see bottom |
Gemini's free tier allocates zero daily requests to every image model and
both Lyria models on this key (confirmed empirically, not a shape bug — see
the bottom of this doc). OpenRouter routes the exact same underlying models
through Google AI Studio as a paid provider and it just works. All three
subcommands are proven end-to-end below with real generated files and real
dollar amounts.
Working: speech (Gemini, free)
python media_gen.py speech --text "الذكاء الاصطناعي وفرص العمل وبناء مصادر الدخل" \
--voice Kore --out ../trials/speech/kore-flash25.wav
Request shape (discovered empirically)
POST /v1beta/models/{model}:generateContent, not :predict and not
the newer /v1beta/interactions endpoint — plain generateContent works.
{
"contents": [{"role": "user", "parts": [{"text": "..."}]}],
"generationConfig": {
"responseModalities": ["AUDIO"],
"speechConfig": {
"voiceConfig": {"prebuiltVoiceConfig": {"voiceName": "Kore"}}
}
}
}
Response shape
{"candidates": [{"content": {"parts": [{"inlineData": {
"mimeType": "audio/L16;codec=pcm;rate=24000", "data": "<base64 PCM>"
}}]}}]}
Audio comes back as raw 16-bit little-endian PCM, mono, 24kHz — not a
WAV file, not MP3, no container at all. media_gen.py decodes the base64
and wraps it in a WAV header via Python's wave module
(_save_audio() in media_gen.py). The two TTS models format mimeType
slightly differently (audio/L16;codec=pcm;rate=24000 vs.audio/l16; rate=24000; channels=1) — the sample-rate regex handles both.
Models tried, both work
gemini-2.5-flash-preview-tts— default, tried first.gemini-3.1-flash-tts-preview— also works, fallback.
Both succeeded on the first try, no 429s during testing. --model overrides
the fallback chain if you want to pin one.
Arabic verification
Generated the exact sentence from the brief, الذكاء الاصطناعي وفرص العمل وبناء, with three voice/model combos, saved under
مصادر الدخلtrials/speech/:
| file | model | voice | duration | peak amplitude | silence ratio |
|---|---|---|---|---|---|
kore-flash25.wav | gemini-2.5-flash-preview-tts | Kore | 4.85s | 28331/32767 | 25.9% |
kore-flash31.wav | gemini-3.1-flash-tts-preview | Kore | 5.52s | 26526/32767 | 29.8% |
puck-flash25.wav | gemini-2.5-flash-preview-tts | Puck | 4.81s | 24744/32767 | 23.8% |
All three: valid RIFF...WAVEfmt header, mono 16-bit PCM @ 24kHz, strong
peak amplitude (not clipped, not silent), and a silence ratio consistent
with natural pauses between words in a ~40-character sentence spoken at a
normal pace (~8 characters/sec). This is strong signal-level evidence the
models produced real speech, not noise or silence.
Caveat — I cannot judge pronunciation quality by ear. I'm a text model;
I verified the audio is valid, non-trivial, and the right duration/shape for
speech, but I did not and cannot listen to confirm the Arabic actually sounds
natural, correctly stressed, or free of mispronunciation on
"الذكاء الاصطناعي" / "مصادر الدخل" specifically. Have a person listen totrials/speech/kore-flash25.wav before relying on this for the workshop.
Gemini's TTS models are documented to support Arabic and auto-detect the
input language (there's no separate "Arabic voice" — all 30 prebuilt voices
are multilingual), so quality risk is about accent/naturalness, not
"garbage output," but that's still a real risk worth a human check.
Recommendation if you need a pick without listening first: Kore viagemini-2.5-flash-preview-tts — it was the first one tried, has the lowest
silence ratio of the three, and matched duration expectations most closely.
Working: image (OpenRouter, ~$0.04/image)
python media_gen.py image --prompt "flat vector icon of a lightbulb, cyan and navy blue color palette, minimal, on white background" \
--aspect 16:9 --ceiling 0.10 --out ../trials/images/lightbulb-icon.png
Real output: trials/images/lightbulb-icon.png — valid PNG (\x89PNG\r\n\x1a\n
header), 1344×768 (confirmed via ffprobe), 262,968 bytes. Cost: $0.0387,
confirmed via the response's own usage.cost field.
Discovery process
The coordinator's tip mattered: GET /v1/images/models (200) lists models
but no generation endpoint. I found the real one by sending an empty POST
body to candidate paths and reading the Zod validation error, which is free
(it 400s before billing):
POST /v1/images/generations {}
POST /v1/images {}
Both return the identical 400:
{"success":false,"error":{"name":"ZodError","message":"[
{\"path\":[\"model\"],\"message\":\"Invalid input: expected string, received undefined\"},
{\"path\":[\"prompt\"],\"message\":\"Invalid input: expected string, received undefined\"}
]"}}
Both routes exist and accept the same body — media_gen.py uses /v1/images
(the name OpenRouter's own docs use). Sending {"n": 0} against the real
model confirmed n's valid range (>=1) the same free way — a second 400,
still no spend.
Request shape (confirmed working)
{
"model": "google/gemini-2.5-flash-image",
"prompt": "<prompt>",
"aspect_ratio": "16:9",
"n": 1
}
Response shape (confirmed working)
{
"data": [{"b64_json": "<base64 PNG>", "media_type": "image/png"}],
"usage": {"prompt_tokens": 20, "completion_tokens": 1290, "total_tokens": 1310,
"cost": 0.038706}
}
usage.cost is the real, authoritative per-call cost — media_gen.py logs
this exact number to trials/spend-log.md as actual, next to the
pre-flight estimate.
Model choice and pricing
google/gemini-2.5-flash-image ("Nano Banana") spends a fixed ~1290
output-image tokens per call (confirmed: 1290 exactly on this test, at
16:9) at $0.00003/token = $0.0387/image, matching the coordinator's
estimate almost exactly. media_gen.py hardcodes $0.04 as the pre-flight
estimate (rounds up slightly, safe). google/gemini-3-pro-image was never
called — the coordinator flagged it as ~4x the price (~$0.15/image) and
not worth it for this volume; it isn't in OR_IMAGE_PRICING at all, so
passing --model google/gemini-3-pro-image fails closed with "no pricing
known" rather than spending blind.
n is capped at 1 server-side for this model (confirmed via/v1/images/models: "n": {"min": 1, "max": 1}), so -n 3 inmedia_gen.py means 3 separate calls, each individually estimated and
guarded — not a batch request.
Gemini-direct backend (--backend gemini) — still blocked
Every Gemini image model — gemini-3.1-flash-image, gemini-3-pro-image,nano-banana-pro-preview, gemini-2.5-flash-image, gemini-3.1-flash-lite-image
— returns the same error on the free tier, on the very first request:
HTTP 429
* Quota exceeded for metric: generate_content_free_tier_requests, limit: 0, model: gemini-3-pro-image
"status": "RESOURCE_EXHAUSTED"
limit: 0 means the free tier is allocated zero requests for that model,
period — not "you used today's quota." This is a billing/tier gate, not a
shape bug: it fires before the request body is even validated.
Confirmed alias: nano-banana-pro-preview and gemini-3-pro-image hit
the identical quota bucket (model: gemini-3-pro-image in both errors) — same
model, two names. The Gemini-side request shape media_gen.py sends (kept
working for whenever billing is turned on):
{
"contents": [{"role": "user", "parts": [{"text": "<prompt>"}]}],
"generationConfig": {"responseModalities": ["IMAGE"], "imageConfig": {"aspectRatio": "16:9"}}
}
Google's docs (ai.google.dev/gemini-api/docs/image-generation) describe a
newer /v1beta/interactions endpoint for these models instead — confirmed to
exist (same 429 limit:0, not a 404) but never validated against a real
success response. If generateContent ever 400s once billing is on, try that
shape next ({"model", "input":[{"type":"text","text":...}], "response_format":, response under
{"type":"image","aspect_ratio":...}}steps[].content[].data).
Working: music (OpenRouter, $0.04/clip or $0.08/song)
python media_gen.py music --prompt "upbeat inspiring instrumental synth pop, hopeful, 120bpm, 30 seconds" \
--seconds 30 --ceiling 0.10 --out ../trials/music/intro-clip.mp3
Real output: trials/music/intro-clip.mp3 — valid MP3 (ID3 header,
confirmed via ffprobe: codec mp3, 44.1kHz stereo, 192kbps, 27.98s),
677,526 bytes. Cost: $0.0400, confirmed via usage.cost.
Discovery process
GET /v1/audio/models and GET /v1/music/models both true 404 (HTML
fallback page, not a JSON error) — no dedicated model-listing surface for
audio. POST /v1/audio/generations and POST /v1/music/generations
also true 404 — no dedicated generation endpoint either. Music only
exists through /v1/chat/completions.
First attempt, non-streaming, cost nothing because it was rejected before
generation:
{"model": "google/lyria-3-clip-preview",
"messages": [{"role": "user", "content": "<prompt>"}],
"modalities": ["text", "audio"]}
{"error": {"message": "Audio output requires stream: true", "code": 400}}
That's the whole discovery cost for this path — one free 400 that states the
exact missing field. Setting "stream": true and parsing the SSE response
worked on the next try (see below).
Request shape (confirmed working)
{
"model": "google/lyria-3-clip-preview",
"messages": [{"role": "user", "content": "<prompt, optionally + \n\nLyrics:\n<lyrics>"}],
"modalities": ["text", "audio"],
"stream": true
}
--lyrics is folded into the prompt text (no separate structured field was
found or needed — Lyria took a plain style/mood prompt fine for an
instrumental test).
Response shape (confirmed working) — SSE, one audio chunk
Content-Type: text/event-stream, lines of data: {...}\n\n, terminated bydata: [DONE]. Unlike a typical token-by-token audio stream, the entire clip
arrived as one chunk in the final event before [DONE]:
{"choices": [{"delta": {"audio": {"data": "<base64 MP3, whole clip>"}}}],
"usage": {"prompt_tokens": 17, "completion_tokens": 4, "total_tokens": 21,
"cost": 0.04}}
media_gen.py concatenates every delta.audio.data seen (correct even if a
future model does chunk it) and decodes once at the end. The decoded bytes
start with ID3 — a real MP3 container, not raw PCM, so no WAV-wrapping is
needed (unlike the Gemini TTS path). usage.cost on the final event is the
real charge, logged verbatim to spend-log.md.
Model choice and pricing
Model card pricing (pricing.prompt/pricing.completion) shows "0" for
both Lyria models — misleading. The real cost is a flat per-generation
fee stated only in the model's free-text description:lyria-3-clip-preview = $0.04 per 30s clip, lyria-3-pro-preview =
$0.08 per full-length song. Confirmed by the actual usage.cost on the
test call: exactly $0.04 for the clip model. media_gen.py defaults tolyria-3-clip-preview (cheaper, matches the coordinator's "do this first").--seconds is accepted for the log/filename only — no duration parameter is
documented or was found; the clip vs. pro model choice is what actually
controls length (fixed ~30s vs. full song), not a numeric field.
Gemini-direct backend (--backend gemini) — still blocked
Both lyria-3-pro-preview and lyria-3-clip-preview return the identical429 limit:0 on Gemini-direct, via both :generateContent and the newer/v1beta/interactions endpoint (confirmed both exist — same error, not a
404). Kept in media_gen.py for whenever billing is turned on:
{"contents": [{"role": "user", "parts": [{"text": "<prompt>"}]}],
"generationConfig": {"responseModalities": ["AUDIO"]}}
OpenRouter operational notes
/v1/keybalance lags real spend by up to ~10 seconds. Checked
immediately after a call, usage can still read the pre-call value; the
response's own usage.cost is the authoritative real-time number.
_or_guard() in media_gen.py reads /v1/key right before a call
(accurate — no calls happened yet that session), never right after one.
- Budget guard:
_or_guard(est, ceiling)mirrorsvideo_gen.py's
generate() exactly — fetch real remaining balance, abort if the estimate
would eat into the $0.50 reserve (OR_RESERVE, same constant value as
video_gen.py's RESERVE) or exceed an explicit --ceiling. Both image
and music refuse to run at all if the chosen model isn't in the hardcoded
pricing dict (OR_IMAGE_PRICING / OR_MUSIC_PRICING) — no blind spend on
an unpriced model.
- Spend actually observed this session: $0.04 (music) + $0.0387 (image)
= $0.1187 for the discovery calls, + $0.04 + $0.0387 for the final
tool-verification calls logged below = $0.157412 total, out of a
$0.60 task budget and a $5 key limit. Key balance after: $4.842588
remaining (GET /v1/key, confirmed).
Credentials & logging (matches video_gen.py)
GEMINI_API_KEYandOPEN_ROUTER_API_KEY, both read from the same
prototypes/.env the same way video_gen.py reads its key — never
printed, never hardcoded. key() reads the Gemini key, or_key() reads
the OpenRouter one.
- Gemini calls (free) log to
trials/gen-log.md: model/params/output/status,
no dollar figure since there isn't one.
- Paid OpenRouter calls log to
trials/spend-log.md, the exact same
table video_gen.py writes to (| when | model | secs | res | est | actual) — all OpenRouter spend for this event lands in one
| output | prompt |
file regardless of which script made the call. For image, secs is -
and res is the aspect ratio; for music, res is - and secs is the
requested duration (unused by the API itself, logged for reference only).
- Model fallback: Gemini subcommands try a short ordered list of models,
skipping to the next on 429/503, mirroring arabic_review.py's
MODELS fallback loop. OpenRouter subcommands default to one specific,
cheap model rather than falling back across a list — there's real money on
the line per attempt, so no automatic retry-a-different-model-and-spend-
again behavior.
sys.stdout.reconfigure(encoding="utf-8")on startup, same as
video_gen.py, so Arabic prints correctly on Windows consoles.
Full command reference
python media_gen.py speech --text "<arabic or english>" [--voice Kore] [--model ...] --out path.wav
python media_gen.py image --prompt "<description>" [-n 1] [--aspect 16:9] \
[--backend openrouter|gemini] [--model ...] [--ceiling 0.10] --out path.png
python media_gen.py music --prompt "<style/mood>" [--lyrics "..."] [--seconds 30] \
[--backend openrouter|gemini] [--model ...] [--ceiling 0.10] --out path.mp3
image and music default to --backend openrouter since that's the only
one that currently works; pass --backend gemini to re-try the free path
(useful once/if billing is enabled on the Gemini project — no code changes
needed, just re-run).