← العودة إلى العروض

الأدوات وتكاليفها

media_gen.py — Gemini + OpenRouter media generation

Companion to video_gen.py (which runs video on OpenRouter, budget-capped).
media_gen.py generates speech, images, and music. Speech runs on the
Gemini API free tier (GEMINI_API_KEY) — free, no spend guard needed.
Images and music default to OpenRouter (OPEN_ROUTER_API_KEY) — the same
$5-capped key video_gen.py uses, real dollars per call, guarded by the same
estimate → check real balance → refuse pattern. Both credentials come from
prototypes/.env.

Verified against the live Gemini and OpenRouter APIs on 2026-08-11. Results
below are real — see trials/speech/, trials/images/, trials/music/,
trials/gen-log.md (every Gemini call), and trials/spend-log.md (every
paid OpenRouter call, same table video_gen.py writes to).

Bottom line

ModalityDefault backendStatusReal cost
Speech (Arabic TTS)Gemini (free)WORKS$0
ImageOpenRouterWORKS~$0.04/image (gemini-2.5-flash-image)
Music (Lyria)OpenRouterWORKS$0.04/30s clip, $0.08/song
Image, MusicGemini (--backend gemini)BLOCKEDfree tier limit: 0 — see bottom

Gemini's free tier allocates zero daily requests to every image model and
both Lyria models on this key (confirmed empirically, not a shape bug — see
the bottom of this doc). OpenRouter routes the exact same underlying models
through Google AI Studio as a paid provider and it just works. All three
subcommands are proven end-to-end below with real generated files and real
dollar amounts.


Working: speech (Gemini, free)

python media_gen.py speech --text "الذكاء الاصطناعي وفرص العمل وبناء مصادر الدخل" \
    --voice Kore --out ../trials/speech/kore-flash25.wav

Request shape (discovered empirically)

POST /v1beta/models/{model}:generateContent, not :predict and not
the newer /v1beta/interactions endpoint — plain generateContent works.

{
  "contents": [{"role": "user", "parts": [{"text": "..."}]}],
  "generationConfig": {
    "responseModalities": ["AUDIO"],
    "speechConfig": {
      "voiceConfig": {"prebuiltVoiceConfig": {"voiceName": "Kore"}}
    }
  }
}

Response shape

{"candidates": [{"content": {"parts": [{"inlineData": {
  "mimeType": "audio/L16;codec=pcm;rate=24000", "data": "<base64 PCM>"
}}]}}]}

Audio comes back as raw 16-bit little-endian PCM, mono, 24kHz — not a
WAV file, not MP3, no container at all. media_gen.py decodes the base64
and wraps it in a WAV header via Python's wave module
(_save_audio() in media_gen.py). The two TTS models format mimeType
slightly differently (audio/L16;codec=pcm;rate=24000 vs.
audio/l16; rate=24000; channels=1) — the sample-rate regex handles both.

Models tried, both work

Both succeeded on the first try, no 429s during testing. --model overrides
the fallback chain if you want to pin one.

Arabic verification

Generated the exact sentence from the brief, الذكاء الاصطناعي وفرص العمل وبناء
مصادر الدخل
, with three voice/model combos, saved under trials/speech/:

filemodelvoicedurationpeak amplitudesilence ratio
kore-flash25.wavgemini-2.5-flash-preview-ttsKore4.85s28331/3276725.9%
kore-flash31.wavgemini-3.1-flash-tts-previewKore5.52s26526/3276729.8%
puck-flash25.wavgemini-2.5-flash-preview-ttsPuck4.81s24744/3276723.8%

All three: valid RIFF...WAVEfmt header, mono 16-bit PCM @ 24kHz, strong
peak amplitude (not clipped, not silent), and a silence ratio consistent
with natural pauses between words in a ~40-character sentence spoken at a
normal pace (~8 characters/sec). This is strong signal-level evidence the
models produced real speech, not noise or silence.

Caveat — I cannot judge pronunciation quality by ear. I'm a text model;
I verified the audio is valid, non-trivial, and the right duration/shape for
speech, but I did not and cannot listen to confirm the Arabic actually sounds
natural, correctly stressed, or free of mispronunciation on
"الذكاء الاصطناعي" / "مصادر الدخل" specifically. Have a person listen to
trials/speech/kore-flash25.wav before relying on this for the workshop.

Gemini's TTS models are documented to support Arabic and auto-detect the
input language (there's no separate "Arabic voice" — all 30 prebuilt voices
are multilingual), so quality risk is about accent/naturalness, not
"garbage output," but that's still a real risk worth a human check.

Recommendation if you need a pick without listening first: Kore via
gemini-2.5-flash-preview-tts — it was the first one tried, has the lowest
silence ratio of the three, and matched duration expectations most closely.


Working: image (OpenRouter, ~$0.04/image)

python media_gen.py image --prompt "flat vector icon of a lightbulb, cyan and navy blue color palette, minimal, on white background" \
    --aspect 16:9 --ceiling 0.10 --out ../trials/images/lightbulb-icon.png

Real output: trials/images/lightbulb-icon.png — valid PNG (\x89PNG\r\n\x1a\n
header), 1344×768 (confirmed via ffprobe), 262,968 bytes. Cost: $0.0387,
confirmed via the response's own usage.cost field.

Discovery process

The coordinator's tip mattered: GET /v1/images/models (200) lists models
but no generation endpoint. I found the real one by sending an empty POST
body
to candidate paths and reading the Zod validation error, which is free
(it 400s before billing):

POST /v1/images/generations  {}
POST /v1/images              {}

Both return the identical 400:

{"success":false,"error":{"name":"ZodError","message":"[
  {\"path\":[\"model\"],\"message\":\"Invalid input: expected string, received undefined\"},
  {\"path\":[\"prompt\"],\"message\":\"Invalid input: expected string, received undefined\"}
]"}}

Both routes exist and accept the same body — media_gen.py uses /v1/images
(the name OpenRouter's own docs use). Sending {"n": 0} against the real
model confirmed n's valid range (>=1) the same free way — a second 400,
still no spend.

Request shape (confirmed working)

{
  "model": "google/gemini-2.5-flash-image",
  "prompt": "<prompt>",
  "aspect_ratio": "16:9",
  "n": 1
}

Response shape (confirmed working)

{
  "data": [{"b64_json": "<base64 PNG>", "media_type": "image/png"}],
  "usage": {"prompt_tokens": 20, "completion_tokens": 1290, "total_tokens": 1310,
            "cost": 0.038706}
}

usage.cost is the real, authoritative per-call cost — media_gen.py logs
this exact number to trials/spend-log.md as actual, next to the
pre-flight estimate.

Model choice and pricing

google/gemini-2.5-flash-image ("Nano Banana") spends a fixed ~1290
output-image tokens per call
(confirmed: 1290 exactly on this test, at
16:9) at $0.00003/token = $0.0387/image, matching the coordinator's
estimate almost exactly. media_gen.py hardcodes $0.04 as the pre-flight
estimate (rounds up slightly, safe). google/gemini-3-pro-image was never
called
— the coordinator flagged it as ~4x the price (~$0.15/image) and
not worth it for this volume; it isn't in OR_IMAGE_PRICING at all, so
passing --model google/gemini-3-pro-image fails closed with "no pricing
known" rather than spending blind.

n is capped at 1 server-side for this model (confirmed via
/v1/images/models: "n": {"min": 1, "max": 1}), so -n 3 in
media_gen.py means 3 separate calls, each individually estimated and
guarded — not a batch request.

Gemini-direct backend (--backend gemini) — still blocked

Every Gemini image model — gemini-3.1-flash-image, gemini-3-pro-image,
nano-banana-pro-preview, gemini-2.5-flash-image, gemini-3.1-flash-lite-image
— returns the same error on the free tier, on the very first request:

HTTP 429
* Quota exceeded for metric: generate_content_free_tier_requests, limit: 0, model: gemini-3-pro-image
"status": "RESOURCE_EXHAUSTED"

limit: 0 means the free tier is allocated zero requests for that model,
period — not "you used today's quota." This is a billing/tier gate, not a
shape bug: it fires before the request body is even validated.
Confirmed alias: nano-banana-pro-preview and gemini-3-pro-image hit
the identical quota bucket (model: gemini-3-pro-image in both errors) — same
model, two names. The Gemini-side request shape media_gen.py sends (kept
working for whenever billing is turned on):

{
  "contents": [{"role": "user", "parts": [{"text": "<prompt>"}]}],
  "generationConfig": {"responseModalities": ["IMAGE"], "imageConfig": {"aspectRatio": "16:9"}}
}

Google's docs (ai.google.dev/gemini-api/docs/image-generation) describe a
newer /v1beta/interactions endpoint for these models instead — confirmed to
exist (same 429 limit:0, not a 404) but never validated against a real
success response. If generateContent ever 400s once billing is on, try that
shape next ({"model", "input":[{"type":"text","text":...}], "response_format":
{"type":"image","aspect_ratio":...}}
, response under steps[].content[].data).


Working: music (OpenRouter, $0.04/clip or $0.08/song)

python media_gen.py music --prompt "upbeat inspiring instrumental synth pop, hopeful, 120bpm, 30 seconds" \
    --seconds 30 --ceiling 0.10 --out ../trials/music/intro-clip.mp3

Real output: trials/music/intro-clip.mp3 — valid MP3 (ID3 header,
confirmed via ffprobe: codec mp3, 44.1kHz stereo, 192kbps, 27.98s),
677,526 bytes. Cost: $0.0400, confirmed via usage.cost.

Discovery process

GET /v1/audio/models and GET /v1/music/models both true 404 (HTML
fallback page, not a JSON error) — no dedicated model-listing surface for
audio. POST /v1/audio/generations and POST /v1/music/generations
also true 404 — no dedicated generation endpoint either. Music only
exists through /v1/chat/completions.

First attempt, non-streaming, cost nothing because it was rejected before
generation:

{"model": "google/lyria-3-clip-preview",
 "messages": [{"role": "user", "content": "<prompt>"}],
 "modalities": ["text", "audio"]}
{"error": {"message": "Audio output requires stream: true", "code": 400}}

That's the whole discovery cost for this path — one free 400 that states the
exact missing field. Setting "stream": true and parsing the SSE response
worked on the next try (see below).

Request shape (confirmed working)

{
  "model": "google/lyria-3-clip-preview",
  "messages": [{"role": "user", "content": "<prompt, optionally + \n\nLyrics:\n<lyrics>"}],
  "modalities": ["text", "audio"],
  "stream": true
}

--lyrics is folded into the prompt text (no separate structured field was
found or needed — Lyria took a plain style/mood prompt fine for an
instrumental test).

Response shape (confirmed working) — SSE, one audio chunk

Content-Type: text/event-stream, lines of data: {...}\n\n, terminated by
data: [DONE]. Unlike a typical token-by-token audio stream, the entire clip
arrived as one chunk in the final event before [DONE]:

{"choices": [{"delta": {"audio": {"data": "<base64 MP3, whole clip>"}}}],
 "usage": {"prompt_tokens": 17, "completion_tokens": 4, "total_tokens": 21,
           "cost": 0.04}}

media_gen.py concatenates every delta.audio.data seen (correct even if a
future model does chunk it) and decodes once at the end. The decoded bytes
start with ID3 — a real MP3 container, not raw PCM, so no WAV-wrapping is
needed (unlike the Gemini TTS path). usage.cost on the final event is the
real charge, logged verbatim to spend-log.md.

Model choice and pricing

Model card pricing (pricing.prompt/pricing.completion) shows "0" for
both Lyria models — misleading. The real cost is a flat per-generation
fee stated only in the model's free-text description:
lyria-3-clip-preview = $0.04 per 30s clip, lyria-3-pro-preview =
$0.08 per full-length song. Confirmed by the actual usage.cost on the
test call: exactly $0.04 for the clip model. media_gen.py defaults to
lyria-3-clip-preview (cheaper, matches the coordinator's "do this first").
--seconds is accepted for the log/filename only — no duration parameter is
documented or was found; the clip vs. pro model choice is what actually
controls length (fixed ~30s vs. full song), not a numeric field.

Gemini-direct backend (--backend gemini) — still blocked

Both lyria-3-pro-preview and lyria-3-clip-preview return the identical
429 limit:0 on Gemini-direct, via both :generateContent and the newer
/v1beta/interactions endpoint (confirmed both exist — same error, not a
404). Kept in media_gen.py for whenever billing is turned on:

{"contents": [{"role": "user", "parts": [{"text": "<prompt>"}]}],
 "generationConfig": {"responseModalities": ["AUDIO"]}}

OpenRouter operational notes

immediately after a call, usage can still read the pre-call value; the
response's own usage.cost is the authoritative real-time number.
_or_guard() in media_gen.py reads /v1/key right before a call
(accurate — no calls happened yet that session), never right after one.

generate() exactly — fetch real remaining balance, abort if the estimate
would eat into the $0.50 reserve (OR_RESERVE, same constant value as
video_gen.py's RESERVE) or exceed an explicit --ceiling. Both image
and music refuse to run at all if the chosen model isn't in the hardcoded
pricing dict (OR_IMAGE_PRICING / OR_MUSIC_PRICING) — no blind spend on
an unpriced model.

= $0.1187 for the discovery calls, + $0.04 + $0.0387 for the final
tool-verification calls logged below = $0.157412 total, out of a
$0.60 task budget and a $5 key limit. Key balance after: $4.842588
remaining
(GET /v1/key, confirmed).


Credentials & logging (matches video_gen.py)

prototypes/.env the same way video_gen.py reads its key — never
printed, never hardcoded. key() reads the Gemini key, or_key() reads
the OpenRouter one.

no dollar figure since there isn't one.

table video_gen.py writes to (| when | model | secs | res | est | actual
| output | prompt |
) — all OpenRouter spend for this event lands in one
file regardless of which script made the call. For image, secs is -
and res is the aspect ratio; for music, res is - and secs is the
requested duration (unused by the API itself, logged for reference only).

skipping to the next on 429/503, mirroring arabic_review.py's
MODELS fallback loop. OpenRouter subcommands default to one specific,
cheap model rather than falling back across a list — there's real money on
the line per attempt, so no automatic retry-a-different-model-and-spend-
again behavior.

video_gen.py, so Arabic prints correctly on Windows consoles.

Full command reference

python media_gen.py speech --text "<arabic or english>" [--voice Kore] [--model ...] --out path.wav

python media_gen.py image --prompt "<description>" [-n 1] [--aspect 16:9] \
    [--backend openrouter|gemini] [--model ...] [--ceiling 0.10] --out path.png

python media_gen.py music --prompt "<style/mood>" [--lyrics "..."] [--seconds 30] \
    [--backend openrouter|gemini] [--model ...] [--ceiling 0.10] --out path.mp3

image and music default to --backend openrouter since that's the only
one that currently works; pass --backend gemini to re-try the free path
(useful once/if billing is enabled on the Gemini project — no code changes
needed, just re-run).