Podcast Agent — Models & TTS

Once Podcast Agent — Ingest, History & Ranking hands over a show’s chosen stories for the day, three things happen: a quick weather/time check, a two-host script, and text-to-speech narration. This page covers all three, plus which AI models and voice providers are behind them — which, unlike ingest, is mostly shared machinery with a couple of deliberate per-show differences.

A quick weather and time check

Before writing the script, every show asks two other small agents in the fleet — one for the local weather, one for the local time — and folds a one-line mention into the intro. If either check fails, that part is simply left out rather than announced as missing, the same way a human host would quietly skip a segment they couldn’t confirm.

Two hosts, not one

Every show is written as a back-and-forth between two hosts, not a single narrator — each show has its own pair of host names and voices, but the format is the same: one host reads each headline as a punchy one-liner, and the other immediately follows with a sentence or two of detail. The script always opens with a short two-host intro and the weather/time check, teases the day’s headlines, covers each story as its own short exchange in priority order, and closes with a two-host sign-off. A short pause is inserted before each story so the pacing doesn’t feel rushed.

The malayalam show works a little differently: rather than a separate ranking step followed by a script-writing step, one AI call both translates that day’s German stories into Malayalam and writes the two-host script directly, since adapting the material into natural spoken Malayalam and structuring it into a script are really the same job for that show.

The whole script runs roughly 400-900 words depending on the show — a few minutes of audio — and every detail comes only from the story’s own headline and description; nothing is invented.

Turning the script into audio

The script is sent to a text-to-speech service, with each host’s lines read in that host’s own voice. If narration fails, most shows re-narrate the entire episode from scratch with a backup voice service rather than patching in just the failed part — mixing two different voice engines mid-episode would sound worse than a short, cheap retry.

Narration is the one place the three shows genuinely diverge, each for its own reason:

Show Narration Notes

Tech

Fish Audio, primary and fallback

Both the primary and backup narration attempts use the same voice provider, so a retry never changes how the show sounds.

Stock

ElevenLabs primary, Fish Audio fallback

The one show with a genuine two-provider fallback — if ElevenLabs fails, the episode is retried against Fish Audio instead.

Malayalam

ElevenLabs only, no fallback

Deliberately has no backup provider: a real narration failure fails that day’s episode outright rather than quietly reading the news in a different voice through a fallback service.

Models behind the show

Role Model Notes

Script writing & ranking (all three shows)

gpt-5.6-luna

A small, low-cost model shared by every show for both writing the script and ranking stories, called directly against the outside provider — see LLM Gateway for how this and every other AI/TTS call in the fleet reach their providers.

See LLM Gateway for how every AI/TTS call in the fleet reaches its provider, and why this was the first app in the fleet to regularly spend real money on a paid model.