Podcast Agent — Ingest, History & Ranking

This page covers everything upstream of script writing: how each show collects stories, how duplicates are filtered out, and how a daily episode’s headlines get chosen. See Podcast Agent — Models & TTS for what happens to the chosen stories afterward.

Collect continuously, publish once a day — independently per show

Diagram

Each show has its own sources

The three shows never share a source list — each pulls from the handful of feeds that actually cover its own beat:

  • Tech — a keyword news search targeting AI/LLM industry topics, plus the official news feeds of OpenAI, Anthropic (via a third-party mirror, since Anthropic doesn’t publish its own feed), and Microsoft’s Azure AI Foundry blog. The only one of the three that uses a keyword search API at all.

  • Stock — three financial-news vendor feeds (CNBC, Yahoo Finance, MarketWatch), RSS only.

  • Malayalam — three German national news vendor feeds (Tagesschau, ZEIT ONLINE, SPIEGEL), RSS only. Source material stays in German all the way through collection; translation into Malayalam happens later, during script writing.

Within a show, each source is checked independently, so one being slow, down, or reshaped only means fewer stories from that one source that hour — it never blocks the others or breaks the collection run.

This collection step is intentionally lightweight for every show: no AI model is involved, nothing is written or narrated, it just gathers and stores. That’s what makes it cheap enough to run continuously throughout the day rather than once.

Never the same story twice

Every incoming story is checked against everything already collected for that same show before it’s kept — matching on its link, its wording, and near-duplicate titles across sources (the same story often gets picked up by more than one outlet with a slightly different headline). Anything that’s already been seen, or already been used in a past episode, is skipped. Stories also age out after a couple of days — a daily show shouldn’t cover stale news just because it’s new to the collection.

Tech and stock also drop anything that isn’t genuinely English content at this stage — a lightweight check, not a translator. Malayalam deliberately skips that check, since its whole source material is German by design and gets carried through to translation rather than filtered out.

Choosing what makes the cut

Once a day, per show, everything collected and not yet used gets ranked by one pass through an AI model. Tech’s ranking uses a fixed priority order — new model releases from any vendor rank highest, followed by Azure AI Foundry news, then Anthropic, then OpenAI, then everything else genuinely newsworthy; stock and malayalam rank by newsworthiness without that vendor-specific ordering. The top handful of stories, in that order, go on to become the episode’s script.

If ranking itself fails for any reason, the show doesn’t stop — it just falls back to the most recent stories instead of the best-ranked ones. And if there’s simply nothing new to cover, that show’s episode for the day is skipped outright rather than publishing something thin.

Each show keeps its own separate story history — a stock story is never a candidate for the tech episode, or vice versa.

Re-running a day’s episode

An episode can be regenerated on demand, per show — useful when iterating on the voice or script. Regenerating releases that day’s stories back into the pool rather than treating them as permanently used, so repeated re-runs don’t quietly burn through the story backlog.