Topic: Training Data

3 chapters across the catalog

Episode 260: Tennessee Trickshot
1:05:22 - 1:10:52

Episode 260: Tennessee Trickshot

Spam Feed Corpus and AI-Generated Podcasts

Efforts are underway to build a training corpus of 50,000 legit and 50,000 spam feeds to improve index filtering. The hosts also examine the rise of AI-generated "deep dive" podcasts on platforms like Apple and Spotify, specifically noting suspicious AI-hosted content regarding public figures like Charlie Kirk.

Episode 258: Perceptron
50:20 - 55:21

Episode 258: Perceptron

LLM Training Data, Spot Checks and Problematic Feed Exports

Dave Jones explains that 90% of model training involves preparing a high-quality dataset. He has developed a new SQL export for the Podcast Index that identifies "problematic" or dead feeds, though it remains private due to DMCA concerns. He describes the confusion of interacting with an LLM that offers to "spot check" data without clear parameters, highlighting the gap between human intuition and machine logic.

Episode 258: Perceptron
1:11:29 - 1:16:10

Episode 258: Perceptron

Training Dataset Diversity, Spreaker Slop and Fexingo Analysis

To create an effective slop detector, the training dataset must include 25,000 "good" podcasts ranging from Joe Rogan to amateur student reports. The challenge lies in differentiating a legitimate human recording from an AI-generated history lesson on Spreaker. An analysis of the "Fexingo" podcast reveals it is a spam farm using two TTS voices to generate hundreds of episodes across 55 languages for ad revenue.