Topic: Model Training

3 chapters across the catalog

Episode 260: Tennessee Trickshot
1:02:31 - 1:05:20

Episode 260: Tennessee Trickshot

Appless Computing and Custom Model Training

The vision of "appless" computing is discussed, where users "vibe code" custom tools for specific tasks rather than using general applications. The DGX Spark server is currently running inference for a 35-billion parameter model and Whisper Turbo, with the ultimate goal of fine-tuning a custom model on the entire Podcast Index database.

Episode 259: SlopJacked
46:05 - 48:40

Episode 259: SlopJacked

Hardware Donation Request, Local AI Model Training

A plea is made for a hardware donation, specifically a high-RAM Mac Mini or a DGX Spark machine costing approximately $3,500, to facilitate local AI model training. The current 16GB Mac Mini M4 is insufficient for the processing required to filter podcast data. Owning the hardware is presented as a more cost-effective long-term solution than paying for expensive cloud-based AI APIs.

Episode 258: Perceptron
50:20 - 55:21

Episode 258: Perceptron

LLM Training Data, Spot Checks and Problematic Feed Exports

Dave Jones explains that 90% of model training involves preparing a high-quality dataset. He has developed a new SQL export for the Podcast Index that identifies "problematic" or dead feeds, though it remains private due to DMCA concerns. He describes the confusion of interacting with an LLM that offers to "spot check" data without clear parameters, highlighting the gap between human intuition and machine logic.