Evaluation metrics for reasoning models
MP3•Episode home
Manage episode 497569197 series 3676690
Content provided by pretrained.fm, Pierce Freeman, and Richard Diehl Martinez. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by pretrained.fm, Pierce Freeman, and Richard Diehl Martinez or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://podcastplayer.com/legal.
Evaluating models on benchmarks, passing a model vibe check, formal reasoning to synthesize datasets, and what type of datasets researchers prefer
8 episodes