Product Metrics Are LLM Evals // Raza Habib CEO Of Humanloop // #320 MLOps.community podcast

Product Metrics are LLM Evals // Raza Habib CEO of Humanloop // #320

6d ago 53:06

Content provided by Demetrios. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by Demetrios or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://podcastplayer.com/legal.

Raza Habib, the CEO of LLM Eval platform Humanloop, talks to us about how to make your AI products more accurate and reliable by shortening the feedback loop of your evals. Quickly iterating on prompts and testing what works, along with some of his favorite Dario from Anthropic AI Quotes.

// Bio

Raza is the CEO and Co-founder at Humanloop. He has a PhD in Machine Learning from UCL, was the founding engineer of Monolith AI, and has built speech systems at Google. For the last 4 years, he has led Humanloop and supported leading technology companies such as Duolingo, Vanta, and Gusto to build products with large language models. Raza was featured in the Forbes 30 Under 30 technology list in 2022, and Sifted recently named him one of the most influential Gen AI founders in Europe.

// Related Links

Websites: https://humanloop.com

~~~~~~~~ ✌️Connect With Us ✌️ ~~~~~~~

Catch all episodes, blogs, newsletters, and more: https://go.mlops.community/TYExplore

MLOps Swag/Merch: [https://shop.mlops.community/]

Connect with Demetrios on LinkedIn: /dpbrinkm

Connect with Raza on LinkedIn: /humanloop-raza

Timestamps:

[00:00] Cracking Open System Failures and How We Fix Them

[05:44] LLMs in the Wild — First Steps and Growing Pains

[08:28] Building the Backbone of Tracing and Observability

[13:02] Tuning the Dials for Peak Model Performance

[13:51] From Growing Pains to Glowing Gains in AI Systems

[17:26] Where Prompts Meet Psychology and Code

[22:40] Why Data Experts Deserve a Seat at the Table

[24:59] Humanloop and the Art of Configuration Taming

[28:23] What Actually Matters in Customer-Facing AI

[33:43] Starting Fresh with Private Models That Deliver

[34:58] How LLM Agents Are Changing the Way We Talk

[39:23] The Secret Lives of Prompts Inside Frameworks

[42:58] Streaming Showdowns — Creativity vs. Convenience

[46:26] Meet Our Auto-Tuning AI Prototype

[49:25] Building the Blueprint for Smarter AI

[51:24] Feedback Isn’t Optional — It’s Everything

441 episodes