This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.
…
continue reading
Tobias Macey Podcasts
The podcast about Python and the people who make it great
…
continue reading

1
From RAG to Relational: How Agentic Patterns Are Reshaping Data Architecture
52:58
52:58
Play later
Play later
Lists
Like
Liked
52:58Summary In this episode of the AI Engineering Podcast Mark Brooker, VP and Distinguished Engineer at AWS, talks about how agentic workflows are transforming database usage and infrastructure design. He discusses the evolving role of data in AI systems, from traditional models to more modern approaches like vectors, RAG, and relational databases. Ma…
…
continue reading

1
Duck Lake: Simplifying the Lakehouse Ecosystem
1:10:41
1:10:41
Play later
Play later
Lists
Like
Liked
1:10:41Summary In this episode of the Data Engineering Podcast Hannes Mühleisen and Mark Raasveldt, the creators of DuckDB, share their work on Duck Lake, a new entrant in the open lakehouse ecosystem. They discuss how Duck Lake, is focused on simplicity, flexibility, and offers a unified catalog and table format compared to other lakehouse formats like I…
…
continue reading

1
Aligning Business and Data: The Essential Role of Data Modeling
1:06:51
1:06:51
Play later
Play later
Lists
Like
Liked
1:06:51Summary In this episode of the Data Engineering Podcast Serge Gershkovich, head of product at SQL DBM, talks about the socio-technical aspects of data modeling. Serge shares his background in data modeling and highlights its importance as a collaborative process between business stakeholders and data teams. He debunks common misconceptions that dat…
…
continue reading

1
From Academia to Industry: Bridging Data Engineering Challenges
50:54
50:54
Play later
Play later
Lists
Like
Liked
50:54Summary In this episode of the Data Engineering Podcast Professor Paul Groth, from the University of Amsterdam, talks about his research on knowledge graphs and data engineering. Paul shares his background in AI and data management, discussing the evolution of data provenance and lineage, as well as the challenges of data integration. He explores t…
…
continue reading

1
High Performance And Low Overhead Graphs With KuzuDB
1:01:29
1:01:29
Play later
Play later
Lists
Like
Liked
1:01:29Summary In this episode of the Data Engineering Podcast Prashanth Rao, an AI engineer at KuzuDB, talks about their embeddable graph database. Prashanth explains how KuzuDB addresses performance shortcomings in existing solutions through columnar storage and novel join algorithms. He discusses the usability and scalability of KuzuDB, emphasizing its…
…
continue reading

1
Bridging Data and Decision-Making: AI's Role in Modern Analytics
1:10:44
1:10:44
Play later
Play later
Lists
Like
Liked
1:10:44Summary In this episode of the Data Engineering Podcast Lucas Thelosen and Drew Gilson from Gravity talk about their development of Orion, an autonomous data analyst that bridges the gap between data availability and business decision-making. Lucas and Drew share their backgrounds in data analytics and how their experiences have shaped their approa…
…
continue reading

1
From Bits to Tables: The Evolution of S3 Storage
50:08
50:08
Play later
Play later
Lists
Like
Liked
50:08Summary In this episode of the Data Engineering Podcast Andy Warfield talks about the innovative functionalities of S3 Tables and Vectors and their integration into modern data stacks. Andy shares his journey through the tech industry and his role at Amazon, where he collaborates to enhance storage capabilities, discussing the evolution of S3 from …
…
continue reading

1
Revolutionizing Python Notebooks with Marimo
51:56
51:56
Play later
Play later
Lists
Like
Liked
51:56Summary In this episode of the Data Engineering Podcast Akshay Agrawal from Marimo discusses the innovative new Python notebook environment, which offers a reactive execution model, full Python integration, and built-in UI elements to enhance the interactive computing experience. He discusses the challenges of traditional Jupyter notebooks, such as…
…
continue reading

1
Warehouse Native Incremental Data Processing With Dynamic Tables And Delayed View Semantics
55:07
55:07
Play later
Play later
Lists
Like
Liked
55:07Summary In this episode of the Data Engineering Podcast Dan Sotolongo from Snowflake talks about the complexities of incremental data processing in warehouse environments. Dan discusses the challenges of handling continuously evolving datasets and the importance of incremental data processing for optimized resource use and reduced latency. He expla…
…
continue reading

1
Streamlining Data Pipelines with MCP Servers and Vector Engines
52:04
52:04
Play later
Play later
Lists
Like
Liked
52:04Summary In this episode of the Data Engineering Podcast Kacper Łukawski from Qdrant about integrating MCP servers with vector databases to process unstructured data. Kacper shares his experience in data engineering, from building big data pipelines in the automotive industry to leveraging large language models (LLMs) for transforming unstructured d…
…
continue reading

1
Foundational Data Engineering At Two Sigma
55:05
55:05
Play later
Play later
Lists
Like
Liked
55:05Summary In this episode of the Data Engineering Podcast Effie Baram, a leader in foundational data engineering at Two Sigma, talks about the complexities and innovations in data engineering within the finance sector. She discusses the critical role of data at Two Sigma, balancing data quality with delivery speed, and the socio-technical challenges …
…
continue reading

1
Enabling Agents In The Enterprise With A Platform Approach
54:18
54:18
Play later
Play later
Lists
Like
Liked
54:18Summary In this episode of the Data Engineering Podcast Arun Joseph talks about developing and implementing agent platforms to empower businesses with agentic capabilities. From leading AI engineering at Deutsche Telekom to his current entrepreneurial venture focused on multi-agent systems, Arun shares insights on building agentic systems at an org…
…
continue reading

1
Dagster's New Era: Modularizing Data Transformation in the Age of AI
1:01:37
1:01:37
Play later
Play later
Lists
Like
Liked
1:01:37Summary In this episode of the Data Engineering Podcast we welcome back Nick Schrock, CTO and founder of Dagster Labs, to discuss the evolving landscape of data engineering in the age of AI. As AI begins to impact data platforms and the role of data engineers, Nick shares his insights on how it will ultimately enhance productivity and expand softwa…
…
continue reading

1
AI and the Lakehouse: How Starburst is Pioneering New Workflows
44:09
44:09
Play later
Play later
Lists
Like
Liked
44:09Summary In this episode of the Data Engineering Podcast Alex Albu, tech lead for AI initiatives at Starburst, talks about integrating AI workloads with the lakehouse architecture. From his software engineering roots to leading data engineering efforts, Alex shares insights on enhancing Starburst's platform to support AI applications, including an A…
…
continue reading

1
Amazon S3: The Backbone of Modern Data Systems
1:01:01
1:01:01
Play later
Play later
Lists
Like
Liked
1:01:01Summary In this episode of the Data Engineering Podcast Mai-Lan Tomsen Bukovec, Vice President of Technology at AWS, talks about the evolution of Amazon S3 and its profound impact on data architecture. From her work on compute systems to leading the development and operations of S3, Mylan shares insights on how S3 has become a foundational element …
…
continue reading

1
Scaling Data Operations With Platform Engineering
42:20
42:20
Play later
Play later
Lists
Like
Liked
42:20Summary In this episode of the Data Engineering Podcast Chakravarthy Kotaru talks about scaling data operations through standardized platform offerings. From his roots as an Oracle developer to leading the data platform at a major online travel company, Chakravarthy shares insights on managing diverse database technologies and providing databases a…
…
continue reading

1
From Data Discovery to AI: The Evolution of Semantic Layers
49:30
49:30
Play later
Play later
Lists
Like
Liked
49:30Summary In this episode of the Data Engineering Podcast, host Tobias Macy welcomes back Shinji Kim to discuss the evolving role of semantic layers in the era of AI. As they explore the challenges of managing vast data ecosystems and providing context to data users, they delve into the significance of semantic layers for AI applications. They dive i…
…
continue reading

1
Balancing Off-the-Shelf and Custom Solutions in Data Engineering
46:05
46:05
Play later
Play later
Lists
Like
Liked
46:05Summary In this episode of the Data Engineering Podcast Tulika Bhatt, a senior software engineer at Netflix, talks about her experiences with large-scale data processing and the future of data engineering technologies. Tulika shares her journey into the data engineering field, discussing her work at BlackRock and Verizon before joining Netflix, and…
…
continue reading

1
StarRocks: Bridging Lakehouse and OLAP for High-Performance Analytics
59:41
59:41
Play later
Play later
Lists
Like
Liked
59:41Summary In this episode of the Data Engineering Podcast Sida Shen, product manager at CelerData, talks about StarRocks, a high-performance analytical database. Sida discusses the inception of StarRocks, which was forked from Apache Doris in 2020 and evolved into a high-performance Lakehouse query engine. He explains the architectural design of Star…
…
continue reading

1
Exploring NATS: A Multi-Paradigm Connectivity Layer for Distributed Applications
1:12:50
1:12:50
Play later
Play later
Lists
Like
Liked
1:12:50Summary In this episode of the Data Engineering Podcast Derek Collison, creator of NATS and CEO of Synadia, talks about the evolution and capabilities of NATS as a multi-paradigm connectivity layer for distributed applications. Derek discusses the challenges and solutions in building distributed systems, and highlights the unique features of NATS t…
…
continue reading

1
Advanced Lakehouse Management With The LakeKeeper Iceberg REST Catalog
57:13
57:13
Play later
Play later
Lists
Like
Liked
57:13Summary In this episode of the Data Engineering Podcast Viktor Kessler, co-founder of Vakmo, talks about the architectural patterns in the lake house enabled by a fast and feature-rich Iceberg catalog. Viktor shares his journey from data warehouses to developing the open-source project, Lakekeeper, an Apache Iceberg REST catalog written in Rust tha…
…
continue reading

1
Simplifying Data Pipelines with Durable Execution
39:49
39:49
Play later
Play later
Lists
Like
Liked
39:49Summary In this episode of the Data Engineering Podcast Jeremy Edberg, CEO of DBOS, about durable execution and its impact on designing and implementing business logic for data systems. Jeremy explains how DBOS's serverless platform and orchestrator provide local resilience and reduce operational overhead, ensuring exactly-once execution in distrib…
…
continue reading

1
Overcoming Redis Limitations: The Dragonfly DB Approach
43:58
43:58
Play later
Play later
Lists
Like
Liked
43:58Summary In this episode of the Data Engineering Podcast Roman Gershman, CTO and founder of Dragonfly DB, explores the development and impact of high-speed in-memory databases. Roman shares his experience creating a more efficient alternative to Redis, focusing on performance gains, scalability, and cost efficiency, while addressing limitations such…
…
continue reading

1
Bringing AI Into The Inner Loop of Data Engineering With Ascend
52:47
52:47
Play later
Play later
Lists
Like
Liked
52:47Summary In this episode of the Data Engineering Podcast Sean Knapp, CEO of Ascend.io, explores the intersection of AI and data engineering. He discusses the evolution of data engineering and the role of AI in automating processes, alleviating burdens on data engineers, and enabling them to focus on complex tasks and innovation. The conversation cov…
…
continue reading

1
Astronomer's Role in the Airflow Ecosystem: A Deep Dive with Pete DeJoy
51:41
51:41
Play later
Play later
Lists
Like
Liked
51:41Summary In this episode of the Data Engineering Podcast Pete DeJoy, co-founder and product lead at Astronomer, talks about building and managing Airflow pipelines on Astronomer and the upcoming improvements in Airflow 3. Pete shares his journey into data engineering, discusses Astronomer's contributions to the Airflow project, and highlights the cr…
…
continue reading

1
Accelerated Computing in Modern Data Centers With Datapelago
55:36
55:36
Play later
Play later
Lists
Like
Liked
55:36Summary In this episode of the Data Engineering Podcast Rajan Goyal, CEO and co-founder of Datapelago, talks about improving efficiencies in data processing by reimagining system architecture. Rajan explains the shift from hyperconverged to disaggregated and composable infrastructure, highlighting the importance of accelerated computing in modern d…
…
continue reading

1
The Future of Data Engineering: AI, LLMs, and Automation
59:39
59:39
Play later
Play later
Lists
Like
Liked
59:39Summary In this episode of the Data Engineering Podcast Gleb Mezhanskiy, CEO and co-founder of DataFold, talks about the intersection of AI and data engineering. He discusses the challenges and opportunities of integrating AI into data engineering, particularly using large language models (LLMs) to enhance productivity and reduce manual toil. The c…
…
continue reading

1
Evolving Responsibilities in AI Data Management
38:57
38:57
Play later
Play later
Lists
Like
Liked
38:57Summary In this episode of the Data Engineering Podcast Bartosz Mikulski talks about preparing data for AI applications. Bartosz shares his journey from data engineering to MLOps and emphasizes the importance of data testing over software development in AI contexts. He discusses the types of data assets required for AI applications, including exten…
…
continue reading

1
CSVs Will Never Die And OneSchema Is Counting On It
54:40
54:40
Play later
Play later
Lists
Like
Liked
54:40Summary In this episode of the Data Engineering Podcast Andrew Luo, CEO of OneSchema, talks about handling CSV data in business operations. Andrew shares his background in data engineering and CRM migration, which led to the creation of OneSchema, a platform designed to automate CSV imports and improve data validation processes. He discusses the ch…
…
continue reading

1
Breaking Down Data Silos: AI and ML in Master Data Management
57:30
57:30
Play later
Play later
Lists
Like
Liked
57:30Summary In this episode of the Data Engineering Podcast Dan Bruckner, co-founder and CTO of Tamr, talks about the application of machine learning (ML) and artificial intelligence (AI) in master data management (MDM). Dan shares his journey from working at CERN to becoming a data expert and discusses the challenges of reconciling large-scale organiz…
…
continue reading

1
Building a Data Vision Board: A Guide to Strategic Planning
49:59
49:59
Play later
Play later
Lists
Like
Liked
49:59Summary In this episode of the Data Engineering Podcast Lior Barak shares his insights on developing a three-year strategic vision for data management. He discusses the importance of having a strategic plan for data, highlighting the need for data teams to focus on impact rather than just enablement. He introduces the concept of a "data vision boar…
…
continue reading

1
How Orchestration Impacts Data Platform Architecture
59:39
59:39
Play later
Play later
Lists
Like
Liked
59:39Summary The core task of data engineering is managing the flows of data through an organization. In order to ensure those flows are executing on schedule and without error is the role of the data orchestrator. Which orchestration engine you choose impacts the ways that you architect the rest of your data platform. In this episode Hugo Lu shares his…
…
continue reading

1
An Exploration Of The Impediments To Reusable Data Pipelines
51:32
51:32
Play later
Play later
Lists
Like
Liked
51:32Summary In this episode of the Data Engineering Podcast the inimitable Max Beauchemin talks about reusability in data pipelines. The conversation explores the "write everything twice" problem, where similar pipelines are built without code reuse, and discusses the challenges of managing different SQL dialects and relational databases. Max also touc…
…
continue reading

1
The Art of Database Selection and Evolution
59:56
59:56
Play later
Play later
Lists
Like
Liked
59:56Summary In this episode of the Data Engineering Podcast Sam Kleinman talks about the pivotal role of databases in software engineering. Sam shares his journey into the world of data and discusses the complexities of database selection, highlighting the trade-offs between different database architectures and how these choices affect system design, q…
…
continue reading

1
Bridging Code and UI in Data Orchestration with Kestra
44:30
44:30
Play later
Play later
Lists
Like
Liked
44:30Summary In this episode of the Data Engineering Podcast, Anna Geller talks about the integration of code and UI-driven interfaces for data orchestration. Anna defines data orchestration as automating the coordination of workflow nodes that interact with data across various business functions, discussing how it goes beyond ETL and analytics to enabl…
…
continue reading

1
Streaming Data Into The Lakehouse With Iceberg And Trino At Going
39:49
39:49
Play later
Play later
Lists
Like
Liked
39:49In this episode, I had the pleasure of speaking with Ken Pickering, VP of Engineering at Going, about the intricacies of streaming data into a Trino and Iceberg lakehouse. Ken shared his journey from product engineering to becoming deeply involved in data-centric roles, highlighting his experiences in ecommerce and InsurTech. At Going, Ken leads th…
…
continue reading

1
An Opinionated Look At End-to-end Code Only Analytical Workflows With Bruin
56:11
56:11
Play later
Play later
Lists
Like
Liked
56:11Summary The challenges of integrating all of the tools in the modern data stack has led to a new generation of tools that focus on a fully integrated workflow. At the same time, there have been many approaches to how much of the workflow is driven by code vs. not. Burak Karakan is of the opinion that a fully integrated workflow that is driven entir…
…
continue reading

1
Feldera: Bridging Batch and Streaming with Incremental Computation
47:36
47:36
Play later
Play later
Lists
Like
Liked
47:36Summary In this episode of the Data Engineering Podcast, the creators of Feldera talk about their incremental compute engine designed for continuous computation of data, machine learning, and AI workloads. The discussion covers the concept of incremental computation, the origins of Feldera, and its unique ability to handle both streaming and batch …
…
continue reading

1
Accelerate Migration Of Your Data Warehouse with Datafold's AI Powered Migration Agent
48:50
48:50
Play later
Play later
Lists
Like
Liked
48:50Summary Gleb Mezhanskiy, CEO and co-founder of DataFold, joins Tobias Macey to discuss the challenges and innovations in data migrations. Gleb shares his experiences building and scaling data platforms at companies like Autodesk and Lyft, and how these experiences inspired the creation of DataFold to address data quality issues across teams. He out…
…
continue reading

1
Bring Vector Search And Storage To The Data Lake With Lance
58:01
58:01
Play later
Play later
Lists
Like
Liked
58:01Summary The rapid growth of generative AI applications has prompted a surge of investment in vector databases. While there are numerous engines available now, Lance is designed to integrate with data lake and lakehouse architectures. In this episode Weston Pace explains the inner workings of the Lance format for table definitions and file storage, …
…
continue reading

1
The Role of Python in Shaping the Future of Data Platforms with DLT
54:08
54:08
Play later
Play later
Lists
Like
Liked
54:08Summary In this episode of the Data Engineering Podcast, Adrian Broderieux and Marcin Rudolph, co-founders of DLT Hub, delve into the principles guiding DLT's development, emphasizing its role as a library rather than a platform, and its integration with lakehouse architectures and AI application frameworks. The episode explores the impact of the P…
…
continue reading

1
Build Your Data Transformations Faster And Safer With SDF
42:36
42:36
Play later
Play later
Lists
Like
Liked
42:36Summary In this episode of the Data Engineering Podcast Lukas Schulte, co-founder and CEO of SDF, explores the development and capabilities of this fast and expressive SQL transformation tool. From its origins as a solution for addressing data privacy, governance, and quality concerns in modern data management, to its unique features like static an…
…
continue reading

1
Scaling Airbyte: Challenges and Milestones on the Road to 1.0
57:11
57:11
Play later
Play later
Lists
Like
Liked
57:11Summary Airbyte is one of the most prominent platforms for data movement. Over the past 4 years they have invested heavily in solutions for scaling the self-hosted and cloud operations, as well as the quality and stability of their connectors. As a result of that hard work, they have declared their commitment to the future of the platform with a 1.…
…
continue reading

1
Enhancing Data Accessibility and Governance with Gravitino
38:41
38:41
Play later
Play later
Lists
Like
Liked
38:41Summary As data architectures become more elaborate and the number of applications of data increases, it becomes increasingly challenging to locate and access the underlying data. Gravitino was created to provide a single interface to locate and query your data. In this episode Junping Du explains how Gravitino works, the capabilities that it unloc…
…
continue reading

1
The Evolution of DataOps: Insights from DataKitchen's CEO
53:30
53:30
Play later
Play later
Lists
Like
Liked
53:30Summary In this episode of the Data Engineering Podcast, host Tobias Macey welcomes back Chris Berg, CEO of DataKitchen, to discuss his ongoing mission to simplify the lives of data engineers. Chris explains the challenges faced by data engineers, such as constant system failures, the need for rapid changes, and high customer demands. Chris delves …
…
continue reading

1
Achieving Data Reliability: The Role of Data Contracts in Modern Data Management
49:26
49:26
Play later
Play later
Lists
Like
Liked
49:26Summary Data contracts are both an enforcement mechanism for data quality, and a promise to downstream consumers. In this episode Tom Baeyens returns to discuss the purpose and scope of data contracts, emphasizing their importance in achieving reliable analytical data and preventing issues before they arise. He explains how data contracts can be us…
…
continue reading

1
How Generative AI Is Impacting Data Engineering Teams
54:45
54:45
Play later
Play later
Lists
Like
Liked
54:45Summary Generative AI has rapidly gained adoption for numerous use cases. To support those applications, organizational data platforms need to add new features and data teams have increased responsibility. In this episode Lior Gavish, co-founder of Monte Carlo, discusses the various ways that data teams are evolving to support AI powered features a…
…
continue reading

1
The Role of Product Managers in Data-Centric Organizations
52:58
52:58
Play later
Play later
Lists
Like
Liked
52:58Summary In this episode Praveen Gujar, Director of Product at LinkedIn, talks about the intricacies of product management for data and analytical platforms. Praveen shares his journey from Amazon to Twitter and now LinkedIn, highlighting his extensive experience in building data products and platforms, digital advertising, AI, and cloud services. H…
…
continue reading

1
Neon: A Serverless And Developer Friendly Postgres
57:43
57:43
Play later
Play later
Lists
Like
Liked
57:43Summary Postgres is one of the most widely respected and liked database engines ever. To make it even easier to use for developers to use, Nikita Shamgunov decided to makee it serverless, so that it can scale from zero to infinity. In this episode he explains the engineering involved to make that possible, as well as the numerous details that he an…
…
continue reading

1
Improve Data Quality Through Engineering Rigor And Business Engagement With Synq
59:48
59:48
Play later
Play later
Lists
Like
Liked
59:48Summary This episode features an insightful conversation with Petr Janda, the CEO and founder of Synq. Petr shares his journey from being an engineer to founding Synq, emphasizing the importance of treating data systems with the same rigor as engineering systems. He discusses the challenges and solutions in data reliability, including the need for …
…
continue reading