TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Software How-To

    Data Engineering Podcast

    This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.

    Advertise

    Copyright: © 2024 Boundless Notions, LLC.

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    Reducing Data Debt with Agile Ledger Architecture Sep 24, 2026
    Show notes

    Summary
    In this episode Christopher Doidge talks about his Agile Ledger Architecture (ALA) approach to data warehousing and how it aims to reduce data debt while shortening the path from raw data to trustworthy business insight. Christopher explained that ALA is not a replacement for existing warehouse patterns like medallion architecture, star schemas, or other modeling approaches, but a complementary discipline focused on pushing business definitions upstream, enforcing cleaner ledger-style transformations, and producing gold-layer tables that stakeholders can actually use without relying on analysts to repeatedly rebuild the same logic. He also discussed his “15-minute litmus test” for time-to-insight, the importance of durable documentation through a data dictionary, and why undocumented business logic living in analyst scripts is often one of the biggest hidden forms of data debt. Overall, this was a thoughtful conversation about designing warehouse systems that serve not just analysts, but the broader business as well.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Today’s episode is sponsored by Parallel - where agents find answers. Most engineers today closely follow new model releases, but don’t pay attention to their agent’s most important tool: web search. Parallel develops enterprise-grade infrastructure for agents to retrieve high-quality context from the web. Their core products are a suite of APIs for retrieving high quality information from the web with Pareto-optimal quality, cost, and speed. Whether you work on voice agents that need 200 millisecond latency, chat bots that balance speed, depth, and quality, or long-horizon agents to do thorough, overnight research for you, Parallel is a single platform for all your agentic research. Get started for free at dataengineeringpodcast.com/parallel
    • Your host is Tobias Macey and today I'm interviewing Christopher Doidge about a data warehousing approach called Agile Ledger Architecture that is designed to drive down data debt
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what the agile ledger architecture is and the story behind it?
    • There are numerous patterns and practices that have been developed for data warehousing over the past 40 years. How does the agile ledger architecture fit in that ecosystem? (e.g. is it compatible with Kimball, Inmon, Data Vault, Anchor Modeling, etc.?)
    • What are some examples of the types of debt that accumulate in current approaches to warehouse implementation, and the impact that it has on the utility of that asset?
    • Digging into the architecture itself, what are the core principles that it is built on?
    • What are the technologies or practices that it is best suited to? (e.g. event streams, lakehouse, ELT workflows, etc.)
    • For someone who has already invested a substantial amount of effort into building a warehouse, what does adoption of the agile ledger architecture look like?
    • Once you have started that adoption, what are the technical controls that you can put in place to ensure that this architecture is maintained and doesn't regress or get sidestepped in another portion of the warehouse?
    • As LLMs and agents grow to become the predominant consumers (and often producers) of a warehouse, what are the benefits that the agile ledger architecture provides to ensure appropriate context and grounding to produce useful insights and accurate answers?
    • What are the most interesting, innovative, or unexpected ways that you have seen the agile ledger architecture used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on data warehouse design and implementation?
    • When is the agile ledger architecture the wrong choice?
    • What do you have planned for the future of this architecture?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Links
    • Agile Ledger Architecture Book (affiliate link)
    • Ralph Kimball
    • Bill Inmon
    • Data Vault
    • Star Schema
    • Anchor Modeling
    • Surrogate Key
    • Data Lakehouse
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    What Context Really Means in Data Engineering and AI Sep 15, 2026
    Show notes

    Summary
    In this episode Soham Mazumdar, co-founder and CEO of Wisdom.ai, talks about what “context” really means in data engineering and AI systems. He explores why context has become such an overloaded term, spanning everything from semantic layers and data catalogs to tribal knowledge, query logs, dashboards, and even agent memory. Soham explained that the big shift is that context is no longer being prepared primarily for human analysts, but for LLMs and agents that can’t reliably fill in missing gaps on their own. That change raises the bar for how context is represented, validated, benchmarked, and maintained so that AI systems can produce trustworthy outcomes.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Your host is Tobias Macey and today I'm interviewing Soham Mazumdar about what "context" actually means in data engineering
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • One of the perennial challenges of engineering in all forms is building a shared understanding of what a given word means. "Context" is one that is being used for an increasing number of purposes with the introduction of AI agents. Can you start by sharing some of the ways that this terminology overload has caused problems in your own experience?
    • Data engineering has arguably always been about context engineering, but at the scale of human consumers. What are the substantive changes that AI/agentic consumers bring to the discipline?
    • While we all understand the notion of "context", turning it into a useful and re-usable component is a different matter entirely. What are some of the ways that "business context" or "technical context" manifests as a tangible artifact?
    • This also brings up the question of data modeling. What are some of the key attributes that are necessary when storing, enriching, evolving, and joining into that context?
    • One could argue that the entire history of data warehousing is about building organizational context. What are the real differences in approach for today's work of capturing and activating that context?
    • How does your work at Wisdom AI address the technical and operational burdens of capturing, modeling, and exposing context at the speed necessary to keep up with organizational demands?
    • What are the most interesting, innovative, or unexpected ways that you have seen Wisdom AI used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Wisdom AI/context engineering?
    • When is Wisdom AI the wrong choice?
    • What do you have planned for the future of Wisdom AI?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
    Links
    • Wisdom AI
    • Context Engineering
    • Knowledge Graph
    • Ontology
    • Snowflake Open Semantic Interchange (OSI)
    • Semantic Layer
    • Palantir Foundry
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    Specialized AI for Data Engineers: Inside Astronomer’s Otto Aug 27, 2026
    Show notes

    Summary
    In this episode Yetunde Dada discusses Otto, Astronomer’s AI agent for Airflow, and the broader challenge of making agentic tooling actually useful for data engineers. She explored why generic coding assistants often fall short in data workflows, how Otto adds the missing context around Airflow, Astro, upgrades, and troubleshooting, and why Astronomer focused first on high-leverage use cases such as DAG authoring, investigation of pipeline failures, version migrations, and legacy scheduler modernization. She also discussed the practical realities of introducing agents into engineering teams: model choice, security boundaries, vendor lock-in concerns, validation of generated code, and the need for agents to fit into existing workflows rather than forcing users into new ones. Overall, this conversation offers a detailed look at how specialized AI agents can support data engineers today, and where Astronomer is headed next with a vision for self-healing pipelines that keep humans in control while automating more of the operational burden.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Your host is Tobias Macey and today I'm interviewing Yetunde Dada about Otto, Astronomer's expert Airflow agent

    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Otto is and the story behind it?
    • What are the core problems that you are trying to solve with Otto and for whom?
    • What was your process for identifying the scope of activities that Otto should be incorporated into?
    • Orchestration engines are a rich source of information. What are the aspects of Airflow that lend themselves to extending with this agentic context?
    • What are the other supporting systems that are necessary to enable Otto to work effectively, especially in mixed orchestration environments? (e.g. metadata platforms)
    • One of the explicit capabilities that you invested in is code review for Airflow DAGs. What are the pain points that you are trying to solve with a specialized review agent?
    • Can you describe the architecture of the Otto system and how you're managing the complex task of context curation?
    • In a production context accuracy and latency are both critical, and often in tension with each other. How do you monitor and optimize for each of those objectives?
      • What are the options for tuning Otto's behavior to bias more toward one direction or another?
    • What are some examples of the type of work that Otto can help automate?
    • How is it measurably different from a generic coding agent that has MCP connections to something like an Open Metadata or DataHub for platform and data context, Airflow documentation, etc.?
    • There are numerous general purpose and specialized agent systems available. What are some of the ways that Otto can work collaboratively with those other products?
    • What are the most interesting, innovative, or unexpected ways that you have seen Otto used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Otto?
    • What do you have planned for the future of Otto?

    Contact Info
    • LinkedIn

    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?

    Links
    • Astronomer
    • Otto
    • Announcement Post
    • Astronomer Cosmo dbt automation
    • Hadoop
    • Spark
    • Otto Automatic Pipeline Failure Investigation
    • Otto Code Review
    • Astro CLI
    • Astro IDE
    • Airflow MCP
    • Kedro
    • Quantum Black
    • Django
    • React
    • Airflow Providers
    • Pi Framework
    • Agent Skills
    • AGENTS.md

    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    Why Multi-Agent Systems Need Shared State, Graph Semantics, and Governance Aug 02, 2026
    Show notes

    Summary
    In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for traversal workloads, and how that perspective shaped OmniGraph’s design on top of object storage, Lance, Arrow, and DataFusion. Ragnor explained the motivation for combining graph semantics with Git-style branching and merging so that teams can manage probabilistic writers such as AI agents with stronger governance, shared context, and safer collaboration patterns. He also dug into the practical tradeoffs of building a graph engine for multi-agent coordination instead of traditional graph analytics use cases. He closed with a look at emerging use cases such as company “brain” systems, software development lifecycle graphs, research workflows, and event-driven agent orchestration, along with a broader conversation about composability, and sovereign AI infrastructure.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Your host is Tobias Macey and today I'm interviewing Ragnor Comerford about OmniGraph, a lakehouse-native graph storage layer with git semantics

    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what OmniGraph is and the story behind it?
    • What was the original problem that you were trying to solve by creating it?
    • There are numerous graph engines available, what are the properties of OmniGraph that differentiate it from the competition?
    • Cypher (GQL) and Gremlin are all established languages with years of examples to work from. What are the benefits of developing a new and more constrained query interface for an agentic audience?
    • How does that change the potential applications of OmniGraph? (e.g. general knowledge graph, fraud detection, SIEM, etc.)
    • Can you describe the architecture of OmniGraph?
    • You have built the system on top of several well-established open source components. What was your process for deciding what to use and how to compose it?
    • What are some examples of systems that can be built with OmniGraph?
    • What are other components/integration points that compose well with OmniGraph?
    • Given the technologies that you are building on top of, what are the automatic benefits/integrations that you benefit from?
    • Given that the underlying storage is Lance, and Lance's interoperability with Parquet/Iceberg, what are the opportunities for modeling graphs on top of existing lakehouse data?
    • What are the most interesting, innovative, or unexpected ways that you have seen OmniGraph used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on OmniGraph?
    • When is OmniGraph the wrong choice?
    • What do you have planned for the future of OmniGraph?

    Contact Info
    • LinkedIn

    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?

    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.

    Links
    • OmniGraph
    • ModernRelay
    • Information Theory
    • Lance
    • Git
    • Neo4J
    • GraphRAG
    • TigerGraph
    • LakeHouse
    • Iceberg
    • Dolt
    • PuppyGraph
    • Podcast Episode
    • Terraform
    • Gremlin
    • Cypher
    • GQL
    • SPARQL
    • In-context Learning
    • BM25 Indexing
    • Data Fusion
    • Adjacency Matrix
    • Predicate Pushdown
    • Agentic Mesh book (affiliate link)
    • witan-council
    • witan-code
    • Web Assembly
    • Clickhouse
    • DSPy
    • MCP-UI

    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    Building the Context Flywheel for AI Data Agents Jul 06, 2026
    Show notes

    Summary
    In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance depends on contextual intelligence: institutional knowledge, semantic meaning, procedural know-how, and access to the right tools. She also dug into how metadata catalogs are evolving into broader context layers that serve both humans and agents, and why agentic systems are changing the economics of metadata and governance work. Prakulpa shared Atlan’s perspective on bootstrapping context from existing systems such as warehouses, BI tools, query logs, and SaaS applications, then using simulation, traces, and human governance loops to improve agent accuracy over time.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Your host is Tobias Macey and today I'm interviewing Prukalpa Sankar about strategies for building a context flywheel for your data agents
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • You have spent several years working in the metadata catalog space with Atlan. What are the notable changes in scope, adoption, and application that you have seen since we last spoke (June 2022)?
    • The recurring theme since the start of 2026 has been agentic augmentation of all engineering workflows, including data. How do you differentiate between data catalogs, semantic layers, agent memory, context layers, etc. when architecting an AI-powered data-oriented system?
    • One of the perennial problems with data catalogs, business glossaries, master data management, etc. is the up-front investment required to get a real-world impact. How can agents help reduce the activation energy needed to get to that return on effort?
    • One of the perennial problems in data engineering is fragmentation and siloing of data. This is exacerbated by AI systems due to the introduction of vector data as a new specialization. What are the forces that you are seeing play into the current set of tensions and the architectural primitives that we need to bring to bear to keep things maintainable?
    • Since the introduction of transformer-based generative models we have been combating hallucinations. While we have made progress, it is still critical to ensure accuracy and trustworthiness when working with business data. What are the policy elements of governance and technical controls to ensure a high degree of confidence in agent-generated context and business semantics?
    • What are the most interesting, innovative, or unexpected ways that you have seen teams build context layers for their agentic data workloads?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on business context engineering?
    • When is agent-managed context the wrong choice?
    • What are your predictions for the next set of architectural shifts that will be driven by the pressures of AI-powered systems?

    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
    Links
    • Atlan
    • Atlan Context Lakehouse
    • Iceberg
    • Business Glossary
    • Master Data Management
    • Semantic Layer
    • Cube.dev
    • MCP == Model Context Protocol
    • A2A == Agent to Agent Protocol
    • Decision Traces
    • Apache Doris
    • StarRocks

    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    Holding Kafka Right: Product-Friendly Streaming with TypeStream Jun 18, 2026
    Show notes

    Summary
    In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka. Jevin shares lessons from roles at Zapier, Humi, and Clio, where real-time synchronization, customer data unification, and document sync at scale highlighted both the strengths and common misuses of Kafka. He digs into using events as the source of truth, materialized views with KTables, and how schema registries and type safety prevent downstream breakage. Jevin explains why teams often reach for heavyweight Kafka clusters without leveraging Streams, Connect, or interactive queries—and how his project, TypeStream, aims to make those capabilities accessible via config-as-code while keeping a thin abstraction and clear escape hatches. He also explore trade-offs across Kafka-compatible alternatives, CDC with Debezium in the real world, and where abstractions should stop so teams can scale responsibility as complexity grows.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • This episode is sponsored by DataDriven.io, the free data engineering interview prep platform built by data engineers for data engineers. Ever walked into a data engineering interview and gotten a question that has nothing to do with real data engineering work? Interviewing is its own skill, separate from the job. Watch your code execute live, inspect Spark internals, and whiteboard your data models and pipelines and defend your decisions. Unlike SQL-only or Python-only practice, DataDriven.io covers the full interview loop: star schemas, slowly changing dimensions, grain and fact table design, idempotency, watermarks, dead letter queues, change data capture, and backpressure. Every question comes from real Data Engineer interview loops at Google, Amazon, Meta, Stripe, Databricks, Netflix, and Airbnb. Go to dataengineeringpodcast.com/datadriven today to start practicing.
    • Your host is Tobias Macey and today I'm interviewing Jevin Maltais about the challenges of building a reliable streaming

    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Typestream is and the story behind it?
    • What are the common challenges that teams encounter when trying to build on top of Kafka?
    • How do those challenges/misconfigurations impact the team's ability to deliver on product goals?
    • What are the fundamental design aspects of Kafka that contribute to the difficulties that teams encounter when using it as an element of their architecture?
    • There have been numerous projects taking aim at Kafka, with varying approaches and degrees of effectiveness (e.g. RedPanda, AutoMQ, Pulsar, etc.). What are the tradeoffs that each of those approaches requires?
    • What makes the original Kafka project so resilient in the face of all of that competition?
    • Can you describe the architecture of Typestream and how each of the core elements contribute to a better user experience?
    • For teams who want to take advantage of streaming capabilities, but don't want to invest in becoming Kafka experts, what does the Typestream workflow look like?
    • If they don't want to manage the operational overhead of a Kafka cluster, how tightly coupled is Typestream to the original Kafka? (can someone use RedPanda or AutoMQ instead?)
    • What are the most interesting, innovative, or unexpected ways that you have seen Typestream used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Typestream?
    • When is Typestream the wrong choice?
    • What do you have planned for the future of Typestream?

    Contact Info
    • Website

    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?

    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.

    Links
    • Typestream
    • Zapier
    • Airflow
    • Kafka
    • KTables
    • KSQL
    • RedPanda
    • Pulsar
    • AutoMQ
    • Kafka Schema Registry
    • Debezium
    • Change Data Capture
    • Kafka Connect
    • Terraform
    • Kafka Compacted Topic

    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    Text to Data Products: Kaarvi’s End-to-End AI for Ingestion, Quality, and Dashboards Jun 08, 2026
    Show notes

    Summary
    In this episode Shravan Gunda, founder and CEO of Kaarvi AI, talks about building an AI-native, agent-driven data platform designed to eliminate the janitorial work that consumes most data teams. He explores Kaarvi’s multi-agent architecture that runs queries across seven LLMs in parallel for reliability, its synthetic data generator that mirrors source schemas for quick testing, and “Hey Kaarvi” chat for text-to-SQL, text-to-transformations, and text-to-dashboard workflows. He also digs into on-prem versus SaaS deployments, domain-specialized agents for privacy and accuracy, code blocks for custom Python/SQL, and the roadmap for a marketplace and desktop assistant. Shravan highlights how Kaarvi compresses weeks of work into hours and bridges the gap between business users and data engineers by turning AI into a dependable force multiplier.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • This episode is sponsored by DataDriven.io, the free data engineering interview prep platform built by data engineers for data engineers. Ever walked into a data engineering interview and gotten a question that has nothing to do with real data engineering work? Interviewing is its own skill, separate from the job. Watch your code execute live, inspect Spark internals, and whiteboard your data models and pipelines and defend your decisions. Unlike SQL-only or Python-only practice, DataDriven.io covers the full interview loop: star schemas, slowly changing dimensions, grain and fact table design, idempotency, watermarks, dead letter queues, change data capture, and backpressure. Every question comes from real Data Engineer interview loops at Google, Amazon, Meta, Stripe, Databricks, Netflix, and Airbnb. Go to dataengineeringpodcast.com/datadriven today to start practicing.
    • Your host is Tobias Macey and today I'm interviewing Shravan Gunda about building an agent-driven data platform at Kaarvi
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Kaarvi is and the story behind it?
    • "AI" is a very broad term that encompasses numerous possible implementations. Can you give some more detail about the different types and applications of AI in Kaarvi's architecture?
    • What are some of the core assumptions of data workflows that need to be reconsidered when AI is embedded in the execution path?
    • What are the most interesting, innovative, or unexpected ways that you have seen Kaarvi used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Kaarvi?
    • When is Kaarvi the wrong choice?
    • What do you have planned for the future of Kaarvi?

    Contact Info
    • LinkedIn

    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?

    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.

    Links
    • Kaarvi
    • Synthetic Data
    • n8n

    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    Scaling Graph Analytics Without ETL: Inside PuppyGraph’s Architecture Jun 01, 2026
    Show notes

    Summary
    In this episode Weimo Liu, co‑founder of PuppyGraph, talks about the engineering behind their “zero-copy” graph querying engine for lakehouse and database sources. He explores how PuppyGraph lets you run Cypher and Gremlin traversals and graph algorithms directly on data in Iceberg, Delta, Hudi, Hive, and even MongoDB—without loading into a separate graph store. Weimo explains their edge-sharded, vectorized, MPP architecture that tackles hub nodes, multi-hop traversals, and shuffle at scale, targeting sub-second to single-digit-second workloads. He digs into practical graph data modeling on top of normalized and denormalized tables, logical views, and flexible mappings; strategies for caching, adaptive reads, and leveraging Iceberg metadata; and how PuppyGraph’s operator-based engine unifies query and algorithms. He also covers real-world applications—from cybersecurity log analysis to entity resolution and agentic workflows—when to choose embedded or transactional graph databases instead, and what’s next for enterprise features and broader warehouse integrations.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • This episode is sponsored by DataDriven.io, the free data engineering interview prep platform built by data engineers for data engineers. Ever walked into a data engineering interview and gotten a question that has nothing to do with real data engineering work? Interviewing is its own skill, separate from the job. Watch your code execute live, inspect Spark internals, and whiteboard your data models and pipelines and defend your decisions. Unlike SQL-only or Python-only practice, DataDriven.io covers the full interview loop: star schemas, slowly changing dimensions, grain and fact table design, idempotency, watermarks, dead letter queues, change data capture, and backpressure. Every question comes from real Data Engineer interview loops at Google, Amazon, Meta, Stripe, Databricks, Netflix, and Airbnb. Go to dataengineeringpodcast.com/datadriven today to start practicing.
    • Your host is Tobias Macey and today I'm interviewing Weimo Liu about the engineering behind PuppyGraph's zero-copy ETL for querying your lakehouse as a graph
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by describing what PuppyGraph is and the story behind it?
    • What are some of the key use cases that people are turning to PuppyGraph and graph data models for?
    • Graph engines have struggled to take off for several years, not least of which is due to the difficulty of scaling them to large data volumes as a result of the topological nature of the data. Can you describe the architecture of PuppyGraph and some of the ways that you are addressing that challenge of data volume for graphs?
    • latency/data exploration
    • types of traversals and limitations
    • lakehouse architecture pros/cons for graphs
    • data modeling/translation
    • shortcomings of zero-ETL and how transforming the underlying representation could provide benefits
    • For someone who is looking for a graph engine to support a connected data use case, what are the guiding questions that you would ask to lead them toward PuppyGraph vs. a dedicated graph database like Memgraph/Neo4J/etc.?
    • What are the most interesting, innovative, or unexpected ways that you have seen PuppyGraph used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on PuppyGraph?
    • When is PuppyGraph the wrong choice?
    • What do you have planned for the future of PuppyGraph and graph data exploration on large data volumes?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
    Links
    • PuppyGraph
    • TigerGraph
    • Google F1
    • Graph Database
    • Google Pregel
    • Iceberg
    • Graph Supernode
    • MPP == Massively Parallel Processing
    • Spark GraphX
    • Trino
    • Ladybug DB
    • lance-graph
    • KuzuDB
    • MemGraph
    • Labelled Property Graph
    • RDF Triples
    • Cypher Query Language
    • Gremlin
    • CDC == Change Data Capture
    • Neo4J
    • JanusGraph
    • NetworkX
    • PyTorch
    • DuckDB
    • Iceberg Array
    • LanceDB
    • Palo Alto Networks
    • Columnar ADBC
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    %

    Maximizing GPU Utilization: Heterogeneous Pipelines with Ray and Kubernetes May 06, 2026
    Show notes

    Summary
    In this episode Robert Nishihara, co-founder of Anyscale and co-creator of Ray, talks about maximizing hardware utilization for AI and data-intensive workloads. He explores Ray’s evolution alongside Kubernetes and PyTorch, and why consolidation at these layers has enabled a new generation of complex, heterogeneous workloads. Robert explains how data preparation has shifted to GPU- and inference-heavy, multimodal pipelines; where Ray fits compared to Spark and workflow orchestrators; and why Ray excels at composing heterogeneous pools of compute, handling failures, and scaling complex systems like multi-node LLM inference and reinforcement learning. He digs into practical strategies for boosting GPU utilization across training and inference, elasticity and prioritization of workloads, topology-aware scheduling, and the importance of fast failure recovery as hardware scales from nodes to racks. If you’re wrestling with expensive GPUs, multimodal data curation, or cross-node LLM inference, this conversation offers concrete mental models and architectural guidance.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Your host is Tobias Macey and today I'm interviewing Robert Nishihara about the challenges of maximizing the utility of your available hardware for AI applications
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by giving an overview of the major contributors to wasted or idle compute?
    • Why does it matter if the available compute isn't being maximized?
    • What are some of the typical ad-hoc methods that teams might use to try to get the most out of their available hardware (especially GPUs)?
    • What are the most interesting, innovative, or unexpected ways that you have seen Ray used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Ray and distributed compute for data and AI?
    • When is Ray the wrong choice?
    • What do you have planned for the future of Ray?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.
    Links
    • AnyScale
    • Ray
    • Deep Learning
    • Computer Vision
    • Kubernetes
    • Cursor
    • Claude Code
    • Kube-Ray
    • PyTorch
    • Tensorflow
    • Theano
    • Caffe
    • vLLM
    • SGLang
    • Ray Tune
    • Neural Network
    • Learning Rates
    • Reinforcement Learning
    • AlphaGo
    • Cursor Composer 2
    • ImageNet
    • Transformer Architecture
    • Stochastic Gradient Descent
    • Airflow
    • Dagster
    • Flyte
    • Mixture of Experts
    • Prefill
    • Temporal
    • Actor Framework
    • RDMA == Remote Direct Memory Access
    • Neoclouds
    • AI Engineering Podcast Episode
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    The AI-First Data Engineer: 10–50x Productivity and What Changes Next Apr 07, 2026
    Show notes

    Summary
    In this episode, I sit down with Gleb Mezhanskiy, CEO and co-founder of Datafold, to explore how agentic AI is reshaping data engineering. We unpack the leap from chat-assisted coding to truly agentic workflows where AI not only writes SQL and dbt models but also executes queries, debugs, runs tests, and ships production-ready outcomes. Gleb explains why teams that master this AI-first loop can see 10–50x gains, how security/compliance concerns can be addressed with platform-native LLM endpoints, and why the role of data engineers is shifting from code authors to operators of autonomous agents. We dig into the consolidation of the modern data stack, the economics driving more data products (Jevons paradox), and why product thinking, domain knowledge, and cross-functional skills will define the next wave of standout data professionals. We also cover practical steps for leaders and ICs: modernizing off legacy platforms, establishing safe AI adoption paths, codifying reusable “skills” and context for agents, and building validation utilities that keep the inner loop fast and trustworthy. Finally, Gleb shares how Datafold moved to fully AI-driven software delivery and why “outcomes over tools” is the emerging model for complex initiatives like data platform migrations—and how this reframes data quality for the AI era, emphasizing broad data access plus rich context over brittle human-centric tests.
    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • If you lead a data team, you know this pain: Every department needs dashboards, reports, custom views, and they all come to you. So you're either the bottleneck slowing everyone down, or you're spending all your time building one-off tools instead of doing actual data work. Retool gives you a way to break that cycle. Their platform lets people build custom apps on your company data—while keeping it all secure. Type a prompt like 'Build me a self-service reporting tool that lets teams query customer metrics from Databricks—and they get a production-ready app with the permissions and governance built in. They can self-serve, and you get your time back. It's data democratization without the chaos. Check out Retool at dataengineeringpodcast.com/retool today and see how other data teams are scaling self-service. Because let's be honest—we all need to Retool how we handle data requests.
    • Your host is Tobias Macey and today I'm bringing back Gleb Mezhanskiy to talk about our predictions for the impact of AI on data engineering for 2026

    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • What are the concrete steps that teams need to be taking today to take advantage of agentic AI capabilities?
    • What are the new guardrails/constraints/workflows that need to be in place before you let AI loose on your data systems?
    • How do you balance the potential cost savings and productivity increases with the up-front investment and variability in inference spend?

    Contact Info
    • LinkedIn

    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?

    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email hosts@dataengineeringpodcast.com with your story.

    Links
    • Blog Post
    • Datafold
    • Claude Opus 4.5
    • Harry Potter - Muggles
    • Jevon's Paradox
    • Modern Data Stack
    • Dagster Compass
    • Gravity Orion
    • MCP == Model Context Protocol
    • Qwen

    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

    1 2 3 52 Next

    Related Podcasts

    Know How… (Video)

    1

    Know How… (Video) Education
    JavaScript Jabber

    2

    JavaScript Jabber Education
    Programming Throwdown

    3

    Programming Throwdown Software How-To
    Software Engineering Radio – the podcast for professional software developers

    4

    Software Engineering Radio – the podcast for professional software developers Software How-To
    The Changelog: Software Development, Open Source

    5

    The Changelog: Software Development, Open Source How To
    Coding Blocks

    6

    Coding Blocks Software How-To
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights