TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    The Data Engineering Show

    The Data Engineering Show is a podcast for data engineering and BI practitioners to go beyond theory. Learn from the biggest influencers in tech about their practical day-to-day data challenges and solutions in a casual and fun setting.

    SEASON 1 DATA BROS
    Eldad and Boaz Farkash shared the same stuffed toys growing up as well as a big passion for data. After founding Sisense and building it to become a high-growth analytics unicorn, they moved on to their next venture, Firebolt, a leading high-performance cloud data warehouse.

    SEASON 2 DATA BROS
    In season 2 Eldad adopted a brilliant new little brother, and with their shared love for query processing, the connection was immediate. After excelling in his MS, Computer Science degree, Benjamin Wagner joined Firebolt to lead its query processing team and is a rising star in the data space.

    For inquiries contact tamar@firebolt.io
    Website: https://www.firebolt.io

    Advertise

    Copyright: © 2024 The Firebolt Data Bros

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    How AI Is Reshaping Modern Data Teams and the Future of Platforms with Xavier Gumara Rigol Aug 18, 2026
    Show notes In this episode of the Data Engineering Show, host Benjamin Wagner sits down with Xavier Gumara Rigol, Head of Data at Manychat, to discuss how artificial intelligence is fundamentally transforming data roles, organizational structures, and the future of analytics platforms while creating new types of work rather than eliminating existing ones.
    What You'll Learn:

    - How to transition data analyst work from query delivery to AI-powered self-service analytics — moving from Tableau dashboards to natural language-to-SQL-to-insight tools that let business users answer their own questions with high accuracy
    - Why cross-functional team structures beat siloed data organizations — embed data analysts and scientists directly within product teams focused on specific business problems rather than centralizing them in separate departments
    - The "Build vs. Buy" framework for AI analytics tools in 2026 — evaluate text-to-SQL solutions based on BI tool consolidation, programmatic manageability, cost competitiveness, and your organization's readiness for AI integration
    - How to structure data platform teams to enable autonomous work — create a central platform team that manages infrastructure, data quality, privacy, and compliance so embedded analysts can operate independently
    - The critical role of context layers for LLM accuracy — well-structured metadata and context on top of your data warehouse is essential for achieving high-accuracy natural language queries at scale
    - Why greenfield projects move faster with AI than legacy platforms — newly built systems with AI require less headcount than legacy systems, but maintaining existing data infrastructure still demands human expertise for years to come
    About the Guest(s)

    Xavier is the Head of Data at Manychat, leading the Machine Learning and Data Platform team with over a decade of experience in the data ecosystem. With a background spanning BI development to data platform management, he has become a thought leader on bridging the divide between data and product functions—a topic he explored in depth through his published book. In this episode, Xavier shares transformative insights on how AI is reshaping data team structures, roles, and workflows, providing actionable strategies for organizations looking to build more versatile, cross-functional data teams. His work implementing natural language-to-SQL analytics agents and customer-facing analytics platforms demonstrates the practical impact of AI adoption in modern data platforms, making this conversation essential for data leaders navigating the rapidly evolving landscape of AI-driven analytics.
    Quotes
    "I think that the most exciting thing that we are working on at the moment is this natural language to SQL to insight type of tool." - Xavier Gumara Rigol
    "The type of work is changing. It's not going anywhere, but it's changing." - Xavier Gumara Rigol
    "Data roles that are value creators, like data analysts and data scientists, need to work in a cross-functional team with a product manager and with a specific problem to solve." - Xavier Gumara Rigol
    "We are repurposing things to do more things at the same time. And this is the benefit of AI and how we are taking advantage today." - Xavier Gumara Rigol
    "To provide high accuracy in the answers, you need to have a very well-structured and built context layer on top of your data warehouse." - Xavier Gumara Rigol
    "The barrier to building has been lowered quite a lot, and we could provide something very quick, very fast to the rest of the organization." - Xavier Gumara Rigol
    "For the current platforms to be maintained and for the current capabilities to be maintained, you will still need humans for quite a few years." - Xavier Gumara Rigol
    "Everything that you create in the system should be able to be created in text because we want to work with this with LLMs and generative AI." - Xavier Gumara Rigol
    "If you start from scratch, you probably can get to a nice place where you need fewer people to maintain in comparison to if you were to build this three years ago." - Xavier Gumara Rigol
    "We still need a data analyst in the company, and we will need them for a long time." - Xavier Gumara Rigol
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review

    Resources

    LinkedIn Profiles:

    • Xavier’s LinkedIn: https://www.linkedin.com/in/xgumara
    • Benjamin's LinkedIn: https://www.linkedin.com/in/wagjamin

    Company Websites:

    • Manychat: manychat.com
    • Snowflake: snowflake.com
    • Firebolt: firebolt.io

    Books:

    • Data As a Product Driver - https://www.amazon.com/Data-Product-Driver-Strategies-Organizations/dp/B0FGNJ8Q3F

    Tools & Platforms:

    • Tableau – Business Intelligence and data visualization
    • Snowflake – Cloud data warehouse and analytical database
    • Claude/Cloud Code – Natural language to SQL query interface with MCP (Model Context Protocol) integration
    • SQL – Stands for Structured Query Language. It is a standardized programming language used to communicate with, manage, and manipulate data stored in relational databases.
    • Firebolt – a high-performance analytical database for real-time analytics and batch processing
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Building Modern Data Platforms Without Legacy Bottlenecks ft Andrew Jones Aug 06, 2026
    Show notes In this episode of The Data Engineering Show, host Benjamin Wagner sits down with Andrew Jones, Staff Data Engineer at LocalStack, to explore how data teams can transition from centralized, bottlenecked operations to scalable, self-serve platforms - while managing real-time reliability demands that now define modern data infrastructure.
    What You'll Learn:

    - How to free up data team capacity by automating support workflows—Andrew's hackathon project automated 15% of access requests using Slack bot agents, reclaiming engineering time for platform-building rather than firefighting.
    - Why data reliability has shifted from a "nice-to-have" to a customer-facing requirement—As companies embed data into product features and AI agents, pipeline failures cascade directly to customers, demanding disciplined software engineering practices like observability, SLOs, and on-call protocols.
    - The framework for segmenting data personas and enabling autonomy without chaos—Technical teams (engineers, product managers) get self-serve tooling to own their data lifecycle; business teams receive curated, pre-built datasets through accessible tools like BI platforms and AI agents.
    - How to shift left on data governance by rethinking upstream data contracts—Rather than relying on manual documentation and checklists, enforce reliability guarantees at the source, ensuring downstream pipelines can't fail due to upstream data quality degradation.
    - Why semantic models and context layers are becoming critical for AI agent grounding—As agents proliferate across organizations, they need structured knowledge of business definitions and data relationships to make contextually accurate decisions and queries.
    - The importance of questioning assumptions during rapid industry transformation—Admit what you don't know yet, stay open-minded about rethinking data ownership and governance, and avoid blindly replicating legacy processes in a fundamentally changed landscape.
    About the Guest(s)

    Andrew Jones is a Staff Data Engineer at LocalStack, bringing a decade of platform-building expertise across companies including GoCardless and Arm. Specializing in data platform architecture and the shift toward data-driven product features, Andrew has pioneered approaches to data reliability and contracts that enable organizations to leverage data as a competitive advantage. In this episode, Andrew shares invaluable insights on transitioning centralized data teams into self-serve platforms, managing the tension between supporting legacy systems while building modern infrastructure, and rethinking data governance in the age of AI agents. His pragmatic approach to solving real-world data challenges makes this conversation essential for data leaders navigating the evolving landscape of analytics and data engineering.
    Quotes
    "I've been building platforms really since I started, and for the last ten years, I was focused on data platforms. That's really my passion." - Andrew Jones
    "The idea there is you can avoid all the cloud costs that you have when you're trying to develop against cloud and work much quicker because it's all local, it's all instant, it's all emulated." - Andrew Jones
    "Data warehouses can be quite expensive if they're up and running, and that's one of the things we're seeing from our customers—they want to make it quicker and also reduce the cost associated with development and CI checking against data warehouses." - Andrew Jones
    "We want to enable people in the business to be more self-serve and make use of the data and add more telemetry to our product without having to involve a person each time." - Andrew Jones
    "One of the biggest challenges is you're trying to build this new world while you're still trying to support the old world—you can't just drop everything and say we're going away for six months." - Andrew Jones
    "With agents now, we can move a lot quicker and build general tooling much quicker than we could before, which means you can automate more than you could before." - Andrew Jones
    "People want to use data either directly or via AI and agents to create product features that are differentiated by the unique data a company has, and that requires a high level of reliability." - Andrew Jones
    "If you want the output to be reliable, then your bits gotta be reliable, but also the bit upstream—which means you have to talk to those people and explain to them how important it is." - Andrew Jones
    "Data contracts are getting more and more traction because more and more companies are trying to use data for customer-facing features and realizing that if you want that to be reliable, you need to go upstream and make sure it's fixed there." - Andrew Jones
    "Things are changing so fast that we're still working a lot of these things out, and I think one of the best things we can do is be open-minded and question some of the assumptions we've had before." - Andrew Jones
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review

    Resources

    LinkedIn Profiles:

    • Andrew’s LinkedIn: https://www.linkedin.com/in/andrewrhysjones/
    • Benjamin's LinkedIn: https://www.linkedin.com/in/wagjamin

    Company Websites:

    • LocalStack: https://www.localstack.cloud
    • Firebolt: firebolt.io

    Tools & Platforms:

    • LocalStack – AWS and Snowflake emulator for local development
    • Snowflake – Cloud data warehouse
    • Tinybird – Real-time analytics platform built on ClickHouse
    • ClickHouse – Open-source columnar database for analytics
    • Slack Bot Agent – Automation for access management requests

    Concepts & Methodologies:

    • Data Contracts – Framework for ensuring data reliability between upstream services and data pipelines
    • Semantic Models – Context layers for grounding agent knowledge
    • Infrastructure as Code – Approach to managing data access and permissions
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Why 99% of BI Tools Get Embedded Analytics Wrong And How Omni Fixed It ft. Chris Merrick Jul 21, 2026
    Show notes In this episode of The Data Engineering Show, host Benjamin Wagner sits down with Chris Merrick, CTO and cofounder of Omni, to explore how AI is transforming the analytics stack and why keeping humans at the center of that transformation is critical to avoiding becoming just another middleware layer.
    What You'll Learn:

    - Why becoming "agentic middleware" would signal failure — and how to design analytics platforms that empower human decision-makers while leveraging AI as an accelerant rather than a replacement
    - How to balance exploratory and governed analytics — the framework for combining ad-hoc AI-driven analysis with consistent, repeatable metrics that organizations depend on for decision-making
    - The semantic layer as the intelligence backbone — why curating datasets and defining business terminology (like "uplift") at the semantic level enables both humans and AI agents to work effectively within guardrails
    - How to avoid the "universal translator" trap — why attempting to abstract away database dialect differences creates more problems than it solves, and how to leverage SQL as the common language instead
    - The feedback loop that makes AI smarter over time — how to automatically capture learnings from user conversations and feed them back into the system so your analytics layer evolves as your business evolves
    - The next frontier: from analytics to true business intelligence — how incorporating unstructured data (Salesforce comments, marketing notes, etc.) alongside structured metrics unlocks deeper "why" questions about your entire organization
    About the Guest(s)

    Chris Merrick is a Co-founder and CTO at Omni, bringing extensive expertise in data engineering and analytics infrastructure. With a distinguished background spanning data pipeline architecture at Stitch (acquired by Talend) and product leadership at Looker, Chris has been instrumental in shaping modern data integration and business intelligence practices. In this episode, Chris discusses how Omni is revolutionizing analytics by combining human-centric design with AI-driven insights, while maintaining governance and semantic rigor across complex data ecosystems. His perspective on balancing agentic workflows with end-user empowerment provides essential guidance for organizations navigating the convergence of AI and analytics, making this conversation invaluable for data leaders, engineers, and analytics practitioners seeking to build intelligent, scalable data platforms.
    Quotes
    "We're building a trusted AI analytics platform for you to do everything from asking questions of your data, getting a trusted response, and then manipulating that, building dashboards from it." - Chris Merrick
    "If Omni becomes agentic middleware, we will have failed." - Chris Merrick
    "There's more than one mode of analytics—there's ad hoc exploratory analysis, and there's the governed version of analytics." - Chris Merrick
    "We have business users logging into this platform with a lot of context about who they are, and that gives us a heck of a lot of context to give them a great data experience." - Chris Merrick
    "You gotta curate the dataset that the human or the LLM is seeing if you want it to be able to use it effectively." - Chris Merrick
    "Semantics exist at every layer of the stack—a database table has semantics, and you go up the stack and get more and more different semantics." - Chris Merrick
    "Don't invent new languages—we don't wanna play that game because we're not language builders and it's not adding a whole lot of value." - Chris Merrick
    "When you define something in Omni, it is relative to the database dialect that it is going to—we don't attempt to be a universal translator." - Chris Merrick
    "The data that used to be the most useless for the BI tool—like long-form AE comments in Salesforce—is now all of a sudden the most interesting stuff." - Chris Merrick
    "How do we pull unstructured data into the fold and make the system intelligent about the entire company, not just the tables in the warehouse?" - Chris Merrick
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review

    Resources

    LinkedIn Profiles:

    • Chris LinkedIn: https://www.linkedin.com/in/merrickchristopher
    • Benjamin's LinkedIn: https://www.linkedin.com/in/wagjamin

    Company Websites:

    • Firebolt: firebolt.io
    • Omni: https://omni.co
    Tools & Platforms:

    • Stitch – ETL data pipeline tool (acquired by Talend)
    • Talend – Data integration platform
    • dbt (data build tool) – Data transformation and modeling tool
    • Looker – Business intelligence and analytics platform
    • Blobby – AI agent for analytics queries at Omni
    • OSI (Open Semantic Interchange) – Standard for semantic model representation
    Mentioned People:

    • Chris Merrick – Co-founder at Omni
    • Jamie – Co-founder at Omni
    • Moshe – CTO (invented MDX), associated with database/semantic model work
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    AI for Data and Data for AI: The Dual Frontier of Modern Data Engineering with Pranav Motarwar Jun 16, 2026
    Show notes In this episode of The Data Engineering Show, host Benjamin Wagner sits down with Pranav Motarwar, a data engineer who worked across major tech companies, and the intersection of AI and data infrastructure, to explore how artificial intelligence is fundamentally reshaping the data engineering landscape not by eliminating roles, but by bifurcating the field into two distinct, equally critical domains.
    What You'll Learn:

    - Why the "data engineering is dying" narrative is clickbait: Data engineers remain essential because 60% of use cases by 2027 will involve providing data to AI agents, while simultaneously human-facing analytics demands continue growing, meaning more work, not less.
    - How to future-proof your career by mastering "AI for Data" AND "Data for AI": Modern AI Data Engineer roles now require both using AI agents to accelerate traditional ETL/DBT workflows AND building entirely new data pipelines (chunking, embedding, vector storage) designed specifically for agent consumption.
    - The transformation framework breaking down how data pipelines for humans differ from pipelines for agents: Human-facing pipelines traditionally handled structured data; agent pipelines now require handling unstructured multimodal inputs (videos, audio, images), demanding completely different architectural approaches.
    - Why individual contributors now own end-to-end pipelines that previously required 7-8 engineers: AI-assisted coding and low-code platforms like Databricks Cortex and Snowflake's GenAI tools reduce traditional pipeline development from one month to 3-4 weeks, freeing engineers to focus on product strategy, governance, and business impact.
    - How the next-gen data stack will evolve: traditional tools (DBT, BI platforms) stay relevant, but new specialized systems emerge: Companies like Vespa handle multimodal retrieval serving, while emerging startups build data warehouses purpose-built for video and complex unstructured data - eventual consolidation will come once larger players (Databricks, Snowflake) evolve their offerings.
    - The exponential data explosion argument that guarantees ongoing demand: Data generated by all humanity through 2008 is now created daily; even single engineers replacing five-person teams will find more work arriving as use cases expand across AI agents, real-time recommendations, robotics, and physical AI systems.
    About the Guest(s)
    Pranav Motarwar is a data engineer with extensive experience across leading tech companies, where he has worked in risk, product, privacy, and core data engineering roles. With a background spanning from traditional data engineering to cloud infrastructure and AI-driven systems, Pranav brings a unique perspective on the industry's rapid evolution. In this episode, he explores how AI is fundamentally transforming data engineering workflows, discussing the emergence of dual pipeline architectures for both human and AI consumption, and the critical skills data engineers need to remain relevant in 2025 and beyond. His insights on the shift from structured data pipelines to multimodal, AI-optimized infrastructure provide actionable guidance for engineers navigating the next generation of data stack technology.
    Quotes
    "I've worked across different product-based companies in different domains like risk and product, as well as privacy, and the core data engineering teams as well." - Pranav Motarwar
    "Data engineering is completely segmented into two different categories: one where the end consumer is human or product, and another where you are building data engineering flow, pipelines, and design for agents to consume." - Pranav Motarwar
    "What used to take one month to create an entire flow with DBT has now been reduced to almost 30% of the time we usually spent three to four years ago." - Pranav Motarwar
    "Data engineers need to be aware of the process of chunking, embedding, and how you are planning the vector store and optimizing the entire process." - Pranav Motarwar
    "The data which was generated by humans from humanity till the year 2008 is currently generated in a day—that's how the volume is exploding." - Pranav Motarwar
    "There are two main aspects to data engineering right now: AI for data and data for AI, and both things are essential for an engineer to plan their future." - Pranav Motarwar
    "You can't say that you should focus on AI for data rather than data for AI because both are going to be very much important for the next couple of years." - Pranav Motarwar
    "Companies like Apple and Tyro are raising relevant job applications in the market known as AI data engineer, with requirements around creating data pipelines for agents and using AI agents in your data engineering flow." - Pranav Motarwar
    "Traditionally, we were consuming and processing data in a very structured format, but now that is getting transformed for agents, where it will be unstructured files, audios, videos—it can be pretty much anything." - Pranav Motarwar
    "If you want to cope with market dynamics, you need to understand the requirements in the market and gauge your skills according to the market dynamics." - Pranav Motarwar
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review

    Resources

    LinkedIn Profiles:

    • Pranav Motarwar's LinkedIn: https://www.linkedin.com/in/pranav-motarwar-648a55169
    • Benjamin's LinkedIn: https://www.linkedin.com/in/wagjamin

    Company Websites:

    • Firebolt: firebolt.io

    Tools & Platforms:

    • DBT – Data transformation and modeling tool for building analytics engineering workflows
    • Fivetran – Data integration platform for automating data pipeline ingestion
    • Snowflake – Cloud-based data warehouse for structured and unstructured data processing
    • Databricks – Unified data analytics platform supporting ETL, data science, and AI workloads
    • BigQuery – Google Cloud's data warehouse for analytics and machine learning
    • Looker – Business intelligence and visualization platform
    • Cortex – Snowflake's AI-powered tool for data pipeline automation
    • LangChain – Framework for building applications with language models and data processing layers
    • Vespa – Retrieval engine for fast vector search and multimodal data serving
    • AdaptDB – Analytical database system for building software products
    Articles & Research Papers:

    • "MIT Technology Review Report on Data Engineering and AI" – Co-published with Snowflake (2023-2025 projections on AI use cases in data engineering)
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    AI Won't Replace Engineers, But This Framework Will Change How They Build with Rohit Girme May 07, 2026
    Show notes Scaling AI from proof-of-concept to production requires more than just deploying models; it demands robust evaluation frameworks, human oversight, and a fundamental shift in how engineering teams approach development.
    In this episode of The Data Engineering Show, host Benjamin Wagner sits down with Rohit Girme, Staff Software Engineer at Airbnb, to explore how Airbnb built a Gen AI evaluation platform to assess LLM outputs across product surfaces, from customer support bots to search and booking experiences. Rohit shares insights into Airbnb's infrastructure choices, evaluation workflows, and lessons learned about leveraging AI tools while maintaining human orchestration.
    What You'll Learn:

    - How to architect a multi-layer Gen AI evaluation platform using Python, VLLM, Kubernetes, and DAG-based workflows to systematically test LLM outputs in production
    - Why splitting monolithic "virtual judges" into specialized LLM-powered metrics (content relevance, hallucination detection, policy adherence) dramatically improves evaluation accuracy and debugging
    - The critical distinction between real-time evaluation (lightweight, sub-second latency) and offline evaluation (comprehensive, human-in-the-loop) and how to route outputs accordingly
    - How to shift from traditional software engineering (deterministic, rule-based testing) to probabilistic AI evaluation where you validate outputs against golden datasets and human judgment benchmarks
    - The framework for breaking down problems into smaller chunks and using AI tools as collaborators rather than end-to-end problem solvers—critical when working with codebases at massive scale
    - Why documentation becomes infrastructure in an AI-driven workflow: LLMs need comprehensive, well-formatted docs to scale tribal knowledge across entire organizations
    - The hard truth about AI and scaling: zero-to-one innovation is now commoditized, but one-to-n execution (the scaling part) still demands human judgment, orchestration, and product sense
    - How to measure AI tool adoption beyond token usage instrument your development workflow to capture whether LLM suggestions actually made it into shipped code and added real value

    About the Guest(s)

    Rohit Girme is a Staff Software Engineer at Airbnb, where he has spent the last seven and a half years building infrastructure and platforms at scale. With deep expertise in search and machine learning infrastructure, Rohit leads efforts in GenAI evaluation and has pioneered Airbnb's approach to ensuring AI-powered features work reliably in production. In this episode, Rohit shares practical insights on building evaluation platforms for large language models, orchestrating AI in product workflows, and leveraging AI tools effectively in software development. His work on integrating LLMs into customer-facing products while maintaining quality and performance provides actionable strategies for engineering teams navigating the rapid adoption of AI, making this conversation essential for data engineers and platform builders looking to scale AI responsibly.

    Quotes

    "Zero to one is easy now, but the one to n, which is a scaling part, I think we still haven't figured that out. You still need humans for that." - Rohit
    "With AI, it's a black box to us as well. We don't know how it's working underneath, so we have to figure out another way to evaluate the surface." - Rohit Girme
    "Humans should be the orchestrators of these tools and not just hand off everything to these tools." - Rohit Girme
    "If we hand off everything to the LLM, it will make a lot of assumptions because context is limited, and it doesn't know the code enough." - Rohit Girme
    "Documentation has become even more relevant because now LLMs need to know everything so everyone can scale up." - Rohit Girme
    "Measuring productivity in LLMs is not just about how many tokens people are using—you need to figure out if they're actually building something on top." - Rohit Girme
    "Internet democratized information, and I think with LLMs, it's capability that would be democratized. If you have a good idea, you can build it very quickly." - Rohit Girme
    "There's always going to be blind spots for every person, but with AI, it'll become even faster because you have this very short cycle of talking to the AI instead of talking to five humans." - Rohit Girme
    "Shipping products or shipping features would become even faster—where earlier it took weeks or months, now it will be days." - Rohit Girme
    "I have supercharged my workflow day to day either at work or at home with access to information that's so easy to get." - Rohit Girme
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review

    Resources

    LinkedIn Profiles:

    • Rohit Girme's LinkedIn: https://www.linkedin.com/in/rohitgirme/
    • Benjamin's LinkedIn: https://www.linkedin.com/in/wagjamin

    Company Websites:

    • Airbnb: airbnb.com
    • Firebolt: firebolt.io

    Tools & Platforms:

    • VLLM – Open source inference framework for hosting and running LLM-based inference engines
    • Kubernetes – Container orchestration platform used for serving infrastructure
    • Apache Airflow – DAG-based workflow orchestration tool (originated from Airbnb)
    • GitHub Copilot – AI-powered code completion tool for software development
    • Claude – LLM tool referenced for code generation and development assistance
    Cloud Services:

    • Azure – Hosted LLM services used at Airbnb
    • AWS – Hosted LLM services used at Airbnb
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    The Framework Canva Uses for 200M+ Designers with Paul Tune Apr 28, 2026
    Show notes AI agents are moving beyond simple automation into collaborative design workflows requiring fundamentally different approaches to user experience, model training, and infrastructure than traditional ML systems.
    In this episode of The Data Engineering Show, host Benjamin Wagner sits down with Paul Tune, Staff Research Scientist at Canva, to explore how the design platform is building agentic workflows, managing multimodal data pipelines, and tackling the unique challenge of teaching machines to understand aesthetic taste alongside functional design.
    What You'll Learn:

    • How to architect user experiences that match intent across expertise levels from seventh graders to professional designers by constraining uncertainty through progressive disclosure rather than forcing upfront specification
    • Why reinforcement learning infrastructure for creative tasks demands different optimization priorities than supervised fine-tuning, with network latency to external API services often dominating compute efficiency
    • The shift in modern ML workflows from "source data → train → deploy" to a verification and evaluation-first paradigm, especially for generative models where training cycles are measured in weeks, not hours
    • How to split ML team responsibilities across data sourcing, supervised fine-tuning, distributed systems tuning, and evaluation with evaluation becoming the critical path as model capabilities scale
    • The difference between LLM inference bottlenecks (token throughput is rarely limiting) versus image-based ML pipelines (where data movement and GPU saturation drive entirely different optimization equations)
    • Why aesthetic evaluation remains harder than mathematical verification and how 2024 will likely see meaningful progress in applying LLMs beyond verifiable domains like coding into subjective areas like design taste
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.

    About the Guest(s)

    Paul is a Staff Research Scientist at Canva, bringing nine years of experience building machine learning systems that empower millions of users worldwide. With deep expertise in large language models, reinforcement learning, and generative AI applications, Paul leads Canva's post-training efforts on LLMs designed for agentic design workflows. In this episode, Paul shares insights into how modern ML teams balance competing priorities - from data efficiency and GPU optimization to evaluation frameworks for subjective tasks like design aesthetics. His work bridging the gap between casual users and professional designers offers valuable lessons for data engineers and ML practitioners looking to scale AI systems across diverse user bases and complex product surfaces.
    Quotes
    "What Canva is is that online graphic design platform for you to be able to design and kinda have this whole end-to-end process from the designing, the brainstorming, and using all sorts of tools in order to create a graphic design." - Paul Tune
    "The whole vision really is to empower the world to design, and what that entails is to then have this entire end-to-end experience of designing on the platform." - Paul Tune
    "I think a lot of it has to do with matching intent, so even for yourself, if you're using cloud code, some folks go down the side of, I want to plan very specific things about my design." - Paul Tune
    "Whereas for a more casual user, they probably do come in without really having an idea of what they actually do want in the first place, and I think having a few options to kind of show, okay, these are kind of like a few designs that you might like as part of that, so that sort of helps to then eventually narrow down the intent." - Paul Tune
    "I think there's a lot of very strong momentum around generative tools right now, and as part of that, Canva is also experimenting with adding generative tools within the product." - Paul Tune
    "I think one of the bigger trends this year is there's been quite a bit of buzz around agents in particular, and Canva is no different in that aspect—we are working towards agentic workflows." - Paul Tune
    "I think for us, the biggest challenge is that every time we do a rollout by an RL algorithm where we do have a sample that needs to be scored and then some level of feedback goes back to the model to then update its weights, we have to heat up specific APIs within different services at Canva." - Paul Tune
    "I think the change has definitely shifted from when we started work on machine learning, where you kind of source the data and then train, to really like how do you evaluate because large language models have so many capabilities." - Paul Tune
    "I think I try to keep focus because I don't think it's very feasible for me to cover every paper out there, even though there are lots and lots of exciting things that happen every day." - Paul Tune
    "I think what I'm particularly excited about is applying these sorts of models into domains that are beyond what is very strongly verifiable, like mathematics and coding, because progress outside these domains has been a bit slower, but I do see at least some progress over time." - Paul Tune
    Resources
    Connect on LinkedIn:

    • Paul Tune - https://www.linkedin.com/in/paul-tune-0ba18116
    • Benjamin Wagner - https://www.linkedin.com/in/wagjamin

    Websites:
    • Canva: https://www.canva.com
    • Canva Engineering Blog: https://www.canva.dev/blog/

    Tools & Platforms:
    • Ray – Distributed training framework for machine learning on Kubernetes clusters
    • Argo – Workflow orchestration tool for managing data pipelines and model training
    • Snowflake – Data warehouse for structured data storage and event management
    • AWS S3 – Object storage for media files and unstructured data
    • Kubernetes – Container orchestration platform for managing distributed training clusters
    • RDS with MySQL – Relational database service for backing services
    • Canva Magic Studio – Generative AI tools suite within Canva, including image generation and LLM-powered writing assistance

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Llama 2 & 3 Safety: Soumya Batra on Agentic AI Training Apr 08, 2026
    Show notes In this episode of The Data Engineering Show, host Benjamin Wagner sits down with Soumya Batra, founder and CEO of WisePort AI and former tech lead at Meta where she led safety efforts for Llama 2 and Llama 3, to explore the evolution of NLP, the complete lifecycle of foundation model training, and why the next AI frontier lies in natively agentic systems rather than simply scaling larger transformers.
    What You'll Learn:

    • Why historical NLP work becomes obsolete with each paradigm shift: Understand how Bayesian networks, RNNs, and LSTMs each dominated until replaced - and why current transformer-scaling dogma will likely face the same fate
    • How to structure the foundation model training lifecycle for safety: Learn the three critical phases - pretraining (data mix optimization), supervised fine-tuning (instruction alignment), and reinforcement learning (human preference integration)—and where safety interventions deliver maximum leverage
    • The counterintuitive data strategy for pretraining safety: Discover why removing all toxic content actually weakens model robustness, and how maintaining a precise balance preserves the model's ability to classify and refuse harmful requests
    • How dual reward models maximize both helpfulness and safety: See why combining helpfulness and safety objectives (as done in Llama 3) ensures every training sample reinforces both capabilities simultaneously rather than creating trade-offs
    • What "natively agentic" means and why it matters more than LLM-powered agents: Learn how foundational agentic models dynamically explore action spaces at inference time instead of relying on fixed developer-defined scaffolding, unlocking domain-agnostic workflows
    • How to build a foundational AI startup without massive training datasets: Understand why synthetic data generation, deterministic task validation, and deep domain expertise can substitute for Internet-scale language corpora in the agentic space
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.

    About the Guest(s)

    Soumya Batra is the Founder and CEO of WisePort AI, a foundational AI company specializing in agentic AI systems. With over twelve years of expertise in NLP and machine learning, she previously served as a Tech Lead and Applied Research Scientist at Meta, where she led safety and controllability efforts for both Llama 2 and Llama 3. Her career spans foundational work at Carnegie Mellon University, Microsoft, and Meta, establishing her as a pioneering voice in conversational AI and foundation model development. In this episode, Soumya demystifies the journey from traditional NLP to large language models, revealing how safety and controllability are embedded across the entire model lifecycle—from pretraining through reinforcement learning. Her insights on the future of agentic AI and the limitations of current scaling-only approaches provide essential perspective for data engineers and ML practitioners navigating the rapidly evolving AI landscape.
    Quotes
    "I did not know then that this would become my career for the next decade." - Soumya
    "Whatever work that I've done in the past becomes irrelevant all of a sudden." - Soumya
    "There is always a notion of, yes, this is the big thing, and then no, it's not anymore." - Soumya
    "I really think that we are going to be proven wrong once again about scaling transformers being the only way to achieve general intelligence." - Soumya
    "Safety was an issue even back then, even though we were training in such controlled settings." - Soumya
    "If you don't put some toxic content there, then it will lose the ability to classify it and it'll be much easier to break the safety later on." - Soumya
    "In the post training phase, we are giving it that ability to be able to answer users' questions." - Soumya
    "The next unlock will now come from foundational agent models that are natively agentic, which will unlock use cases that look unimaginable to us right now." - Soumya
    "Natively agentic means the foundational model itself needs to dynamically explore the action space, rather than scaffolding around existing LLMs." - Soumya
    "The real unlock comes from creating your own use cases, creating your own synthetic data, and going deep into a few workflows." - Soumya
    Resources
    Connect on LinkedIn:

    • Soumya Batra - https://in.linkedin.com/in/soumyabatra
    • Benjamin Wagner - https://www.linkedin.com/in/wagjamin

    Websites:
    • WisePort AI – https://www.wiseport.ai
    • Firebolt - https://www.firebolt.io

    Articles & Research Papers:
    • LLaMA: Open and Efficient Foundation Language Models – Meta AI Research
    • Lima: Less Is More for Alignment – Stanford & Meta AI Research

    Educational Institutions:
    • Carnegie Mellon University - Language Technologies Institute (ATI)

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    The Data Fusion Secret & Why Custom Query Engines Fail with Nikita Lapkov Mar 24, 2026
    Show notes In this episode of The Data Engineering Show, host Benjamin Wagner sits down with Nikita Lapkov, Senior Software Engineer at Cloudflare, to explore the architecture, design decisions, and future roadmap of R2 SQL- Cloudflare's new R2-based distributed query engine launched in September 2024.
    What You'll Learn:

    • How to leverage existing query engines strategically: Why Cloudflare chose Apache Data Fusion for single-node query processing rather than building an analytical engine from scratch, freeing engineering resources for distributed orchestration challenges.
    • The stateless architecture pattern for global infrastructure: How to design compute nodes that hold zero persistent state by storing all metadata in a distributed catalog (Iceberg), enabling per-query worker provisioning across 300+ geographically dispersed data centers.
    • Why filter pushdown and metadata-driven pruning are non-negotiable optimizations: How to reduce data scanned from object storage before query execution begins by leveraging catalog statistics and range filtering - the foundation of R2 SQL's performance gains.
    • How to solve version compatibility at infrastructure scale: Why backward compatibility matters more than cross-version support when you can't control individual node upgrade timing, and how this constraint drives architectural decisions.
    • The shuffle strategy for point-to-point distributed joins: How to implement in-memory and disk-based shuffles within ephemeral worker clusters using network-addressable worker IDs, allowing stateless workers to forget completely after query completion.
    • Why adaptive query execution is the next frontier for petabyte-scale analytics: How collecting runtime data distribution statistics mid-query execution enables mid-flight plan reconfiguration - a technique worth the overhead investment when queries run for minutes or hours rather than milliseconds.
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review

    About the Guest(s)

    Nikita is a Senior Software Engineer at Cloudflare, specializing in distributed query engines and data platform architecture. With extensive experience in database internals gained through roles at ClickHouse, Yandex, and MongoDB, Nikita has developed deep expertise in query optimization and system design at scale. At Cloudflare, he leads the development of R2 SQL, a distributed analytical query engine built on Apache Data Fusion, serving as a critical component of Cloudflare's data platform. In this episode, Nikita discusses the architecture, design decisions, and technical challenges of building a stateless, distributed SQL engine across Cloudflare's unique 300-location infrastructure, offering valuable insights for engineers working on large-scale data systems. Their work demonstrates how thoughtful architectural choices and infrastructure constraints drive innovation in distributed database systems.
    Quotes

    "It was my crash course into OS engineering. We encouraged every possible bug in this project. It was very painful and very hard." - Nikita Lapkov
    "Collecting a stack trace is very hidden, especially if you're not writing in C or C++. It is actually a very complicated and involved process." - Nikita Lapkov
    "What excites me is that it has free egress. Usually, you would pay per gigabyte to load your data. You don't have that with R2." - Nikita Lapkov
    "What we explicitly wanted to avoid when building R2 SQL is building an analytical query engine again. We would much rather use something off the shelf and work on the interesting distributed parts." - Nikita Lapkov
    "No matter how complex the query is, you can make a case that, with extreme cases, the throughput for a single load operation is relatively constant, no matter how complex the query is." - Nikita Lapkov
    "We try to be as stateless as possible. All our state lives in the catalog itself, so we only need what's in the catalog and the query that comes from the request." - Nikita Lapkov
    "The shuffles cannot really be reused unless you do some very fancy heuristics. Once we have picked the workers for a particular query, we can think of them as our little cluster." - Nikita Lapkov
    "Joins consume your entire roadmap, and this is pretty much what will be happening with us at some point. We need to make sure that distributed joins work really well, no matter what your data distribution is like." - Nikita Lapkov
    "We have potentially minutes to spare, and optimizing some even subparts of the query is worthy investigation because it could shave hours or something like that." - Nikita Lapkov
    "Finding the safe points for replanning and doing this distributed coordination while we have 50 different workers working on different parts of the query is definitely the area we want to look at in the coming year." - Nikita Lapkov
    Resources

    Connect on LinkedIn:

    • Nikita Lapkov - https://www.linkedin.com/in/nikitalapkov
    • Benjamin Wagner - https://www.linkedin.com/in/wagjamin

    Websites:

    • Firebolt – firebolt.io
    • Cloudflare –cloudflare.com
    • Apache Arrow DataFusion –datafusion.apache.org

    Tools & Platforms:

    • R2 SQL – Cloudflare's R2-based query engine for analytical queries
    • Apache Arrow DataFusion – Analytical query engine used for single-node number crunching
    • Arroyo – Rust-based streaming solution built on DataFusion
    • R2 – S3-compatible object storage with free egress
    • Apache Iceberg – Catalog system for state management
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    How Zipline AI Turns Weeks of Engineering Into Minutes of SQL Queries ft. Nikhil Simha Mar 10, 2026
    Show notes In this episode of The Data Engineering Show, host Benjamin sits down with Nikhil Simha, CTO of Zipline AI and co-author of Chronon, to explore how a declarative feature platform solves the speed-vs-scale paradox in modern ML infrastructure, from fraud detection at Airbnb to powering OpenAI's recommendation systems.
    What You'll Learn:

    • How to eliminate the data scientist-to-ML engineer bottleneck by generating Spark, Flink, and orchestration pipelines automatically from simple SQL queries, enabling data scientists to ship features independently without waiting for engineering resources

    • Why fraud detection demands real-time feature iteration: The adversarial nature of fraud requires companies to build and deploy new detection models in days, not months- a timeline impossible with manual pipeline engineering

    • The "precompute everything" optimization principle for serving latency: Chronon minimizes query response time by batching feature computation upstream through stream and batch processing, then delivering pre-aggregated signals to models in milliseconds

    • How to safely ship feature versions in production using dual-write strategies that keep old and new feature versions running simultaneously, enabling A/B testing and instant rollbacks without service disruption

    • Why context engineering, not just RAG, powers modern LLM applications: ML model predictions (fraud risk scores, user signals, embeddings) feed directly into LLM prompts as structured context, improving decision quality for both human and AI agents

    • The critical gap in open-source data infrastructure: Modern systems need query engines that scale seamlessly from single-machine to distributed clusters - today's choice between lightweight tools (DuckDB) and heavyweight platforms (Spark) leaves mid-scale and product-embedded analytics underserved

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review

    About the Guest(s)

    Nikhil Simha is the CTO at Zipline AI, bringing extensive experience from leadership roles at Airbnb and Facebook. He is a co-author of Chronon, an open-source feature engineering platform that automates the generation of ML infrastructure from declarative queries. With deep expertise in real-time data systems, fraud detection, and feature engineering at scale, Nikhil has architected solutions powering recommendation systems and risk detection across billions of user interactions. In this episode, he shares insights on building scalable ML infrastructure, integrating LLMs with real-time feature contexts, and the evolving data engineering landscape. His work has directly impacted how organizations from early-stage startups to Fortune 500 companies approach feature engineering and real-time ML serving, making this conversation essential for engineers building production AI systems.

    Quotes

    "Fraud is adversarial. Right? Like, someone comes up with a new way to do fraud somewhere around the world, and people at Airbnb need to react to it very quickly." - Nikhil
    "Chronon, at its core, generates these systems from queries. So users write queries on Chronon, and we generate all of these under the hood." - Nikhil
    "Chronon allows data scientists to operate independently." - Nikhil
    "The main problem there was that the traditional model of data scientists writing some logic and ML engineers going and billing system out for that logic, that was too slow for fraud detection." - Nikhil
    "They have to come up with a new model in a matter of days. They don't have, like, this three to five month period where they can sit and create the new model, build all of these pipelines." - Nikhil
    "There is a real gap in the industry for an engine that goes all the way from single machine scale to thousands of machine scale seamlessly." - Nikhil
    "Most people, for ninety-five percent of their queries, don't need Spark in RPA. Right? But there is that 5% usually, like, a lot of ML falls into that." - Nikhil
    "We are handling query fragments. Right? We take query fragments, generate very specialized logic for that, and run that through Spark's distributed processing topologies." - Nikhil
    "The new trend in the industry would be, like, towards these engines that can work at any scale and be useful for interactive and large processing workloads." - Nikhil
    "I think Iceberg is great that way because you're not fragmenting to different proprietary data formats, different proprietary engines." - Nikhil

    Resources


    Connect on LinkedIn:

    • Nikhil Simha - https://www.linkedin.com/in/nikhilsimha
    • Benjamin Wagner - https://www.linkedin.com/in/wagjamin


    Websites:

    • Zipline AI – zipline.ai
    • Firebolt – firebolt.io


    Tools & Platforms:

    • Chronon – Feature engineering and real-time ML infrastructure platform for generating data pipelines from queries
    • Apache Spark – Distributed data processing engine for batch and large-scale processing workloads
    • Apache Flink – Stream processing engine for real-time data transformations
    • Redis – In-memory key-value store for feature serving
    • Apache Iceberg – Open table format for data lake storage
    • Airflow – Workflow orchestration platform for pipeline scheduling
    • DuckDB – Open-source analytical database for single-machine to moderate-scale processing
    • BigQuery – Google Cloud data warehouse
    • Snowflake – Cloud-based data warehouse platform
    • Kubernetes – Container orchestration platform

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    The Geo-Data Problem Nobody Talks About And How Voi Solved It ft. Magnus Dahlbäck Feb 19, 2026
    Show notes
    In this episode of The Data Engineering Show, host Benjamin sits down with Magnus Dahlbäck, Senior Director of Data and Platform at Voi, to explore how a rapidly scaling European e-scooter company transformed its data infrastructure, adopted a metrics-first approach to analytics, and is now leveraging AI to solve real-time operational challenges across 150 cities and 150,000 vehicles.
    What You'll Learn:
    • How to escape the "dashboard chaos" trap by adopting a metrics-first architecture with a semantic layer, reducing confusion from hundreds of conflicting dashboards to a single source of truth across the organization

    • Why replacing Tableau with Steep (a metrics-centric BI tool) unlocked self-service analytics for non-technical users, empowering teams to answer their own data questions without waiting months for custom dashboard builds

    • The real-world cost optimization challenge of managing Snowflake expenses that scale 1:1 with ride volume—and why data leaders must constantly rethink architecture to control FinOps in high-growth environments

    • How to architect for IoT at scale: processing billions of daily events from connected vehicles using micro-batch pipelines (5-minute intervals) while keeping real-time machine learning inference separate through cross-functional product teams

    • The decision framework for choosing traditional ML vs. LLMs: use traditional methods for accuracy-critical workloads (supply-demand forecasting for vehicle positioning) and LLMs for pattern discovery where 100% precision isn't required (analyzing rider feedback)

    • How to build proactive customer support powered by data and AI: leverage sensor data and ride telemetry to detect poor user experiences and reach out before customers complain, rather than waiting for refund requests

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here: https://www.fame.so/follow-rate-review.

    About the Guest(s)

    Magnus Dahlbäck is Senior Director of Data and Platform at Voi, a leading European micro-mobility company, where he oversees the data analytics team, platform infrastructure, and AI initiatives. With over four years at Voi, Magnus has scaled the data organization from three people to a comprehensive team of platform engineers, data analysts, and data scientists while architecting a modern data stack centered on metrics-first analytics and semantic layers. In this episode, Magnus shares insights on building scalable data platforms for IoT-heavy, real-world products, including strategies for managing billions of daily events, implementing self-service analytics, and balancing traditional machine learning with large language models. His work at Voi—where the data platform powers both internal analytics and customer-facing product features—demonstrates how thoughtful data architecture drives measurable business impact, making this conversation essential for data leaders navigating AI integration and data democratization.

    Quotes

    "There are hundreds of dashboards, and I'm looking for some data, some metrics, and there are 10 dashboards that contain that, and they all show different numbers." - Magnus
    "Metrics is a very natural way of interacting with data rather than dashboards that are named something randomly." - Magnus
    "We're basically throwing man hours on slicing and dicing data, trying to find patterns, anomalies that we often miss, right, because it just takes too much time." - Magnus
    "The way we work with data hasn't really changed that much in the last ten, twenty years to be completely fair, but now we're seeing new technologies, new approaches to it." - Magnus
    "It comes down to the use case. What's the accuracy we need?" - Magnus
    "We can see from the sensor data, from the IoT, from other data points during your ride if it was a good or bad experience, so why don't we reach out to you?" - Magnus
    "Building software around physical objects is really cool when you're a techie guy like me, working at a company where it's a combination of software, B to C, hardware, IoT." - Magnus
    "The biggest dataset that we process is IoT data—billions of events every day, basically, that we process." - Magnus
    "We have cross functional teams where all the product teams have everything from back end to front end to data people, designers, and so on." - Magnus
    "Metrics is kind of the business language that we use—we talk about rides, average ride charge, active vehicles—so metrics is a very natural way of interacting with data." - Magnus

    Resources


    Connect on LinkedIn:

    • Magnus Dahlbäck - https://www.linkedin.com/in/magnusdahlback/
    • Benjamin Wagner - https://www.linkedin.com/in/wagjamin/


    Websites:
    • Guest's Company: Voi Technologies Website (voi.com)
    • Host's Company: Firebolt Website (firebolt.io)

    Tools & Platforms:
    • Snowflake – Data warehouse for analytics and machine learning workloads
    • DBT (Data Build Tool) – Data transformation and modeling
    • Apache Airflow – Workflow orchestration
    • Steep – Metrics-first BI tool with semantic layer (Swedish startup)
    • GCP Vertex AI – Machine learning platform for model training and deployment

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    1 2 3 7 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights