TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    The Data Engineering Show

    The Data Engineering Show is a podcast for data engineering and BI practitioners to go beyond theory. Learn from the biggest influencers in tech about their practical day-to-day data challenges and solutions in a casual and fun setting.

    SEASON 1 DATA BROS
    Eldad and Boaz Farkash shared the same stuffed toys growing up as well as a big passion for data. After founding Sisense and building it to become a high-growth analytics unicorn, they moved on to their next venture, Firebolt, a leading high-performance cloud data warehouse.

    SEASON 2 DATA BROS
    In season 2 Eldad adopted a brilliant new little brother, and with their shared love for query processing, the connection was immediate. After excelling in his MS, Computer Science degree, Benjamin Wagner joined Firebolt to lead its query processing team and is a rising star in the data space.

    For inquiries contact tamar@firebolt.io
    Website: https://www.firebolt.io

    Advertise

    Copyright: © 2024 The Firebolt Data Bros

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    Why 99% of Data Teams Give Up on Real-Time And How Artie Changes That Feb 03, 2026
    Show notes In this episode of The Data Engineering Show, Benjamin sits down with Artie CTO and co-founder Robin Tang, to explore the complexities of high-performance data movement. Robin shares his journey from building Maxwell at Zendesk to scaling data systems at Open Door, highlighting the gap between business-oriented SaaS connectors and the rigorous demands of production database replication.
    Robin dives deep into Artie’s architecture, explaining how they leverage a split-plane model (Control Plane and Data Plane) to provide a "Bring Your Own Cloud" (BYOC) experience that engineering teams actually trust. You’ll hear about the technical nuances of CDC, from handling Postgres TOAST columns to the "economy of scale" challenges of processing billions of rows for Substack, Artie’s first customer. Whether you're struggling with real-time ingestion costs or curious about the future of platform-agnostic partitioning, this conversation provides a masterclass in modern data movement.
    What You'll Learn:

    • Why the data movement market is bifurcating: Managed vendors like Fivetran excel at SaaS integrations (hundreds of connectors), while specialized vendors like Artie focus on production databases at high volume - a fundamentally different job to be done requiring expertise in failure recovery, observability, and advanced use cases.
    • How to design CDC architecture that doesn't break production databases: Use online backfill strategies (DB log framework) instead of long-running transactions that hold write locks; implement table-level parallelism so a single table error doesn't halt the entire pipeline.
    • The split-plane architecture pattern for flexible deployment models: Build control plane and data plane separation from day one, allowing customers to choose between fully managed cloud deployments or bring-your-own-cloud (BYOC) without compromising UX or architecture.
    • Why database-specific expertise matters more than breadth: SQL Server CDC requires reverse engineering undocumented code; Postgres has TOAST columns; MongoDB allows invalid timestamp values - each data source has hidden complexity that justifies deep specialization over connector sprawl.
    • How to build trust with early-stage customers on mission-critical workloads: Walk prospects through architecture and failure modes before implementation; encourage them to stress-test with real data volumes; establish deep engineering partnerships where both teams debug problems together (not sales-driven relationships).
    • The platform-specific optimization trap and how to solve it: Instead of requiring customers to understand nuances of BigQuery time partitioning vs. Snowflake's lack thereof, build platform-agnostic features (like soft partitioning) that work consistently across destinations while handling platform-specific optimizations under the hood.

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.

    About the Guest(s)

    Robin is the CTO and cofounder of Artie, a data movement platform built for high-volume, low-latency production database replication. With over a decade of experience building large-scale data systems, including early work on Maxwell (an open-source CDC framework at Zendesk) and database architecture at venture-backed startups, Robin identified a critical gap: existing tools optimize for SaaS integrations, not production databases at scale. In this episode, Robin shares hard-won lessons from building mission-critical infrastructure, including architectural innovations that prevent data loss and failure modes that only surface under real-world production load. His work at Artie has powered reliable data replication for companies like Substack, making this conversation essential for engineering teams building or evaluating real-time data movement solutions.

    Quotes

    “Artie helps companies make data streaming accessible." - Robin
    "I didn't want to make any sort of compromises and it just turned out to be a really hard problem, so then we started a company around this." - Robin
    "The complexity is not just at the destination level, the complexity is also at the source level." - Robin
    "Every pipeline that we touch is mission critical for customers, or else they would just use either their existing pipeline or a managed vendor that's out there." - Robin
    "We handle the whole thing, whereas other vendors more or less provide a component and expect engineers to either build or attach additional pieces." - Robin
    "I think the biggest bottleneck for real time right now is accessibility. When people think about real time, they immediately think it's not worth it because they implicitly have a cost associated with it." - Robin
    "We use Kafka transactions, so we do not commit offsets until the destination tells us the data has actually been flushed." - Robin
    "There's so much nuance with every single data source that it becomes a whack-a-mole problem." - Robin
    "When there's sufficient pain on the other side and they buy into your vision, it's easier to overcome obstacles during technical implementation." - Robin
    "We're spending more time developing platform-agnostic solutions so customers don't have to understand platform nuances." - Robin
    Resources

    Connect on LinkedIn:
    • Robin Tang - https://www.linkedin.com/in/tang8330/
    • Benjamin Wagner - https://www.linkedin.com/in/wagjamin/


    Websites:
    • Artie: https://www.artie.com/
    • Fivetran: https://www.fivetran.com
    • Estuary: https://www.estuary.dev
    • Airbyte: https://airbyte.com
    • Debezium: https://debezium.io

    Tools & Platforms:
    • Maxwell – Open source CDC framework for MySQL to read binlog into Kafka
    • Kafka – Distributed event streaming platform for data movement
    • WarpStream – Cost-optimized Kafka alternative using object storage
    • Streamsy – Kubernetes-native Kafka deployment tool
    • Apache Iceberg – Open table format for data lakehouse architecture
    • Delta Live Tables – Databricks' data movement and transformation tool
    • ClickPipes – ClickHouse's native data ingestion platform
    • Snowpipe Streaming – Snowflake's real-time data ingestion service
    • Google Datastream – Google Cloud's CDC and data movement service
    • AWS MSK Tiered Storage – Amazon managed Kafka with tiered storage capabilities

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    The $100M Problem: How Lyft's Data Platform Prevents ML Failures with Ritesh Varyani at Lyft Dec 16, 2025
    Show notes In this episode of the Data Engineering Show, host Benjamin Wagner sits down with Ritesh Varyani, Staff Software Engineer at Lyft, to explore how the company manages a sophisticated multi-engine data stack serving thousands of engineers, while simultaneously integrating AI across infrastructure and user-facing analytics.
    What You'll Learn:

    • How to architect a polyglot data platform that serves fundamentally different workloads, Spark for ML training and massive parallel processing, Trino for dashboarding and medium-scale ETL, and ClickHouse for sub-second OLAP queries without creating operational chaos
    • Why unification matters more than expansion: Lyft's 2026 strategy prioritizes consolidating and simplifying the data stack rather than adding new tools, reducing maintenance burden and improving reliability for end users
    • The dual-layer AI strategy that simultaneously enhances user analytics (semantic layer v2 with AI-native support) while automating platform operations (intelligent job failure diagnosis, adaptive resource allocation, and agentic workflow optimization)
    • How to fund innovation from the bottom-up: Lyft's model encourages individual engineers to experiment with AI on their own time, prove business value through POCs, and secure leadership buy-in through demonstrated alignment with company strategy
    • Why vendor selection now includes AI explainability and debuggability as standard RFP requirements, even when AI isn't the primary driver of a purchasing decision
    • The framework for deciding open-source investment vs. managed services: Prioritize business-critical goals first, then determine whether in-house ownership or vendor solutions accelerate that mission, AI becomes the accelerant, not the decision driver
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    About the Guest(s)

    Ritesh is a Staff Software Engineer at Lyft, bringing six years of experience architecting and scaling the company's data platform. With a background spanning Microsoft's data and cloud infrastructure, including work on Hadoop, Azure, and SaaS products. Ritesh leads Lyft's critical data systems including Trino, Spark, and ClickHouse. In this episode, Ritesh shares insights on building scalable, AI-native data platforms that serve diverse organizational needs, from batch processing and analytics to real-time marketplace operations. His strategic approach to unifying complex data stacks while integrating AI-driven reliability and user experience improvements provides actionable guidance for data engineers and platform leaders navigating infrastructure modernization at scale.
    Quotes

    "The goal of our platform is to give our users access to the data as fast as possible so that they can drive the meaning from the data that they are getting and take better data driven decisions." - Ritesh
    "We are a Hive format shop. We are going to be moving to other open table formats in the future, but at this point, we are a hive table format." - Ritesh
    "Our main goal at this point is primarily understanding how we see the data platform running five years from now, three years from now, and how we are able to future proof it." - Ritesh
    "In this world of AI, we should not be falling behind in any way, and bringing AI in the right places within our platform." - Ritesh
    "We want to make our semantic layer ready for the AI native side of things so that our teams are able to drive the best meaning possible from the data that they see." - Ritesh
    "Big data systems are distributed systems by nature, and where AI can help you is very clearly understand how the patterns are changing and what is a good action to take." - Ritesh
    "Rather than thinking of this as an AI versus an open source thing, it's about a question of what work is the most business critical and how do you go 100% behind it." - Ritesh
    "Not everybody is working on AI initiatives at this point, but where it makes sense according to our business strategy, if it aligns with it, then obviously we go and invest." - Ritesh
    "If you are the one who's going to take on the initiative, probably spend a few hours outside of what you're already working on, and that is how you will discover AI and the tooling for it." - Ritesh
    "We are trying to consolidate into a single direction of providing different kinds of models so that you are easily able to integrate and focus on the value you want to provide to your customers." - Ritesh
    Resources
    Connect on LinkedIn:

    • Ritesh Varyani - https://www.linkedin.com/in/riteshvaryani/
    • Benjamin Wagner - https://www.linkedin.com/in/wagjamin/
    • Eldad Farkash - https://www.linkedin.com/in/eldadfarkash/

    Websites:
    • Lyft - https://www.lyft.com

    Tools & Platforms:
    • Apache Spark – Batch processing engine for ML training jobs, large-scale data processing, and GDPR operations
    • Trino – Query engine for BI dashboarding, ETL workflows, and SQL-based data access
    • ClickHouse – Columnar database for sub-second query latency and real-time analytics
    • Amazon S3 – Data lake storage for parquet tables and offline data processing
    • AWS EKS (Elastic Kubernetes Service) – Kubernetes infrastructure for hosting Spark and Trino
    • ClickHouse Cloud – Managed ClickHouse offering used by Lyft
    • Hive Table Format – Current table format for organizing parquet files in S3
    • Kubernetes Operators – Infrastructure for managing ClickHouse deployments

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    60 Billion Predictions Daily: Inside Credit Karma’s Agentic Data Layer with Maddie Daianu Nov 19, 2025
    Show notes What does MLOps look like when you are deploying 60 billion machine learning predictions a day?
    Maddie Daianu, Head of Data and AI at Intuit Credit Karma, joins the Data Bros to pull back the curtain on one of the most high-volume data environments in FinTech. With a 100-person team serving 140 million members, standard data practices break down.
    Maddie shares how her team manages terabytes of daily data on Google Cloud and explains the massive strategic pivot they are undertaking right now: The move from "Information" to "Agency."
    What You'll Learn:

    • Extreme Scale: How to architect a system that handles 80 billion daily predictions without latency.
    • The Unified Consumer Profile: The hackathon project that unlocked real-time personalization across Credit Karma and TurboTax.
    • The "Done-For-You" Future: Why they are building an "Agentic Data Layer" to move from recommending financial products to actively managing them for the user.
    If you want to know what the future of high-scale AI infrastructure looks like, this is the blueprint.
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    About the Guest(s)

    Maddie Daianu is the Head of Data and AI at Intuit Credit Karma, where she leads the teams responsible for AI science, machine learning engineering, data engineering, and the experimentation platform. She brings a background that spans academic research in biomedical engineering and machine learning, and experience at both smaller companies and Meta. Her current focus is on building the data and AI infrastructure that drives highly personalized financial experiences for Credit Karma's 140 million members and contributes to Intuit's broader consumer ecosystem.
    Quotes

    "The key elements and ingredients of making this app successful is data and AI." - Maddie
    "We have and we process and transform multiple terabytes of information daily for our 140,000,000 members every single day." - Maddie
    "We have our models that essentially, lead to almost 60,000,000,000 daily predictions for our 140,000,000 member base every single day." - Maddie
    "We want to take this to the next level. So Intuit as a whole believes... in creating done for you experiences for our users." - Maddie
    "If you don't structure your data in a semantically, well structured way, you are not likely able to provide the most highly relevant and personalized experiences for users." - Maddie
    "One thing that we've been building, in the last year or so it's called the unified consumer profile." - Maddie
    "Intuit has been investing in tremendously over the last, few years... the generative AI operating system... to move fast and continuously disrupt ourselves, especially in the age of AI." - Maddie
    ResourcesConnect on LinkedIn:

    • Maddie Daianu
    • Benjamin Wagner
    • Eldad Farkash

    Websites:

    • Credit Karma

    Tools & Platforms:

    • BigQuery – Data warehouse for processing multiple terabytes of information daily
    • Bigtable – Operational serving layer for real-time data access
    • Vertex AI – Machine learning platform for model training and deployment
    • Alchemy – Feature online feature store for real-time transformations and aggregations
    • Generative AI Operating System – Centralized platform for democratizing Gen AI adoption across Intuit products

    Products & Services Mentioned:

    • TurboTax – Tax preparation and filing software
    • Debt Agent – AI-powered tool for debt consolidation and management assistance
    • Unified Consumer Profile – Semantic graph depicting financial journey across Credit Karma and TurboTax

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Block Bad Data Before the Write with Nike’s Ashok Singamaneni Oct 07, 2025
    Show notes In this episode of The Data Engineering Show, Benjamin and Eldad are joined by Ashok Singamaneni, a Principal Data Engineer at Nike. Ashok dives deep into his work on the open-source projects BrickFlow and Spark Expectations. He shares his journey from mechanical engineering to data engineering and the lessons learned over a decade of tackling production data quality issues that lead to costly recomputes.
    Ashok explains the philosophy behind Spark Expectations: treating the ingestion and transformation layers of a data pipeline (Bronze/Silver) as a software product rather than just a data engineering product. This means implementing rigorous checks like data quality, unit testing, and integration testing before the data is written to the final layer. He details the implementation using a Python decorator pattern within Spark jobs, allowing engineers to define rules that check for everything from basic column validation to complex referential integrity and aggregation consistency. The discussion also covers the trade-offs of using generative AI tools like Cursor for data engineering and the growing industry trend of prioritizing upfront data quality due to the rise of AI-powered analytics and direct leadership access to data.
    What You'll Learn:

    • Why the ingestion and transformation layers (Bronze/Silver) of a data pipeline should be treated as a software product with rigorous testing.
    • How Spark Expectations moves data quality checks to before data is written to the final tables to prevent mission-critical failures and recomputes.
    • The three types of checks in Spark Expectations: row-level, aggregation-level, and query DQ (for referential integrity).
    • How the tool handles failures with options to ignore, drop the record, or fail the entire job.
    • Why big data quality is becoming a prime focus across the industry due to AI integrations and direct executive-level access to data.
    • Ashok’s lessons on using Generative AI tools (like Cursor/Cloud Code) in data engineering projects and the necessity of restrictive permissions.
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    About the Guest(s)

    Ashok Singamaneni is a Principal Data Engineer at Nike, with over twelve years of experience in the data space across the banking, healthcare, and retail domains. He is the creator of the popular open-source frameworks Spark Expectations and BrickFlow, which focus on improving data quality and pipeline reliability. Ashok advocates for treating data ingestion and transformation as a software product, ensuring checks and balances are in place early in the pipeline. He holds a background in mechanical engineering.
    Quotes

    "DLT expectations gave an idea to the industry that you can do data quality before actually writing the data into your final tables." - Ashok
    "I think over the time, in my experience, what I learned is this ingestion layer and the transformation layer, you should treat that as a software product, not like a data engineering product." - Ashok
    "If it's mission critical, then you fail the job, not process the data, and don't put that data into the final table so that you don't need to recompute that again." - Ashok
    "As the scale of the product increases, it becomes even more difficult for us to find exactly where the issue went wrong... it takes time for you to debug and see, like, lot of human effort also involved." - Ashok
    "Data observability and quality is becoming prime because of AI integrations that are happening." - Ashok
    "Ultimately, at the end of the day, you are responsible when you're checking in the code. It's not Claude or Karsar that will be blamed if something goes wrong." - Ashok
    "The leadership is directly looking at the data and if there is something wrong in the data, then there can be some serious repercussions happening on the business decisions." - Ashok
    "Rather than having bad data in the tables and then recomputing or reclarifying things, let's not put that data first in the first place." - Ashok
    "You can drop the record and put that in an error table and give that alert to the engineering team that there is some error in the error table you can look at." - Ashok
    "The road eq checks that happens are very fast. It should happen as a pretty standard checks that happens on the scale." - Ashok
    Resources

    Projects:

    • Spark Expectations - Data quality framework
    • BrickFlow - Open source project for data pipelines
    Tools & Technologies:

    • Apache Spark
    • Databricks DLT (Delta Live Tables)
    • Great Expectations - Post-processing data quality tool
    • Cursor / Cloud Code - Generative AI coding tools
    • SQLMesh
    For Feedback & Discussions on Firebolt Core:

    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository
    • Benjamin@Firebolt.io
    • Eldad@Firebolt.io

    Primary Speakers:

    • Ashok Singamaneni
    • Benjamin Wagner
    • Eldad Farkash

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Postgres vs. Elasticsearch: The Unexpected Winner in High-Stakes Search for Instacart with Ankit Mittal Sep 17, 2025
    Show notes In this episode of The Data Engineering Show, Benjamin Wagner sits down with Ankit Mittal, former Senior Engineer at Instacart, to explore how they revolutionized their search infrastructure by transitioning from Elasticsearch to PostgreSQL. Learn how Instacart tackled the unique challenges of fast-moving grocery inventory, achieved high-performance search capabilities, and leveraged PostgreSQL extensions for complex retrieval operations. Whether you're scaling search functionality or optimizing database performance, this deep dive offers valuable insights into building robust, production-ready search systems using PostgreSQL.
    • Discover why Instacart moved from Elasticsearch to PostgreSQL for retailer search
    • Learn about handling real-time inventory updates and search optimization
    • Explore PostgreSQL extensions, sharding strategies, and data flow architecture
    • Understand the trade-offs between different search infrastructure approaches

    What You'll Learn:

    • How Instacart managed fast-moving grocery inventory data by consolidating search, ranking, and filtering into a single PostgreSQL cluster
    • Why pushing compute closer to the data layer can significantly improve search performance and reduce network calls
    • The architecture decisions behind using PostgreSQL extensions like PG Vector and custom solutions for search functionality
    • How to implement efficient data ingestion through S3-based pipelines and bulk writes instead of real-time updates
    • Why table maintenance operations like PGD pack are crucial for optimizing read throughput in production environments
    • The trade-offs between traditional search engines and relational databases for complex search implementations
    • The challenges of maintaining self-hosted PostgreSQL in a predominantly cloud-managed environment
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    About the Guest(s)

    Ankit is a Software Engineer at ParadeDB and former Senior Engineer at Instacart, where he specialized in PostgreSQL infrastructure and search systems. With extensive experience in database optimization and search architecture, he played a key role in modernizing Instacart's search infrastructure by transitioning from Elasticsearch to a custom PostgreSQL solution. In this episode, Ankit shares deep insights into building and scaling high-performance search systems for e-commerce, particularly focusing on the unique challenges of grocery retail's fast-moving inventory. His work at Instacart revolutionized their single-retailer search functionality, demonstrating how traditional relational databases can be adapted for complex search operations. His expertise in database systems and their practical applications in high-scale environments makes this conversation particularly valuable for engineers interested in modern search architecture and database optimization.
    Quotes

    "Think about it. If there's a lot of things that you can get the database to do, then the applications become simpler." - Ankit
    "My non-Instacart experience has largely been in pre-PMF startups where the approach of abuse your database to its absolute limits works wonders." - Ankit
    "Almost everything that we got retrieved had to be filtered out. So we go back to Elasticsearch again." - Ankit
    "We traded off the quality of retrieval, hardcore core retrieval, with the whole system reducing the network calls." - Ankit
    "It's a place to go to find what item is available, in what store, what item is available, at what price, including full product taxonomy graph and product and ontology." - Ankit
    "The grand theme here is that we wanted more control over the cluster, how to spin it off, what kind of disks it would have." - Ankit
    "We tell teams who want to have their data in this cluster, create an s3 home, create either a bucket or a home, whatever they want to do, and tell us that we would sync ourselves." - Ankit
    "What we found is that the read throughput, we can throw more data if the tables are repacked nicely." - Ankit
    "Most engineers who want to work on search, they are more used to the Elasticsearch shape of the query." - Ankit
    "The relevance is better because they could join more things in the database. They also saw the cost of the normalized data reduced." - Ankit
    Resources

    Company Websites:
    - Instacart - Grocery delivery platform
    - ParadeDB - Database technology company
    - Firebolt - Cloud data warehouse (firebolt.io)
    Tools & Technologies:

    - PostgreSQL - Database system
    - Elasticsearch - Search engine
    - PG Cat/PG Dog - PostgreSQL proxy tools
    - PG Vector - PostgreSQL vector extension
    - PG Repack - PostgreSQL table repacking tool
    - ClickHouse - Column-oriented DBMS
    - TantiVy - Rust-based search engine library
    Articles:

    - Instacart Search Modernization Blog Posts (Series on hybrid retrieval)
    - Target's AlloyDB Migration Blog Post
    For Feedback & Discussions on Firebolt Core:

    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository
    • Benjamin@Firebolt.io

    Primary Speakers:

    • Ankit Mittal
    • Benjamin Wagner
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Is Self-Service BI a False Promise? Lei Tang of Fabi.ai Thinks So Aug 28, 2025
    Show notes Explore the future of AI-powered business intelligence with Lei Tang, CTO and Co-founder of Fabi.ai, as he discusses the evolution from traditional self-service BI to "Vibe-analytics." Learn how AI is transforming data accessibility, enabling anyone to perform sophisticated analytics without deep technical expertise. From building trust in AI-generated insights to creating intelligent semantic layers, discover how modern BI platforms are bridging the gap between data teams and business stakeholders. Tune in to understand why static dashboards are becoming obsolete and how AI agents will soon proactively surface business opportunities and insights.
    Key points:
    • The limitations of traditional self-service BI and how AI is addressing them
    • Building secure, context-aware AI systems for data analysis
    • The future of human-AI interaction in business intelligence
    • Technical insights into modern BI platform architecture
    • Vision for proactive, AI-driven business insights

    What You'll Learn:

    • Why traditional self-service BI has failed to deliver on its promises and how AI can bridge the gap
    • How to build an AI-native BI platform that combines SQL, Python, and natural language processing
    • The framework for implementing "Vibe-analytics" - a new paradigm of AI-powered visual analytics
    • Why context engineering and semantic understanding are crucial for accurate AI-driven analysis
    • How to balance security and accessibility when deploying AI-powered analytics tools
    • The future of BI platforms as proactive insight generators rather than passive dashboards
    • Why caching and stateful environments are essential for responsive AI-powered analytics
    • How to leverage AI to translate business questions into accurate technical queries while maintaining data integrity

    About the Guest(s)
    Lei is the Co-founder and CTO of Fabi.ai, where he leads the development of AI-native business intelligence solutions. With a PhD in machine learning and over a decade of experience in the data domain, Lei has held significant roles, including positions at Yahoo, Walmart, Lyft (as Director of Data Science), and Clari (as Chief Data Scientist). His expertise spans machine learning, data engineering, and business analytics, with a particular focus on making data analysis more accessible and efficient. In this episode, Lei shares insights on the evolution of self-service BI and how AI is transforming business intelligence, drawing from his experience building Fabi.ai, a platform that combines SQL, Python, and AI to democratize data analysis. His work in developing "Vibe AI" (AI-powered BI) represents a significant advancement in making complex data analysis accessible to non-technical users while maintaining data accuracy and trust.
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    Quotes
    "For the past decade, it's really difficult to make sure the self-service BI can work. And then now with AI, the worst part is that it can run properly, but the numbers are wrong." - Lei
    "If you talk to anybody working in the BI space, like self-service BI, that has been termed for maybe for the past decade. But I have to say that is a false promise." - Lei
    "We're saying that we really want those data team to be able to, like, say, what type of data is exposed to, like, say, less technical folks." - Lei
    "In order to build AI native BI, I would say the focus should be how human interact with AI." - Lei
    "We believe that, essentially, this BI system or, like, AI BI system would be more like a agent, and then it'll actually looking for, like, business opportunities and insight and surface to you." - Lei
    "The one common theme I have been experiencing is that normally would work with other business stakeholders, could be marketing, could be operations, could be sales." - Lei
    "We strongly believe that BI should be stored as code." - Lei
    "Enterprise data tends to be very noisy, very complex." - Lei
    "The semantics of itself becomes part of the context for the AI engine." - Lei
    "Most organizations, the data, like the schema, the kind of business, like metrics and logic, has been constantly evolving." - Lei
    Resources

    • Fabi.ai - AI-native BI platform
    • Firebolt (firebolt.io) - Cloud data warehouse platform
    Tools & Technologies:

    • Firebolt Core - Free self-hosted query engine
    • Looker - BI Platform
    • Tableau - BI Platform
    • Sisense - BI Platform
    • Snowflake - Data Warehouse
    • BigQuery - Data Warehouse
    • PostgreSQL - Database
    • SQL Alchemy - Database toolkit
    • Pandas - Data analysis library
    For Feedback & Discussions on Firebolt Core:

    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository
    • Benjamin@Firebolt.io

    Primary Speakers:

    • Lei Tang
    • Benjamin Wagner
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Building Uber's AI Assistant: How Genie Revolutionizes On-Call Support with Paarth Chothani from Uber Jul 22, 2025
    Show notes Journey inside Uber's innovative AI assistant "Genie" with Paarth Chotani, Staff Engineer at Uber, as he shares how they're revolutionizing on-call support using LLMs and vector search. From processing massive amounts of internal documentation to building scalable RAG pipelines, discover how Uber tackles the challenges of implementing AI assistants at scale. Get insights into the evolution from traditional chatbots to agent-based solutions, and learn practical lessons about staying current in the rapidly evolving AI landscape. Whether you're building AI-powered tools or scaling data infrastructure, this episode offers valuable perspectives on balancing innovation with real-world implementation.
    • Building and scaling RAG pipelines at enterprise scale• Evolution from traditional chatbots to AI agents• Practical insights on data processing and vector search implementation• Leveraging open-source technologies in production environments• Navigating rapid technological changes in AI development
    What You'll Learn:

    • How Uber transformed its on-call support system by building an AI assistant that searches across internal documentation, wikis, and code
    • Why combining multiple data sources with vector databases creates more accurate and contextual responses for enterprise support
    • The evolution from basic RAG implementation to agent-based architecture for handling complex support scenarios
    • How to scale AI processing pipelines using Apache Spark for large-scale data chunking and embedding generation
    • Why customization and internal data sources are crucial for enterprise AI assistant effectiveness
    • The future of AI assistants: moving from documentation lookup to automated problem resolution through multi-agent systems
    • How to balance rapid AI innovation with setting realistic customer expectations in fast-moving tech environments
    Paarth is a Staff Engineer at Uber, where he works on Michelangelo, Uber's machine learning platform. With over four years at Uber, he specializes in feature store development, online serving at scale, and GenAI implementations. He has been instrumental in developing Genie, an AI-powered on-call assistant that revolutionizes how Uber's engineering teams handle support requests and documentation access. In this episode, Paarth shares valuable insights on building and scaling RAG-based systems, vector search implementations, and the evolution of AI assistants from traditional chatbots to sophisticated agent-based solutions. His experience spanning both AWS chatbot development and current GenAI innovations at Uber offers listeners a unique perspective on the rapid advancement of AI-powered enterprise solutions.
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    Quotes
    "Think of Genie as your on-call assistant. Different infra teams have their Slack channels, and because these technologies are widely used, you have to wait a lot." - Paarth
    "What we realized is for our engineers to really get help, data sources really should be internal only because we customize lot of these open source engines for making it work at Uber scale." - Paarth
    "Instead of building a mega scale pipeline that just ingest all data sources and then keeps a central data source solution, we instead are giving users the flexibility to ingest what data sources they want." - Paarth
    "We had to scale our you can say the whole infrared layer to chunk data faster to be able to create embedding set scale." - Paarth
    "It almost felt like they're doing what EMR was doing. You have your Hadoop and big data technology, and we needed these pipelines to basically process all this data quickly." - Paarth
    "We've even evolved from just giving you the right documentation to starting to evolve into a situation where we'll also start taking actions on your behalf." - Paarth
    "That intuition that comes from building this kind of bot, I feel like that intuition came again as we were starting to see this technology come, and we're like, hey, this looks like where you can pretty much fit all these pieces together." - Paarth
    "What we have seen with several use cases is agentic genie works well when designed well, when you've analyzed the problem of which type of subproblems the bot should resolve per channel, per use case." - Paarth
    "I think having a problem in mind always helps that way, the energy is little bit focused and directed." - Paarth
    "Whatever you're building is not enough because the expectation has already gone to the next level, so the pace is too fast right now." - Paarth
    Resources
    • Companies & Platforms:
    • Uber - ML Platform & Engineering
    • Firebolt - Cloud Data Warehouse (firebolt.io)
    Tools & Technologies:

    • Michelangelo - Uber's ML Platform
    • Genie - Uber's On-Call Assistant Bot
    • Cursor - Developer IDE
    • OpenSearch - Vector Database
    • LangGraph - Agent Framework
    Notable Projects Mentioned:

    • MetaMate (Meta)
    • Query Copilot (Uber)
    • Scale at AI (Meta Meetup)
    Company Blogs:

    • Uber Engineering Blog - Genie and Query Optimization articles
    Primary Speakers:

    • Paarth Chotani - Staff Engineer, Uber
    • Benjamin - Firebolt
    • Eldad - Firebolt

    For Feedback & Discussions on Firebolt Core:
    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository
    • Benjamin@Firebolt.io

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    From Zero to 100M Users: Inside Notion’s Data Stack and AI Strategy with Sumit Gupta Jun 10, 2025
    Show notes AI's transformative impact on data engineering and analytics is reshaping how professionals create value, shifting focus from technical skills to strategic thinking and communication.
    In this episode of The Data Engineering Show, the bros talk with Sumit Gupta, Lead BI Engineer at Notion, about his journey through prominent tech companies, modern data stacks, and how AI is revolutionizing data workflows and professional development.
    What You'll Learn:
    • How modern data stacks are evolving with tools like Snowflake, dbt, Iceberg, and Hex
    • Why transferable skills are becoming more crucial than technical expertise in the AI era
    • How to leverage AI tools strategically
    • The framework for automating content creation workflows using AI tools and APIs
    • Why this is "the worst AI will ever be" and how to prepare for accelerating change
    • How to balance AI automation with authentic human connection in content creation
    • Why modern data professionals must embrace AI while maintaining ethical considerations
    • How companies like Notion are implementing AI for improved customer insights and engagement

    This episode offers valuable insights into the practical application of AI in data workflows, content creation, and professional development, while addressing both the opportunities and challenges in the evolving tech landscape.
    Highlights:
    [03:19] - Modern Data Stack Evolution for Scale
    [10:05] - AI-Powered Customer Intelligence Platform
    [15:17] - Future of Data Careers in the AI Era
    [18:49] - Automated Content Creation Workflow
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    About the Guest
    Sumit Gupta is a Lead BI Engineer at Notion, where he spearheads reporting and dashboarding initiatives for marketing and sales teams. With over a decade of experience in data and analytics, including notable roles at industry leaders like Snowflake and Dropbox, he brings deep expertise in modern data stack implementation and AI integration. In this episode, Sumit shares valuable insights on the evolution of data engineering, the impact of AI on analytics workflows, and how he leverages various AI tools to enhance productivity both professionally and as a content creator with 21,000+ Instagram followers. His unique perspective on balancing technical expertise with transferable skills in the age of AI, combined with his experience at "Bay Area Darlings" like Notion, Snowflake, and Dropbox, makes this conversation particularly relevant for data professionals navigating the rapidly evolving tech landscape.
    Quotes
    "The scariest part about the whole AI boom is this is the worst AI will ever be." - Sumit
    "If you are someone who's starting new in data field, the value of your technical skills that used to be very valuable until 2021 is not as much - your transferable skills or soft skills comes into picture." - Sumit
    "Every bit is expensive - all the servers are cheap, but when you're dealing with hundred million users and trillions of rows of data a day, you have to find that one percent saving." - Sumit
    "AI has made me a lot more productive, but at the same time, it has also made me dumber." - Sumit
    "If you are someone new, especially in data, scared of AI or skeptic of AI, I would say jump in - if you don't jump onto the bandwagon right now, you might be left out in a year or so." - Sumit
    Resources
    • Sumit Gupta LinkedIn
    • Notion Website
    • Sumit Gupta Instagram
    • Firebolt Website

    For Feedback & Discussions on Firebolt Core:
    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository
    • Benjamin@Firebolt.io

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    How Rising Wave Is Redefining Real-Time Data with Postgres Power May 07, 2025
    Show notes
    In this episode of The Data Engineering Show, host Benjamin and co-host Eldad sit with Yingjun Wu, founder and CEO of Rising Wave, to explore the evolution of stream processing systems and the innovations his company is bringing to the space.
    What you’ll learn:
    • Yingjun's journey from academic research in stream processing to founding Rising Wave, and the challenges of building trust in a new database system.
    • How Rising Wave's architecture, using S3 as primary storage, delivers second-level scalability, while other systems can take hours to scale.
    • The competitive landscape of stream processing, with Rising Wave's Postgres compatibility providing a significant advantage in ease of use.
    • How one major company reduced its CPU requirements from 20,000 to just 600 by switching from a traditional stream processing system to Rising Wave.
    • The rising importance of Apache Iceberg as a destination for stream processing output, helping companies avoid vendor lock-in.
    • How streaming systems fit into modern data stacks, especially as companies seek to avoid being locked into proprietary systems.
    Yingjun Wu is the founder and CEO of Rising Wave, a stream processing system built in Rust and designed with a cloud-native architecture. With a PhD focused on stream processing and database systems, Yingjun previously worked at Redshift and IBM Research before founding Rising Wave. His company has developed a system that achieves significant performance and resource efficiency advantages over traditional stream processing solutions, while maintaining Postgres compatibility for ease of use.
    Episode Highlights:
    The Origins of Rising Wave (00:30)

    Yingjun shares his background in stream processing from his PhD days and explains how his experience at Redshift revealed the need for better stream processing solutions, especially since many data warehouse workloads involve data ingested from streaming sources like Kinesis or Kafka.
    Building a System from Scratch (04:10)

    Yingjun describes the challenging first 2-3 years of developing Rising Wave without customers, highlighting how trust is a major barrier for new database systems. After 2.5 years, they secured their first customers, including a startup and several larger companies, which helped establish Rising Wave's credibility.
    The Current Stream Processing Landscape (07:47)

    Benjamin asks about the current stream processing space, with Yingjun positioning Rising Wave as a leader, particularly for SQL-based workloads. He highlights several key advantages of Rising Wave, including its Rust-based implementation and S3-based storage architecture.
    S3 as Primary Storage (10:27)

    Yingjun explains their decision to use S3 as primary storage from day one, despite its slowness and expense. He discusses how they've optimized for these challenges and would still make the same architectural choice today due to benefits like simplified state management and superior elastic scaling.
    The Business Model (13:52)

    Rising Wave offers open-source, cloud, and on-premise versions of its product. Yingjun notes that many highly regulated industries require on-premise deployment, including customers in the banking and aerospace sectors.
    Typical Users and Competitive Advantages (15:01)

    When asked about their typical users, Yingjun explains they directly compete with Flink but have advantages in ease of use due to Postgres compatibility. Their users are either new to stream processing or are migrating from systems like Spark Streaming or Flink due to performance issues or development complexity.
    Apache Iceberg Integration (19:25)

    Yingjun discusses how Apache Iceberg is emerging as an important destination for Rising Wave output, as companies seek to avoid vendor lock-in with proprietary data warehouses. He explains how Rising Wave typically performs ETL functions before data is sent to Iceberg tables.
    The Future of Data Management (32:06)

    The conversation concludes with a discussion about Iceberg becoming a "single source of truth" for data, with multiple specialized query engines potentially accessing the same data. Yingjun and Eldad share perspectives on how this shift away from proprietary data lock-in is changing the data ecosystem.
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    Episode Resources:

    • Rising Wave Website
    • Yingjun Wu LinkedIn

    For Feedback & Discussions on Firebolt Core:
    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository
    • Benjamin@Firebolt.io

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Revolutionizing Data Governance with DataStrato’s Unified Open Source Approach Apr 08, 2025
    Show notes In this episode of The Data Engineering Show, the bros sit with Lisa Cao, Product Manager at DataStrato, to explore data catalogs and Apache Gravitino, a unified metadata lake used to manage access and perform data governance for all data sources.
    What You’ll Learn:
    • How Apache Gravitino differs from others like Unity catalog and Polaris by being able to support multiple catalog systems.
    • What the “Push-Down Permission Management” security model is and how to implement it across different data systems.
    • How to maintain consistent governance across various query engines like Spark, Trino, and Flink.
    • Why interoperability, flexibility and open source ecosystem are becoming an important dynamics of data infrastructure rather than performance benchmarking.
    • How to evaluate new data tools based on their real-world adoption rather than the social media hype.

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts instructions on how to do this here [insert link].
    Lisa Cao is a Product Manager at DataStrato, specializing in AI/ML product partnerships and developer relations. With deep expertise in data catalog technologies and open-source ecosystems, she plays a key role in developing Apache Gravitino, an ASF incubating project that provides a unified governance and security layer for diverse data systems. Her work in developing extensible catalog frameworks has helped organizations manage complex data environments across multiple platforms.
    Episode Highlights:
    • What is Apache Gravitino? (01:24)
    Apache Gravitino is a meta-catalog that serves as a unified data governance and security layer used to manage different data systems. Lisa shares that Gravitino was the first to release an iceberg rest catalog and ended up open sourcing for the general community to use and as time passed, Polaris and Unity Catalog were also announced in open source. She highlights that although Gravitino, Polaris and Unity Catalog are very similar, Gravitino differs in that it is able to support multiple catalogs.
    • Unifying AI/ML and Big Data Stack (03:15)
    One of the interesting things about Gravitino is that it offers more than just a catalog of data models and these model catalogs are the first step into looking at how to merge two worlds of AI and ML catalogs. Lisa shares the goal of effective management, that is, creating a system that can store and manage different types of data models, track changes to the models, and control access to the models.
    • Simplifying Data Governance (10:49)
    Think of Gravitino as a “traffic cop” that helps to manage and secure data from multiple sources. It is crucial to have a system that provides unified access control across all data sources, allowing teams to manage access and data governance so that ML teams don't have to worry about access. Lisa says that Apache Gravitino is the system that makes data accessible to different teams and users while making sure that it is secure and governed appropriately.
    • The Gravitino’s Query Engine Solution (21:34)
    Every query engine has its own way of managing data, which makes it difficult to switch between engines - you have to reconfigure everything. Lisa highlights that Gravitino solves the problem by providing a single layer of data governance that works across multiple query engines.
    • Navigating the Fast-Paced World of Data Engineering (24:41)

    Lisa talks about how fast the data engineering space is moving and shares some insights to catching up;
    • Don’t try to learn everything at once.
    • Don't get too deep into every tool
    • Look for real-world adoption

    She warns against the social media hype that can amplify the messaging around new tools, making it seem everyone is using it, when in reality, that can’t be easily seen.
    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    Episode Resources:
    • Apache Gravitino website

    For Feedback & Discussions on Firebolt Core:
    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository
    • Benjamin@Firebolt.io

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so
    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.
    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing

    Previous 1 2 3 4 7 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights