TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    News

    The Python Podcast.__init__

    The podcast about Python and the people who make it great

    Advertise

    Copyright: © 2023 Boundless Notions, LLC.

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    Build Composable And Reusable Feature Engineering Pipelines with Feature-Engine Oct 31, 2021
    Show notes

    Summary

    Every machine learning model has to start with feature engineering. This is the process of combining input variables into a more meaningful signal for the problem that you are trying to solve. Many times this process can lead to duplicating code from previous projects, or introducing technical debt in the form of poorly maintained feature pipelines. In order to make the practice more manageable Soledad Galli created the feature-engine library. In this episode she explains how it has helped her and others build reusable transformations that can be applied in a composable manner with your scikit-learn projects. She also discusses the importance of understanding the data that you are working with and the domain in which your model will be used to ensure that you are selecting the right features.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Soledad Galli about feature-engine, a Python library to engineer features for use in machine learning models

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you describe what feature-engine is and the story behind it?
    • What are the complexities that are inherent to feature engineering?
      • What are the problems that are introduced due to incidental complexity and technical debt?
    • What was missing in the available set of libraries/frameworks/toolkits for feature engineering that you are solving for with feature-engine?
    • What are some examples of the types of domain knowledge that are needed to effectively build features for an ML model?
    • Given the fact that features are constructed through methods such as normalizing data distributions, imputing missing values, combining attributes, etc. what are some of the potential risks that are introduced by incorrectly applied transformations or invalid assumptions about the impact of these manipulations?
    • Can you describe how feature-engine is implemented?
      • How have the design and goals of the project changed or evolved since you started working on it?
    • What (if any) difference exists in the feature engineering process for frameworks like scikit-learn as compared to deep learning approaches using PyTorch, Tensorflow, etc.?
    • Can you describe the workflow of identifying and generating useful features during model development?
      • What are the tools that are available for testing and debugging of the feature pipelines?
    • What do you see as the potential benefits or drawbacks of integrating feature-engine with a feature store such as Feast or Tecton?
    • What are the most interesting, innovative, or unexpected ways that you have seen feature-engine used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on feature-engine?
    • When is feature-engine the wrong choice?
    • What do you have planned for the future of feature-engine?

    Keep In Touch

    • LinkedIn
    • @Soledad_Galli on Twitter
    • solegalli on GitHub

    Picks

    • Tobias
      • Dune Movie
      • Dune Series
    • Soledad
      • The Social Dilemma
      • Don’t Be Evil by Rana Foroohar

    Closing Announcements

    • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
    • To help other people find the show please leave a review on iTunes and tell your friends and co-workers

    Links

    • feature-engine
    • Feature Engineering
    • Python Feature Engineering Cookbook
    • scikit-learn
    • Feature Stores
      • Podcast Episode
    • Pandas
      • Podcast Episode
    • PyTorch
      • Podcast Episode
    • Tensorflow
    • Feast
    • Tecton
      • Data Engineering Podcast Episode
    • Kaggle
    • Dask
      • Data Engineering Podcast Episode

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Speed Up Your Python Data Applications By Parallelizing Them With Bodo Oct 25, 2021
    Show notes

    Summary

    The speed of Python is a subject of constant debate, but there is no denying that for compute heavy work it is not the optimal tool. Rather than rewriting your data oriented applications, or having to rearchitect them, the team at Bodo wrote a compiler that will do the optimization for you. In this episode Ehsan Totoni explains how they are able to translate pure Python into massively parallel processes that are optimized for high performance compute systems.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Ehsan Totoni about Bodo, an inferential compiler for Python that automatically parallelizes your data oriented projects

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you describe what Bodo is and the story behind it?
    • What are some of the use cases that it is being applied to?
    • What are the motivating factors for something like Dask or Ray as compared to Bodo?
    • What are the software patterns that contribute to slowdowns in data processing code?
      • What are some of the ways that the compiler is able to optimize those operations?
    • Can you describe how Bodo is implemented?
    • How does Bodo process the Python code for compiling to the optimized form?
      • What are the compilation techniques for understanding the semantics of the code being processed?
      • How do you manage packages that rely on C extensions?
      • What do you use as an intermediate representation for translating into the optimized output?
    • What is the workflow for applying Bodo to a Python project?
      • What debugging utilities does it provide for identifying any errors that occur due to the added parallelism?
    • What kind of support does Bodo have for optimizing a machine learning project with Bodo? (e.g. using PyTorch/Tensorflow/MxNet/etc.)
    • When working with a workflow orchestrator such as Dagster for Airflow, what would the integration process look like for being able to take advantage of the optimized Bodo output?
    • What are the most interesting, innovative, or unexpected ways that you have seen Bodo used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Bodo?
    • When is Bodo the wrong choice?
    • What do you have planned for the future of Bodo?

    Keep In Touch

    • LinkedIn
    • @EhsanTn on Twitter
    • ehsantn on GitHub

    Picks

    • Tobias
      • Paracord Crafts
    • Ehsan
      • [

    Closing Announcements

    • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
    • To help other people find the show please leave a review on iTunes and tell your friends and co-workers

    Links

    • Bodo
      • Data Engineering Podcast Episode
    • University of Illinois Urbana-Champaign
    • HPC
    • MPI
    • Elastic Fabric Adapter
    • All-to-All Communication
    • Dask
      • Data Engineering Podcast Episode
    • Ray
      • Podcast Episode
    • Pandas Extension Arrays
      • Podcast Episode
    • GeoPandas
    • Numba
    • LLVM
    • scikit-learn
    • Horovod
    • Dagster
      • Podcast.__init__ Episode
      • Data Engineering Podcast Episode
    • Airflow
      • Podcast Episode
    • IPython Parallel
    • Parquet

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    An Exploration Of Financial Exchange Risk Management Strategies Oct 16, 2021
    Show notes

    Summary

    The world of finance has driven the development of many sophisticated techniques for data analysis. In this episode Paul Stafford shares his experiences working in the realm of risk management for financial exchanges. He discusses the types of risk that are involved, the statistical methods that he has found most useful for identifying strategies to mitigate that risk, and the software libraries that have helped him most in his work.

    Announcements

    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With their managed Kubernetes platform it’s now even easier to deploy and scale your workflows, or try out the latest Helm charts from tools like Pulsar and Pachyderm. With simple pricing, fast networking, object storage, and worldwide data centers, you’ve got everything you need to run a bulletproof data platform. Go to dataengineeringpodcast.com/linode today and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Paul Stafford about building risk models to guard against financial exchange rate volatility

    Interview

    • Introductions
    • How did you get introduced to Python?
    • What are the principles involved in risk management, and how are statistical methods used?
    • How did you get involved in financial markets?
      • In what ways did your background in science and engineering prepare you for work in finance and risk management?
    • What are the tools that you have found most useful in your career in finance?
    • How have recent trends such as the widespread adoption of deep learning impacted the capabilities and risks present in foreign exchange strategies?
    • What are the challenges that you face in obtaining and validating the input data that you are relying on for building financial and statistical models?
      • How has the volatility of the pandemic impacted the robustness and resilience of your predictive capabilities?
    • What are the areas where the available tools are typically insufficient?
    • What are the most interesting, innovative, or unexpected strategies or techniques that you have seen applied to risk management?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working in risk management?
    • What are the economic and industry trends that you are keeping a close eye on for your work at Deaglo and your own personal projects?

    Keep In Touch

    • LinkedIn

    Picks

    • Tobias
      • The Vault (movie)
    • Paul
      • Motorcycle Trip of the Grand Canyon

    Links

    • Deaglo Partners, LLC.
    • Value At Risk (VaR)
    • Black-Scholes Equation
    • Linear Algebra
    • Principal Component Analysis
    • Eigenvectors and Eigenvalues
    • Markov Chain Monte Carlo
    • Violin Plot
    • Kurtosis
    • PyMC3
      • Podcast Episode
    • Bayesian Regression
    • Constrained Optimization
    • Ethereum
    • Smart Contracts
    • Behavioral Finance
    • Black Swan by Nassim Nicholas Taleb (affiliate link)
    • SciPy Convention
    • RealPython
    • 3Blue1Brown
    • Sentiment Analysis

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Build Better Machine Learning Models By Understanding Their Decisions With SHAP Oct 09, 2021
    Show notes

    Summary

    Machine learning and deep learning techniques are powerful tools for a large and growing number of applications. Unfortunately, it is difficult or impossible to understand the reasons for the answers that they give to the questions they are asked. In order to help shine some light on what information is being used to provide the outputs to your machine learning models Scott Lundberg created the SHAP project. In this episode he explains how it can be used to provide insight into which features are most impactful when generating an output, and how that insight can be applied to make more useful and informed design choices. This is a fascinating and important subject and this episode is an excellent exploration of how to start addressing the challenge of explainability.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Scott Lundberg about SHAP, a library that implements a game theoretic approach to explain the output of any machine learning model

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you describe what SHAP is and the story behind it?
    • What are some of the contexts that create the need to explain the reasoning behind the outputs of an ML model?
    • How do different types of models (deep learning, CNN/RNN, bayesian vs. frequentist, etc.) and different categories of ML (e.g. NLP, computer vision) influence the challenge of understanding the meaningful signals in their reasoning?
    • Taking a step back, how do you define "explainability" when discussing inferences produced by ML models?
      • What are the degrees of specificity/accuracy when seeking to understand the decision processes involved?
    • Can you describe how SHAP is implemented?
      • What are the signals that you are tracking to understand what features are being used to determine a given output?
      • What are the assumptions that you had as you started this project that have been challenged or updated as you explored the problem in greater depth?
    • Can you describe the workflow for someone using SHAP?
      • What are the challenges faced by practitioners in interpreting the visualizations generated from SHAP?
    • How much domain knowledge and context is necessary to use SHAP effectively?
    • What are the ongoing areas of research around tracking of ML decision processes?
    • How are you using SHAP in your own work?
    • What are the most interesting, innovative, or unexpected ways that you have seen SHAP used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on SHAP?
    • When is SHAP the wrong choice?
    • What do you have planned for the future of SHAP?

    Keep In Touch

    • slundberg on GitHub
    • Website
    • LinkedIn

    Picks

    • Tobias
      • Reminiscence
    • Scott
      • Augustine’s Confessions

    Links

    • SHAP
    • Microsoft Research
    • Matlab
    • Game Theory
    • Computational Biology
    • LIME
    • Shapley Values
    • Julia Language
    • ResNet
    • CNN == Convolutional Neural Network
    • RNN == Recurrent Neural Network
    • A* Algorithm
    • CFPB == Consumer Financial Protection Bureau
    • NP Hard
    • Huggingface
    • Right for the Right Reasons: Training Differentiable Models by Constraining their Explanations
    • Numba
    • Log Odds
    • InterpretML
    • Polyjuice

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Accelerating Drug Discovery Using Machine Learning With TorchDrug Sep 30, 2021
    Show notes

    Summary

    Finding new and effective treatments for disease is a complex and time consuming endeavor, requiring a high degree of domain knowledge and specialized equipment. Combining his expertise in machine learning and graph algorithms with is interest in drug discovery Jian Tang created the TorchDrug project to help reduce the amount of time needed to find new candidate molecules for testing. In this episode he explains how the project is being used by machine learning researchers and biochemists to collaborate on finding effective treatments for real-world diseases.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Jian Tang about TorchDrug

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you describe what TorchDrug is and the story behind it?
    • What are the goals of the TorchDrug project?
      • Who are the target users of the project?
      • What are the main ways that it is being used?
    • What are the challenges faced by biologists and chemists working on development and discovery of pharmaceuticals?
      • What are some of the other tools/techniques that they would use (in isolation or combination with TorchDrug)?
    • Can you describe how TorchDrug is implemented?
      • How have you approached the design of the project and its APIs to make it accessible to engineers that don’t possess domain expertise in drug discovery research?
    • How do graph structures help when modeling and experimenting with chemical structures for drug discovery?
    • What are the formats and sources of data that you are working with?
      • What are some of the complexities/challenges that you have had to deal with to integrate with up or downstream systems to fit into the overall research process?
    • Can you talk through the workflow of using TorchDrug to build and validate a model?
      • What is involved in determining and codifying a goal state for the model to optimize for?
    • What are the biggest open questions in the area of drug discovery and research?
      • How is TorchDrug being used to assist in the exploration of those problems?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on TorchDrug?
    • When is TorchDrug the wrong choice?
    • What do you have planned for the future of TorchDrug?

    Keep In Touch

    • tangjianpku on GitHub
    • @tangjianpku on Twitter
    • Website
    • LinkedIn

    Picks

    • Tobias
      • Rope refactoring library
    • Jian
      • Attending conferences once the pandemic is over

    Closing Announcements

    • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
    • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
    • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat

    Links

    • TorchDrug
    • Mila
    • Yoshua Bengio
    • Alphafold
    • Few-shot learning
    • Metalearning
    • PyTorch Geometric
    • DeepGraph Library
    • NetworKit
      • Podcast Episode
    • graph-tool
      • Podcast Episode

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    An Exploration Of Automated Speech Recognition Sep 26, 2021
    Show notes

    Summary

    The overwhelming growth of smartphones, smart speakers, and spoken word content has corresponded with increasingly sophisticated machine learning models for recognizing speech content in audio data. Dylan Fox founded Assembly to provide access to the most advanced automated speech recognition models for developers to incorporate into their own products. In this episode he gives an overview of the current state of the art for automated speech recognition, the varying requirements for accuracy and speed of models depending on the context in which they are used, and what is required to build a special purpose model for your own ASR applications.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Dylan Fox about the challenges of training and deploying large models for automated speech recognition

    Interview

    • Introductions
    • How did you get introduced to Python?
    • What is involved in building an ASR model?
      • How does the complexity/difficulty compare to models for other data formats? (e.g. computer vision, NLP, NER, etc.)
    • How have ASR models changed over the last 5, 10, 15 years?
    • What are some other categories of ML applications that work with audio data?
      • How does the level of complexity compare to ASR applications?
    • What is the typical size of an ASR model that you are deploying at Assembly?
      • What are the factors that contribute to the overall size of a given model?
    • How does accuracy compare with model size?
    • How does the size of a model contribute to the overall challenge of deploying/monitoring/scaling it in a production environment?
    • How can startups effectively manage the time/cost that comes with training large models?
    • What are some techniques that you use/attributes that you focus on for feature definitions in the source audio data?
    • Can you describe the lifecycle stages of an ASR model at Assembly?
    • What are the aspects of ASR which are still intractable or impractical to productionize?
    • What are the most interesting, innovative, or unexpected ways that you have seen ASR technology used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on ASR?
    • What are the trends in research or industry that you are keeping an eye on?

    Keep In Touch

    • LinkedIn
    • @YouveGotFox on Twitter

    Picks

    • Tobias
      • The Hitman’s Wife’s Bodyguard
    • Dylan
      • Inspiration 4 Documentary

    Closing Announcements

    • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
    • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
    • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat

    Links

    • Learn Python The Hard Way
    • DeepSpeech
    • Wav2Letter
    • BERT
    • GPT-3
    • Convolutional Neural Network (CNN)
    • Recurrent Neural Network (RNN)
    • Mycroft
      • Podcast Episode
    • CMU Sphinx
    • Pocket Sphinx
    • Gaussian Mixture Model (GMM)
    • Hidden Markov Model (HMM)
    • DeepSpeech Paper
    • Transformer Architecture
    • Audio Analytic Sound Recognition Podcast Episode
    • Horovod distributed training library
    • Knowledge Distillation
    • Libre Speech Data Set
    • Lambda Labs
    • Wav2Vec

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Experimenting With Reinforcement Learning Using MushroomRL Sep 19, 2021
    Show notes

    Summary

    Reinforcement learning is a branch of machine learning and AI that has a lot of promise for applications that need to evolve with changes to their inputs. To support the research happening in the field, including applications for robotics, Carlo D’Eramo and Davide Tateo created MushroomRL. In this episode they share how they have designed the project to be easy to work with, so that students can use it in their study, as well as extensible so that it can be used by businesses and industry professionals. They also discuss the strengths of reinforcement learning, how to design problems that can leverage its capabilities, and how to get started with MushroomRL for your own work.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Davide Tateo and Carlo D’Eramo about MushroomRL, a library for building reinforcement learning experiments

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you start by describing what reinforcement learning is and how it differs from other approaches for machine learning?
    • What are some example use cases where reinforcement learning might be necessary?
    • Can you describe what MushroomRL is and the story behind it?
      • Who are the target users of the project?
      • What are its main goals?
    • What are your suggestions to other developers for implementing a succesful library?
    • What are some of the core concepts that researchers and/or engineers need to understand to be able to effectively use reinforcement learning techniques?
    • Can you describe how MushroomRL is architected?
      • How have the goals and design of the project changed or evolved since you began working on it?
    • What is the workflow for building and executing an experiment with MushroomRL?
      • How do you track the states and outcomes of experiments?
    • What are some of the considerations involved in designing an environment and reward functions for an agent to interact with?
    • What are some of the open questions that are being explored in reinforcement learning?
    • How are you using MushroomRL in your own research?
    • What are the most interesting, innovative, or unexpected ways that you have seen MushroomRL used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on MushroomRL?
    • When is MushroomRL the wrong choice?
    • What do you have planned for the future of MushroomRL?
    • How can the open-source community contribute to MushroomRL?
    • What kind of support you are willing to provide to users?

    Keep In Touch

    • Davide
      • boris-il-forte on GitHub
      • Website
    • Carlo
      • carloderamo on GitHub
      • Website

    Picks

    • Tobias
      • Britannia TV Series
    • Davide
      • 1984 by George Orwell
    • Carlo
      • Twin Peaks TV Series

    Closing Announcements

    • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
    • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
    • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat

    Links

    • MushroomRL
    • TU Darmstadt
    • MuJoCo
    • PyBullet
    • iGibson
    • Habitat
    • OpenAI Gym
    • PyTorch
      • Podcast Episode
    • RLLib
    • Ray
      • Podcast Episode
    • OpenAI Baselines
    • Stable Baselines
    • ROS

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Doing Dask Powered Data Science In The Saturn Cloud Sep 10, 2021
    Show notes

    Summary

    A perennial problem of doing data science is that it works great on your laptop, until it doesn’t. Another problem is being able to recreate your environment to collaborate on a problem with colleagues. Saturn Cloud aims to help with both of those problems by providing an easy to use platform for creating reproducible environments that you can use to build data science workflows and scale them easily with a managed Dask service. In this episode Julia Signall, head of open source at Saturn Cloud, explains how she is working with the product team and PyData community to reduce the points of friction that data scientists encounter as they are getting their work done.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Julia Signell about building distributed processing workflows in Python through the power of Dask

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you describe what you are building at Saturn Cloud?
      • Who are your target users and how does that inform the features and priorities that you build into your platform?
    • What are the road blocks that data scientists typically encounter when working on their laptop/workstation?
    • How does open source factor into the Saturn product?
      • What are some of the projects that you are collaborating with/contributing to as part of your work at Saturn?
      • How has your experience at Anaconda informed your work at Saturn?
    • Can you describe how the Saturn Cloud platform is architected?
      • How has it changed or evolved since it was first launched?
    • Can you describe the learning curve that data scientists go through when adopting Dask?
    • What are some examples of projects or workflows that Dask enables which are not possible/practical to do locally?
    • How would you characterize the overall awareness/adoption of Dask in the Python data science community?
    • What are the most interesting, innovative, or unexpected ways that you have seen Saturn Cloud used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Saturn Cloud?
    • When is Saturn Cloud the wrong choice?
    • What do you have planned for the future of Saturn Cloud?

    Keep In Touch

    • @jsignell on Twitter
    • jsignell on GitHub

    Picks

    • Tobias
      • Peter Rabbit 2
    • Julia
      • PawPaw Fruit

    Closing Announcements

    • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
    • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
    • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat

    Links

    • Saturn Cloud
    • Dask
      • Podcast Episode
    • Pangeo
    • XArray
    • Conda
    • Mamba
    • Holoviz
    • Dash
    • Anaconda
      • Podcast Episode
    • Kubernetes
    • Tornado
      • Podcast Episode
    • Prefect
      • Podcast Episode
    • Dagster
      • Podcast Episode
    • Airflow
    • Ray
      • Podcast Episode

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Monitor The Health Of Your Machine Learning Products In Production With Evidently Sep 03, 2021
    Show notes

    Summary

    You’ve got a machine learning model trained and running in production, but that’s only half of the battle. Are you certain that it is still serving the predictions that you tested? Are the inputs within the range of tolerance that you designed? Monitoring machine learning products is an essential step of the story so that you know when it needs to be retrained against new data, or parameters need to be adjusted. In this episode Emeli Dral shares the work that she and her team at Evidently are doing to build an open source system for tracking and alerting on the health of your ML products in production. She discusses the ways that model drift can occur, the types of metrics that you need to track, and what to do when the health of your system is suffering. This is an important and complex aspect of the machine learning lifecycle, so give it a listen and then try out Evidently for your own projects.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Emeli Dral about monitoring machine learning models in production with Evidently

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you describe what Evidently is and the story behind it?
    • What are the metrics that are useful for determining the performance and health of a machine learning model?
      • What are the questions that you are trying to answer with those metrics?
    • How does monitoring of machine learning models compare to monitoring of infrastructure or "traditional" software projects?
    • What are the failure modes for a model?
    • Can you describe the design and implementation of Evidently?
      • How has the architecture changed or evolved since you started working on it?
    • What categories of model is Evidently designed to work with?
      • What are some strategies for making models conducive to monitoring?
    • What is involved in monitoring a model on a continuous basis?
    • What are some considerations when establishing useful thresholds for metrics to alert on?
      • Once an alert has been triggered what is the process for resolving it?
      • If the training process takes a long time, how can you mitigate the impact of a model failure until the new/updated version is deployed?
    • What are the most interesting, innovative, or unexpected ways that you have seen Evidently used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Evidently?
    • When is Evidently the wrong choice?
    • What do you have planned for the future of Evidently?

    Keep In Touch

    • LinkedIn
    • @EmeliDral on Twitter
    • emeli-dral on GitHub

    Picks

    • Tobias
      • The Suicide Squad
    • Emeli
      • Airflow

    Links

    • Evidently AI
      • Open Source
    • Yandex
    • Grafana

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Making Automated Machine Learning More Accessible With EvalML Aug 25, 2021
    Show notes

    Summary

    Building a machine learning model is a process that requires a lot of iteration and trial and error. For certain classes of problem a large portion of the searching and tuning can be automated. This allows data scientists to focus their time on more complex or valuable projects, as well as opening the door for non-specialists to experiment with machine learning. Frustrated with some of the awkward or difficult to use tools for AutoML, Angela Lin and Jeremy Shih helped to create the EvalML framework. In this episode they share the use cases for automated machine learning, how they have designed the EvalML project to be approachable, and how you can use it for building and training your own models.

    Announcements

    • Hello and welcome to Podcast.__init__, the podcast about Python’s role in data and science.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $100 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • Your host as usual is Tobias Macey and today I’m interviewing Angela Lin, Jeremy Shih about EvalML, an AutoML library which builds, optimizes, and evaluates machine learning pipelines

    Interview

    • Introductions
    • How did you get introduced to Python?
    • Can you describe what EvalML is and the story behind it?
    • What do we mean by the term AutoML?
    • What are the kinds of problems that are best suited to applications of automated ML?
    • What does the landscape for AutoML tools look like?
      • What was missing in the available offerings that motivated you and your team to create EvalML?
    • Who is the target audience for EvalML?
    • How is the EvalML project implemented?
      • How has the project changed or evolved since you first began working on it?
    • What is the workflow for building a model with EvalML?
      • Can you describe the preprocessing steps that are necessary and the input formats that it is expecting?
    • What are the supported algorithms/model architectures?
    • How does EvalML explore the search space for an optimal model?
      • What decision functions does it employ to determine an appropriate stopping point?
    • What is involved in operationalizing an AutoML pipeline?
    • What are some challenges or edge cases that you see users of EvalML run into?
    • What are the most interesting, innovative, or unexpected ways that you have seen EvalML used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on EvalML?
    • When is EvalML the wrong choice?
    • When is auto ML the wrong approach?
    • What do you have planned for the future of EvalML?

    Keep In Touch

    • Angela
      • angela97lin on GitHub
      • LinkedIn
    • Jeremy
      • jeremyliweishih on GitHub
      • LinkedIn

    Picks

    • Tobias
      • Gloryhammer
    • Angela
      • Sarma mediterranean restaurant
    • Jeremy
      • Crucial Conversations by Stephen Covey (affiliate link)

    Closing Announcements

    • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you’ve learned something or tried out a project from the show then tell us about it! Email hosts@podcastinit.com) with your story.
    • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
    • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat

    Links

    • EvalML
    • FeatureLabs
    • Alteryx
    • Scheme
    • NetLogo
    • Flask
    • AutoML
    • Woodwork
    • FeatureTools
    • Compose
    • Random Forest
    • XGBoost
    • Prophet
    • GreyKite
    • Shap

    The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA


    Previous 1 4 5 6 7 8 39 Next

    Related Podcasts

    Inside Strategic Coach: Connecting Entrepreneurs With What Really Matters

    1

    Inside Strategic Coach: Connecting Entrepreneurs With What Really Matters Business
    WSJ Your Money Briefing

    2

    WSJ Your Money Briefing Business
    FORTUNE Unfiltered with Aaron Task

    3

    FORTUNE Unfiltered with Aaron Task Business
    FORTUNE OnStage Presents: The Most Powerful Women

    4

    FORTUNE OnStage Presents: The Most Powerful Women Business
    Slate Money

    5

    Slate Money Business
    In The Dark – The New Yorker

    6

    In The Dark – The New Yorker Business News
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights