TopPodcast.com
Menu
  • Home
  • Top Charts
  • Top Networks
  • Top Apps
  • Top Independents
  • Top Podfluencers
  • Top Picks
    • Top Business Podcasts
    • Top True Crime Podcasts
    • Top Finance Podcasts
    • Top Comedy Podcasts
    • Top Music Podcasts
    • Top Womens Podcasts
    • Top Kids Podcasts
    • Top Sports Podcasts
    • Top News Podcasts
    • Top Tech Podcasts
    • Top Crypto Podcasts
    • Top Entrepreneurial Podcasts
    • Top Fantasy Sports Podcasts
    • Top Political Podcasts
    • Top Science Podcasts
    • Top Self Help Podcasts
    • Top Sports Betting Podcasts
    • Top Stocks Podcasts
  • Podcast News
  • About Us
  • Podcast Advertising
  • Contact
Not in our directory?
Add Show Here
Podcast Equipment
Center

toppodcastlogoOur TOPPODCAST Picks

  • Comedy
  • Crypto
  • Sports
  • News
  • Politics
  • True Crime
  • Business
  • Finance

Follow Us

toppodcastlogoStay Connected

    View Top 200 Chart
    Back to Rankings Page
    Technology

    The Data Life Podcast

    This is a podcast where we talk all-about real life experiences of dealing with data and machine learning tools, techniques and personalities. We cover not just the technical aspects but also the “life” aspects of working in the field.

    Note: Opinions expressed are my own and do not express the views or opinions of my employer.

    Advertise

    Copyright: © 917496

    • Apple Podcasts
    • Google Play
    • Spotify

    Latest Episodes:
    17: Why Pandas is the new Excel Oct 25, 2019
    Show notes

    The Data Life Podcast is a podcast where we talk all-about real life experiences with data and data science science tools, techniques, models and personalities.

    In this episode, we will talk about how Pandas is becoming a tool of choice for many data scientists for doing their data analysis work. We will explore how Pandas wins over Excel in several key areas that are important for businesses today:

    1) Large dataset sizes
    2) Different kinds of input formats such as JSON, CSV, HTML, SQL etc
    3) Complex business logic
    4) Linking data analysis work to websites and databases
    5) Cost

    Pandas has lots of helpful functions such as read_csv, read_json, read_sql that allow easy input of data into dataframes. DataFrames have several useful methods like "describe", "value_counts", "groupby", "loc" and more that allow easy understanding of your dataset. It also supports plotting out of the box with "plot" method.
    We also cover how Pandas differs from SQL in things like ease of handling time series data, visualizations and more.
    Tune in to the episode to learn more about how Pandas might be the tool for your data analysis needs to take your business to next level!

    Fantastic Resources:
    1) Book by Pandas creator Wes McKinney: https://www.amazon.com/dp/1491957662/?tag=omnilence-20
    2) Great workshop video by Kevin Markham in PyCon: https://www.youtube.com/watch?v=0hsKLYfyQZc
    3) Input output methods for Pandas: https://pandas.pydata.org/pandas-docs/stable/user_guide/io.html
    4) Comparison of some operations of Pandas with SQL https://pandas.pydata.org/pandas-docs/stable/getting_started/comparison/comparison_with_sql.html

    Thanks for listening! Please consider supporting this podcast from the link in the end.


    16: Getting Started with Natural Language Processing Oct 05, 2019
    Show notes

    So many tweets and news articles and unstructured text surrounds us. How do we make sense of all of these? Natural language processing or NLP can help. NLP refers to algorithms that process, understand and generate aspects of natural language either in text or in spoken voice. In this episode we will cover some of the common techniques in NLP to help get started in this exciting field!

    We cover several tasks in a NLP pipeline:
    1. Tokenization and punctuation removal
    2. Stemming and Lemmatization
    3. One hot vectors
    4. Word embeddings including Word2Vec and Glove
    5. Recurrent Neural Networks and LSTMs
    6. tf and tf-idf approaches - when to use word embeddings, when to use tf / tf-idf approaches?
    7. Generating text using encoder-decoder or sequence to sequence models
    Some resources:
    1. Sequence Models - course by Andrew Ng on Coursera - one of the best courses I have seen on this topic! https://www.coursera.org/learn/nlp-sequence-models
    2. Awesome collection of resources for NLP for Python, C++, Scala etc. and popular resource: https://github.com/keon/awesome-nlp
    3. Overview of Text Similarity Metrics (a blog written by me on Medium): https://towardsdatascience.com/overview-of-text-similarity-metrics-3397c4601f50
    4. How to train custom word embeddings on a GPU https://towardsdatascience.com/how-to-train-custom-word-embeddings-using-gpu-on-aws-f62727a1e3f6
    Thanks for listening, please support this podcast by following the link in the end.



    15: Using Flask, REST API and Vue.js to build a Single Page Web Application Sep 16, 2019
    Show notes

    As a data scientist, you will work on machine learning models that are deployed on websites - usually wrapped around a REST API, these days they also call this approach a “micro-service”. It is for this reason it is important to know how backends and front ends work and how to build them. In this episode, we talk about building a note app which is a Single Page Application or SPA using Pythons flask library for backend and Vue.js for frontend. We use REST API to communicate between them.

    We cover following topics in Q and A format:

    1. Why should data scientists care about building frontend and backend and rest api?

    2. What is a single page application?

    3. Why Vue.js?

    4. Why do we need server side code?

    5. What is REST API?

    6. How does Flask help with building rest api?


    Then we go into the exact mechanics of building the SPA:

    Step 1: Database setup

    Step 2: Write REST API in flask

    Step 3: Postman setup and testing of the API

    Step 4: Build frontend and write forms to get information

    Step 5: Build routing and login pages

    Step 6: Front end design and UI/UX

    Finally you can deploy both the server and client separately on AWS or Heroku so that other users can see it and use it.


    Dependencies:

    1) Flask to build server side REST APIs

    2) Sqlalchemy which is ORM to access database

    3) Bcrypt for hashing user passwords to store in your database

    4) Vue for building frontend

    5) Bootstrap-Vue for using bootstrap with Vue.js

    6) Axios to communicate via AJAX between client and server

    7) Vue CLI 3 to manage the tooling of the client


    Really awesome resources:

    1) Learn Vue.JS from scratch by the awesome teacher Net Ninja - YouTube https://www.youtube.com/watch?v=5LYrN_cAJoA&list=PL4cUxeGkcC9gQcYgjhBoeQH7wiAyZNrYa&index=1

    2) Building book recording app using Vue and Flask https://testdriven.io/blog/developing-a-single-page-app-with-flask-and-vuejs/#bootstrap-vue

    3) Managing state in Vue.js including Vuex and simple global store: https://medium.com/fullstackio/managing-state-in-vue-js-23a0352b1c87

    4) Authenticating a Flask API Using JSON Web Tokens - YouTube https://www.youtube.com/watch?v=J5bIPtEbS0Q

    5) Really nice tutorial for using databases with Flask by Corey Schafer - YouTube https://www.youtube.com/watch?v=cYWiDiIUxQc&list=PL-osiE80TeTs4UjLw5MM6OjgkjFeUxCYH&index=4


    If this has been of value please consider supporting me by buying me a coffee at the Anchor link at the end. If you support, I will provide extra bonus content for you. Thanks for listening!


    14: Building a Character-Based Text Classifier Aug 07, 2019
    Show notes

    Ever wonder how to automatically detect language from a script? How does Google do it?

    Ever wonder how Amazon knows whether you are searching for a product or a SKU on its search bar?

    We look into character-based text classifiers in this episode. We cover 2 types of models. First is the bag-of-words models such as Naive Bayes, logistic regression and vanilla neural network. Second we cover sequence models such as LSTMs and how to prepare your characters for the LSTMs including things like one-hot encoding, padding, creating character embeddings and then feeding these into LSTMs. We also cover how to set up and compile these sequence models.

    Thanks for listening, and if you find this content useful, please leave a review and consider supporting this podcast from the link below.


    13: Statistics of A/B Testing Jul 17, 2019
    Show notes

    You and your team might spend a lot of time building a new feature. But how do you know if this feature will be liked by the users? One of the ways to statistically prove this is by using A/B testing. Listen to this episode to get tips, tricks and intuition behind hypothesis testing, alpha, beta, p-values, two-sample t-tests and more.

    These understandings have been learnt from experiences deploying A/B tests in the field, and talking to experts.

    These ideas are typically not covered in traditional A/B testing texts which tend to focus a lot on math without the intuition, and that's why I really wanted to cover it in this podcast episode. Thanks for listening! I'd really appreciate your support for this podcast. Follow the link below.


    12: Your Users Don't Care How Smart You Are Jun 25, 2019
    Show notes

    In this episode, we will talk about the importance of business impact in data science. "Your users don't care how smart you are" was a quote I read that got me started in thinking about this.

    The right way to do data science is to think of users, revenue impact, business value and go for the simplest solution possible. The wrong way to do data science is to just find a nail to hit the hammer with rather than the other way around.

    We will cover about all this and more!

    Amazon link of Inspired by Marty Cagan (a great read to get better at product thinking): https://www.amazon.com/dp/1119387507/?tag=omnilence-20

    Please consider buying me a coffee if you find this content useful. Refer to link at the bottom.


    11: The Ten Essential Machine Learning Questions Jun 21, 2019
    Show notes

    This episode covers the ten essential machine learning questions. Disclaimer: Baseline answers have been provided in the episode for guidance. For complete accuracy, please refer to textbooks or to courses by Andrew Ng on Coursera.

    If this content is useful, please consider buying me a coffee via the link https://anchor.fm/the-data-life-podcast/support

    Resources:
    1. Machine Learning Course by Andrew Ng: https://www.coursera.org/learn/machine-learning
    2. Deep Learning Course by Andrew Ng: https://www.coursera.org/specializations/deep-learning
    Questions:
    1. What is underfitting and overfitting? How to avoid it?
    2. What is the difference between batch, SGD and mini-batch gradient descents? When will you use each?
    3. How to choose a machine learning model?
    4. How to improve the latency of a machine learning model in production?
    5. If your training and cross validation accuracies are high, but testing accuracy is less - how would you debug this?
    6. Name 3 hyper-parameters. Why can’t we train them as hyper-parameters, why should only humans set them?
    7. Which metric should be used to evaluate a classifier? How do you connect it to business value?
    8. What prevents someone to select deep learning model for everything?
    9. Say you have to classify a lot of data, but you don’t have labelled training examples. How would you begin to solve the problem? How many training data points are needed?
    10. Say you have a perfectly working machine learning model. How do you deploy this in production? How do you check if users will actually like it?


    Please leave a review on Apple Podcasts or wherever you listen to this.
    Thanks for listening!



    Mining Twitter Data for Sentiment Analysis of Events Jun 01, 2019
    Show notes

    Twitter is a rich source of live information. Is it possible to run sentiment analysis on what the world is thinking as an event unfolds over time? Could we track Twitter data and see if it correlates to news that affects stock market movements? These are some of the questions that we will answer in this podcast episode.

    There are 6 steps for mining Twitter data for sentiment analysis of events that we will cover:

    1) Get Twitter API Credentials
    2) Setup API Credentials in Python
    3) Get Tweet Data via Streaming API using Tweepy
    4) Use out-of-the-box sentiment analysis libraries to get sentiment information
    5) Plot sentiment information to see trends for events
    6) Set this up on AWS or Google Cloud Platform
    This episode covers information about saving the tweets in a database, and using them to plot sentiment information.

    Corresponding Blog Post With Code: https://towardsdatascience.com/mining-live-twitter-data-for-sentiment-analysis-of-events-d69aa2d136a1?source=friends_link&sk=e06ae49f4ce6fb52157ea0eaee72f4c4
    Tweepy: https://github.com/tweepy/tweepy
    TextBlob: https://textblob.readthedocs.io/en/dev/
    Vader Sentiment: https://github.com/cjhutto/vaderSentiment
    Set up AWS instance: https://aws.amazon.com/ec2/getting-started/
    Set up GCP instance: https://cloud.google.com/compute/docs/quickstart-linux

    My Twitter Profile: https://twitter.com/sanket107
    Thanks for listening!


    Don't Be Shy To Pursue Your Interest May 19, 2019
    Show notes

    In this episode, we will talk about things like Maslow's Hierarchy of Needs, and focussing on higher level needs such as satisfaction and achieving full potential. In the area of tech, data science and software development, admitting your interest could involve "shyness" as the next shiny cool thing is pursued by everyone. But if your interest is in a niche, don't let others stop you from putting in an effort to become great at it.

    Thanks for listening, and please show your support to keep this podcast going!


    Review of Udacity Nanodegrees - are they worth it? May 03, 2019
    Show notes

    Udacity has become a popular platform for learning about various things in data science, machine learning and programming in general. In this episode, we will discuss the good, bad and ugly of the Udacity nanodegrees. I will also cover my experiences with Deep Learning and NLP Nanodegrees.

    We will cover things like how Udacity has great production quality and has nice intro courses, but due to their lack of depth and low community engagement, the high costs might not be justified (most of their nanodegrees are around $1,000 currently) But if cost is not a concern, then Udacity could be a good way to get into a new area. If you prefer a structured approach with timelines, they could be good too but if you don't mind doing your own research, reading of blogs and watching free videos online, then again Udacity nanodegrees may not be worth the cost.

    Resources:

    1) Deep Learning Nanodegree: https://www.udacity.com/course/deep-learning-nanodegree--nd101

    2) NLP Nanodegree: https://www.udacity.com/course/natural-language-processing-nanodegree--nd892

    3) DeepLearning.AI by Andrew Ng: https://www.coursera.org/deeplearning-ai

    Please support the podcast by rating it in Apple Podcasts, and also leaving a review :) Thanks for listening!



    Previous 1 2 3 Next

    Related Podcasts

    Reply All

    1

    Reply All Games & Hobbies
    Inside VR & AR

    2

    Inside VR & AR Gadgets
    Note to Self

    3

    Note to Self News
    BrainStuff

    4

    BrainStuff Natural Sciences
    This Week in Tech (Audio)

    5

    This Week in Tech (Audio) News
    Hands-On Tech (Audio)

    6

    Hands-On Tech (Audio) Technology
    footer-logo

    Contact Us

    Toll Free: 844-670-7747

    Links

    • Home
    • Top Charts
    • Networks
    • Apps
    • Independents Podcasts
    • Podcast Advertising
    • Podcast News
    • Contact Us
    • About Us
    • Analytics & Insights

    Stay Connected

      Privacy, Terms of Use & Our Code of Ethics Protecting Content Creators Copyrights