Sample space is a podcast about tools, thoughts and techniques from machine learning practitioners. We talk to toolmakers and practitioners about interesting problems in the real world to find out how great ideas in our field actually manifest.
Time for some (extreme) distillation with Thomas van Dongen - founder of the Minish Lab
Thomas van Dongen, founder of the Minish Lab, distills the power of word embeddings, revealing how to get highly performant word embeddings with the right distillation technique.
6 Dec 2024
Imbalanced learn: regrets and onwards
Guillaume Lemaitre, maintainer of Imbalanced learn, shares lessons learned over the last decade on rethinking resampling techniques for imbalanced classification use-cases.
6 Nov 2024
You want to be in control of your own Copilot
There are many LLMs that you can use for programming these days. Some of them even go into your IDE like Cursor or Github Copilot. But what if you want to tweak these LLMs do to what you want? Instead of being stuck with the tools that a vendor gives you, the goal of Continue.dev (http://Continue.dev) is to allow you to customise this yourself. In this podcast we talk to Ty Dunn, co-founder of the project to learn more about this.
If you are curious to learn more about this effort, please check out https://continue.dev (https://continue.dev). You may always want to read the manifesto over at https://amplified.dev/ (https://amplified.dev/).
We have a Discord these days, feel free to discuss the podcast with us there! https://discord.probabl.ai (https://discord.probabl.ai)
This podcast is part of the open efforts over at probabl. To learn more you can check out website or reach out to us on social media.
Website: https://probabl.ai/ (https://probabl.ai/)
Bluesky: https://bsky.app/profile/probabl.bsky.social (https://bsky.app/profile/probabl.bsky.social)
LinkedIn: https://www.linkedin.com/company/probabl (https://www.linkedin.com/company/probabl)
Twitter: https://x.com/probabl_ai (https://x.com/probabl_ai)
31 Oct 2024
What it is like to maintain the scikit-learn docs
Scikit-learn's documentation pages are celebrated. But not everyone is aware that the project actually has somebody on payroll to take care of it. In this episode we talk to Arturo about stories from the scikit-learn documentation. In particular, the docs have a recommender that few folks are aware of. People just assume that it is manually curated, but there are a few base scikit-learn tools under the hood there.
Link to the official scikit-learn MOOC: https://inria.github.io/scikit-learn-mooc/ (https://inria.github.io/scikit-learn-mooc/)
We have a Discord these days, feel free to discuss the podcast with us there! https://discord.probabl.ai (https://discord.probabl.ai)
You can follow the podcast on most podcast players including apple podcasts, spotify and rss.com (http://rss.com).
- https://podcasts.apple.com/us/podcast/sample-space/id1739598572 (https://podcasts.apple.com/us/podcast/sample-space/id1739598572)
- https://open.spotify.com/show/0BnwEHuyOlHgeZfselpn1n (https://open.spotify.com/show/0BnwEHuyOlHgeZfselpn1n)
- https://rss.com/podcasts/sample-space/ (https://rss.com/podcasts/sample-space/)
This podcast is part of the open efforts over at probabl. To learn more you can check out website or reach out to us on social media.
Website: https://probabl.ai/ (https://probabl.ai/)
Bluesky: https://bsky.app/profile/probabl.bsky.social (https://bsky.app/profile/probabl.bsky.social)
LinkedIn: https://www.linkedin.com/company/probabl (https://www.linkedin.com/company/probabl)
Twitter: https://x.com/probabl_ai (https://x.com/probabl_ai)
23 Oct 2024
Sqlite can totally do embeddings now
Vector databases are kind of everywhere these days. There is a big pool of VC's that are pooring money into the ecosystem too. But while all of that is happening, sqlite has also gotten support for it. In this episode we talk the Alex Garcia, the maintainer of this project, and discuss how the project got created on what the future has in store.
Sqlite-vec Github repo:
https://github.com/asg017/sqlite-vec (https://github.com/asg017/sqlite-vec)
Alex Garcia blog:
https://alexgarcia.xyz/blog/2024/sqlite-vec-hybrid-search/index.html (https://alexgarcia.xyz/blog/2024/sqlite-vec-hybrid-search/index.html)
Datasette discord:
https://discord.com/invite/ktd74dm5mw (https://discord.com/invite/ktd74dm5mw)
Sqlite-vec channel on Mozilla Discord:
https://discord.gg/Ve7WeCJFXk (https://discord.gg/Ve7WeCJFXk)
16 Oct 2024
How to rethink the notebook - with Akshay Agrawal, co-creator of Marimo
Jupyter has been a great environment to explore computational ideas, but that doesn't mean that it can be the only environment for interactive coding in Python. It also comes with some downsides, which led Akshay Agrawal to create an alternative called Marimo. We discussed it in a previous livestream and figured that it was time to sit down with the creator to learn what led to the development of this exciting new too.
You can learn more about Marimo by going to their website over at https://marimo.io (https://marimo.io)
To learn more you can check out website or reach out to us on social media.
Website: https://probabl.ai/ (https://probabl.ai/)
LinkedIn: https://www.linkedin.com/company/probabl (https://www.linkedin.com/company/probabl)
Twitter: https://x.com/probabl_ai (https://x.com/probabl_ai)
10 Sept 2024
You are always dealing with many tables - with Madelon Hulsebos
When you are working on a data pipeline for ML ... you are never dealing with a single table. It always demands different tables for different reasons that all have to be mashed together in order to have something that you can learn from. But if that is the case, why do we spend so much time talking about ML pipelines that only work on a single table? Madelon Hulsebos has a Phd on the topic and so we figured that we might ask her.
As mentioned in the podcast, here is the link to Madelon's homepage. https://www.madelonhulsebos.com/ (https://www.madelonhulsebos.com/)
Some links to interesting articles from Madelon, as well as her homepage, can be found below. https://www.madelonhulsebos.com/assets/dataset_search_survey.pdf (https://www.madelonhulsebos.com/assets/dataset_search_survey.pdf)
https://dl.acm.org/doi/pdf/10.1145/3654975 (https://dl.acm.org/doi/pdf/10.1145/3654975)
https://dl.acm.org/doi/pdf/10.1145/3588710 (https://dl.acm.org/doi/pdf/10.1145/3588710)
21 Aug 2024
How Narwhals has many end users ... that never use it directly with Marco Gorelli
When you pip install a package you will for sure end up using it later. But often you will also install a bunch of dependencies and it is very likely that you won't directly interact with all of them. That does not mean that such a package is not useful, it merely means that the package might be directly used by a maintainer instead. This is interesting, because recently one such tool came into existence. It is called Narwhals and it seems to be on track to become critical infrastructure for data science projects. We have the maintainer of Narwhals on the show this week to talk about it.
To learn more about Narwhals, you can check the repository here: https://github.com/narwhals-dev/narwhals (https://github.com/narwhals-dev/narwhals)
This podcast is part of the open efforts over at probabl. To learn more you can check out website or reach out to us on social media.
Website: https://probabl.ai/ (https://probabl.ai/)
LinkedIn: https://www.linkedin.com/company/probabl (https://www.linkedin.com/company/probabl)
Twitter: https://x.com/probabl_ai (https://x.com/probabl_ai)
17 Jul 2024
Pragmatic data science checklists with Peter Bull - cofounder Drivendata
A lot of things can (and have) gone wrong when folks tried to apply data science projects. So how might we prevent that? Maybe what we need to do is to look at the medical profession and their practice of checklists before surgery.
27 Jun 2024
Model safety, that's a pickle! with Adrin Jalali - scikit-learn maintainer
Historically it's always been the case that you would use a pickle file to store a trained scikit-learn model on disk for deployment. Pickles make sense because these are so flexible, but they do carry a security concern. Adrin has been working on a remedy called skops, which is the main topic of this podcast.
To learn more about skops, make sure to check the documentation: https://skops.readthedocs.io/en/stable/ (https://skops.readthedocs.io/en/stable/)
30 May 2024
Moving Towards KDearestNeighbors with Leland McInnes - creator of UMAP
Leland McInnes is known for a lot of packages. There's UMAP, but also PyNNDescent and HDBScan. Recently he's also been working on tools to help visualise clusters of data and he's also cooking up something new that's related to nearest neighbor algorithms. This interview touches all of these topics.
If you're interested in learning more about the MoMA exhibition, it was by Refik Anadol: https://refikanadol.com/ (https://refikanadol.com/) and this was the work at MoMA: https://refikanadol.com/works/unsupervised/. (https://refikanadol.com/works/unsupervised/)
The other artist was Kyle McDonald: https://kylemcdonald.net/ (https://kylemcdonald.net/) and the piece we mentioned was this one: https://www.youtube.com/watch?v=04DqdT0-NtI (https://www.youtube.com/watch?v=04DqdT0-NtI).
2 May 2024
Talk like a DataFrame, run like SQL with Phillip Cloud - core-committer on Ibis
Ibis is a Python library that offers a single data-frame API, from Python, which can run your queries on many different backends. These include databases like Postgres, but also commercial vendors like BigQuery and Snowflake. This ability to control multiple backends from a single API has a lot of use-cases, as well as maintainer challenges, all of which are discussed in this episode.
To learn more about Ibis, check out the docs here: https://ibis-project.org/ (https://ibis-project.org/)
If you're attending PyCon US this year, you may be interested in Philip's talk: https://us.pycon.org/2024/schedule/presentation/55/ (https://us.pycon.org/2024/schedule/presentation/55/)
During the podcast, Philip also mentioned a blogpost about DuckDB, here: https://ibis-project.org/posts/why-duckdb/ (https://ibis-project.org/posts/why-duckdb/)
There was also a dogfooding blogpost, which is this one: https://ibis-project.org/posts/ci-analysis/ (https://ibis-project.org/posts/ci-analysis/)
2 May 2024
Talk like a DataFrame, run like SQL with Philip Cloud - core-committer on Ibis
Ibis is a Python library that offers a single data-frame API, from Python, which can run your queries on many different backends. These include databases like Postgres, but also commercial vendors like BigQuery and Snowflake. This ability to control multiple backends from a single API has a lot of use-cases, as well as maintainer challenges, all of which are discussed in this episode.
To learn more about Ibis, check out the docs here: https://ibis-project.org/ (https://ibis-project.org/)
If you're attending PyCon US this year, you may be interested in Philip's talk: https://us.pycon.org/2024/schedule/presentation/55/ (https://us.pycon.org/2024/schedule/presentation/55/)
During the podcast, Philip also mentioned a blogpost about DuckDB, here: https://ibis-project.org/posts/why-duckdb/ (https://ibis-project.org/posts/why-duckdb/)
There was also a dogfooding blogpost, which is this one: https://ibis-project.org/posts/ci-analysis/ (https://ibis-project.org/posts/ci-analysis/)
11 Apr 2024
Enhancing Jupyter with Widgets with Trevor Manz - creator of anywidget.
In this (first!) episode of Sample Space we talk to Trevor Mantz, the creator of anywidget. It's a (neat!) tool to help you build more interactive notebooks by giving you tools to apply just enough Javascript to get directional communication working in your favorite notebook environment. That means that Python can talk to widgets, but also that widgets can talk to Python. There's a lot to like about these widgets and we're doing a proper deep dive in this first episode.
To learn more about anywidget, check out the docs (https://anywidget.dev/). In particular you may want to glance at the gallery (https://anywidget.dev/en/community/) first, it has loads of nice examples.
You can also find the project on Github (https://github.com/manzt/anywidget) and if you're eager to talk to folks involved with the project, consider joining the discord here (https://discord.gg/W5h4vPMbDQ).
3 Apr 2024
Introducing Sample Space
We're starting a new podcast!
Reach and audience
Public platform figures. Ratings count people who left a rating, not total listeners.
Apple Podcasts (US)
5.0 / 5
3 ratings
Spotify
4.2 / 5
5 ratings
Podcast Authority Score: 16 / 100
A composite of feed quality, social presence, YouTube performance and engagement. Read the methodology.
Quality
9
Social presence
0
YouTube
0
Engagement
60
Contact Sample Space
Guest appearances
Does not typically book guests
Based on episode analysis; this does not confirm that the show is currently accepting guests.
Host of Sample Space?
Claim your podcast to manage its listing and keep your show details accurate.
Pod Engine is an independent podcast discovery and analytics service and is not affiliated with or endorsed by this podcast. Artwork and show content belong to their owners. Full legal notice.
Explore this show Podcast research with Pod Engine