We discuss seminal mathematical papers (sometimes really old đ ) that have shaped and established the fields of machine learning and data science as we know them today. The goal of the podcast is to introduce you to the evolution of these fields from a mathematical and slightly philosophical perspective.
We will discuss the contribution of these papers, not just from pure a math aspect but also how they influenced the discourse in the field, which areas were opened up as a result, and so on.
Our podcast episodes are also available on our youtube:
https://youtu.be/wThcXx_vXjQ?si=vnMfs
Data Science #34 - The deep learning original paper review, Hinton, Rumelhard & Williams (1985)
On the 34th episode, we review the 1986 paper, "Learning representations by back-propagating errors" , which was pivotal because it provided a clear, generalized framework for training neural networks with internal 'hidden' units. The core of the procedure, back-propagation, repeatedly adjusts the weights of connections in the network to minimize the error between the actual and desired output vectors. Crucially, this process forces the hidden units, whose desired states aren't specified, to develop distributed internal representations of the task domain's important features.This capability to construct useful new features distinguishes back-propagation from earlier, simpler methods like the perceptron-convergence procedure. The authors demonstrate its power on non-trivial problems, such as detecting mirror symmetry in an input vector and storing information about isomorphic family trees. By showing how the network generalizes correctly from one family tree to its Italian equivalent, the paper illustrated the algorithm's ability to capture the underlying structure of the task domain.Despite recognizing that the procedure was not guaranteed to find a global minimum due to local minima in the error-surface , the paper's clear formulation (using equations 1-9 ) and its successful demonstration of learning complex, non-linear representations served as a powerful catalyst.
It fundamentally advanced the field of connectionism and became the standard, foundational algorithm used today to train multi-layered networks, or deep learning models, despite the earlier, lesser-known work by Werbos
3 Nov 2025
Data Science #33 - The Backpropagation method, Paul Werbos (1980)
On the 33rd episdoe we review Paul Werbosâs âApplications of Advances in Nonlinear Sensitivity Analysisâ which presents efficient methods for computing derivatives in nonlinear systems, drastically reducing computational costs for large-scale models. Werbos, Paul J. "Applications of advances in nonlinear sensitivity analysis." System Modeling and Optimization: Proceedings of the 10th IFIP Conference New York City, USA, August 31âSeptember 4, 1981These methods, especially the backward differentiation technique, enable better sensitivity analysis, optimization, and stochastic modeling across economics, engineering, and artificial intelligence. The paper also introduces Generalized Dynamic Heuristic Programming (GDHP) for adaptive decision-making in uncertain environments.Its importance to modern data science lies in laying the foundation for backpropagation, the core algorithm behind training neural networks. Werbosâs work bridged traditional optimization and todayâs AI, influencing machine learning, reinforcement learning, and data-driven modeling.
19 Sept 2025
Data Science #32 - A Markovian Decision Process, Richard Bellman (1957)
We reviewed Richard Bellmanâs âA Markovian Decision Processâ (1957), which introduced a mathematical framework for sequential decision-making under uncertainty.
By connecting recurrence relations to Markov processes, Bellman showed how current choices shape future outcomes and formalized the principle of optimality, laying the groundwork for dynamic programming and the Bellman equationThis paper is directly relevant to reinforcement learning and modern AI: it defines the structure of Markov Decision Processes (MDPs), which underpin algorithms like value iteration, policy iteration, and Q-learning.
From robotics to large-scale systems like AlphaGo, nearly all of RL traces back to the foundations Bellman set in 1957
Reach and audience
Public platform figures. Ratings count people who left a rating, not total listeners.
YouTube views
23,047
Score snapshot 15 Sept 2025
Apple Podcasts (US)
4.0 / 5
6 ratings
Podcast Authority Score: 49 / 100
A composite of feed quality, social presence, YouTube performance and engagement. Read the methodology.
Quality
52
Social presence
0
YouTube
85
Engagement
31
Host of Data Science Decoded?
Claim your podcast to manage its listing and keep your show details accurate.
Pod Engine is an independent podcast discovery and analytics service and is not affiliated with or endorsed by this podcast. Artwork and show content belong to their owners. Full legal notice.
Explore this show Podcast research with Pod Engine