The Kubernetes AI Show dives deep into the real-world challenges of adopting AI on Kubernetes platforms.

The AI Kubernetes Show
Claim This Podcastby The AI Kubernetes Show
Podcast Overview
The Kubernetes AI Show dives deep into the real-world challenges of adopting AI on Kubernetes platforms.
Language
🇺🇲
Publishing Since
11/28/2025
1 verified contact email on file for The AI Kubernetes Show
Pitch yourself as a guest, propose sponsorships, or reach out directly to the host.
Recent Episodes

August 12, 2026
Disaggregated Inference on Kubernetes
<p>Every LLM inference request splits into two phases with completely different hardware profiles: pre-fill and decode. But most Kubernetes clusters still run them on the same pod. Morgan Foster, who works on the llm-d project and the AI Gateway Working Group out of Red Hat's Office of the CTO, joins William Morgan to explain what that means. </p><p>Foster and Morgan get into why real-time inference workloads fit the primitives Kubernetes already has, while training often runs on Slurm instead. They walk through disaggregated inference: how pre-fill and decode get scheduled on separate pods, how KV cache blocks move between them over RDMA, and what happens today when a pre-fill or decode pod dies mid-request (nothing handles the failover cleanly, and the computed KV blocks just get thrown away). They also cover why faster token generation buys agentic systems more capability, not just quicker responses, and why Kubernetes proxies, built around header-based ingress routing, are struggling to govern AI traffic that depends on streamed JSON request bodies and inference happening outside the cluster. </p><p>FIND AND FOLLOW US ON: </p><p>✦ Spotify: <a href="https://open.spotify.com/show/4LLTNjMbd7wNCYhSI9pkQU?nd=1&dlsi=3baba6bcc6b44cae"><u>https://open.spotify.com/show/4LLTNjMbd7wNCYhSI9pkQU?nd=1&dlsi=3baba6bcc6b44cae</u></a> </p><p>✦ Apple Music: <a href="https://podcasts.apple.com/us/podcast/the-ai-kubernetes-show/id1895650894"><u>https://podcasts.apple.com/us/podcast/the-ai-kubernetes-show/id1895650894</u></a> </p><p>TAKEAWAYS</p><p>✓ Why real-time inference workloads fit Kubernetes' existing primitives, while training often runs on Slurm instead </p><p>✓ How disaggregated inference splits the pre-fill and decode phases of an LLM request across separate pods </p><p>✓ Why KV cache blocks move between pods over RDMA, and what happens if the pre-fill or decode pod dies mid-transfer </p><p>✓ Why more tokens per second buys agentic systems more capability, not just faster responses </p><p>✓ Why Kubernetes proxies built for header-based routing struggle with AI traffic that depends on streamed request bodies </p><p>✓ What the AI Gateway Working Group is building to fix it</p><p>Read the blog post at: <a href="https://www.buoyant.io/ai-kubernetes-episode/disaggregated-inference-on-kubernetes" target="_blank" rel="noopener noreferer">https://www.buoyant.io/ai-kubernetes-episode/disaggregated-inference-on-kubernetes</a></p><p><br></p>

July 29, 2026
The State of Cloud Native AI and What’s Still Evolving
<p>There is no agreed upon definition of what cloud native AI means yet, but if you're trying to figure out where Kubernetes AI standards are really headed, check out this episode with someone who's worked on the CNCF's first attempt to define cloud native AI. </p><p>William Morgan talks with Ron Petty (RXM, CNCF founding contributor, and member of the CNCF's TCG AI) about what's actually settled in cloud native AI and what isn't. They cover the CNCF's shift from working groups to technical community groups (TCGs), Petty's own framing of cloud native AI as a two-way street between building models and using AI to run cloud native systems, and the 15-criteria self-reported conformance program (mostly GPU management) that gets you an "AI certified" stamp. </p><p>TAKEAWAYS </p><p>✓ Why Ron Petty defines cloud native AI as a two-way street: building models on cloud native infrastructure, and using AI to run cloud native systems </p><p>✓ What the CNCF's self-reported AI conformance program actually checks, including GPU management and having an AI router on your cluster </p><p>✓ Why certificate management and rotation is still unsolved, and how that connects to why Linkerd defaults to mutual TLS </p><p>✓ What's actually wrong with Model Context Protocol's design, according to a proxy engineer's read on it </p><p>✓ Why long-running AI agents don't need fundamentally different guardrails than short-running ones </p><p>✓ How K8sGPT uses existing Kubernetes checks plus an LLM to help junior engineers diagnose cluster issues faster</p><p>Read the blog post at: www.buoyant.io/ai-kubernetes-episode/the-state-of-cloud-native-ai-and-whats-still-evolving</p><p><br></p>

July 15, 2026
Treat Testing as a Platform Service on Kubernetes
<p>Most testing tools bolt onto a CI pipeline and hope for the best. Ole Lensmar, CTO of Testkube, built a company that changes that script. Pull test execution out of CI/CD entirely and run tests as Kubernetes workloads instead.</p><p>In this episode, Lensmar walks through why Testkube uses Kubernetes as its test execution engine, running Selenium, Playwright, K6, JMeter, and Postman tests as cluster workloads instead of asking teams to adopt new tools. He and podcast host William Morgan dig into who should own testing (developers own functional tests, while the platform team owns performance, security, and chaos testing). Lensmar's core recommendation is to treat testing and quality as a platform capability, the same way you already treat CI and CD. </p><p>More AI-generated code means more tests are needed, intelligent test selection can keep pipelines fast as test counts grow, and AI is already useful for triaging failed test logs using Kubernetes and Grafana MCP servers. </p><p>TAKEAWAYS: </p><p>✓ Where testing ownership splits between developers and the platform team, and where it doesn't </p><p>✓ Why "testing as a platform service" should get the same priority as CI and CD </p><p>✓ How intelligent test selection works, and why you should never rely on it alone </p><p>✓ Where AI already adds value today: triaging failed test logs with Kubernetes and Grafana MCP servers </p><p>✓ Why testing the AI components of your application (agents, evals) is becoming its own discipline.</p><p>Read the blog post at: <br></p>
27 total episodes available
Deep-dive analytics for The AI Kubernetes Show
Frequently asked questions
Have a different question and can't find the answer you're looking for? Reach out to our support team by sending us an email and we'll get back to you as soon as we can.
- What is The AI Kubernetes Show?
- How often does this podcast release new episodes?
This podcast updates daily.
- Where can I listen to this podcast?
This podcast is available on 4 platforms including Apple Podcasts, Spotify, and more. You can also use the RSS feed directly.
- Does this podcast accept guests?
Yes, this podcast regularly features guests.
Legal Disclaimer
Pod Engine is not affiliated with, endorsed by, or officially connected with any of the podcasts displayed on this platform. We operate independently as a podcast discovery and analytics service.
All podcast artwork, thumbnails, and content displayed on this page are the property of their respective owners and are protected by applicable copyright laws. This includes, but is not limited to, podcast cover art, episode artwork, show descriptions, episode titles, transcripts, audio snippets, and any other content originating from the podcast creators or their licensors.
We display this content under fair use principles and/or implied license for the purpose of podcast discovery, information, and commentary. We make no claim of ownership over any podcast content, artwork, or related materials shown on this platform. All trademarks, service marks, and trade names are the property of their respective owners.
While we strive to ensure all content usage is properly authorized, if you are a rights holder and believe your content is being used inappropriately or without proper authorization, please contact us immediately at hey@podengine.ai for prompt review and appropriate action, which may include content removal or proper attribution.
By accessing and using this platform, you acknowledge and agree to respect all applicable copyright laws and intellectual property rights of content owners. Any unauthorized reproduction, distribution, or commercial use of the content displayed on this platform is strictly prohibited.
