MBA Training AI. The daily podcast for everyone learning to use AI well at work. Foundations of LLMs, prompting, AI agents, responsible AI, and getting the most out of ChatGPT, Claude and Gemini. New episode every day at mba-training.com.
GPT-6 Intelligent UI: why ChatGPT calculators need auditing
OpenAI rolled out GPT-6 to all ChatGPT tiers with Intelligent UI, which answers with charts, forms, buttons and small working calculators. Is it a real upgrade? Independent Artificial Analysis scores show GPT-6 Sol and Luna about level with GPT-5.6 at half the cost, with a lower hallucination rate but more non-answers. The episode argues the polished tools can look more trustworthy than they are.
You will be able to read the 44% time to first token claim correctly, weigh vendor benchmarks against independent ones, compare Luna with Anthropic's Claude Haiku 5.5 pricing and its 100,000 token threshold, check Enterprise admin settings, and rebuild a ChatGPT calculator in your own spreadsheet to verify its formula and inputs.
0:00 GPT-6 Sol and Luna rollout by tier
1:08 Why interactive calculators are harder to audit
2:30 Time to first token speed claims
3:25 Artificial Analysis benchmarks and hallucination rate
5:59 Claude Haiku 5.5 pricing versus Luna
7:12 Interface hype and a spreadsheet audit
8 Oct 2026
Fintech AI timelines: the evidence clock beats the build clock
How long does it really take to put an AI agent live in regulated finance? Vendors such as Gradient Labs quote days, and Nous Research just raised $90 million for Hermes for Businesses. The episode argues the build was never the delay. The real timeline is the evidence clock: the time needed to observe enough rare worst-case outcomes, with SR 26-2 excluding agentic AI and EU AI Act credit scoring rules arriving December 2027.
Listeners learn to separate value metrics from guardrail metrics, apply the rule of three to size a shadow mode run, assign guardrail ownership outside the project team, read the Klarna case, and cross out go-live dates that ignore base rates.
0:00 Nous Research funding and agent hype
2:06 SR 26-2 and EU AI Act deadlines
3:43 Guardrail metrics and Klarna's lesson
5:13 Rule of three shadow mode math
6:35 FCA testing and vendor go-live claims
8:25 Guardrail swaps and pilot traps
6 Oct 2026
ChatGPT Ads: the measurement stack behind image ads
Did OpenAI launch ads inside ChatGPT image generation, or announce a test? The episode separates the two, then argues the ad format is the smaller story. The measurement stack shipped alongside it, attribution partners like AppsFlyer, Triple Whale, Adjust, Branch and Northbeam, plus geo-based incrementality work with Haus, Measured and WorkMagic, is what turns a placement into an advertising business.
You come away able to read partner-reported results with the right scepticism, including WorkMagic's Dose study and Triple Whale's Portland Leather figure, understand view-through conversions and why 52.7% landing inside an hour matters, and judge the DoubleVerify and Integral Ad Science pilots that exclude user conversations. Practical steps: check the partner list against your analytics stack and write your own evidence standard.
0:00 What OpenAI actually announced on October 5
0:46 Attribution and incrementality behind the ad unit
1:48 Partner-reported lift numbers and view-through conversions
3:17 Can anyone audit the ads-don't-influence-answers claim
4:45 The case for ads and the inventory shortage
6:46 What enterprise AI teams should do this week
Contact Leaders Insights — AI
Guest appearances
Does not typically book guests
Based on episode analysis; this does not confirm that the show is currently accepting guests.
Host of Leaders Insights — AI?
Claim your podcast to manage its listing and keep your show details accurate.
Pod Engine is an independent podcast discovery and analytics service and is not affiliated with or endorsed by this podcast. Artwork and show content belong to their owners. Full legal notice.
Explore this show Podcast research with Pod Engine