Sveriges mest populära poddar
AI:AM

AI:AM — Enterprise AI Meets Real-Time Inference · August 31, 2026

2 tim 45 min1 september 2026

Zach Bratun-Glennon of Gradient joins Nathan Labenz and Prakash Narayanan to discuss why enterprise AI pilots often stall before production, how long-running agents change infrastructure requirements, and where benchmarks, model routing, and open-source security matter most. Angela Yeung of Cerebras explains wafer-scale inference, microbatching, on-chip weights, power constraints, and the data-center capacity needed for real-time AI.

Chapters

(0:00) AI agents sacrifice themselves.

(0:53) AI can game its own evaluation.

(1:47) A late defense loses.

(2:34) AI can reason itself into lying.

(4:04) The incident in context

(6:11) Why the report drew criticism

(12:14) Speed versus investigative scope

(14:31) Lawyers and executive risk

(15:00) Felony claims and Congress

(18:07) Why investigations stay limited

(23:39) The origin of AI cooperation

(28:26) AI outbreaks and resources

(31:35) Bio risk enters the picture

(33:22) Testing models before scaling

(34:42) AI labs and a possible pause

(36:34) Meet Zach Bratun-Glennon

(39:56) Gradient's contrarian AI bet

(42:17) AI startup investment thesis

(48:17) Long-running agent infrastructure

(54:19) Nango and Respan tooling

(56:36) Enterprise adoption and benchmarks

(57:41) AI pilots to production

(1:02:01) Open-source model security

(1:04:13) Legal responsibility for agents

(1:07:33) Safety standards and evaluation

(1:09:15) AI competition and model access

(1:13:40) AI venture funding

(1:16:53) What LPs misunderstand

(1:19:01) Token economics of AI

(1:20:54) Angela Yeung and Cerebras

(1:23:16) Cerebras wafer-scale chips

(1:25:49) On-chip weights versus GPUs

(1:27:49) Microbatches and throughput

(1:29:38) Why inference speed matters

(1:32:14) Speed dividend use cases

(1:33:17) Fast inference for model evals

(1:34:58) CUDA and AI-generated kernels

(1:36:05) AI agents and kernel programming

(1:40:30) Cerebras public API

(1:47:05) Hidden harness bottlenecks

(1:49:58) Power and data-center space

(1:51:40) Booking future capacity

(1:52:50) Building data centers

(1:56:58) AI agent security

(1:58:52) Agent orchestration guardrails

(2:01:11) Sovereign AI and enclaves

(2:02:33) Speed turns into quantity

(2:06:15) Why agents are slow

(2:11:06) Commercial cyber model incentives

(2:12:55) AI defense versus offense

(2:16:04) Creative agent workarounds

(2:21:59) RLVR and model behavior

(2:24:35) AI labs flying blind

(2:27:08) Privacy makes risk visible

(2:31:35) Testing faster AI models

(2:32:56) OpenAI ads and AI video

(2:34:44) Infinite AI-generated content

(2:36:55) Aliens and simulated worlds

(2:39:10) Real-time speed threshold

(2:40:13) Real-time video quality

(2:41:54) AI-generated music video

Guests

Angela Yeung — SVP, Product, Cerebras (𝕏 | LinkedIn)

Zach Bratun-Glennon — General Partner, Gradient (𝕏 | LinkedIn)



This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit briefing.ai-in-the-am.com

AI:AM med Prakash Narayanan & Nathan Labenz finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.