Sveriges mest populära poddar

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

Speculative Decoding and Efficient LLM Inference with Chris Lott - #717

1 tim 17 min•4 februari 2025

Today, we're joined by Chris Lott, senior director of engineering at Qualcomm AI Research to discuss accelerating large language model inference. We explore the challenges presented by the LLM encoding and decoding (aka generation) and how these interact with various hardware constraints such as FLOPS, memory footprint and memory bandwidth to limit key inference metrics such as time-to-first-token, tokens per second, and tokens per joule. We then dig into a variety of techniques that can be used to accelerate inference such as KV compression, quantization, pruning, speculative decoding, and leveraging small language models (SLMs). We also discuss future directions for enabling on-device agentic experiences such as parallel generation and software tools like Qualcomm AI Orchestrator.

The complete show notes for this episode can be found at https://twimlai.com/go/717.

Fler avsnitt av The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

Why AI Agents Break the GenAI Security Model with Devvret Rishi - #770

16 juni•56 min

Is RAG Dead? Lessons from Building AI for Tax Law with Alex Bowcut - #769

9 juni•52 min

Relational Foundation Models for Enterprise Data with Jure Leskovec - #768

21 maj•1 tim 6 min

How to Find the Agent Failures Your Evals Miss with Scott Clark - #767

7 maj•53 min

How to Engineer AI Inference Systems with Philip Kiely - #766

30 apr.•55 min

How Capital One Delivers Multi-Agent Systems with Rashmi Shetty - #765

16 apr.•54 min

The Race to Production-Grade Diffusion LLMs with Stefano Ermon - #764

26 mars•1 tim 3 min

Agent Swarms and Knowledge Graphs for Autonomous Software Development with Siddhant Pardeshi - #763

10 mars•1 tim 16 min

AI Trends 2026: OpenClaw Agents, Reasoning LLMs, and More with Sebastian Raschka - #762

26 feb.•1 tim 19 min

The Evolution of Reasoning in Small Language Models with Yejin Choi - #761

29 jan.•1 tim 6 min

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) med Sam Charrington finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.