Best reads for PMs & Designers
Get 5 personalized best reads each week, with TL;DR and clear next steps.
One free email every Tuesday · No sponsored posts · See a sample email
Topics
Sources
First seen 24 January 2026
All right. Um, welcome everyone to yet another inference talk. I hope you have had a good conference so far. And u, so in this session, I mean I'm sure you people who have been in the room uh must have heard these terms many times by now. So we're going to do a little bit more deep dive into the challenges of LLM deployments for agentic workloads and uh in this session we'll focus specifically on KV cache away routing and uh PD disagregation um and also you know when you when you look at public inference uh benchmark results you are typically looking at very steady state isolated highly sanitized numbers and what those benchmarks actually don't show you u is the chaotic reality of multi-turn interactions, massive context fluctuations which are very typical of agentic workloads. So we'll also try to pull the curtain back on some of those complexities. Um by by way of introduction uh my name is Ashish Kamra. I'm a senior manager of performance engineering at Red Hat. And with me >> hi I'm Yuch Chen. I'm the product manager at Red Hat Inference working closely with VLM and AMD core mainta
YouTube: Sequoia Capitalyoutube.com · 25 August 2026
Parag Agrawal is making a bet that goes against two decades of web search: agents will query the web a thousand times more than humans ever have, and the infrastructure built around human clicks is wrong for them. The former Twitter CEO, now founder and CEO of Parallel Web Systems, explains why Parallel treats human click data as a bug and trains on agent feedback instead. He unpacks the counterintuitive choice to ship a search agent before a search engine, building an index incrementally, and how the new Turbo product cut agentic search to 200 milliseconds. But the problem Parag keeps returni