@ryanshrout
The most anticipated announcement of the keynote: NVIDIA is announcing a new LPU, the @NVIDIA @GroqInc 3 LPU, that pairs with Vera Rubin NVL72 via a dedicated LPX rack connected over Direct C2C. Claims are 35x inference throughput over Blackwell for trillion-parameter models. The idea is to combine GPU throughput with LPU latency to push the Pareto curve on high-interactivity workloads. Lots of questions still on cost, capacity tradeoffs, and real-world deployment, but the architecture concept is fascinating.