@NehallAgrawal
π[LIVE] on day 2 of @aiDotEngineer πΈπ¬ Daria Soboleva from @cerebras on building mixture of experts at scale. βGPU scaling hits a wall when communication becomes the bottleneck. The real breakthrough is keeping massive expert networks on-chip without model partitioning.β https://t.co/HYZpBeWtIZ