@Modular
What does AI-assisted GPU kernel development actually look like? We used Cursor and Claude to port NVIDIA's CUTLASS Blackwell conv2d to Mojo 🔥 in one session. 90% matmul reuse, ~770 lines, 6.6x faster than cuDNN on B200 GPUs. All kernels in the Modular repo: https://t.co/Hv6x9rKTeS