@charles_irl
Inference isn't everything, but it does require a new stack -- not Kubernetes, not SLURM. At @modal, we dove deep to build that stack. In this blog post we explain how, from compute management & cloud-native cacheing to CRIU & GPU checkpointing. https://t.co/DQ4wvuXjre https://t.co/iF0ZYJQWFL