@charles_irl
There was a flippening in the last few months: you can run your own LLM inference with rates and performance that match or beat LLM inference APIs. We wrote up the techniques to do so in a new guide, along with code samples. https://t.co/vpNGPZ0vgM https://t.co/wCaWIXwBf6