@binsquares
omg, GPU acceleration on smolvm works way better than I thought. can run llama.cpp inside the smol machine with close to ~90% of performance of the host using vulkan backend. ~127 t/s on qwen-0.5b. Min hassle, max ease of setup. https://t.co/SQzj7PTs9P