@PyTorch
TorchAO Quantized Gemma-3 Models Now Available on HuggingFace π TorchAO quantized Gemma-3 models (4b, 12b and 27b) are now available in multiple quantized formats (FP8, INT4, AWQ-INT4, HQQ-INT8-INT4, QAT-INT4) on HuggingFace Hub! Our models: * Run efficiently on A100/H100 GPUs (with vLLM) & mobile devices (with ExecuTorch) * Deliver faster inference, big memory savings, and minimal accuracy loss compared to BF16 model * Include reproducible quantization recipes for your own models, including algorithms like AWQ, HQQ and QAT ποΈ Check out the collection: https://t.co/hAEeQuqSIL #LLM #PyTorch #TorchAO #ExecuTorch #VLLM #Quantization #Gemma3 #HuggingFace #EdgeAI