@PyTorch
PyTorch 2.10 is now optimized for @Intel Core Ultra Series 3 processors to bring high-performance AI to the PC and edge. This release leverages the new Xe3 architecture and Arc B-series GPUs to deliver up to 120 XMX TOPs. Native TorchAO integration enables seamless int4 weight-only quantization, allowing a Llama 3.2 3B Instruct model to run locally with a first token latency of 242.29 ms and a 2+ token latency of 27.24 ms. For edge scenarios, developers can achieve 1.4x to 1.7x faster training for vision models via Anomalib, with WinClip seeing gains up to 2.5x compared to previous generations. Read our latest blog from the Intel PyTorch and Client AI SW teams for the full technical deep dive and benchmarks: https://t.co/pEGpmwtROi #PyTorch #Intel #AIPC #EdgeAI #OpenSourceAI