Your curated collection of saved posts and media
Arm has released a set of hands-on labs designed to help developers deploy PyTorch models using ExecuTorch across real edge devices. The labs walk-through: โข Exporting models into lightweight .pte artifacts โข Running optimized inference on Arm CPUs (XNNPACK + KleidiAI) โข Offloading workloads to Ethos-U NPUs using TOSA and Vela โข Visualizing model partitioning and performance behavior Itโs a practical way to understand how models are executed across heterogeneous computeโand how to optimize for it. If youโre working on edge AI or exploring on-device inference, this is worth a look: https://t.co/cml4cBVF3Q #PyTorch #ExecuTorch #OpenSourceAI
๐จ WARNING: The self-spreading โMini Shai-Huludโ worm compromised npm & PyPI packages tied to TanStack, Mistral AI, Guardrails AI, OpenSearch & more. The attack used GitHub OIDC token hijacking and cache poisoning to spread credential-stealing malware across 42 TanStack packages and 84 versions. Check your dependencies immediately โ https://t.co/33fxlrOPzz
Don't tell me Reachy Mini isn't the most adorable robot ever! https://t.co/BOGd95ann4
Due to rising RAM prices and tariff-related costs, Reachy Mini pricing will change on June 1st. Good news: you still have a bit more than 2 weeks to order at the Early Bird prices. Order here: https://t.co/UcMxyyfYO9 https://t.co/Oq4LjBlWUl

Meet physics-intern๐งโ๐, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics is hard for humans and LLMs alike. But physics-intern decomposes problems and dispatches them to a team of specialized agents, solving research-level questions far more effectively than the base model alone.
Need document parsing that stays fully local and private? ๐ Meet liteparse-server, a self-hostable, open-source HTTP server for parsing documents and generating screenshots from PDFs, Office files, and images. โ 100% self-hosted โ Private by default โ Open source โ Built for production deployments Deploy it as: ๐ณ a @Docker container โก or a serverless Express.js API It also integrates easily with: - @Redisinc for caching and rate limiting - @opentelemetry-compatible collectors for traces and metrics - observability tools like @JaegerTracing, @PrometheusIO and @grafana Read the full breakdown here: https://t.co/E3y2ZHvURm GitHub repo: https://t.co/K0d8XVEFGK
For our free newsletter this week, we discuss how AI-powered cyberattacks are a growing risk. And this morning a ton of hacks are being disclosed, as @emollick points out here. โจ@IrenaCronin and I write this newsletter every week. ย AI-powered cyberattacks are a growing risk because AI helps hackers move faster and create more convincing attacks. They can use AI to find software flaws, build attack tools, automate parts of cyber operations, and write phishing messages that look more realistic. Businesses need to respond with stronger security basics, including faster patching, better login protections, employee training, software visibility, and AI-aware defenses. Subscribe and read for free: https://t.co/HHwYy7NoAl
Expect your feed to look more and more like this in the coming weeks and months. https://t.co/kcnHnnCVya
Cool idea from Nous Research. What if you could speed up long-context pretraining with a subquadratic wrapper that you remove before deployment? That is the idea behind Lighthouse Attention. The method wraps ordinary SDPA with a hierarchical, gradient-free selection layer that compresses and decompresses queries, keys, and values symmetrically, preserving left-to-right causality. Crucially, it can be removed near the end of training in a short recovery phase, so the deployed model still runs vanilla attention with no architectural cost at inference. Preliminary LLM experiments report faster total training time and lower final loss than full-attention baselines. Why does it matter? Most efficient-attention work either changes the deployment-time architecture or pays a quality tax to do so. A training-only wrapper that survives a clean recovery phase sidesteps both. If it scales, this becomes an important training-time speedup for long-context pretraining. Paper: https://t.co/9g5Ldnb1rV Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
We just crossed 1,000,000 public datasets on Hugging Face! That's petabytes of data available that millions of AI builders are downloading, analyzing, and training AI models on every day! What's interesting is that we see a clear acceleration since agents started to be good as the number of datasets doubled over the past 8 months (it took 4 years to reach the first 500k). It's becoming easier and faster to build, share and use your own datasets! Many are saying the next bottleneck for more people to build AI themselves (instead of relying on APIs) is better data so we're just getting started! Thanks everyone for your amazing contributions, we couldn't do it without you!
Don't tell me Reachy Mini isn't the most adorable robot ever! https://t.co/BOGd95ann4
Don't tell me Reachy Mini isn't the most adorable robot ever! https://t.co/BOGd95ann4
NEW paper from Google DeepMind. (bookmark it) AI Co-Mathematician is an agentic research workbench for mathematicians, and it just hit 48% on FrontierMath Tier 4, a new high score among AI systems evaluated. The system is an asynchronous, stateful environment that supports ideation, literature discovery, computational analysis, theorem verification, and knowledge development. It manages uncertainty, clarifies intentions, records unsuccessful attempts, and emits formal mathematical outputs. Early applications yielded solved open problems, fresh research angles, and recovered overlooked citations during active research sessions. This is one of the cleanest demonstrations that agentic AI moves the frontier on genuinely hard mathematical research, not just problem-solving but discovery support. The asynchronous stateful workbench design is interesting to adopt if you are building agents for any expert workflow. Paper: https://t.co/C1ro3mGPQi Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c
https://t.co/a0sKQPZSz3
We've just hit 1M open datasets on the Hugging Face Hub ๐ Open models need open data. Today we hit that milestone, together with the most incredible community in AI! ๐ค Onwards to the next million ๐ https://t.co/PV6knP3XlJ
Expect your feed to look more and more like this in the coming weeks and months. https://t.co/kcnHnnCVya
Planning your trip to San Jose? ๐ซ Join the #PyTorch community at #PyTorchCon North America, Oct 20-21. Register by July 31 & keep $400 in your pocket with early bird pricing. Secure your pass: https://t.co/AVHdaIFT20 https://t.co/JpGdQYzDDb
Reachy is mad, but RAM costs + tariffs are forcing our hand. Prices will go up on June 1st! Still at the early bird price until then though if you were looking for an excuse to get one now: https://t.co/veqPEwFIaP! https://t.co/UP45svdMr8
We just hit #1 trending on @huggingface Spaces ๐ โThe Ultimate Guide to RL Environmentsโ dives into building & scaling RL environments for LLMs. If you're exploring RL + agents, this might be useful https://t.co/2bbwtic6xN
Excited to release the Ultimate guide to RL environments! Definitions of RL environments differ wildly in the LLM era, so we spent the last month building several RL environments across 6 different frameworks, domains and complexities to map out which are easiest to build with a
Unsloth just published MTP-enabled quantized GGUFs for Qwen3.6-35B-A3B. https://t.co/9iuepdo5AW
Unsloth just published MTP-enabled quantized GGUFs for Qwen3.6-35B-A3B. https://t.co/9iuepdo5AW
M3 Max users really got local AGI before GTA VI https://t.co/AfaFukk6jR
M3 Max users really got local AGI before GTA VI https://t.co/AfaFukk6jR
Meta silently dropped Sapiens2 last week ๐ฅ a family of high-res models trained on 1B human images > for pose estimation, body-part segmentation, surface normals, pointmaps (sota) > 6 sizes: 0.1B โ 5B params (all ViT patch 16) > high-res: 1024ร768 and 4K https://t.co/e2Nl7zAhKM
Reason-ModernColBERT nearly solved BrowseComp-Plus, smashing SOTA and outperforming models models 54ร bigger Not bad for a 1 year old model not optimized for deep research What if we actually tried? Introducing Agent-ModernColBERT: adding another 10% on top with a 5 min training
@IntCyberDigest POV: you are downloading packages in 2026 https://t.co/EHqMTy4wzl
We published new research on how we serve post-trained Qwen3 235B models on NVIDIA GB200 NVL72 Blackwell racks. GB200 is a major step up over Hopper for high-throughput inference on large MoE models, not just a training platform. https://t.co/yYZuPRXWzr
The battle over AI is increasingly becoming a battle over intellectual property. Elsevier joining the lawsuit against Meta shows that publishers, authors and researchers are pushing back against how AI models were trained. One thing is becoming clear. Data ownership may become one of the defining issues of the AI era. https://t.co/TovYOET7FN @nature
Itโs time! Get your Reachy mini at https://t.co/TVtTsxloZf https://t.co/rg4s4JtYzN
Itโs time! Get your Reachy mini at https://t.co/TVtTsxloZf https://t.co/rg4s4JtYzN
TMAS Scaling Test-Time Compute via Multi-Agent Synergy https://t.co/SVQfgTSiTK
paper: https://t.co/4XxBgCVgZG
Rebellious Student Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR https://t.co/JL6MJ3Txum