Your curated collection of saved posts and media
HeartMuLa A Family of Open Sourced Music Foundation Models https://t.co/TJzg6eMEXZ
paper: https://t.co/2ALsoeOEji
model: https://t.co/W7HamtcMzs
STEP3-VL-10B Technical Report https://t.co/TiSlflEB58
SeedFold Scaling Biomolecular Structure Prediction https://t.co/XFVd625BEW
Transition Matching Distillation for Fast Video Generation https://t.co/YXFty6ul0W
discuss: https://t.co/YWnc9zFtNO
RigMo Unifying Rig and Motion Learning for Generative Animation https://t.co/FkionIZZLr
discuss: https://t.co/doYeEa9MO3
What makes preference data truly effective for LLM alignment? π€ Introducing AIR: A systematic framework that deconstructs preference datasets into 3 core components (Annotations, Instructions, Response Pairs) and reveals evidence-based optimization principles. No more trial and error! π€ Paper: https://t.co/dFxRHSux2W π arXiv: https://t.co/QJl1T1uGyi Why it matters: 1οΈβ£ Simplicity Wins in Annotations: Basic point-wise scoring with generative models (like Llama-3.1-70B-Instruct) + greedy decoding outperforms complex methods. Less is moreβexcessive design introduces noise rather than clarity. 2οΈβ£ Smart Instruction Filtering: Select instructions with low response variance across LLMs. This forces models to learn fine-grained preferences (like logical rigor) rather than relying on obvious differences. 3οΈβ£ Balanced Response Pairs: Optimal pairs combine moderate score gaps (Ξ=2-3), high absolute quality (β₯8), and 1:1 On/Off-Policy mixingβachieving clear contrast without overfitting. The results? +5.3 average gain across 6 benchmarks (WildBench, Arenahard, etc) with just 14k curated pairs from 17 open-source LLMs covering coding, math, and chat tasks. AIR transforms preference learning from "scale blindly" to component-aware designβa blueprint for building smarter, more aligned AI systems. π #AI #LLM #RLHF #PreferenceLearning #Alignment
Alterbute Editing Intrinsic Attributes of Objects in Images https://t.co/BGBFK5w6eE
Tiny model, big orbit. https://t.co/zssZcL2T8C
Tiny model, big orbit. https://t.co/zssZcL2T8C
Download weights: @kaggle β https://t.co/DfOtwrIBiS @huggingface β https://t.co/iteHS5eolu

Download weights: @kaggle β https://t.co/DfOtwrIBiS @huggingface β https://t.co/iteHS5eolu

Built on Gemma 3, TranslateGemma was trained on data generated by Gemini β effectively transferring its intelligence into a smaller package. This means developers can build low-latency translation tools that run entirely on-device. Try it now on @huggingface and @Kaggle β https://t.co/M7wJmngSCk
Today we're releasing our first official LoRA: INFL8 Trained on the incredible @Alibaba_Qwen Image Edit 2511 model, it can inflate anything! Links to the model and @huggingface space below so you can try it out for yourself β¬οΈ https://t.co/sT2E83vsXi
Finally! We (the community + @OpenAIDevs + @huggingface ) bring you an open standard for inference. It's called 'Open Responses' it's based on Responses and it's perfect for agent workloads. Fewer special cases, more consistency, faster shipping. Excited for what this unlocks. Below is a deep dive blog post, weβll look at how Open Responses works and why the open source community should use Open Responses.
Crazy ! π€― I was sceptical so I had to check the dataset See for yourself, the truth is here: there are indeed white spaces indicating the true answer in some examples (thanks to the AI assistant on @huggingface for the SQL query) https://t.co/5nXLgHCn3f
The presence of a leading whitespace leaks the correct choice selection in the MMLU-Pro benchmark. Am I missing something? Seems to impact Chemistry, Physics, and Math. HF Issue in reply. https://t.co/FdtKNvIX9m
TranslateGemma is out https://t.co/jgoHfqjiJH
TranslateGemma is out https://t.co/jgoHfqjiJH
Benchmark Leaderboards seen in the wild on @huggingface ππ https://t.co/nsHvwDCHwy
Benchmark Leaderboards seen in the wild on @huggingface ππ https://t.co/nsHvwDCHwy
Introducing FLUX.2 [klein]. Blazing fast. Beautiful. Generate stunning images in under a second while maintaining exceptional quality. Great for fast editing, changing styles, and developing ideas from 0 β 1. Available via API, or run it locally - Klein 4B under Apache 2.0, Klein 9B as open weights. Try it for free in our demo app (link in the thread).
New gated models are rolling into Microsoft Foundryβstarting with vision, language, and multilingual models from Hugging Face. Explore whatβs available and how access works here: https://t.co/r08oUVmk51 https://t.co/RT8mObCOsx

Welcome π https://t.co/tQM1xp4SS6
Pro https://t.co/qNmlCkzd48
Welcome π https://t.co/tQM1xp4SS6
β¨ One year after the debut of "Project Digits," DGX Spark has transformed from concept to reality and is now available from NVIDIA and OEM partners. See it in action lighting up the show floor at #CES2026. #SparkSomethingBig https://t.co/6UcfxWqsw6
π Introducing OctoCodingBench, a new benchmark for aligned coding agents: https://t.co/oKaF7jjagb Passing tests β aligned behavior. An agent can produce code that aces every unit test while ignoring system guidelines, violating project conventions, or misusing tools. In real-world coding, how you solve matters as much as what you solve. Nobody wants an agent that ships perfect code while deleting your README, reformatting every file, and mass-commenting in LLM-ese. Don't let your coding agent paperclip-max your repo!
π Introducing OctoCodingBench, a new benchmark for aligned coding agents: https://t.co/oKaF7jjagb Passing tests β aligned behavior. An agent can produce code that aces every unit test while ignoring system guidelines, violating project conventions, or misusing tools. In real-world coding, how you solve matters as much as what you solve. Nobody wants an agent that ships perfect code while deleting your README, reformatting every file, and mass-commenting in LLM-ese. Don't let your coding agent paperclip-max your repo!
FLUX.2-klein-9B https://t.co/5vaiIKiSot
FLUX.2-klein-9B https://t.co/5vaiIKiSot