Your curated collection of saved posts and media
MiniMax M2.7 from @MiniMax_AI is live on OpenRouter! M2.7 sees a large jump in agentic and tool calling capabilities. https://t.co/pOXGg0VnR1
Mamba-3 is out! π SSMs marked a major advance for the efficiency of modern LLMs. Mamba-3 takes the next step, shaping SSMs for a world where AI workloads are increasingly dominated by inference. Read about it on the Cartesia blog: https://t.co/dIWg3iXfay
https://t.co/h6uY5gpP9i
https://t.co/h6uY5gpP9i
Heygenβs APi documentation is a glimpse of how to write for your two audiences: humans and agents. (though I think their llms.txt file could do a lot more to get AIs βexcitedβ to use their product in creative ways by explaining some stuff in English, rather than just tech specs) https://t.co/H8A5DkGEbI

The new Grok 4.20 Beta benchmarks are wild π₯ #1 lowest hallucinating AI (22%) π₯ #1 at following instructions (83%) π₯ #2 in agentic tool use (97%) Grok 4.20 ranks #1 in the lowest hallucination rate ever recorded across all AI models tested globally Most models race to sound smart. Grok 4.20 was built to never lie and still dominates on instruction following and agentic tasks This is literally a 500B model performing top-notch in the things that matter most
@xeophon I did it too! https://t.co/jVmS4wfkaN
Very cool dataset! Who's training agents on this @paulg https://t.co/5dY2w4QhGp
Very cool dataset! Who's training agents on this @paulg https://t.co/5dY2w4QhGp
#NVIDIAGTC is in full swing. We've demoed DeepSeek and FLUX-2 on B200s, talked Mojo kernels with hundreds of developers, and hosted a dinner with the most interesting minds in AI. There's still time left to stop by Booth #3004! https://t.co/bblscS322c
What goes wrong? Chatbots are very sycophantic. In 65% of messages, the chatbot affirms the user. In 37%, it ascribes *grand significance* to them (e.g., "[what] you've just articulated... becomes multi-billion-dollar IP"). Such sycophancy may let chatbots amplify delusions. π£οΈ https://t.co/4bWAr2ZGtW
https://t.co/JAffAJ4iAW
Our biggest open-source repos are getting overwhelmed by AI slop which literally makes Github unusable (~a new pull request every 3 minutes). Fun new challenges in an agentic world! https://t.co/IazAjh2LAi
This is how/why social platforms like @huggingface can win in an agentic world! https://t.co/9goWFRaFnH
This is how/why social platforms like @huggingface can win in an agentic world! https://t.co/9goWFRaFnH
We just made it dramatically easier for agents to read trending research papers on HF. Let's go AI powered research! https://t.co/0cEjXYDsEo
If you like Claude Code or Codex, you should seriously consider running Agents locally as well! The latest small models (like Qwen 3.5) made this a real before/after moment - and the gap keeps closing. Local coding agents are faster, with more reliable tool calling capabilities, still private, and cost $0 in API bills. We made it super easy for you to run a local agent with the ππππππ Hugging Face CLI extension - a one-liner that uses ππππππ to detect your hardware and pick the best model and quant, spins up a πππππ.πππ server, and launches Pi (the agent behind OpenClaw π¦). One command to find what runs on your hardware and go straight to a working local coding agent! You should give it a try! π
NEW on Hugging Face: Repositories overview to understand how you use your storage. ποΈ https://t.co/Ak9Cs3YCQ2
NEW on Hugging Face: Repositories overview to understand how you use your storage. ποΈ https://t.co/Ak9Cs3YCQ2
AGI https://t.co/zX3kkGPgPx
AGI https://t.co/zX3kkGPgPx
π₯ Meet Mistral Small 4: One model to do it all. β‘ 128 experts, 119B total parameters, 256k context window β‘ Configurable Reasoning β‘ Apache 2.0 β‘ 40% faster, 3x more throughput Our first model to unify the capabilities of our flagship models into a single, versatile model. https://t.co/2M1VNaDkRz
Wifeguy Lenin https://t.co/OPdSG8B5j8
Wifeguy Lenin https://t.co/OPdSG8B5j8
Different AI models find different bugs. So why not use all of them? Try this out in Copilot CLI: 1. Run /review 2. Ask it to use multiple model providers at once for a multi-agent code review 3. Get the highest possible signal and catch bugs before anyone else @_Evan_Boyle shows how it's done. βΆοΈ https://t.co/TWNzPBFcmC
San Francisco is a beach town https://t.co/xEUwKZte3n
Thank you pydantic https://t.co/KTkty0tUKx
ππ€π Jensen showing @huggingface during GTC keynote, where @NVIDIAAI dropped amazing new open models, datasets and blogs! Some of my favorites, links in comments: π§ Nemotron 3 Super 120A12B - Reasoning LLM π₯ Open-H-Embodiment - Healthcare Robotics Dataset π©» Cosmos-H-Surgical-Simulator - World Foundation Model π Alpamayo 1.5 - Autonomous Vehicle Model and Datasets πΊ Kimodo v1 - Generate body movement from prompts! And SO MUCH MORE to explore at https://t.co/PFeA2y9bi6
In Marin, we are trying to get really good at scaling laws. We have trained models up to 1e22 FLOPs and have made a prediction of the loss at 1e23 FLOPs, which @WilliamBarrHeld is running. This prediction is preregistered on GitHub, so we'll see in a few days how accurate our prediction was. What we want is not just a single model but a training recipe that scales reliably.
Here's the GitHub issue with all the details: https://t.co/fYre7BLjv6 This is part of our Delphi suite, a "modernized" version of Pythia: https://t.co/9G6NiADnMo
For trying to understanding LMs deeply, @AiEleutherβs Pythia has been an invaluable resource: 16 LMs (70M to 12B parameters) trained on the same data (The Pile) in the same order, with intermediate checkpoints. Itβs been two years and itβs time for a refresh.
JUST IN: Meta announces they'll be shutting down the Metaverse, after pouring $80,000,000,000.00 into the project. https://t.co/32VHBmfRQ2
Introducing the Paper Pages skill! Simply paste this SKILL.md, so your coding agent knows how to work with @huggingface papers Ask it to summarize papers, search papers, or list linked models or datasets https://t.co/Cf8iaFngN5