Your curated collection of saved posts and media
Hermes 一丢 Agent,全网程序员集体进化了! Nous Research 扔出 hermes-agent(90k+ stars),核心就一个词:自我进化。它不是玩具,而是带持久记忆、自动提炼技能、跨会话成长的底层骨架。 结果?社区直接把它当 DNA,短短几周卷出 80+ 进化体,生态总星 10 万+。这才是开源的最高境界:一个 Agent,变成全世界的进化树。 我挑了 4 个正在 X 上刷屏的 “Hermes 系进化体”,AI 玩家看了会沉默,爱好者看了会狂喜: 1️⃣ Hermes Ecosystem Map / Hermes Atlas(https://t.co/3wVPdY9cFg + https://t.co/L5UNrkxP1N) Kevin 用 Hermes Agent 自己刮全网,做出 80+ 项目安全审查地图 + RAG 聊天机器人。 还能生成《State of Hermes》报告,社区 PR 狂冲,官方都点赞了。 2️⃣ Hermes-Wiki(https://t.co/gTncSju9b0) 喂它自己源码,它就自动生成并维护完整项目 Wiki。 RAG 哭晕,Agent 自己给自己立传、太 meta 了。 3️⃣ hermes-hud(https://t.co/DivT37Yf9r) TUI 仪表盘进化成浏览器 Web HUD,实时监控 Agent 灵魂状态 + 持久记忆。 生态地图直接收录,玩家直呼终于看到 Agent 在思考。 4️⃣ hermes-control-interface (https://t.co/khcp2KHgly) 自托管 Dashboard,一键管多 Agent、进程、记忆和调度。 “把玩具 Agent 进化成生产基础设施”的代表作。 🟢 为什么 Hermes 系这么爆? 它从不给你黑箱,它给的是可读、可 hack、可自我改进的底层循环。你 fork 它,不是在抄,而是在和 Agent 一起递归自我改进。 这才是 2026 年开源 Agent 的正确打开方式。

@omnivaughn My AI can: https://t.co/kiuZ7QXLzb
Introducing the MLX-Benchmark Suite!! https://t.co/sp4ZMIBxov The first comprehensive benchmark for evaluating LLMs on Apple's MLX framework. 🎯 What is this? MLX Benchmark is a CLI tool and dataset that measures how well large language models understand, write, and debug code for Apple's MLX machine learning framework — covering everything from core array operations to LoRA fine-tuning with mlx-lm, mlx-vlm, and mlx-embeddings. 📊 Dataset https://t.co/5b04a7PKAp - 520 questions across 6 task types: knowledge QA, multiple choice, true/false, fill-in-the-blank, code generation, and debugging - 11 categories spanning the full MLX ecosystem: mlx_core, mlx_nn, mlx_lm, mlx_lm_lora, mlx_vlm, mlx_embeddings, mlx_embeddings_lora, mlx_optimizers, coding, debugging, conceptual - 4 difficulty levels: easy → medium → hard → very-hard - 90+ subcategories covering everything from array_creation to lora_finetuning ✨ Features - 🏃 Multi-provider benchmarking — Ollama, Anthropic, OpenAI, Groq, OpenRouter - ⚖️ LLM-as-judge evaluation — strict scoring with an independent judge model - 🔍 Fine-grained filtering — by type, difficulty, and category - 📝 LaTeX export — --latex generates publication-ready booktabs tables - 📈 PNG chart export — --plot generates grouped bar charts comparing models A detailed paper will be coming as well!!!

Introducing the MLX-Benchmark Suite!! https://t.co/sp4ZMIBxov The first comprehensive benchmark for evaluating LLMs on Apple's MLX framework. 🎯 What is this? MLX Benchmark is a CLI tool and dataset that measures how well large language models understand, write, and debug code for Apple's MLX machine learning framework — covering everything from core array operations to LoRA fine-tuning with mlx-lm, mlx-vlm, and mlx-embeddings. 📊 Dataset https://t.co/5b04a7PKAp - 520 questions across 6 task types: knowledge QA, multiple choice, true/false, fill-in-the-blank, code generation, and debugging - 11 categories spanning the full MLX ecosystem: mlx_core, mlx_nn, mlx_lm, mlx_lm_lora, mlx_vlm, mlx_embeddings, mlx_embeddings_lora, mlx_optimizers, coding, debugging, conceptual - 4 difficulty levels: easy → medium → hard → very-hard - 90+ subcategories covering everything from array_creation to lora_finetuning ✨ Features - 🏃 Multi-provider benchmarking — Ollama, Anthropic, OpenAI, Groq, OpenRouter - ⚖️ LLM-as-judge evaluation — strict scoring with an independent judge model - 🔍 Fine-grained filtering — by type, difficulty, and category - 📝 LaTeX export — --latex generates publication-ready booktabs tables - 📈 PNG chart export — --plot generates grouped bar charts comparing models A detailed paper will be coming as well!!!
Stop what you’re doing. This is worth 18 minutes. https://t.co/033UgU9Gm1
Stop what you’re doing. This is worth 18 minutes. https://t.co/033UgU9Gm1
@togethercompute @realDanFu But looking at the reported metrics, the looped 770M model isn’t really close to the 1.3B model. https://t.co/NkFfNjj8eR
For those who made it this far, use code "HERMESAGENT0010" for $10 off a valid Nous Portal Subscription (new or existing). Code valid for the first 250 users. Subscribe to Nous Portal: https://t.co/R6H1nIpuzO Learn more about Hermes Agent: https://t.co/iHmhvv5McL
// Self-Evolving Agent Protocol // One of the more interesting papers I read this week. (bookmark it if you are an AI dev) The paper introduces Autogenesis, a self-evolving agent protocol where agents identify their own capability gaps, generate candidate improvements, validate them through testing, and integrate what works back into their own operational framework. No retraining, no human patching, just an ongoing loop of assessment, proposal, validation, and integration. Why it's worth reading this paper: Static agents age quickly. As deployment environments change and new tools arrive, the agents that survive will be the ones that can safely rewrite themselves. Autogenesis is part of a growing wave of self-improving agent systems, alongside work like Meta-Harness and the Darwin Gödel Machine line, and it's one of the cleaner protocol-level takes on continual self-improvement so far. Paper: https://t.co/3aj9LLjSbk Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
This is nuts y'all, we've got new tornado warnings popping up everywhere across Iowa and Minnesota. The storms are moving fast and dropping warnings back to back. We are covering all of this live right now, hop on the stream and stay safe. WATCH LIVE: https://t.co/zA4VX48Zhe https://t.co/7LDc79Xjjd
Proud to announce Spatial Desk for Vision Pro available for pre-order with official release 8th May. After a Vision Pro onboarding session with a client - this is the App I wish was available. It's the 'blank canvas' concept. @AppleEDU @Apple https://t.co/xsywkCLWVw @Scobleizer https://t.co/LGcx2dvH6j

The Hermes Agent Creative Hackathon starts now 16 Days, $25k in Prizes Presented by @Kimi_Moonshot & @NousResearch For the tinkerers pushing Hermes Agent into creative domains: video, image, audio, 3D, long-form writing, creative software, interactive media and more. Show us what your Hermes Agent can do. Details Below ↓
@NousResearch @Kimi_Moonshot Hackathon time https://t.co/0W4NUREjQS
I'm a slow learner. And a minor in economics. I interviewed @VitalikButerin back when he was just starting. He developed Ethereum. If I was fast I would have bought some. When I interviewed him an Ethereum cost $2. Today it costs a LOT more. Same with understanding the new economic model that's coming. People in San Francisco say "don't worry about jobs, we are headed toward an age of abundance." Hard to get on that train as I see family and friends struggling economically after getting laid off. But this message from @elonmusk made it all click for me. We are, indeed, headed into a world where AI and robots do most of our current jobs. Anthropic's own engineers believe that AI will be able to do half of all jobs within two years, I heard from an investor who recently gave a speech there and asked them. We all feel it and know it. AI is taking away our jobs. Been seeing many messages along these lines on X's AI community. And normies are hating AI for it. We all hate change, even me. The only real answer is not to slow down, but accelerate through this moment of pain. I recently had dinner with a woman who is running a new construction robot company. Her robot can tie rebar 6x faster than a human. And the humans who do that job? They have the highest injury rate in construction, she told me. Robots don't get hurt. If we accelerate we get to the world where AI agents and robots do all of our shitty jobs. Which, let's be honest, almost all jobs suck. I've had jobs all my life and had some of the best ones that humans can ever have (I got paid to walk around Microsoft with a video camera, remember). Even that job sucked. A robot could do that even better. Here is proof of that. Asked Grok to simulate a conversation on this topic between a simulated Robert Scoble and a simulated Elon Musk: https://t.co/OjU9hJcxXE In 10 years when there are millions of Optimus, or other humanoids, walking around, they will be able to do even our most "fun" jobs. Which will leave us to do new jobs. Like setting up shop in a Holodeck that most people don't yet see. Or building new things for robots to do. Like entertainment most people can't yet see either. Grok lays it out. Acceleration is the right answer. Get us through this period of maximum pain fast so we can get on the other side where there is a lot of fun to be had.

@boardyai If I can build this, https://t.co/kiuZ7QXLzb which is a new way to read the AI community here on X, then someone can build a new LinkedIn as well. Get your Hermes or OpenClaw cooking!
@pmitu f I can build this, https://t.co/kiuZ7QXLzb which is a new way to read the AI community here on X, then someone can build a new LinkedIn as well. Get your Hermes or OpenClaw cooking!
Recent attacks on the open source supply chain are targeting CI/CD workflows to exfiltrate secrets. 👀 Here's what you can do right now to protect your projects. 🔒 https://t.co/fhLuOCo2vK
This is evidence of a new "headless" Web being developed by the nerds. A Web just for people's agents to come and visit and do business with. Web 1.0: sharing info. Web 2.0: making it dynamic. Web 3.0: making it trustworthy. Web 4.0: making it for agents. :-) I once gave @Benioff shit for not having enough AI (I'm an investor in Salesforce, albeit a tiny one). I hate most enterprise software. Once had to apologize to the CEO of Workday when I gave them shit in public back when I worked at Rackspace and did it in an improper way. Now that we can create our own software with AI I expect this "headless Web" to rapidly expand. It's why I added a feed for OpenClaws to my news site: https://t.co/8L5xphk0qQ
Welcome Salesforce Headless 360: No Browser Required! Our API is the UI. Entire Salesforce & Agentforce & Slack platforms are now exposed as APIs, MCP, & CLI. All AI agents can access data, workflows, and tasks directly in Slack, Voice, or anywhere else with Salesforce Headless
That's fuckin' right! https://t.co/RALKAISD1m
@ilyamiskov That's why I created https://t.co/kiuZ7QXLzb so you can keep up with all the announcements on X without having to scroll through 40,000 posts a day (my AI reads that many every day to make that page).
It's never been easier for your agent to harness the power of Hugging Face, from account creation to getting your own finetuned model, with everything in between, without any manual action on your end! 🤗🤖 The future is now! 🚀 https://t.co/NSbri6Sijy
Just launched some great new providers in Stripe Projects (https://t.co/QwQF6d2TIz): @huggingface @Cloudflare @OpenRouter @firecrawl @flydotio @Amplitude_HQ @mixpanel @inngest Many more coming later this month.
https://t.co/XmSdspq5zP
[ACL 2026✨] FinWorkBench! Can AI agents truly handle real world finance & accounting 💰 work? ✨ Built from real enterprise data, not synthetic tasks ✨ Long-horizon finance workflows ✨ Messy multimodal & cross-file reasoning ✨ Expert annotated (700 hours) https://t.co/g2gvY
https://t.co/XmSdspq5zP
Thank you @_akhaliq for sharing our work! We also have a 24/7 live stream for GameWorld: https://t.co/eFfD4a437G. Watch the agents play games in real time.🕹️
GameWorld Towards Standardized and Verifiable Evaluation of Multimodal Game Agents paper: https://t.co/IfbTgfNnSM https://t.co/gL3BURxzkV
Thank you @_akhaliq for sharing our work! We also have a 24/7 live stream for GameWorld: https://t.co/eFfD4a437G. Watch the agents play games in real time.🕹️
Guess what model is being fed today https://t.co/QZAaBaSzus
RSVP here if you haven't : https://t.co/EXOuswdiqh
grok 4.3 beta can use an ubuntu shell and a persistent file layer to generate artifacts grok wrote python to encode the xai / grok logo into audio, i gave it the script back and had it render a spectrogram video from that signal, and use the grok_files tool to save the mp4 into the product's files layer i opened the file from the files panel and played it myself this is getting crazy
Grok 4.3 can Take in Video and Extract Audio files https://t.co/5uprx2dM85
NEWS: Dubai Police just admitted spying on a PRIVATE WhatsApp chat using electronic monitoring operations. They arrested an airline crew member over a photo shared ONLY in a closed group (never posted publicly). Meta & WhatsApp have been LYING to billions with their “fully end-to-end encrypted” ads. Your private messages aren’t private on WhatsApp. (Source: Detained in Dubai)
Offline-first AI agent for Raspberry Pi https://t.co/iapUnKRhXI https://t.co/FtE8vK8kSu
