Your curated collection of saved posts and media
Google just released TIPS on Hugging Face A vision-language model with spatial awareness, built for dense understanding tasks like segmentation and depth estimation. https://t.co/7aNlJcrGXE
okay this is too cute 😭 i saw Microduck and immediately asked Qwen3.8-27B to build a tiny browser gym where i could teach it a treat-delivery trick. 250 episodes later, the little guy learned it in 9 steps. https://t.co/vc6aZ7Hsx4
BIG ANNOUNCEMENT FROM HUGGING FACE TODAY: We're unveiling Microduck 🐥🤖 It's a tiny $399 open-source robot you can teach new tricks with reinforcement learning. It can walk, pick things up, get back up when it falls, and even roller-skate. Welcome to the era of open-source af
My /dev-pair Hermes Skill works. Same model marking its own homework is how silent bugs ship. This skill refuses that by construction. https://t.co/xZktDbG6X9 The driver writes. A reviewer from a different model family reads. The reviewer is read-only — no tools, no edits, no commands. Verdicts only: SHIP, BLOCKER, MAJOR. Same-family review is blocked. Sessions persist so the follow-up still knows the last fight. This morning’s pass didn’t just rank ten changes. It found a live bug and we shipped that first (commit 89d2c8767). main() always returned 0. The docstring promised exit 1 on critical failure. So a CRITICAL finding and a total email-delivery miss both recorded cron last_status='ok'. The watchdog could scream. The cron still called the run healthy. Same false-negative as the fitness bug earlier that day. Fixed: exit 1 on critical-or-undelivered, plus a Telegram fallback that does not share the Google OAuth path. Then the ten, in the order the pair actually agreed: Crash-isolate all 22 checks — one bad check was killing the rest and the email. Start + completion heartbeat — “end of run” can’t see a hang. Blocker. Revised Watch the watchman off-box Decouple cooldown latches per alert class Parse jobs.json for the window — Luna called it over-engineering. I pushed back. A second copy of the schedule was the root of this morning’s bug Tests — 1,529 lines, 23 checks, zero tests. Narrowed to crash / hang / delivery / overnight / corrupt-state Auto-repair only if four gates hold: idempotent, postcondition checked in the same run, failure leaves things no worse, no credential / consent / billing boundaryVerify advisory commands exist 9–10. Structured history and chronic-alert detection — deferred. Real value, not outage prevention Honest bit from the thread: the critique is Luna’s. The concessions are Kimi’s. Dev-pair rotated reviewers between turns. Worth knowing before you attribute the win. Screenshot is the discussion. #AIAgents @NousResearch @Teknium

While I think about longer-term plans, I do have one short + urgent project I'd love to pursue if I can find funding: putting together an eval on embedded device/hardware hacking capabilities. With talk of cyber being 'defense dominant' in the long term, I'm worried that many are focused on software like Chrome that can be easily updated, but we live surrounded by mice and keyboards and hard drives and garage door openers that run software too. How vulnerable are these to AI-assisted attacks? Testing is harder than pure software - it needs all the hardware in question as well as support circuitry for cycling power, reading out voltages, and so on. I'd love to build up a wall of common devices and put models to the test, but I don't have the money to spend hundreds or thousands of dollars on hardware and I don't have the runway to work on this unpaid. Ideal outcome of this tweet: you tell me that someone is already working on this and I can let it go. But failing that I hope this either inspires work on this to start somewhere. Or reaches someone willing to fund me to give this a bash! I'd tried applying for a grant on this with no luck (https://t.co/ahSSp7JRqt), but I also don't know much about how one goes about seeking funding for something like this. If you would like to make this happen, please reach out :)
I've officially left https://t.co/MD2Xc5heDO I've had a good rest, with plenty of time for travel & tinkering, and now I'm thinking about what to do next. I've got a few ideas to share soon, but I'm also open to suggestions - feel free to reach out :)
the industry’s preferred way to obscure things is to conflate LLMs with AI, and to ignore the fact that recent contributions have come from adding things - mostly (neuro)symbolic - to those LLMs (precisely as my famous 2022 paper forcecast) and to ignore the fact that the new techniques for doing so still at present work best in limited domains such as math and coding. keep your eye on the ball; they sure as hell won’t be candid about what’s happening.
chatting with @HamelHusain tomorrow about * why "it's hard to eval" is a product smell * how agents have changed the eval landscape and how @sh_reya and he have updated their course * whether data science is dead in the age of AI agents or not Register to join live or get the recording afterwards: https://t.co/NxIERZAEak