Your curated collection of saved posts and media

Showing 32 posts ยท last 7 days ยท newest first
O
omarsar0
@omarsar0
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”25659668

When you build AI agents, don't treat prompts like config strings. Treat them like executable business logic. Because that's what they really are. @arshdilbagi's blog and this Stanford CS 224G lecture lay out one of the clearest mental models I have seen for LLM evaluation. Stop treating evals like unit tests. That works for deterministic software. For LLM products, it creates false confidence because real-world usage changes over time. Example: an insurance prompt passed 20 eval cases. The team shipped. In production, a new class of requests showed up and failed quietly. No crash, no alert, just wrong answers at scale. The fix is not "write more eval cases," which is what many teams do. It is building evals as a living feedback loop. Start with a small set, ship, watch what breaks in production, add those failures back, and re-run on every prompt or model change. What eval failure caught your team off guard? Blog: https://t.co/HCVhcow5rA Stanford CS 224G lecture: https://t.co/q667gGwckt

Media 1Media 2
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”25659668

When you build AI agents, don't treat prompts like config strings. Treat them like executable business logic. Because that's what they really are. @arshdilbagi's blog and this Stanford CS 224G lecture lay out one of the clearest mental models I have seen for LLM evaluation. Stop treating evals like unit tests. That works for deterministic software. For LLM products, it creates false confidence because real-world usage changes over time. Example: an insurance prompt passed 20 eval cases. The team shipped. In production, a new class of requests showed up and failed quietly. No crash, no alert, just wrong answers at scale. The fix is not "write more eval cases," which is what many teams do. It is building evals as a living feedback loop. Start with a small set, ship, watch what breaks in production, add those failures back, and re-run on every prompt or model change. What eval failure caught your team off guard? Blog: https://t.co/HCVhcow5rA Stanford CS 224G lecture: https://t.co/q667gGwckt

Media 1Media 2
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Mar 05, 2026
160d ago
๐Ÿ†”12277409

Google Workspace CLI: https://t.co/Rg229zYsoA

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Mar 05, 2026
160d ago
๐Ÿ†”12277409

Google Workspace CLI: https://t.co/Rg229zYsoA

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”89368331

Read on for more: https://t.co/bLxqT3yUTl

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”89368331

Read on for more: https://t.co/bLxqT3yUTl

Media 1
๐Ÿ–ผ๏ธ Media
S
SIGKITTEN
@SIGKITTEN
๐Ÿ“…
Mar 06, 2026
159d ago
๐Ÿ†”32368826

thanks codex team for supporting the radare2 project with some Pro subs!๐Ÿ™โค๏ธ https://t.co/0fKEt16AeI

Media 1
๐Ÿ–ผ๏ธ Media
๐Ÿ”jxnlco retweeted
S
SIGKITTEN
@SIGKITTEN
๐Ÿ“…
Mar 06, 2026
159d ago
๐Ÿ†”32368826

thanks codex team for supporting the radare2 project with some Pro subs!๐Ÿ™โค๏ธ https://t.co/0fKEt16AeI

Media 1
โค๏ธ61
likes
๐Ÿ”1
retweets
๐Ÿ–ผ๏ธ Media
Z
ZeffMax
@ZeffMax
๐Ÿ“…
Mar 06, 2026
159d ago
๐Ÿ†”25134380

that was the old me, the six days ago me https://t.co/btAJYAwAjL

@ โ€ข

Media 1
๐Ÿ–ผ๏ธ Media
J
jxnlco
@jxnlco
๐Ÿ“…
Mar 06, 2026
158d ago
๐Ÿ†”50911745

amazing work team https://t.co/9d1cEx7vy2

Media 1
๐Ÿ–ผ๏ธ Media
C
chongdashu
@chongdashu
๐Ÿ“…
Mar 06, 2026
159d ago
๐Ÿ†”73340254

Couldn't help it! Had to give GPT 5.4 (High) + /fast mode a try. โ†’ Added height terrains to the level โ†’ Animation tweens for the jumps Used xHigh to solve a gnarly bug with the controls successfully ๐Ÿ’ช This Final Fantasy Tactics-inspired game was completely vibe coded! https://t.co/q2K7PovU62

@chongdashu โ€ข Fri Mar 06 00:41

Just recorded a step by step walkthrough of how I vibe code games using Codex and Claude Code. I implement new features 'live' in the recording, showing how I get the most out of GPT and Opus. Full video hopefully landing tomorrow! It's going to be a good one... don't miss it!

๐Ÿ–ผ๏ธ Media
B
bertgodel
@bertgodel
๐Ÿ“…
Mar 03, 2026
161d ago
๐Ÿ†”11940087

Weโ€™re announcing Kos-1 Lite, a medical model that achieves SOTA on HealthBench Hard at 46.6%. As a medium sized language model (~100B), it achieves these results at a fraction of the serving cost of frontier trillion-parameter models. https://t.co/27sxAHPgZM

Media 1
๐Ÿ–ผ๏ธ Media
B
bertgodel
@bertgodel
๐Ÿ“…
Mar 03, 2026
161d ago
๐Ÿ†”11940087

Weโ€™re announcing Kos-1 Lite, a medical model that achieves SOTA on HealthBench Hard at 46.6%. As a medium sized language model (~100B), it achieves these results at a fraction of the serving cost of frontier trillion-parameter models. https://t.co/27sxAHPgZM

Media 1
๐Ÿ–ผ๏ธ Media
H
HuggingPapers
@HuggingPapers
๐Ÿ“…
Mar 04, 2026
161d ago
๐Ÿ†”28876865

SWE-rebench V2 A language-agnostic pipeline that automatically harvests 32,000+ executable real-world software engineering tasks across 20 programming languages. Built for large-scale RL training of code agents with reproducible Docker environments. https://t.co/JJ0vLH5N7B

Media 1
๐Ÿ–ผ๏ธ Media
H
HuggingPapers
@HuggingPapers
๐Ÿ“…
Mar 04, 2026
161d ago
๐Ÿ†”28876865

SWE-rebench V2 A language-agnostic pipeline that automatically harvests 32,000+ executable real-world software engineering tasks across 20 programming languages. Built for large-scale RL training of code agents with reproducible Docker environments. https://t.co/JJ0vLH5N7B

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”91333805

Image Generation with a Sphere Encoder https://t.co/6I2FbpogaC

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”27662834

Utonia Toward One Encoder for All Point Clouds paper: https://t.co/AJFPivgBm9 https://t.co/Xbux4iY1QV

Media 2
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”36665410

BeyondSWE Can Current Code Agent Survive Beyond Single-Repo Bug Fixing? paper: https://t.co/IrLgJJomQU

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”50449052

Beyond Language Modeling An Exploration of Multimodal Pretraining paper: https://t.co/GmtPAQDo8T

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”19687332

Beyond Length Scaling Synergizing Breadth and Depth for Generative Reward Models https://t.co/25QhR93OKK

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”36016044

Kiwi-Edit Versatile Video Editing via Instruction and Reference Guidance https://t.co/s9xlDgXhfc

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”67577372

BBQ-to-Image Numeric Bounding Box and Qolor Control in Large-Scale Text-to-Image Models paper: https://t.co/54U6zmx2ZA https://t.co/fW8zbIrE19

Media 2
๐Ÿ–ผ๏ธ Media
G
GordonWetzstein
@GordonWetzstein
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”59376026

Video world models today have a very limited context length. Mode Seeking meets Mean Seeking (MMM) unlocks long-context, persistent video world models through a unified representation. 1/8 ๐Ÿงต https://t.co/XXMic82qoc

๐Ÿ–ผ๏ธ Media
A
andimarafioti
@andimarafioti
๐Ÿ“…
Mar 04, 2026
160d ago
๐Ÿ†”53775745

The Faster-Qwen3-TTS demo just passed the official Qwen3-TTS demo in Hugging Face trending (last 7 days). Now the #5 most trending Space ๐Ÿš€ https://t.co/8Ar622nKU9

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”69962587

Helios Real Real-Time Long Video Generation Model paper: https://t.co/ae0ZH4zPzn https://t.co/kCnNfF3ImI

Media 2
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”08415662

Heterogeneous Agent Collaborative Reinforcement Learning https://t.co/ASb1VwtCeK

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”75057561

Proact-VL A Proactive VideoLLM for Real-Time AI Companions https://t.co/GkHdSKxSvi

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”72206297

CubeComposer Spatio-Temporal Autoregressive 4K 360ยฐ Video Generation from Perspective Video paper: https://t.co/mnDM1VrYn7 https://t.co/iHtlZJCo1w

Media 2
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”64405355

LTX-2.3 is out on Hugging Face model: https://t.co/te5nwPL1LE https://t.co/biO7szxFGz

Media 2
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Mar 05, 2026
159d ago
๐Ÿ†”52917888

Tencent released HY-WU on Hugging Face An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing model: https://t.co/jAnic8Z9i1 https://t.co/LsLpyjMVQT

Media 2
๐Ÿ–ผ๏ธ Media
X
Xianbao_QIAN
@Xianbao_QIAN
๐Ÿ“…
Mar 02, 2026
162d ago
๐Ÿ†”61966034

New model updates from iquestlab. If you're trying to find an inference model that you can run offline, this is probably the one you're looking for. - 7B and 14B coding models - Optimized for tool use, CLI agents and HTML generation - 128k context length - Explicit and detailed prompting works best - MiT license with requirement of display logo - available on @huggingface

Media 1
๐Ÿ–ผ๏ธ Media
A
ariG23498
@ariG23498
๐Ÿ“…
Mar 02, 2026
162d ago
๐Ÿ†”67511857

With the help of @huggingface we (/w @RisingSayak) are building a ML Club India ๐Ÿ‡ฎ๐Ÿ‡ณ What we want to do: 1. Online talks 2. IST compatible timing 2. Open to all More to come in this week! Watch this space. ๐Ÿค— Special thanks to @LysandreJik who motivated me to keep working on this. ๐Ÿ”ฅ

Media 1
๐Ÿ–ผ๏ธ Media