Your curated collection of saved posts and media

Showing 32 posts ยท last 7 days ยท newest first
I
illustrata_ai
@illustrata_ai
๐Ÿ“…
Jan 15, 2026
210d ago
๐Ÿ†”15625396

๐Ÿ–ค https://t.co/bXGfPfQtlE

Media 1
๐Ÿ–ผ๏ธ Media
P
perrymetzger
@perrymetzger
๐Ÿ“…
Jan 09, 2026
217d ago
๐Ÿ†”48188541

Stoicism 101: You are not required to have an opinion on every controversy. In fact, youโ€™ll be much happier if you donโ€™t. https://t.co/KXkPgZByU3

Media 1
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
Jan 09, 2026
216d ago
๐Ÿ†”99261893

#reallifedataviz https://t.co/l415d8BEdB

Media 1
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
Jan 13, 2026
213d ago
๐Ÿ†”40696270

'How Much Grams?' A mini eval inspired by that viral guy who would ask ChatGPT voice+video to guess the weight of things. Flash crushes it - both speed and accuracy. Post: https://t.co/HNJWhGQIEb https://t.co/191EqCtbnt

Media 1
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
Jan 14, 2026
211d ago
๐Ÿ†”51710851

Data from my sleep tracking came in handy recently: can you see which 5 days my wife tried a new feather pillow? (Graph shows my snoring duration) A similarly stark positive change came from running an air purifier. Probably lots more places I'm unknowingly sub-optimal... :) https://t.co/PdjHZrxMyI

Media 1
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
Jan 17, 2026
208d ago
๐Ÿ†”79659169

@ATinyGreenCell I'm trying a https://t.co/ZWqHLe1GOr for similar reasons

Media 1
๐Ÿ–ผ๏ธ Media
C
cloneofsimo
@cloneofsimo
๐Ÿ“…
Jan 18, 2026
208d ago
๐Ÿ†”51537590

It is November 2022, you yap "LLM cant reason because they are autoregressive!! We need cat level intelligence, neuro-symbolic AI! Chatgpt is fun toy product!" It is year 2026, they solve conjectures in 40 min. You literally share proof of open problems with chatgpt links. Actually insane timeline.

@neelsomani โ€ข Sun Jan 18 01:17

I've solved a second Erdos problem (#281) using only GPT 5.2 Pro - no prior solutions found. Terence Tao calls it "perhaps the most unambiguous instance" of AI solving an open problem: https://t.co/TBiCwiSFzl

Media 1
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
Jan 18, 2026
208d ago
๐Ÿ†”36878904

Saturday hobby fun: biolistics tests with different nozzles, sticking DNA to carriers, and making a computer-controlled pipette for moving microliters of liquid around ๐Ÿ˜ (Also breakfast date, happy walks with toddler niece, fireside reading - perfect Saturday) https://t.co/eNM9l57uiZ

Media 1
+2 more
๐Ÿ–ผ๏ธ Media
J
johnowhitaker
@johnowhitaker
๐Ÿ“…
Jan 18, 2026
208d ago
๐Ÿ†”75570734

Going to make so much agar art with this bad boy ๐Ÿ˜ https://t.co/iFYdZq64SG

๐Ÿ–ผ๏ธ Media
L
LiorOnAI
@LiorOnAI
๐Ÿ“…
Jan 16, 2026
210d ago
๐Ÿ†”26642286

Repo: https://t.co/PRh999W0kp

Media 1
๐Ÿ–ผ๏ธ Media
L
LiorOnAI
@LiorOnAI
๐Ÿ“…
Jan 16, 2026
210d ago
๐Ÿ†”30575956

Must watch. Why Transformers are taking over CNNs in computer vision. https://t.co/rH6CV0qdAK

Media 1
๐Ÿ–ผ๏ธ Media
L
LiorOnAI
@LiorOnAI
๐Ÿ“…
Jan 16, 2026
210d ago
๐Ÿ†”57191786

Video: https://t.co/JySWpaKtvF

Media 1
๐Ÿ–ผ๏ธ Media
L
LiorOnAI
@LiorOnAI
๐Ÿ“…
Jan 17, 2026
209d ago
๐Ÿ†”65950481

> Be Maor Shlomo. > Fail for a decade. > See Lovable take off. > Grab Claude 3.5. > Build a competitor in weeks. > Add a twist to it. > Hit $230k MRR in 90 days. > Sell it for $80M in 4 months. https://t.co/iOiAqRIP5a

Media 1
๐Ÿ–ผ๏ธ Media
L
LiorOnAI
@LiorOnAI
๐Ÿ“…
Jan 17, 2026
209d ago
๐Ÿ†”87073541

Google just gave language models real long-term memory. A new architecture learns during inference and keeps context across millions of tokens. It holds ~70 percent accuracy at 10 million tokens. ๐—ง๐—ต๐—ถ๐˜€ ๐—ฎ๐—ฟ๐—ฐ๐—ต๐—ถ๐˜๐—ฒ๐—ฐ๐˜๐˜‚๐—ฟ๐—ฒ ๐—น๐—ฒ๐—ฎ๐—ฟ๐—ป๐˜€ ๐˜„๐—ต๐—ถ๐—น๐—ฒ ๐—ถ๐˜ ๐—ฟ๐˜‚๐—ป๐˜€ Titans adds a neural long-term memory that updates during generation. Not weights. Not retraining. Live learning. โ€ข A small neural network stores long-range context โ€ข It updates only when something unexpected appears โ€ข Routine tokens get ignored to stay fast This lets the model remember facts from far earlier text without scanning everything again. ๐—œ๐˜ ๐—ธ๐—ฒ๐—ฒ๐—ฝ๐˜€ ๐˜€๐—ฝ๐—ฒ๐—ฒ๐—ฑ ๐˜„๐—ต๐—ถ๐—น๐—ฒ ๐˜€๐—ฐ๐—ฎ๐—น๐—ถ๐—ป๐—ด ๐—ฐ๐—ผ๐—ป๐˜๐—ฒ๐˜…๐˜ Attention stays local. Memory handles the past. โ€ข Linear inference cost โ€ข No quadratic attention blowups โ€ข Stable accuracy past two million tokens ๐—œ๐˜ ๐—ฎ๐—น๐—น๐—ผ๐˜„๐˜€ ๐˜†๐—ผ๐˜‚ ๐˜๐—ผ ๐—ฏ๐˜‚๐—ถ๐—น๐—ฑ ๐—ป๐—ฒ๐˜„ ๐—ธ๐—ถ๐—ป๐—ฑ๐˜€ ๐—ผ๐—ณ ๐—ฎ๐—ฝ๐—ฝ๐˜€ You can process full books, logs, or genomes in one pass. You can keep state across long sessions. You can stop chunking context just to survive limits.

Media 1
๐Ÿ–ผ๏ธ Media
L
LiorOnAI
@LiorOnAI
๐Ÿ“…
Jan 17, 2026
209d ago
๐Ÿ†”87443872

Project: https://t.co/8aKlYQ2gb6

Media 1
๐Ÿ–ผ๏ธ Media
L
LiorOnAI
@LiorOnAI
๐Ÿ“…
Jan 18, 2026
208d ago
๐Ÿ†”07875549

Small models just beat giant LLM agents at their own job. Not by thinking harder, but by coordinating better. A new system just outscored GPT-5 on Humanityโ€™s Last Exam, using far less compute. ๐—ง๐—ต๐—ถ๐˜€ ๐˜€๐˜†๐˜€๐˜๐—ฒ๐—บ ๐—ฟ๐—ฒ๐—ฝ๐—น๐—ฎ๐—ฐ๐—ฒ๐˜€ ๐—ผ๐—ป๐—ฒ ๐—ฏ๐—ถ๐—ด ๐—ฏ๐—ฟ๐—ฎ๐—ถ๐—ป ๐˜„๐—ถ๐˜๐—ต ๐—ฎ ๐—ฐ๐—ผ๐—ป๐—ฑ๐˜‚๐—ฐ๐˜๐—ผ๐—ฟ Instead of one model doing everything, it assigns roles. โ€ข Large models handle hard reasoning. โ€ข Small models handle routine steps. โ€ข A controller decides what to call, when. That controller is trained only to make decisions. ๐—œ๐˜ ๐—น๐—ฒ๐—ฎ๐—ฟ๐—ป๐˜€ ๐—ฐ๐—ผ๐—ผ๐—ฟ๐—ฑ๐—ถ๐—ป๐—ฎ๐˜๐—ถ๐—ผ๐—ป, ๐—ป๐—ผ๐˜ ๐—ฝ๐—ฟ๐—ผ๐—บ๐—ฝ๐˜ ๐˜๐—ฟ๐—ถ๐—ฐ๐—ธ๐˜€ It uses reinforcement learning, not hand rules. Rewards optimize three things at once: - Task success - Latency - Compute cost ๐—ง๐—ต๐—ฒ ๐—ฟ๐—ฒ๐˜€๐˜‚๐—น๐˜๐˜€ It scores 37.1% on HLE versus 35.1%. Runs about 2.5ร— faster. Uses roughly 70% less cost. This lets you build agents that scale by coordination, not parameters.

Media 1
๐Ÿ–ผ๏ธ Media
C
ChuanmingLiu
@ChuanmingLiu
๐Ÿ“…
Jan 14, 2026
212d ago
๐Ÿ†”44212466

https://t.co/GTBUA1RzVy https://t.co/sB5dE5xoiY

Media 1
๐Ÿ–ผ๏ธ Media
๐Ÿ”arnicas retweeted
C
Chuanming
@ChuanmingLiu
๐Ÿ“…
Jan 14, 2026
212d ago
๐Ÿ†”44212466

https://t.co/GTBUA1RzVy https://t.co/sB5dE5xoiY

Media 1
โค๏ธ341
likes
๐Ÿ”74
retweets
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 12, 2026
214d ago
๐Ÿ†”33730234

Great paper on Agentic Memory. LLM agents need both long-term and short-term memory to handle complex tasks. However, the default approach today treats these as separate components, each with its own heuristics, controllers, and optimization strategies. But memory isn't two independent systems. It's one cognitive process that decides what to store, retrieve, summarize, and forget. This new research introduces AgeMem, a unified framework that integrates long-term and short-term memory management directly into the agent's policy through tool-based actions. Instead of relying on trigger-based rules or auxiliary memory managers, the agent learns when and how to invoke memory operations: ADD, UPDATE, DELETE for long-term storage, and RETRIEVE, SUMMARY, FILTER for context management. It uses a three-stage progressive RL strategy. First, the model learns long-term memory storage. Then it masters short-term context management. Finally, it coordinates both under full task settings. To handle the fragmented experiences from memory operations, they design a step-wise GRPO (Group Relative Policy Optimization) that transforms cross-stage dependencies into learnable signals. The results across five long-horizon benchmarks: > On Qwen2.5-7B, AgeMem achieves 41.96 average score compared to 37.14 for Mem0, a 13% improvement. > On Qwen3-4B, the gap widens: 54.31 vs 44.70. Adding long-term memory alone provides +10-14% gains. > Adding RL training adds another +6%. > The full unified system with both memory types achieves up to +21.7% improvement over no-memory baselines. The unified memory management through learnable tool-based actions outperforms fragmented heuristic pipelines, enabling agents to adaptively decide what to remember and forget based on task demands. Paper: https://t.co/twhfiEsnho Learn to build effective AI agents in our academy: https://t.co/JBU5beIoD0

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 12, 2026
214d ago
๐Ÿ†”08455881

During my live cohort this week, I will share more on my process for quickly building tools like this in Claude Code: https://t.co/lmI57QkT5U It's an intensive training, but you only need to learn this stuff once and build from there. There are a few seats left!

Media 1
๐Ÿ–ผ๏ธ Media
D
dair_ai
@dair_ai
๐Ÿ“…
Jan 12, 2026
214d ago
๐Ÿ†”86348593

Efficient Lifelong Memory for LLM Agents LLM agents need memory to handle long conversations. The way this is handled today is that memory either retains full interaction histories, leading to massive redundancy, or relies on iterative reasoning to filter noise, consuming excessive tokens. This new research introduces SimpleMem, an efficient memory framework based on semantic lossless compression that maximizes information density while minimizing token consumption. The framework operates through a three-stage pipeline. 1) First, Semantic Structured Compression applies entropy-aware filtering to distill raw dialogue into compact memory units, resolving coreferences and converting relative time expressions ("last Friday") into absolute timestamps. 2) Second, Recursive Memory Consolidation incrementally integrates related memories into higher-level abstractions, turning repetitive entries like "ordered a latte at 8 AM" into patterns like "regularly drinks coffee in the morning." 3) Third, Adaptive Query-Aware Retrieval dynamically adjusts the retrieval scope based on query complexity. The results: On the LoCoMo benchmark with GPT-4.1-mini, SimpleMem achieves 43.24 F1, outperforming the strongest baseline Mem0 (34.20) by 26.4%, while reducing token consumption to just 531 tokens per query compared to 16,910 for full-context approaches, a 30x reduction. They claim that memory construction is 14x faster than Mem0 (92.6s vs 1350.9s per sample) and 50x faster than A-Mem. Even a 3B parameter model with SimpleMem outperforms larger models using inferior memory strategies. This work shows that structured semantic compression and adaptive retrieval enable LLM agents to maintain reliable long-term memory without drowning in tokens or sacrificing accuracy.

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 13, 2026
213d ago
๐Ÿ†”64542639

New research from Meta and collaborators. This is a good paper showing what's possible with proper world models. World models need actions to predict consequences. The default approach today requires labeled action data, which is expensive to obtain and limited to narrow domains like video games or robotic manipulation. But the vast majority of video data online has no action labels at all. This new research tackles learning latent action world models directly from in-the-wild videos, expanding beyond the controlled settings of previous work to capture the full diversity of real-world actions. The challenge is significant. In-the-wild videos contain actions far beyond simple navigation or manipulation: people entering frames, objects appearing and disappearing, dancers moving, fingers forming guitar chords. There's also no consistent embodiment across videos, unlike robotics datasets, where the same arm appears throughout. So how do the authors address this? Continuous but constrained latent actions, using sparse or noisy regularization, effectively capture this action complexity. Discrete quantization, the common approach in prior work, struggles to adapt. Without a shared embodiment, the model learns spatially-localized, camera-relative transformations. The results demonstrate genuine action transfer. Motion from a walking person can be applied to a flying ball. Actions like "someone entering the frame" transfer across completely different videos. By training a small controller to map known actions to latent ones, the world model trained purely on natural videos can solve robotic manipulation and navigation tasks with performance close to models trained on domain-specific, action-labeled data. Latent action spaces learned from unlabeled internet videos can serve as a universal interface for planning, removing the bottleneck of action annotation. Paper: https://t.co/BL6mpuLZGD Learn to build effective AI agents in our academy: https://t.co/JBU5beHQNs

Media 1
๐Ÿ–ผ๏ธ Media
D
dair_ai
@dair_ai
๐Ÿ“…
Jan 13, 2026
213d ago
๐Ÿ†”86443905

On building more powerful self-evolving agents. LLM agents struggle to learn from experience after deployment. Fine-tuning is expensive and causes catastrophic forgetting. RAG retrieves based on semantic similarity alone, often pulling noise instead of what actually works. Similarity and utility are not the same thing. This new research introduces MemRL, a framework that enables agents to self-evolve through non-parametric reinforcement learning on episodic memory, keeping the LLM completely frozen. The core idea is to treat memory retrieval as a decision-making problem, not a matching problem. Each memory stores an Intent-Experience-Utility triplet. The utility is a learned Q-value representing expected returns, continuously refined through environmental feedback. MemRL implements Two-Phase Retrieval. First, filter candidates by semantic similarity to ensure relevance. Then, rank by learned Q-values to select what actually works. This distinguishes high-value strategies from semantically similar noise. When the agent succeeds or fails, it updates the Q-values of retrieved memories using Bellman-style backups. No gradient updates to model weights. The frozen LLM provides stable reasoning while the memory evolves plastically. Results across four benchmarks: On HLE (knowledge frontier tasks), MemRL significantly outperforms both RAG and existing memory systems like MemP. The pattern holds on BigCodeBench for code generation, ALFWorld for exploration tasks, and Lifelong Agent Bench for OS and database operations. Analysis confirms a strong correlation between learned utility scores and actual task success, validating that Q-values capture genuine functional value rather than superficial similarity. Why does it matter? Decoupling stable reasoning from plastic memory enables continuous runtime improvement without the catastrophic forgetting or computational costs of fine-tuning. Paper: https://t.co/HvLUnXW2Jd Learn to build effective AI agents in our academy: https://t.co/zQXQt0PMbG

Media 1Media 2
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 13, 2026
212d ago
๐Ÿ†”77973851

I applaud Anthropic for relentlessly making Claude Code easier to use. You can now leverage modes within Code in Claude Desktop. Ask, Plan, and Execute are some of the most important components to make an agent work. And they are now available at the press of a button. https://t.co/SV5MGSYJeY

Media 1
๐Ÿ–ผ๏ธ Media
D
dair_ai
@dair_ai
๐Ÿ“…
Jan 14, 2026
212d ago
๐Ÿ†”43994185

Super interesting paper from Meta Superintelligence Labs. This work suggests complex reasoning and search capabilities can emerge solely through self-evolution, challenging the assumption that human supervision is necessary for advanced agent abilities. Let's break down the paper: Self-evolving LLMs can improve without human-curated data by generating their own training problems. However, existing data-free frameworks focus on narrow domains like math and coding. They struggle with open-domain search agents due to limited question diversity and the massive compute required for multi-step reasoning with tools. But what if search agents could evolve from scratch using only an external search engine? This new research introduces Dr. Zero (DeepResearch-Zero), a framework enabling search agents to self-evolve without any training data, demonstrations, or human annotations. The core design: a proposer-solver feedback loop where both models initialize from the same base LLM. The proposer generates diverse questions to train the solver. As the solver improves, it pushes the proposer to create harder yet still solvable queries, establishing an automated curriculum. Standard GRPO requires nested sampling, generating multiple queries each with multiple responses. This becomes computationally prohibitive for multi-turn search agents. Dr. Zero introduces Hop-Grouped Relative Policy Optimization (HRPO), which clusters structurally similar questions by their cross-hop complexity to construct group-level baselines. This eliminates nested sampling while maintaining stable training. The proposer reward balances verifiability and difficulty. If the solver gets everything right, the question is too easy. If it fails completely, too hard. The sweet spot maximizes the learning signal. Results: The data-free Dr. Zero matches or surpasses fully supervised search agents by up to 14.1% on complex QA benchmarks, including HotpotQA, 2WikiMQA, and MuSiQue. On Qwen2.5-7B, Dr. Zero achieves 0.372 average score compared to 0.347 for supervised Search-R1. Paper: https://t.co/CjkbRQNQIl Learn to build effective AI agents in our academy: https://t.co/zQXQt0PMbG

Media 1Media 2
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 14, 2026
212d ago
๐Ÿ†”34754243

UniversalRAG RAG systems retrieve knowledge to ground model responses. However, most existing approaches are limited to a single modality, typically text. For many real-world RAG systems, some queries need images. Others need videos. Many need combinations. This new research introduces UniversalRAG, a framework that retrieves and integrates knowledge from heterogeneous sources across diverse modalities and granularities. Real-world queries vary widely in what knowledge they need. A universal RAG framework that dynamically routes to the right modality and granularity serves diverse information needs that no single-corpus approach can address. Instead of forcing everything into one embedding space, UniversalRAG uses modality-aware routing. A router dynamically predicts which modality-specific corpus best matches the query, then performs targeted retrieval within it. This sidesteps the modality gap entirely by avoiding cross-modal comparisons. Beyond modality, the framework also handles granularity. Complex analytical questions may need full documents or complete videos. Simple factoid questions are better served with paragraphs or short clips. UniversalRAG organizes each modality into multiple granularity levels: paragraphs and documents for text, clips and full videos for video, plus tables and images. The router can be trained or training-free. The trained version uses inductive biases from existing benchmarks. The training-free version prompts frontier models like Gemini to predict the best modality-granularity pairs directly. Validation across 10 benchmarks spanning text, images, tables, and videos shows UniversalRAG outperforms both unimodal RAG baselines and unified embedding approaches by large margins on average. Paper: https://t.co/OAfR65bEm2 Learn to build effective Agentic RAG systems in our academy: https://t.co/JBU5beIoD0

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 15, 2026
211d ago
๐Ÿ†”68379782

This is great insight from Cursor on long-running agents. It turns out planning is all you need. On a serious note, planning is critical to be productive and effective with AI Agents. It's aligned with how I get Claude Code to effectively work on long-running tasks. (More of my thoughts on Claude Code towards the end of the post) First, let's discuss the insights from the Cursor article. The big problem with multi-agent systems today is the coordination/communication. The solution Cursor proposes is careful planning. I agree that the best way to deal with this challenge today is to do careful planning. Planners can explore the codebase and create these tasks. Then, subplanners are spawned to address specific categories of tasks. The great thing about this is that it enables parallelization and recursive loops, ideal for this kind of work. From here, subagents can focus on assigned subtasks once they are completed (and push changes). One important aspect of this work is that subagents don't coordinate at all and are oblivious to the bigger picture. But they don't need to be to produce high-quality code that doesn't conflict. The issue with having subagents talk to each other is that this can lead to communication bottlenecks, duplicate work, and potential drift. This can operate in cycles, which are all verified using a judge agent. The judge agent determines if work can continue or if there is an issue to address on every cycle. Cursor managed to build a web browser from scratch with this approach. The agent ran for a week, writing over 1M+ lines of code across 1K files. Cursor found that GPT-5.2 is better for this set up. They find that Opus 4.5 tends to stop earlier, take shortcuts, and quickly yield back control. The simpler system worked best. "Too little structure and agents conflict, duplicate work, and drift. Too much structure creates fragility." An interesting finding: designing an effective system prompt to focus over long periods was more important than the harness and models themselves. Why this resonated with me, even though I am not a Cursor user: I have been testing Claude Code on long-running tasks. And what Cursor reports is aligned with my own findings. However, better planning and tuning of the system prompt, including tuning CLAUDE MD, has allowed me to leverage Claude Code more effectively for these long-running tasks. Here are a few notes on planning and how you get something like this to work in Claude Code: You can do effective planning in many different ways. You can create an initial plan and complete it with Claude Code (in plan mode). Or you can brainstorm the plan with Claude Code directly (in plan mode). Claude Code is excellent at managing plans for you in case you don't want lots of moving parts. This, together with subagents works extremely well in Claude Code already. However, you can also get more creative with how planning is done to mimic the subplanners proposed by Cursor. Claude Code is extremely flexible with all its functionalities (Skills, Slash Commands, Subagents, Hooks, etc.). I will share more on this later after I finish with some experiments I am currently working on. When planning, it helps if you are also involved in the process. If you are a Cluade Code user, you can trigger the AskUserQuestion tool to inject inputs that will help with making the plan robust. From here, you can offload individual work to subagents (in parallel if you want). The great part about this is that in Claude Code subagents manage their own context, which keeps the main orchestrator's context clean and only for the high-level stuff. You can customize your subagents with models and tools. The planning is core for the coordination to work. The system prompt helps to maintain stability and better manage context. The subagents are just in charge of executing the work. I will be sharing more on my setup in the coming weeks. I am fascinated by how far we can push agent harnesses for long-horizon tasks. Stay tuned!

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 16, 2026
209d ago
๐Ÿ†”63249167

The reality is that we should all be trying to build our own ideal agentic coworker. Anthropic's Cowork signals a new wave of agent orchestration tools on the horizon. It's not just about making it easy to use Claude Code. IMO, it's more about building intuitive interfaces to interact, manage, and orchestrate agents to be more productive with AI. There is room for innovation in how human-to-agent (and vice versa) communication happens. And I mean everywhere and every domain you can imagine. With the rise of Skills and plugins, anyone can build powerful experiences with these agents and tools. You don't need to be technical to disrupt and build creative and insanely useful skills (either for work, a personal project, or even a startup). You need to have good taste in the domain you are operating, pay close attention to emerging AI technology, experiment relentlessly, build context, and build with a compounding mindset. Exciting times ahead. It's time to build!

Media 1
๐Ÿ–ผ๏ธ Media
D
dair_ai
@dair_ai
๐Ÿ“…
Jan 15, 2026
211d ago
๐Ÿ†”37827206

Autonomous Memory Management in LLM Agents LLM agents struggle with long-horizon tasks due to context bloat. As interaction history grows, computational costs explode, latency increases, and reasoning degrades from distraction by irrelevant past errors. The standard approach is append-only: every thought, tool call, and response permanently accumulates. This works for short tasks but guarantees failure for complex exploration. This research introduces Focus, an agent-centric architecture inspired by slime mold (Physarum polycephalum). The biological insight: organisms do not retain perfect records of every movement through a maze. They retain the learned map. Focus gives agents two new primitives: start_focus and complete_focus. The agent autonomously decides when to consolidate learnings into a persistent Knowledge block and actively prunes the raw interaction history. No external timers or heuristics forcing compression. It declares what you are investigating, explores using standard tools, and then consolidates by summarizing what was attempted, what was learned, and the outcome. The system appends this to a persistent Knowledge block and deletes everything between the checkpoint and the current step. This converts monotonically increasing context into a sawtooth pattern: growth during exploration, collapse during consolidation. Evaluation on SWE-bench Lite with Claude Haiku 4.5 shows Focus achieves 22.7% token reduction (14.9M to 11.5M tokens) while maintaining identical accuracy (60% for both baseline and Focus). Individual instances showed savings up to 57%. Aggressive prompting matters. Passive prompting yielded only 6% savings. Explicit instructions to compress every 10-15 tool calls, with system reminders, increased compressions from 2.0 to 6.0 per task. Capable models can autonomously self-regulate their context when given appropriate tools and prompting, opening pathways for cost-aware agentic systems without sacrificing task performance. Paper: https://t.co/bVkeQlrvGJ Learn to build effective AI agents in our academy: https://t.co/zQXQt0PMbG

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 17, 2026
208d ago
๐Ÿ†”54414130

Ralph Research now has a dashboard to monitor progress. While the ralph-research loop cooks inside of Claude Code, I can now visually track code implementation and experiments. Claude Code hooks are very cool! Now building a control center to orchestrate & scale experiments. https://t.co/qXmxSmfVxN

๐Ÿ–ผ๏ธ Media
D
dair_ai
@dair_ai
๐Ÿ“…
Jan 17, 2026
208d ago
๐Ÿ†”91865973

A huge claim from this paper on the end of reward engineering. Reward engineering remains a persistent bottleneck in multi-agent RL. This paper argues that LLMs enable a fundamental shift: from hand-crafted reward functions to natural language objectives. If language can specify what we want and LLMs can translate that into working rewards, the end of reward engineering may mark the beginning of truly scalable multi-agent coordination. Instead of translating human intent into numbers, a lossy and error-prone process, we can describe it in the same language we use with each other. EUREKA demonstrates GPT-4 can generate reward functions achieving human-level performance from language descriptions alone, It outperforms human-designed rewards on 83% of robotics tasks. CARD enables autonomous reward refinement without human intervention. RLVR (as in DeepSeek-R1) shows that language-based training produces emergent reasoning capabilities. What enables all of this? First, semantic reward specification: language preserves intent that numerical functions lose. "Collaborate efficiently" carries rich meaning about task division, smooth handoffs, and failure recovery that no weighted sum captures. Second, dynamic adaptation: when reward hacking occurs, an LLM can observe trajectories, generate feedback in natural language, and refine rewards automatically. No more weeks of manual debugging. Third, inherent human alignment: language objectives are interpretable. Debugging "minimize delivery time while avoiding collisions" is far easier than debugging opaque weight vectors. A few challenges remain: computational cost of LLM inference, hallucination risks in safety-critical systems, language ambiguity, and scaling to hundreds of agents. Paper: https://t.co/czW7QPVML1 Learn to build effective AI Agents in our academy: https://t.co/Y5kVy5iKiQ

Media 1Media 2
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Jan 18, 2026
208d ago
๐Ÿ†”60297408

The better the plan, the better your agents perform. Long-horizon agents break not because they can't plan. They break because they plan over entangled contexts. The default approach to LLM agent planning falls into two camps. Step-wise planning (like ReAct) interleaves reasoning and acting but makes short-sighted decisions. One-shot planning generates complete plans upfront but becomes brittle when execution errors occur. But both share the same flaw: a single, growing execution history that mixes information across multiple sub-tasks. This new research introduces Task-Decoupled Planning (TDP), a training-free framework that replaces entangled reasoning with explicit task decoupling. How does it work? A Supervisor decomposes tasks into a directed acyclic graph (DAG) of sub-goals. A Planner and Executor then operate with scoped contexts, reasoning only over the active sub-task. This reminds me of the new blog published by Cursor, which uses a similar tactic where planning is decoupled. When something goes wrong, replanning stays local. Independent decisions remain untouched. Isolating context, decisions, and error correction at the sub-task level prevents local failures from cascading across the entire workflow. On TravelPlanner, TDP achieves the highest hard-constraint micro pass rate (32.5%) under DeepSeek-V3.2. On HotpotQA, it reaches 85.88% delivery accuracy. On ScienceWorld, it matches or exceeds strong baselines across both GPT-4o and DeepSeek models. TDP reduces token consumption by up to 82% compared to Plan-and-Act while improving task outcomes. On HotpotQA, it uses just 1,747 output tokens versus 9,929 for the baseline. Fewer tokens, better results. Sub-task decoupling offers a unified mechanism that works across heterogeneous demands like multi-hop reasoning, interactive environments, and constraint-heavy tool planning. You get all of this without sacrificing performance or efficiency. Paper: https://t.co/0hOsV3wFsZ Learn to build effective AI agents in our academy: https://t.co/JBU5beIoD0

Media 1
๐Ÿ–ผ๏ธ Media