🐦 Twitter Post Details

Viewing enriched Twitter post

@dair_ai

On building more powerful self-evolving agents. LLM agents struggle to learn from experience after deployment. Fine-tuning is expensive and causes catastrophic forgetting. RAG retrieves based on semantic similarity alone, often pulling noise instead of what actually works. Similarity and utility are not the same thing. This new research introduces MemRL, a framework that enables agents to self-evolve through non-parametric reinforcement learning on episodic memory, keeping the LLM completely frozen. The core idea is to treat memory retrieval as a decision-making problem, not a matching problem. Each memory stores an Intent-Experience-Utility triplet. The utility is a learned Q-value representing expected returns, continuously refined through environmental feedback. MemRL implements Two-Phase Retrieval. First, filter candidates by semantic similarity to ensure relevance. Then, rank by learned Q-values to select what actually works. This distinguishes high-value strategies from semantically similar noise. When the agent succeeds or fails, it updates the Q-values of retrieved memories using Bellman-style backups. No gradient updates to model weights. The frozen LLM provides stable reasoning while the memory evolves plastically. Results across four benchmarks: On HLE (knowledge frontier tasks), MemRL significantly outperforms both RAG and existing memory systems like MemP. The pattern holds on BigCodeBench for code generation, ALFWorld for exploration tasks, and Lifelong Agent Bench for OS and database operations. Analysis confirms a strong correlation between learned utility scores and actual task success, validating that Q-values capture genuine functional value rather than superficial similarity. Why does it matter? Decoupling stable reasoning from plastic memory enables continuous runtime improvement without the catastrophic forgetting or computational costs of fine-tuning. Paper: https://t.co/HvLUnXW2Jd Learn to build effective AI agents in our academy: https://t.co/zQXQt0PMbG

Media 1
Media 2

📊 Media Metadata

{
  "media": [
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2011086096986443905/media_0.jpg?",
      "filename": "media_0.jpg"
    },
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2011086096986443905/media_1.png?",
      "filename": "media_1.png"
    }
  ],
  "processed_at": "2026-01-18T17:31:45.108434",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2011086096986443905",
  "url": "https://x.com/dair_ai/status/2011086096986443905",
  "twitterUrl": "https://twitter.com/dair_ai/status/2011086096986443905",
  "text": "On building more powerful self-evolving agents.\n\nLLM agents struggle to learn from experience after deployment. Fine-tuning is expensive and causes catastrophic forgetting. RAG retrieves based on semantic similarity alone, often pulling noise instead of what actually works.\n\nSimilarity and utility are not the same thing.\n\nThis new research introduces MemRL, a framework that enables agents to self-evolve through non-parametric reinforcement learning on episodic memory, keeping the LLM completely frozen.\n\nThe core idea is to treat memory retrieval as a decision-making problem, not a matching problem. Each memory stores an Intent-Experience-Utility triplet. The utility is a learned Q-value representing expected returns, continuously refined through environmental feedback.\n\nMemRL implements Two-Phase Retrieval. First, filter candidates by semantic similarity to ensure relevance. Then, rank by learned Q-values to select what actually works. This distinguishes high-value strategies from semantically similar noise.\n\nWhen the agent succeeds or fails, it updates the Q-values of retrieved memories using Bellman-style backups. No gradient updates to model weights. The frozen LLM provides stable reasoning while the memory evolves plastically.\n\nResults across four benchmarks: On HLE (knowledge frontier tasks), MemRL significantly outperforms both RAG and existing memory systems like MemP. The pattern holds on BigCodeBench for code generation, ALFWorld for exploration tasks, and Lifelong Agent Bench for OS and database operations.\n\nAnalysis confirms a strong correlation between learned utility scores and actual task success, validating that Q-values capture genuine functional value rather than superficial similarity.\n\nWhy does it matter? Decoupling stable reasoning from plastic memory enables continuous runtime improvement without the catastrophic forgetting or computational costs of fine-tuning.\n\nPaper: https://t.co/HvLUnXW2Jd\n\nLearn to build effective AI agents in our academy: https://t.co/zQXQt0PMbG",
  "source": "Twitter for iPhone",
  "retweetCount": 52,
  "replyCount": 14,
  "likeCount": 281,
  "quoteCount": 5,
  "viewCount": 25960,
  "createdAt": "Tue Jan 13 14:41:04 +0000 2026",
  "lang": "en",
  "bookmarkCount": 281,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2011086096986443905",
  "displayTextRange": [
    0,
    299
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "dair_ai",
    "url": "https://x.com/dair_ai",
    "twitterUrl": "https://twitter.com/dair_ai",
    "id": "889050642903293953",
    "name": "DAIR.AI",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1643277398522187778/31dedbLo_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/889050642903293953/1742055232",
    "description": "",
    "location": "",
    "followers": 85514,
    "following": 1,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Sun Jul 23 09:12:45 +0000 2017",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 4016,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 122,
    "statusesCount": 2821,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2012903315890225220"
    ],
    "profile_bio": {
      "description": "Democratizing AI research, education, and technologies. New Claude Code cohort: https://t.co/XCpQTjR9hg",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "dair-ai.thinkific.com/courses/claude…",
              "expanded_url": "https://dair-ai.thinkific.com/courses/claude-code-for-everyone-cohort-3",
              "indices": [
                80,
                103
              ],
              "url": "https://t.co/XCpQTjR9hg"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "dair.ai",
              "expanded_url": "https://www.dair.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/lkqPZtMmfU"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/WjQuzIn80T",
        "expanded_url": "https://twitter.com/dair_ai/status/2011086096986443905/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {},
          "orig": {}
        },
        "id_str": "2011086093966610432",
        "indices": [
          300,
          323
        ],
        "media_key": "3_2011086093966610432",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARvo0CWzm5AACgACG+jQJmeakIEAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABG+jQJbObkAAKAAIb6NAmZ5qQgQAA",
            "media_key": "3_2011086093966610432"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/G-jQJbObkAAWJ3o.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 904,
              "w": 1614,
              "x": 0,
              "y": 0
            },
            {
              "h": 1614,
              "w": 1614,
              "x": 0,
              "y": 0
            },
            {
              "h": 1800,
              "w": 1579,
              "x": 0,
              "y": 0
            },
            {
              "h": 1800,
              "w": 900,
              "x": 0,
              "y": 0
            },
            {
              "h": 1800,
              "w": 1614,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1800,
          "width": 1614
        },
        "sizes": {
          "large": {
            "h": 1800,
            "w": 1614
          }
        },
        "type": "photo",
        "url": "https://t.co/WjQuzIn80T"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "urls": [
      {
        "display_url": "arxiv.org/abs/2601.03192",
        "expanded_url": "https://arxiv.org/abs/2601.03192",
        "indices": [
          1924,
          1947
        ],
        "url": "https://t.co/HvLUnXW2Jd"
      },
      {
        "display_url": "dair-ai.thinkific.com",
        "expanded_url": "https://dair-ai.thinkific.com/",
        "indices": [
          2000,
          2023
        ],
        "url": "https://t.co/zQXQt0PMbG"
      }
    ]
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "article": null
}