🐦 Twitter Post Details

Viewing enriched Twitter post

@dair_ai

Interesting new paper on online RL for agents. Most agent training still treats deployment and learning as separate phases. Serve the model first, collect data later, fine-tune offline. But every agent interaction already contains a learning signal. This paper introduces OpenClaw-RL, a framework that trains agents from the next state that follows each action: user replies, tool outputs, terminal traces, GUI changes, and test results. The key idea is to recover two signals at once. Evaluative signals become scalar rewards through a PRM judge. Directive signals become token-level supervision through hindsight-guided on-policy distillation. In their personalization setup, the combined method improves the score from 0.17 to 0.81 after 16 update steps, outperforming binary RL or OPD alone. They also show gains across tool-call and GUI agents when combining process and outcome rewards. Why it matters? Agents should improve simply by being used. This is a practical step toward online, always-learning agent systems instead of static deployments. Paper: https://t.co/bzLdREoet8 Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c

Media 1
Media 2

📊 Media Metadata

{
  "media": [
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2032261877284446344/media_0.jpg",
      "filename": "media_0.jpg"
    },
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2032261877284446344/media_1.png",
      "filename": "media_1.png"
    }
  ],
  "processed_at": "2026-03-13T01:16:31.154517",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2032261877284446344",
  "url": "https://x.com/dair_ai/status/2032261877284446344",
  "twitterUrl": "https://twitter.com/dair_ai/status/2032261877284446344",
  "text": "Interesting new paper on online RL for agents.\n\nMost agent training still treats deployment and learning as separate phases. Serve the model first, collect data later, fine-tune offline.\n\nBut every agent interaction already contains a learning signal.\n\nThis paper introduces OpenClaw-RL, a framework that trains agents from the next state that follows each action: user replies, tool outputs, terminal traces, GUI changes, and test results.\n\nThe key idea is to recover two signals at once. Evaluative signals become scalar rewards through a PRM judge. Directive signals become token-level supervision through hindsight-guided on-policy distillation.\n\nIn their personalization setup, the combined method improves the score from 0.17 to 0.81 after 16 update steps, outperforming binary RL or OPD alone. They also show gains across tool-call and GUI agents when combining process and outcome rewards.\n\nWhy it matters?\n\nAgents should improve simply by being used. This is a practical step toward online, always-learning agent systems instead of static deployments.\n\nPaper: https://t.co/bzLdREoet8\n\nLearn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c",
  "source": "Twitter for iPhone",
  "retweetCount": 2,
  "replyCount": 1,
  "likeCount": 4,
  "quoteCount": 0,
  "viewCount": 213,
  "createdAt": "Fri Mar 13 01:06:03 +0000 2026",
  "lang": "en",
  "bookmarkCount": 9,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2032261877284446344",
  "displayTextRange": [
    0,
    274
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "dair_ai",
    "url": "https://x.com/dair_ai",
    "twitterUrl": "https://twitter.com/dair_ai",
    "id": "889050642903293953",
    "name": "DAIR.AI",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1643277398522187778/31dedbLo_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/889050642903293953/1773242460",
    "description": "",
    "location": "",
    "followers": 91678,
    "following": 1,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Sun Jul 23 09:12:45 +0000 2017",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 4256,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 174,
    "statusesCount": 3013,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2032107624007876781"
    ],
    "profile_bio": {
      "description": "Democratizing AI research, education, and technologies. New AI learning portal: https://t.co/LRnpZN7L4c",
      "entities": {
        "description": {
          "hashtags": [],
          "symbols": [],
          "urls": [
            {
              "display_url": "academy.dair.ai",
              "expanded_url": "https://academy.dair.ai/",
              "indices": [
                80,
                103
              ],
              "url": "https://t.co/LRnpZN7L4c"
            }
          ],
          "user_mentions": []
        },
        "url": {
          "urls": [
            {
              "display_url": "dair.ai",
              "expanded_url": "https://www.dair.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/lkqPZtMU5s"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/qG90fnEFru",
        "expanded_url": "https://twitter.com/dair_ai/status/2032261877284446344/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {
            "faces": [
              {
                "h": 80,
                "w": 80,
                "x": 149,
                "y": 1236
              }
            ]
          },
          "orig": {
            "faces": [
              {
                "h": 80,
                "w": 80,
                "x": 149,
                "y": 1236
              }
            ]
          }
        },
        "id_str": "2032261874059034624",
        "indices": [
          275,
          298
        ],
        "media_key": "3_2032261874059034624",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARw0C2g128AACgACHDQLaPYboIgAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABHDQLaDXbwAAKAAIcNAto9hugiAAA",
            "media_key": "3_2032261874059034624"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HDQLaDXbwAAXUC3.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 750,
              "w": 1340,
              "x": 0,
              "y": 0
            },
            {
              "h": 1340,
              "w": 1340,
              "x": 0,
              "y": 0
            },
            {
              "h": 1528,
              "w": 1340,
              "x": 0,
              "y": 0
            },
            {
              "h": 1720,
              "w": 860,
              "x": 240,
              "y": 0
            },
            {
              "h": 1720,
              "w": 1340,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1720,
          "width": 1340
        },
        "sizes": {
          "large": {
            "h": 1720,
            "w": 1340
          }
        },
        "type": "photo",
        "url": "https://t.co/qG90fnEFru"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "arxiv.org/abs/2603.10165",
        "expanded_url": "https://arxiv.org/abs/2603.10165",
        "indices": [
          1069,
          1092
        ],
        "url": "https://t.co/bzLdREoet8"
      },
      {
        "display_url": "academy.dair.ai",
        "expanded_url": "https://academy.dair.ai/",
        "indices": [
          1145,
          1168
        ],
        "url": "https://t.co/LRnpZN7L4c"
      }
    ],
    "user_mentions": []
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "article": null
}