🐦 Twitter Post Details

Viewing enriched Twitter post

@dair_ai

Interesting technical work from Microsoft. Provides a better understanding on SFT and how to leverage it better for RL. Microsoft researchers asked whether a standard SFT pipeline actually produces the model you want to run RL on. Their answer is no. Standard SFT keeps spending gradient on sequences the model has already fit, which narrows the distribution RL later needs to explore. TailSFT filters those sequences out during training and concentrates learning on the under-modeled tail of the data. That is the only modification they implement. Results: On OLMo-3 7B, pass@16 improves by up to 16.8 points absolute on coding and 3.1 on math. Those higher-coverage checkpoints then lift final pass@1 after GRPO by up to 3.9 points, and in some settings early reward climbs 2.5x faster than the matched standard SFT run. Paper: https://t.co/QocPRtNhjH Chat with Paper: https://t.co/wfTytUm5jp

Media 1

📊 Media Metadata

{
  "media": [
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2094138107314753938/media_0.jpg",
      "filename": "media_0.jpg"
    }
  ],
  "processed_at": "2026-08-30T21:02:15.584174",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2094138107314753938",
  "url": "https://x.com/dair_ai/status/2094138107314753938",
  "twitterUrl": "https://twitter.com/dair_ai/status/2094138107314753938",
  "text": "Interesting technical work from Microsoft.\n\nProvides a better understanding on SFT and how to leverage it better for RL.\n\nMicrosoft researchers asked whether a standard SFT pipeline actually produces the model you want to run RL on.\n\nTheir answer is no.\n\nStandard SFT keeps spending gradient on sequences the model has already fit, which narrows the distribution RL later needs to explore.\n\nTailSFT filters those sequences out during training and concentrates learning on the under-modeled tail of the data.\n\nThat is the only modification they implement.\n\nResults:\n\nOn OLMo-3 7B, pass@16 improves by up to 16.8 points absolute on coding and 3.1 on math. Those higher-coverage checkpoints then lift final pass@1 after GRPO by up to 3.9 points, and in some settings early reward climbs 2.5x faster than the matched standard SFT run.\n\nPaper: https://t.co/QocPRtNhjH\n\nChat with Paper: https://t.co/wfTytUm5jp",
  "source": "Twitter for iPhone",
  "retweetCount": 7,
  "replyCount": 3,
  "likeCount": 37,
  "quoteCount": 1,
  "viewCount": 2953,
  "createdAt": "Sun Aug 30 19:00:06 +0000 2026",
  "lang": "en",
  "bookmarkCount": 25,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2094138107314753938",
  "displayTextRange": [
    0,
    273
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "dair_ai",
    "url": "https://x.com/dair_ai",
    "twitterUrl": "https://twitter.com/dair_ai",
    "id": "889050642903293953",
    "name": "DAIR.AI",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1643277398522187778/31dedbLo_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/889050642903293953/1773242460",
    "description": "",
    "location": "",
    "followers": 130805,
    "following": 1,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Sun Jul 23 09:12:45 +0000 2017",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 5147,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 337,
    "statusesCount": 3668,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2093324233158045788"
    ],
    "profile_bio": {
      "description": "Democratizing AI research, education, and technologies. Learn about AI Agents for FREE at https://t.co/HHXg8rryu4",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "academy.dair.ai/courses/elemen…",
              "expanded_url": "https://academy.dair.ai/courses/elements-of-ai-agents",
              "indices": [
                90,
                113
              ],
              "url": "https://t.co/HHXg8rryu4"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "dair.ai",
              "expanded_url": "https://www.dair.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/lkqPZtMU5s"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.x.com/1LnePyS9YN",
        "expanded_url": "https://x.com/dair_ai/status/2094138107314753938/photo/1",
        "ext_master_playlist_only": [],
        "ext_media_availability": {
          "status": "Available"
        },
        "ext_playlists": [],
        "features": {
          "large": {
            "faces": []
          },
          "orig": {
            "faces": []
          }
        },
        "id_str": "2094138103485308928",
        "indices": [
          274,
          297
        ],
        "media_key": "3_2094138103485308928",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAAR0P34KI2iAACgACHQ/fg20a0ZIAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABHQ/fgojaIAAKAAIdD9+DbRrRkgAA",
            "media_key": "3_2094138103485308928"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HQ_fgojaIAAe4EC.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 987,
              "w": 1762,
              "x": 0,
              "y": 0
            },
            {
              "h": 1762,
              "w": 1762,
              "x": 0,
              "y": 0
            },
            {
              "h": 1858,
              "w": 1630,
              "x": 0,
              "y": 0
            },
            {
              "h": 1858,
              "w": 929,
              "x": 325,
              "y": 0
            },
            {
              "h": 1858,
              "w": 1762,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1858,
          "width": 1762
        },
        "sizes": {
          "large": {
            "h": 1858,
            "w": 1762
          }
        },
        "type": "photo",
        "url": "https://t.co/1LnePyS9YN"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "arxiv.org/abs/2608.25756",
        "expanded_url": "https://arxiv.org/abs/2608.25756",
        "indices": [
          839,
          862
        ],
        "url": "https://t.co/QocPRtNhjH"
      },
      {
        "display_url": "academy.dair.ai/papers/tailsft…",
        "expanded_url": "https://academy.dair.ai/papers/tailsft-filtered-fine-tuning-improves-post-training-performance-2608.25756",
        "indices": [
          881,
          904
        ],
        "url": "https://t.co/wfTytUm5jp"
      }
    ],
    "user_mentions": []
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}