🐦 Twitter Post Details

Viewing enriched Twitter post

@rasbt

I have been pretty heads-down this year to finish Chapter 6 on implementing reinforcement learning with verifiable rewards from scratch (using GRPO). I just finished it this weekend, and I'd say it's the best (or at least my favorite) chapter yet! The goal of this chapter is to explain and implement GRPO from the bottom up. This means coding and walking through each GRPO step one by one (advantages, rewards, logprobs, and loss) and then training a 0.6B base model on the 12k examples from the MATH training set. (This takes the model from 15% to 47% accuracy on the MATH-500 test set, which is about as good as the official Qwen3 reasoning model of similar size.) The focus is on readability and understanding GRPO, but the supplementary materials also contain scripts to run it in a multi-GPU setting. The code notebook is already available on GitHub if you want to take a look: https://t.co/SM58MXjf8V. (And the full chapter should make it to the early access version of the book at https://t.co/vzCr5sTjrf soon!) PS: The next chapter will introduce additional tips and tricks to improve the GRPO algorithm for better and more stable training behavior.

Media 1
Media 2

📊 Media Metadata

{
  "media": [
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2012897755916579278/media_0.jpg?",
      "filename": "media_0.jpg"
    },
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2012897755916579278/media_1.jpg?",
      "filename": "media_1.jpg"
    }
  ],
  "processed_at": "2026-01-18T17:51:07.712774",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2012897755916579278",
  "url": "https://x.com/rasbt/status/2012897755916579278",
  "twitterUrl": "https://twitter.com/rasbt/status/2012897755916579278",
  "text": "I have been pretty heads-down this year to finish Chapter 6 on implementing reinforcement learning with verifiable rewards from scratch (using GRPO). I just finished it this weekend, and I'd say it's the best (or at least my favorite) chapter yet!\n\nThe goal of this chapter is to explain and implement GRPO from the bottom up. This means coding and walking through each GRPO step one by one (advantages, rewards, logprobs, and loss) and then training a 0.6B base model on the 12k examples from the MATH training set. \n(This takes the model from 15% to 47% accuracy on the MATH-500 test set, which is about as good as the official Qwen3 reasoning model of similar size.)\n\nThe focus is on readability and understanding GRPO, but the supplementary materials also contain scripts to run it in a multi-GPU setting.\n\nThe code notebook is already available on GitHub if you want to take a look: https://t.co/SM58MXjf8V. \n\n(And the full chapter should make it to the early access version of the book at https://t.co/vzCr5sTjrf soon!)\n\nPS: The next chapter will introduce additional tips and tricks to improve the GRPO algorithm for better and more stable training behavior.",
  "source": "Twitter for iPhone",
  "retweetCount": 91,
  "replyCount": 21,
  "likeCount": 660,
  "quoteCount": 2,
  "viewCount": 21020,
  "createdAt": "Sun Jan 18 14:39:57 +0000 2026",
  "lang": "en",
  "bookmarkCount": 481,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2012897755916579278",
  "displayTextRange": [
    0,
    304
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "rasbt",
    "url": "https://x.com/rasbt",
    "twitterUrl": "https://twitter.com/rasbt",
    "id": "865622395",
    "name": "Sebastian Raschka",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1661187442043486209/a3E4t1eV_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/865622395/1742309979",
    "description": "",
    "location": "",
    "followers": 380682,
    "following": 1116,
    "status": "",
    "canDm": false,
    "canMediaTag": true,
    "createdAt": "Sun Oct 07 02:06:16 +0000 2012",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 23824,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 2044,
    "statusesCount": 19230,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2006015301717028989"
    ],
    "profile_bio": {
      "description": "ML/AI research engineer. Ex stats professor.\nAuthor of \"Build a Large Language Model From Scratch\" (https://t.co/O8LAAMRzzW) & reasoning (https://t.co/5TueQKx2Fk)",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "amzn.to/4fqvn0D",
              "expanded_url": "https://amzn.to/4fqvn0D",
              "indices": [
                100,
                123
              ],
              "url": "https://t.co/O8LAAMRzzW"
            },
            {
              "display_url": "mng.bz/lZ5B",
              "expanded_url": "https://mng.bz/lZ5B",
              "indices": [
                138,
                161
              ],
              "url": "https://t.co/5TueQKx2Fk"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "sebastianraschka.com",
              "expanded_url": "https://sebastianraschka.com",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/HrtQQ5tgJl"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "allow_download_status": {
          "allow_download": true
        },
        "display_url": "pic.twitter.com/HBJDq5TzoE",
        "expanded_url": "https://twitter.com/rasbt/status/2012897755916579278/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {},
          "orig": {}
        },
        "id_str": "2012897342144368640",
        "indices": [
          305,
          328
        ],
        "media_key": "3_2012897342144368640",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARvvP3fH19AACgACG+8/2B6Wsc4AAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABG+8/d8fX0AAKAAIb7z/YHpaxzgAA",
            "media_key": "3_2012897342144368640"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/G-8_d8fX0AAfki4.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 2294,
              "w": 4096,
              "x": 0,
              "y": 0
            },
            {
              "h": 3235,
              "w": 3235,
              "x": 531,
              "y": 0
            },
            {
              "h": 3235,
              "w": 2838,
              "x": 729,
              "y": 0
            },
            {
              "h": 3235,
              "w": 1618,
              "x": 1339,
              "y": 0
            },
            {
              "h": 3235,
              "w": 4096,
              "x": 0,
              "y": 0
            }
          ],
          "height": 3235,
          "width": 4096
        },
        "sizes": {
          "large": {
            "h": 1618,
            "w": 2048
          }
        },
        "type": "photo",
        "url": "https://t.co/HBJDq5TzoE"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "urls": [
      {
        "display_url": "github.com/rasbt/reasonin…",
        "expanded_url": "https://github.com/rasbt/reasoning-from-scratch/blob/main/ch06/01_main-chapter-code/ch06_main.ipynb",
        "indices": [
          888,
          911
        ],
        "url": "https://t.co/SM58MXjf8V"
      },
      {
        "display_url": "mng.bz/Nwr7",
        "expanded_url": "https://mng.bz/Nwr7",
        "indices": [
          995,
          1018
        ],
        "url": "https://t.co/vzCr5sTjrf"
      }
    ]
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "article": null
}