Your curated collection of saved posts and media
I have been pretty heads-down this year to finish Chapter 6 on implementing reinforcement learning with verifiable rewards from scratch (using GRPO). I just finished it this weekend, and I'd say it's the best (or at least my favorite) chapter yet! The goal of this chapter is to explain and implement GRPO from the bottom up. This means coding and walking through each GRPO step one by one (advantages, rewards, logprobs, and loss) and then training a 0.6B base model on the 12k examples from the MATH training set. (This takes the model from 15% to 47% accuracy on the MATH-500 test set, which is about as good as the official Qwen3 reasoning model of similar size.) The focus is on readability and understanding GRPO, but the supplementary materials also contain scripts to run it in a multi-GPU setting. The code notebook is already available on GitHub if you want to take a look: https://t.co/SM58MXjf8V. (And the full chapter should make it to the early access version of the book at https://t.co/vzCr5sTjrf soon!) PS: The next chapter will introduce additional tips and tricks to improve the GRPO algorithm for better and more stable training behavior.

@StronglyAI I tested all these below. Regarding evals: yes, I am testing against MATH-500 (non-overlapping with the 12k examples in the training set) https://t.co/zXXsxYpw3f
Btw I also ran several additional ablation studies on improving GRPO. I summarized them in my latest blog a few weeks ago. The goal here is to have the GRPO foundation (this chapter) to which we can then apply these tweaks (next chapter) https://t.co/HLRhabyW0A
@aina_oluwa Build a Reasoning Model From Scratch, https://t.co/fQndtsmUJv (sequel to Build A Large Language Model From Scratch)
using @GoogleDeepMind's Nano Banana to create a 3D version of @ivanleomk for our hack and roll project! @thorwebdev might be next hehe Can you guess what we're building!? CC: @miinnong @dongkiatt https://t.co/jH4HFtcTWe
my last hackathon as a full-time student: stickers, food and one hell of a time building at @nushackers's annual hack n roll! happy to have built something i've always wanted to try for the longest time! hearing your code! See what we built: - Strudel Your Code: "Turn your coding sessions into live generative music performances" https://t.co/yTnCC2UJFq (featuring 3d chibis of @gabrielchua @agrimsingh @ivanleomk @thorwebdev and @GeoffreyHuntley's ralph wiggum) Strudel is "a new live coding platform to write dynamic music pieces in the browser!" - We built Strudel Your Code with @cursor_ai, @openai, @AnthropicAI and more. It runs primarily on OpenAI's responses, with a single 2D Nano Banana Pro from @GeminiApp to @TencentHunyuan's 2d-3d model for our 3D models. It was insanely fun to build on @threejs, i'll write a short post on how we got the 3D models into fruition as this seemed to be the most asked question on the board. See our full video here: https://t.co/hTkzJIYKED but of course hackathons arent always about the tech! it's also about connecting to people from various walks and sharing more from each other! also big ups to the sponsors!! for sponsoring @cursor_ai @ExaAILabs @GoogleDeepMind @AnthropicAI @elevenlabsio @v0 @ManusAI @jigsawstack @OpenAI

introducing Clipmorph. an agentic layer between cmd+c and cmd+v. copy data, speak what you want (reformat table, generate chart, extract columns), paste the result No app switching. Your clipboard, now intelligent. https://t.co/cUGgCDBNql
Hack&Roll 2026 is over! What a crazy and exciting day with 809 participants, 156 judges and one really happy duck (*quack) ππ·π₯ https://t.co/cpuBL99KZA
About to let @ManusAI choose my next haircut hahahaha https://t.co/QZgtwh2N7N
age of empire ii mod to visualize @cursor_ai agents running in parallel, all working on separate worktrees ralph-style? surely not... π @ericzakariasson @leerob https://t.co/EVhf5exYmE
https://t.co/9SmVP9plVZ
https://t.co/9SmVP9plVZ
https://t.co/Dk4Cw3BGyz
https://t.co/Dk4Cw3BGyz
Donald Trump Bought at Least $1 Million in Bonds in Netflix, Warner Bros. Discovery Following Their Deal Announcement https://t.co/IvLCIPygcT
Imagine getting called Sug by her https://t.co/OlhcgiDg2Q
Imagine getting called Sug by her https://t.co/OlhcgiDg2Q
https://t.co/tNWdI4MbXC
not his best work https://t.co/kd4B5lEluo
not his best work https://t.co/kd4B5lEluo
Too much hate in this world and people spreading hate. Letβs spread some love. Love conquers all. Bosh. https://t.co/7ZRtBWPE0l
More than 200,000 Danish citizens have signed a petition to buy California as a response to Trumpβs attempt to take Greenland. They say they will provide Californians with βrule of law, universal health care, fact-based politics, and a lifetime supply of Danish pastries.β https://t.co/02hVzcEsi2

I FOUND THE CYBER KIRK https://t.co/1WHmMMnTHe
I FOUND THE CYBER KIRK https://t.co/1WHmMMnTHe
Sinkhornβs Descent credit: @tokenbender https://t.co/br2pkIK2I2
Growing Pains Across the Journey of Adolescence credit: @NousResearch https://t.co/GCUS2qNOYP
Stranger Descents, Eleven Enters The Upside Down credit: @karpathy https://t.co/zNRGnWmkM2
You can now use your GitHub Copilot subscription with @opencode π https://t.co/hhgy2GxbiL
GitHub Copilot now remembers what it learns across your workflow. β‘οΈ Copilot coding agent, CLI, and code review share knowledge automatically β‘οΈ Memories include citations that get verified in real-time before use β‘οΈ Outdated information self-heals as agents discover contradictions β‘οΈ 7% higher PR merge rates in testing, 2% better code review feedback Available now in public preview for all paid Copilot plans. Deep dive below β¬οΈ https://t.co/gJ88y5rr5i
Today on Open Source Friday ποΈ @dennisivy11 from @appwrite joins us to chat: β’ building in public with open source β’ full-stack development workflows β’ community-driven product development β’ migration strategies between platforms Plus, tune in for hints at an exciting AI announcement. π Set a reminder for when we go live.π https://t.co/PczQFA31EU

Imagine if 94% of your compilation errors just β¦ vanished. Your skin would clear. The laundry would be folded. β¨ A new study shows thatβs exactly the percentage of LLM errors that are just type-check failures. @cassidoo explains why AI is settling the "typed vs. untyped" debate. β¬οΈ https://t.co/kSYQfVCkRX
Have you tried /delegate yet in Copilot CLI? Watch the magic β¨ https://t.co/YPCpl7X1rA