Your curated collection of saved posts and media
β¨ One year after the debut of "Project Digits," DGX Spark has transformed from concept to reality and is now available from NVIDIA and OEM partners. See it in action lighting up the show floor at #CES2026. #SparkSomethingBig https://t.co/6UcfxWqsw6
π Introducing OctoCodingBench, a new benchmark for aligned coding agents: https://t.co/oKaF7jjagb Passing tests β aligned behavior. An agent can produce code that aces every unit test while ignoring system guidelines, violating project conventions, or misusing tools. In real-world coding, how you solve matters as much as what you solve. Nobody wants an agent that ships perfect code while deleting your README, reformatting every file, and mass-commenting in LLM-ese. Don't let your coding agent paperclip-max your repo!
π Introducing OctoCodingBench, a new benchmark for aligned coding agents: https://t.co/oKaF7jjagb Passing tests β aligned behavior. An agent can produce code that aces every unit test while ignoring system guidelines, violating project conventions, or misusing tools. In real-world coding, how you solve matters as much as what you solve. Nobody wants an agent that ships perfect code while deleting your README, reformatting every file, and mass-commenting in LLM-ese. Don't let your coding agent paperclip-max your repo!
FLUX.2-klein-9B https://t.co/5vaiIKiSot
FLUX.2-klein-9B https://t.co/5vaiIKiSot
Meta just released Action100M on Hugging Face A massive video dataset with 100M+ hierarchical action annotations. Every video includes tree-of-captions with action labels, brief and detailed summaries. https://t.co/ns9d13Zbfj
voyage-4-nano is now available. Itβs a lightweight, open-weights embedding model designed for fast, cost-efficient retrieval. Available on @huggingface. https://t.co/Vm3W0ytgoc
In case anyone missed -- new models were shipped in Diffusers this week. 1οΈβ£ Flux.2 Klein - significantly consumer-friendlier than Flux.2 2οΈβ£ GLM Image - AR + Diffusion Decoder Check'em out! https://t.co/TufRTTrVKg
Happy 2026! Will this be the year we finally achieve AGI? Iβd like to propose a new version of the Turing Test, which Iβll call the Turing-AGI Test, to see if weβve achieved this. Iβll explain in a moment why having a new test is important. The public thinks achieving AGI means computers will be as intelligent as people and be able to do most or all knowledge work. Iβd like to propose a new test. The test subject β either a computer or a skilled professional human β is given access to a computer that has internet access and software such as a web browser and Zoom. The judge will design a multi-day experience for the test subject, mediated through the computer, to carry out work tasks. For example, an experience might consist of a period of training (say, as a call center operator), followed by being asked to carry out the task (taking calls), with ongoing feedback. This mirrors what a remote worker with a fully working computer (but no webcam) might be expected to do. A computer passes the Turing-AGI Test if it can carry out the work task as well as a skilled human. Most members of the public likely believe a real AGI system will pass this test. Surely, if computers are as intelligent as humans, they should be able to perform work tasks as well as a human one might hire. Thus, the Turing-AGI Test aligns with the popular notion of what AGI means. Hereβs why we need a new test: βAGIβ has turned into a term of hype rather than a term with a precise meaning. A reasonable definition of AGI is AI that can do any intellectual task that a human can. When businesses hype up that they might achieve AGI within a few quarters, they usually try to justify these statements by setting a much lower bar. This mismatch in definitions is harmful because it makes people think AI is becoming more powerful than it actually is. Iβm seeing this mislead everyone from high-school students (who avoid certain fields of study because they think itβs pointless with AGIβs imminent arrival) to CEOs (who are deciding what projects to invest in, sometimes assuming AI will be more capable in 1-2 years than any likely reality). The original Turing Test, which required a computer to fool a human judge, via text chat, into being unable to distinguish it from a human, has been insufficient to indicate human-level intelligence. The Loebner Prize competition actually ran the Turing Test and found that being able to simulate human typing errors β perhaps even more than actually demonstrating intelligence β was needed to fool judges. A main goal of AI development today is to build systems that can do economically useful work, not fool judges. Thus a modified test that measures ability to do work would be more useful than a test that measures the ability to fool humans. For almost all AI benchmarks today (such as GPQA, AIME, SWE-bench, etc.), a test set is determined in advance. This means AI teams end up at least indirectly tuning their models to the published test sets. Further, any fixed test set measures only one narrow sliver of intelligence. In contrast, in the Turing Test, judges are free to ask any question to probe the model as they please. This lets a judge test how βgeneralβ the knowledge of the computer or human really is. Similarly, in the Turing-AGI Test, the judge can design any experience β which is not revealed in advance to the AI (or human subject) being tested. This is a better way to measure generality of AI than a predetermined test set. AI is on an amazing trajectory of progress. In previous decades, overhyped expectations led to AI winters, when disappointment about AI capabilities caused reductions in interest and funding, which picked up again when the field made more progress. One of the few things that could get in the way of AIβs tremendous momentum is unrealistic hype that creates an investment bubble, risking disappointment and a collapse of interest. To avoid this, we need to recalibrate societyβs expectations on AI. A test will help. If we run a Turing-AGI Test competition and every AI system falls short, that will be a good thing! By defusing hype around AGI and reducing the chance of a bubble, we will create a more reliable path to continued investment in AI. This will let us keep on driving forward real technological progress and building valuable applications β even ones that fall well short of AGI. And if this test sets a clear target that teams can aim toward to claim the mantle of achieving AGI, that would be wonderful, too. And we can be confident that if a company passes this test, they will have created more than just a marketing release β it will be something incredibly valuable. [Original text: https://t.co/mGAmoOGga7 ]
If youβve never written code before, this is for you. Iβve just launched a course that shows you, in less than 30 minutes, how to describe an idea for an app and build it with AI. In this course, you'll build a working web application - a funny interactive birthday message generator that runs in your browser and can be shared with friends. You'll customize it by telling AI how you want it changed, and tweak it until it works the way you want. By the end, you'll have a repeatable process you can apply to build a wide variety of applications. If you want to try vibe coding, this will be the best place to start! Further, you'll be able to use these techniques with whatever tool you're most comfortable with (like ChatGPT, Gemini, Claude, or others) -- we're vendor neutral. Skills you'll gain: - How to build web apps with AI - zero coding skills needed - How to fix and improve your creations by chatting with AI - A simple process you can use to build other things you can dream up Building with AI is one of the most fun things in the world. Please join me and take your first step! I think you will be surprised at what you can build. And if you're an experienced engineer, please share this with someone in your life who's been curious about building with AI. Come build with me! https://t.co/q6gyzlxWFS
New course: Document AI: From OCR to Agentic Doc Extraction, built with @LandingAI, where I'm executive chairman, and taught by David Park and Andrea Kropp. Much of the world's data is locked in PDFs, JPEGs, and other documents. This short course shows you how to build agentic workflows that process documents accurately: breaking them into parts, examining each piece carefully, and extracting information through multiple iterations. Traditional Optical Character Recognition (OCR) captures text but loses context from table headers, chart captions, or reading order of columns. After exploring OCR's limitations, youβll use LandingAI's Agentic Document Extraction (ADE) framework to process documents. ADE treats pages as visually -- as images -- to parse information and extract fields. Skills you'll gain: - Build agents to convert unstructured files into structured Markdown/HTML and JSON - Use ADE to parse complex data like forms, handwriting, or equations - Map extracted information to named fields using a specified schema, with bounding boxes for grounding and validation - Deploy RAG applications with event-driven document processing Come learn about the best tools for processing documents like financial invoices, medical records, or academic papers intelligently: https://t.co/PYjgnoaD2K
I am proud to announce that I have successfully completed the worldβs first USA coast to coast fully autonomous drive! I left the Tesla Diner in Los Angeles 2 days & 20 hours ago, and now have ended in Myrtle Beach, SC (2,732.4 miles) This was accomplished with Tesla FSD V14.2 with absolutely 0 disengagements of any kind even for all parking including at Tesla Superchargers.
@DanielleFong Well there was this thing once... :) I'm sure it's coming. https://t.co/fSSFmtgSwo
@sergeykarayev Hah I was just thinking about the same analogy. How I suddenly feel about all of the code I've written so far https://t.co/s5DRHFBviZ
I came across this video of Justin Bieber using a full orchestra to turn the rhythms in his mind into music. Soon, anyone will be able to hum a tune, describe the vibe youβre aiming for, and have an AI-powered producer turn your ideas into reality. https://t.co/91Sqlte3MK
New post: nanochat miniseries v1 The correct way to think about LLMs is that you are not optimizing for a single specific model but for a family models controlled by a single dial (the compute you wish to spend) to achieve monotonically better results. This allows you to do careful science of scaling laws and ultimately this is what gives you the confidence that when you pay for "the big run", the extrapolation will work and your money will be well spent. For the first public release of nanochat my focus was on end-to-end pipeline that runs the whole LLM pipeline with all of its stages. Now after YOLOing a few runs earlier, I'm coming back around to flesh out some of the parts that I sped through, starting of course with pretraining, which is both computationally heavy and critical as the foundation of intelligence and knowledge in these models. After locally tuning some of the hyperparameters, I swept out a number of models fixing the FLOPs budget. (For every FLOPs target you can train a small model a long time, or a big model for a short time.) It turns out that nanochat obeys very nice scaling laws, basically reproducing the Chinchilla paper plots: Which is just a baby version of this plot from Chinchilla: Very importantly and encouragingly, the exponent on N (parameters) and D (tokens) is equal at ~=0.5, so just like Chinchilla we get a single (compute-independent) constant that relates the model size to token training horizons. In Chinchilla, this was measured to be 20. In nanochat it seems to be 8! Once we can train compute optimal models, I swept out a miniseries from d10 to d20, which are nanochat sizes that can do 2**19 ~= 0.5M batch sizes on 8XH100 node without gradient accumulation. We get pretty, non-itersecting training plots for each model size. Then the fun part is relating this miniseries v1 to the GPT-2 and GPT-3 miniseries so that we know we're on the right track. Validation loss has many issues and is not comparable, so instead I use the CORE score (from DCLM paper). I calculated it for GPT-2 and estimated it for GPT-3, which allows us to finally put nanochat nicely and on the same scale: The total cost of this miniseries is only ~$100 (~4 hours on 8XH100). These experiments give us confidence that everything is working fairly nicely and that if we pay more (turn the dial), we get increasingly better models. TLDR: we can train compute optimal miniseries and relate them to GPT-2/3 via objective CORE scores, but further improvements are desirable and needed. E.g., matching GPT-2 currently needs ~$500, but imo should be possible to do <$100 with more work. Full post with a lot more detail is here: https://t.co/na8zVLqWLf And all of the tuning and code is pushed to master and people can reproduce these with scaling_laws .sh and miniseries .sh bash scripts.

@andrew_n_carr Yeah, $10B is the difference in finding it first and ~5 years ago. :) I just love reproducing landmark results for much cheaper, it's so fun! Reproducing LeCun 1989 was super fun too: https://t.co/oOZcQW3Y9H What runs unoptimized on a consumer laptop in 1 minute was a state of the art neural net trained for days in 1989. Another favorite example: CIFAR-10. In 2011 state of the art was 77%. I estimated human accuracy to be ~94% but said that performance might go up to 85-90%. https://t.co/KJl0V4T0ei Now you can speedrun to 94% accuracy in 1.98 seconds on a single GPU (yes, <2 seconds). https://t.co/wHUQs6htdV So e.g. right now GPT-2 (imo the landmark result that launched LLMs and where the modern stack is basically in full form) is ~$500, but I'm unreasonably obsessed with how much that can be brought down.
@patrickc This repo shows a way that works well for me: https://t.co/L3K42MU4wF Basically I use epub (not pdf), the code then parses it into text. I usually go chapter by chapter, manually copy paste the chapter text around, get a summary, do a Q&A and read alongside.
You no longer need to leave Python to write high-performance hardware kernels. Learn how to use Pallas in Keras to author custom ops that lower to Mosaic for TPUs or Triton for GPUs: https://t.co/oeV4cmV4M0
If youβre building with JAX, Keras should be your default choice. Talked about why this works so well today at JAX DevLab with @mattdangerw and Fabien. We also covered KerasHub (LLMs) and KerasRS (RecSys). Feel free to shoot me a DM if you have any questions! https://t.co/k747byAaVE
See post here: https://t.co/XiLiUI78v4
https://t.co/IAJYrXOpjP
https://t.co/IAJYrXOpjP
The Kaleidoscope Hypothesis by @fchollet https://t.co/zwlzKu91gr
The Kaleidoscope Hypothesis by @fchollet https://t.co/zwlzKu91gr
I have been pretty heads-down this year to finish Chapter 6 on implementing reinforcement learning with verifiable rewards from scratch (using GRPO). I just finished it this weekend, and I'd say it's the best (or at least my favorite) chapter yet! The goal of this chapter is to explain and implement GRPO from the bottom up. This means coding and walking through each GRPO step one by one (advantages, rewards, logprobs, and loss) and then training a 0.6B base model on the 12k examples from the MATH training set. (This takes the model from 15% to 47% accuracy on the MATH-500 test set, which is about as good as the official Qwen3 reasoning model of similar size.) The focus is on readability and understanding GRPO, but the supplementary materials also contain scripts to run it in a multi-GPU setting. The code notebook is already available on GitHub if you want to take a look: https://t.co/SM58MXjf8V. (And the full chapter should make it to the early access version of the book at https://t.co/vzCr5sTjrf soon!) PS: The next chapter will introduce additional tips and tricks to improve the GRPO algorithm for better and more stable training behavior.

@StronglyAI I tested all these below. Regarding evals: yes, I am testing against MATH-500 (non-overlapping with the 12k examples in the training set) https://t.co/zXXsxYpw3f
Btw I also ran several additional ablation studies on improving GRPO. I summarized them in my latest blog a few weeks ago. The goal here is to have the GRPO foundation (this chapter) to which we can then apply these tweaks (next chapter) https://t.co/HLRhabyW0A
@aina_oluwa Build a Reasoning Model From Scratch, https://t.co/fQndtsmUJv (sequel to Build A Large Language Model From Scratch)
using @GoogleDeepMind's Nano Banana to create a 3D version of @ivanleomk for our hack and roll project! @thorwebdev might be next hehe Can you guess what we're building!? CC: @miinnong @dongkiatt https://t.co/jH4HFtcTWe
my last hackathon as a full-time student: stickers, food and one hell of a time building at @nushackers's annual hack n roll! happy to have built something i've always wanted to try for the longest time! hearing your code! See what we built: - Strudel Your Code: "Turn your coding sessions into live generative music performances" https://t.co/yTnCC2UJFq (featuring 3d chibis of @gabrielchua @agrimsingh @ivanleomk @thorwebdev and @GeoffreyHuntley's ralph wiggum) Strudel is "a new live coding platform to write dynamic music pieces in the browser!" - We built Strudel Your Code with @cursor_ai, @openai, @AnthropicAI and more. It runs primarily on OpenAI's responses, with a single 2D Nano Banana Pro from @GeminiApp to @TencentHunyuan's 2d-3d model for our 3D models. It was insanely fun to build on @threejs, i'll write a short post on how we got the 3D models into fruition as this seemed to be the most asked question on the board. See our full video here: https://t.co/hTkzJIYKED but of course hackathons arent always about the tech! it's also about connecting to people from various walks and sharing more from each other! also big ups to the sponsors!! for sponsoring @cursor_ai @ExaAILabs @GoogleDeepMind @AnthropicAI @elevenlabsio @v0 @ManusAI @jigsawstack @OpenAI

introducing Clipmorph. an agentic layer between cmd+c and cmd+v. copy data, speak what you want (reformat table, generate chart, extract columns), paste the result No app switching. Your clipboard, now intelligent. https://t.co/cUGgCDBNql