🐦 Twitter Post Details

Viewing enriched Twitter post

@peteflorence

Ever since I started working on robot foundation models a handful of years ago, the broad ability to one-shot in-context learn has been the single most vivid goal in my mind along the long road ahead. It’s something @andyzengineer and I have especially been thinking together on for several years now, and something shared by the whole team at Generalist as a major goal. But even before anybody started talking about robot foundation models, too, the capability this enables, i.e. generalized one-shot learning, has been both incredibly concrete and elusive. In terminology that I think many other people have come up with as well, in grad school we used to talk about a “Ctrl-C-Ctrl-V” type of capability. See something, and have the robot do it too. This same type of idea inspired the name of the 1970 MIT copy demo (https://t.co/qs08JRYJGl). The thing is, it sounds simple, but is incredibly hard since the world is never quite the same when it got copied and where you want to paste it. The real world can be hard to predict and is full of variation. To do this, you need strong generalization, and it needs to be acquired in a single example. Doing this over a wide range of tasks, especially for dexterous tasks, is hard mode. Lots of the components of the idea of making this all happen have been there for a long time. As an example, this 2017 NeurIPS paper “one-shot imitation learning” https://t.co/IWr9KWeeGi has excellent vision, with ambition well beyond what was achievable at the time, and although it’s not referred to as “in-context learning” since it was pre-Transformer, it actually uses attention to condition on a single demonstration. And now, many things have happened since early 2017, including the broadly celebrated arrival of one/few-shot in-context learning in language models in 2020. This new model GEN-1.5 takes in everything we have built and learned over the past couple years at Generalist. It has been training for 8 months. It has taken an incredible amount of commitment and grit from the whole team to get here. The level to which this model has survived many surgeries has continued to surprise me. And its capabilities have continued to surprise as well. We found compositional generalization on Friday. We found sim2real prompting earlier last week. We filmed the contiguous uncut videos of live prompting yesterday. To be clear, the success rates are modest, and there’s still a long way to go. But now I have definitely seen a ~decade-long imagination come into the real world.

📊 Media Metadata

{
  "score": 0.46,
  "score_components": {
    "author": 0.09,
    "engagement": 0.0,
    "quality": 0.16000000000000003,
    "source": 0.135,
    "nlp": 0.05,
    "recency": 0.025
  },
  "scored_at": "2026-08-19T21:02:39.630770",
  "import_source": "api_import",
  "source_tagged_at": "2026-08-19T21:02:39.630779",
  "enriched": true,
  "enriched_at": "2026-08-19T21:02:39.630781"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2090165901509411299",
  "url": "https://x.com/peteflorence/status/2090165901509411299",
  "twitterUrl": "https://twitter.com/peteflorence/status/2090165901509411299",
  "text": "Ever since I started working on robot foundation models a handful of years ago, the broad ability to one-shot in-context learn has been the single most vivid goal in my mind along the long road ahead. It’s something @andyzengineer  and I have especially been thinking together on for several years now, and something shared by the whole team at Generalist as a major goal.\n\nBut even before anybody started talking about robot foundation models, too, the capability this enables, i.e. generalized one-shot learning, has been both incredibly concrete and elusive. In terminology that I think many other people have come up with as well, in grad school we used to talk about a “Ctrl-C-Ctrl-V” type of capability. See something, and have the robot do it too. This same type of idea inspired the name of the 1970 MIT copy demo (https://t.co/qs08JRYJGl).\n\nThe thing is, it sounds simple, but is incredibly hard since the world is never quite the same when it got copied and where you want to paste it. The real world can be hard to predict and is full of variation. To do this, you need strong generalization, and it needs to be acquired in a single example. Doing this over a wide range of tasks, especially for dexterous tasks, is hard mode.\n\nLots of the components of the idea of making this all happen have been there for a long time. As an example, this 2017 NeurIPS paper “one-shot imitation learning” https://t.co/IWr9KWeeGi has excellent vision, with ambition well beyond what was achievable at the time, and although it’s not referred to as “in-context learning” since it was pre-Transformer, it actually uses attention to condition on a single demonstration.\n\nAnd now, many things have happened since early 2017, including the broadly celebrated arrival of one/few-shot in-context learning in language models in 2020.\n\nThis new model GEN-1.5 takes in everything we have built and learned over the past couple years at Generalist. It has been training for 8 months. It has taken an incredible amount of commitment and grit from the whole team to get here. The level to which this model has survived many surgeries has continued to surprise me. And its capabilities have continued to surprise as well. We found compositional generalization on Friday. We found sim2real prompting earlier last week. We filmed the contiguous uncut videos of live prompting yesterday.\n\nTo be clear, the success rates are modest, and there’s still a long way to go. But now I have definitely seen a ~decade-long imagination come into the real world.",
  "source": "Twitter for iPhone",
  "retweetCount": 9,
  "replyCount": 5,
  "likeCount": 63,
  "quoteCount": 2,
  "viewCount": 4436,
  "createdAt": "Wed Aug 19 19:55:58 +0000 2026",
  "lang": "en",
  "bookmarkCount": 15,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2090165901509411299",
  "displayTextRange": [
    0,
    279
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "peteflorence",
    "url": "https://x.com/peteflorence",
    "twitterUrl": "https://twitter.com/peteflorence",
    "id": "577537524",
    "name": "Pete Florence",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1789006541699567616/lnV0BkPh_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/577537524/1705784496",
    "description": "",
    "location": "San Francisco, CA",
    "followers": 7618,
    "following": 381,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Fri May 11 20:59:26 +0000 2012",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 1898,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 98,
    "statusesCount": 695,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2041529286562402804"
    ],
    "profile_bio": {
      "description": "Co-Founder & CEO @GeneralistAI",
      "entities": {
        "description": {
          "user_mentions": [
            {
              "id_str": "",
              "indices": [
                17,
                30
              ],
              "name": "",
              "screen_name": "GeneralistAI"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "peteflorence.com",
              "expanded_url": "http://peteflorence.com",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/lucjHdxLRN"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {},
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "people.csail.mit.edu/bkph/phw_copy_…",
        "expanded_url": "https://people.csail.mit.edu/bkph/phw_copy_demo.shtml",
        "indices": [
          823,
          846
        ],
        "url": "https://t.co/qs08JRYJGl"
      },
      {
        "display_url": "arxiv.org/pdf/1703.07326",
        "expanded_url": "https://arxiv.org/pdf/1703.07326",
        "indices": [
          1402,
          1425
        ],
        "url": "https://t.co/IWr9KWeeGi"
      }
    ],
    "user_mentions": [
      {
        "id_str": "907346058207936512",
        "indices": [
          216,
          230
        ],
        "name": "Andy Zeng",
        "screen_name": "andyzengineer"
      }
    ]
  },
  "quoted_tweet": {
    "type": "tweet",
    "id": "2090161945307664621",
    "url": "",
    "twitterUrl": "",
    "text": "",
    "source": "Twitter for iPhone",
    "retweetCount": 0,
    "replyCount": 0,
    "likeCount": 0,
    "quoteCount": 0,
    "viewCount": 0,
    "createdAt": "",
    "lang": "",
    "bookmarkCount": 0,
    "isReply": false,
    "inReplyToId": null,
    "conversationId": "",
    "displayTextRange": [],
    "inReplyToUserId": null,
    "inReplyToUsername": null,
    "author": {},
    "extendedEntities": {},
    "card": null,
    "place": {},
    "entities": {},
    "quoted_tweet": null,
    "retweeted_tweet": null,
    "isLimitedReply": false,
    "communityInfo": null,
    "article": null
  },
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}