🐦 Twitter Post Details

Viewing enriched Twitter post

@omarsar0

New research from Meta and collaborators. This is a good paper showing what's possible with proper world models. World models need actions to predict consequences. The default approach today requires labeled action data, which is expensive to obtain and limited to narrow domains like video games or robotic manipulation. But the vast majority of video data online has no action labels at all. This new research tackles learning latent action world models directly from in-the-wild videos, expanding beyond the controlled settings of previous work to capture the full diversity of real-world actions. The challenge is significant. In-the-wild videos contain actions far beyond simple navigation or manipulation: people entering frames, objects appearing and disappearing, dancers moving, fingers forming guitar chords. There's also no consistent embodiment across videos, unlike robotics datasets, where the same arm appears throughout. So how do the authors address this? Continuous but constrained latent actions, using sparse or noisy regularization, effectively capture this action complexity. Discrete quantization, the common approach in prior work, struggles to adapt. Without a shared embodiment, the model learns spatially-localized, camera-relative transformations. The results demonstrate genuine action transfer. Motion from a walking person can be applied to a flying ball. Actions like "someone entering the frame" transfer across completely different videos. By training a small controller to map known actions to latent ones, the world model trained purely on natural videos can solve robotic manipulation and navigation tasks with performance close to models trained on domain-specific, action-labeled data. Latent action spaces learned from unlabeled internet videos can serve as a universal interface for planning, removing the bottleneck of action annotation. Paper: https://t.co/BL6mpuLZGD Learn to build effective AI agents in our academy: https://t.co/JBU5beHQNs

Media 1

📊 Media Metadata

{
  "media": [
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2011076284164542639/media_0.jpg?",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2011076284164542639/media_0.jpg?",
      "type": "photo",
      "filename": "media_0.jpg"
    }
  ],
  "processed_at": "2026-01-18T17:31:54.949758",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2011076284164542639",
  "url": "https://x.com/omarsar0/status/2011076284164542639",
  "twitterUrl": "https://twitter.com/omarsar0/status/2011076284164542639",
  "text": "New research from Meta and collaborators.\n\nThis is a good paper showing what's possible with proper world models.\n\nWorld models need actions to predict consequences. The default approach today requires labeled action data, which is expensive to obtain and limited to narrow domains like video games or robotic manipulation.\n\nBut the vast majority of video data online has no action labels at all.\n\nThis new research tackles learning latent action world models directly from in-the-wild videos, expanding beyond the controlled settings of previous work to capture the full diversity of real-world actions.\n\nThe challenge is significant. In-the-wild videos contain actions far beyond simple navigation or manipulation: people entering frames, objects appearing and disappearing, dancers moving, fingers forming guitar chords. There's also no consistent embodiment across videos, unlike robotics datasets, where the same arm appears throughout.\n\nSo how do the authors address this?\n\nContinuous but constrained latent actions, using sparse or noisy regularization, effectively capture this action complexity. Discrete quantization, the common approach in prior work, struggles to adapt. Without a shared embodiment, the model learns spatially-localized, camera-relative transformations.\n\nThe results demonstrate genuine action transfer.\n\nMotion from a walking person can be applied to a flying ball. Actions like \"someone entering the frame\" transfer across completely different videos.\n\nBy training a small controller to map known actions to latent ones, the world model trained purely on natural videos can solve robotic manipulation and navigation tasks with performance close to models trained on domain-specific, action-labeled data.\n\nLatent action spaces learned from unlabeled internet videos can serve as a universal interface for planning, removing the bottleneck of action annotation.\n\nPaper: https://t.co/BL6mpuLZGD\n\nLearn to build effective AI agents in our academy: https://t.co/JBU5beHQNs",
  "source": "Twitter for iPhone",
  "retweetCount": 112,
  "replyCount": 23,
  "likeCount": 609,
  "quoteCount": 3,
  "viewCount": 70042,
  "createdAt": "Tue Jan 13 14:02:04 +0000 2026",
  "lang": "en",
  "bookmarkCount": 494,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2011076284164542639",
  "displayTextRange": [
    0,
    298
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "omarsar0",
    "url": "https://x.com/omarsar0",
    "twitterUrl": "https://twitter.com/omarsar0",
    "id": "3448284313",
    "name": "elvis",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/939313677647282181/vZjFWtAn_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/3448284313/1565974901",
    "description": "",
    "location": "DAIR.AI Academy",
    "followers": 285337,
    "following": 758,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Fri Sep 04 12:59:26 +0000 2015",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 34412,
    "hasCustomTimelines": true,
    "isTranslator": true,
    "mediaCount": 4448,
    "statusesCount": 17028,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2012877899360297408"
    ],
    "profile_bio": {
      "description": "Building @dair_ai • Prev: Meta AI, Elastic, PhD • New cohort: https://t.co/AqEvlVuYdM",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "dair-ai.thinkific.com/courses/claude…",
              "expanded_url": "https://dair-ai.thinkific.com/courses/claude-code-for-everyone-cohort-3",
              "indices": [
                62,
                85
              ],
              "url": "https://t.co/AqEvlVuYdM"
            }
          ],
          "user_mentions": [
            {
              "id_str": "0",
              "indices": [
                9,
                17
              ],
              "name": "",
              "screen_name": "dair_ai"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "dair.ai",
              "expanded_url": "https://www.dair.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/XQto5ypkSM"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/MWDgCscdCC",
        "expanded_url": "https://twitter.com/omarsar0/status/2011076284164542639/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {},
          "orig": {}
        },
        "id_str": "2011076278653247488",
        "indices": [
          299,
          322
        ],
        "media_key": "3_2011076278653247488",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARvoxzhlV2AACgACG+jHOa3XEK8AAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABG+jHOGVXYAAKAAIb6Mc5rdcQrwAA",
            "media_key": "3_2011076278653247488"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/G-jHOGVXYAAXdg8.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 904,
              "w": 1614,
              "x": 0,
              "y": 0
            },
            {
              "h": 1614,
              "w": 1614,
              "x": 0,
              "y": 0
            },
            {
              "h": 1802,
              "w": 1581,
              "x": 33,
              "y": 0
            },
            {
              "h": 1802,
              "w": 901,
              "x": 404,
              "y": 0
            },
            {
              "h": 1802,
              "w": 1614,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1802,
          "width": 1614
        },
        "sizes": {
          "large": {
            "h": 1802,
            "w": 1614
          }
        },
        "type": "photo",
        "url": "https://t.co/MWDgCscdCC"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "urls": [
      {
        "display_url": "arxiv.org/abs/2601.05230",
        "expanded_url": "https://arxiv.org/abs/2601.05230",
        "indices": [
          1899,
          1922
        ],
        "url": "https://t.co/BL6mpuLZGD"
      },
      {
        "display_url": "dair-ai.thinkific.com",
        "expanded_url": "https://dair-ai.thinkific.com/",
        "indices": [
          1975,
          1998
        ],
        "url": "https://t.co/JBU5beHQNs"
      }
    ]
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "article": null
}