@Scobleizer
RT @EmbodiedAIRead: 1X World Model | From Video to Action: A New Way Robots Learn Blog: https://t.co/RsW7Zl2cai 1X describes and shows in…
Viewing enriched Twitter post
RT @EmbodiedAIRead: 1X World Model | From Video to Action: A New Way Robots Learn Blog: https://t.co/RsW7Zl2cai 1X describes and shows in…
{
"media": [
{
"type": "photo",
"url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2012790326726156631/media_0.jpg?",
"filename": "media_0.jpg"
}
],
"processed_at": "2026-01-18T19:36:11.117037",
"pipeline_version": "2.0"
} {
"type": "tweet",
"id": "2012790326726156631",
"url": "https://x.com/Scobleizer/status/2012790326726156631",
"twitterUrl": "https://twitter.com/Scobleizer/status/2012790326726156631",
"text": "RT @EmbodiedAIRead: 1X World Model | From Video to Action: A New Way Robots Learn\n\nBlog: https://t.co/RsW7Zl2cai\n\n1X describes and shows in…",
"source": "Twitter for iPhone",
"retweetCount": 28,
"replyCount": 3,
"likeCount": 226,
"quoteCount": 5,
"viewCount": 11495,
"createdAt": "Sun Jan 18 07:33:04 +0000 2026",
"lang": "en",
"bookmarkCount": 154,
"isReply": false,
"inReplyToId": null,
"conversationId": "2012790326726156631",
"displayTextRange": [
0,
140
],
"inReplyToUserId": null,
"inReplyToUsername": null,
"author": {
"type": "user",
"userName": "Scobleizer",
"url": "https://x.com/Scobleizer",
"twitterUrl": "https://twitter.com/Scobleizer",
"id": "13348",
"name": "Robert Scoble",
"isVerified": false,
"isBlueVerified": true,
"verifiedType": null,
"profilePicture": "https://pbs.twimg.com/profile_images/1915614118876504066/zVnfpAMf_normal.jpg",
"coverPicture": "https://pbs.twimg.com/profile_banners/13348/1765416783",
"description": "",
"location": "My Free Newsletter 👉",
"followers": 554066,
"following": 35595,
"status": "",
"canDm": true,
"canMediaTag": false,
"createdAt": "Mon Nov 20 23:43:44 +0000 2006",
"entities": {
"description": {
"urls": []
},
"url": {}
},
"fastFollowersCount": 0,
"favouritesCount": 560395,
"hasCustomTimelines": true,
"isTranslator": false,
"mediaCount": 6124,
"statusesCount": 235033,
"withheldInCountries": [],
"affiliatesHighlightedLabel": {},
"possiblySensitive": false,
"pinnedTweetIds": [
"2012934334546862425"
],
"profile_bio": {
"description": "San Francisco/Silicon Valley AI | Robots, holodecks, BCIs, analysis of new things | Ex-Microsoft, Rackspace, Fast Company | Wrote eight books about the future.",
"entities": {
"description": {},
"url": {
"urls": [
{
"display_url": "unaligned.io/subscribe",
"expanded_url": "http://unaligned.io/subscribe",
"indices": [
0,
23
],
"url": "https://t.co/YVCwrHlZ0x"
}
]
}
}
},
"isAutomated": false,
"automatedBy": null
},
"extendedEntities": {},
"card": null,
"place": {},
"entities": {
"urls": [
{
"display_url": "1x.tech/discover/world…",
"expanded_url": "https://www.1x.tech/discover/world-model-self-learning",
"indices": [
89,
112
],
"url": "https://t.co/RsW7Zl2cai"
}
],
"user_mentions": [
{
"id_str": "1941983406176481280",
"indices": [
3,
18
],
"name": "Embodied AI Reading Notes",
"screen_name": "EmbodiedAIRead"
}
]
},
"quoted_tweet": null,
"retweeted_tweet": {
"type": "tweet",
"id": "2012770636855418983",
"url": "https://x.com/EmbodiedAIRead/status/2012770636855418983",
"twitterUrl": "https://twitter.com/EmbodiedAIRead/status/2012770636855418983",
"text": "1X World Model | From Video to Action: A New Way Robots Learn\n\nBlog: https://t.co/1sPpUJBcrF\n\n1X describes and shows initial results for a new potential way of learning robot policy using video generation based world modeling, compared to VLA which is based on VLM.\n\n- How it works: at inference time, the system receives a text prompt and a starting frame. The World Model rolls out the intended future image frames, the Inverse Dynamics Model extracts the trajectory, and the robot executes the sequence in the real world.\n\n- The World Model backbone: A text-conditioned diffusion model trained on web-scale video, mid-trained on 900 hours of egocentric human data of first-person manipulation tasks for capturing general manipulation behaviors, and fine-tuned on 70 hours of NEO-specific sensorimotor logs for adapting to NEO’s visual appearance and kinematics.\n\n- The Inverse Dynamics Model: similar to architecure used in DreamGen, and trained on 400 hours of robot data on random play and motions.\n\n- Results: The model can generate videos aligning well with real-world execution, and the robot can perform object grasping, manipulation with some degree of generalization.\n\n- Current limitations: The pipeline latency is high and it’s not lose-loop. Currently the WM takes 11 second to generate 5 second video on a multi-GPU server and IDM takes another 1 second to extract actions.",
"source": "Twitter for iPhone",
"retweetCount": 28,
"replyCount": 3,
"likeCount": 226,
"quoteCount": 5,
"viewCount": 11495,
"createdAt": "Sun Jan 18 06:14:49 +0000 2026",
"lang": "en",
"bookmarkCount": 154,
"isReply": false,
"inReplyToId": null,
"conversationId": "2012770636855418983",
"displayTextRange": [
0,
300
],
"inReplyToUserId": null,
"inReplyToUsername": null,
"author": {
"type": "user",
"userName": "EmbodiedAIRead",
"url": "https://x.com/EmbodiedAIRead",
"twitterUrl": "https://twitter.com/EmbodiedAIRead",
"id": "1941983406176481280",
"name": "Embodied AI Reading Notes",
"isVerified": false,
"isBlueVerified": true,
"verifiedType": null,
"profilePicture": "https://pbs.twimg.com/profile_images/1941985661059239936/yPCcVMKL_normal.jpg",
"coverPicture": "",
"description": "",
"location": "California, USA",
"followers": 2813,
"following": 1,
"status": "",
"canDm": false,
"canMediaTag": true,
"createdAt": "Sun Jul 06 22:12:07 +0000 2025",
"entities": {
"description": {
"urls": []
},
"url": {}
},
"fastFollowersCount": 0,
"favouritesCount": 19,
"hasCustomTimelines": true,
"isTranslator": false,
"mediaCount": 122,
"statusesCount": 138,
"withheldInCountries": [],
"affiliatesHighlightedLabel": {},
"possiblySensitive": false,
"pinnedTweetIds": [],
"profile_bio": {
"description": "Sharing daily personal notes on selected interesting Embodied AI papers, blogs and talks | Maintained by @yilun_chen_ | Opinions are my own.",
"entities": {
"description": {
"user_mentions": [
{
"id_str": "0",
"indices": [
105,
117
],
"name": "",
"screen_name": "yilun_chen_"
}
]
}
}
},
"isAutomated": false,
"automatedBy": null
},
"extendedEntities": {
"media": [
{
"allow_download_status": {
"allow_download": true
},
"display_url": "pic.twitter.com/VPokzisOnY",
"expanded_url": "https://twitter.com/EmbodiedAIRead/status/2012770636855418983/photo/1",
"ext_media_availability": {
"status": "Available"
},
"features": {
"large": {},
"orig": {}
},
"id_str": "2012770545582972928",
"indices": [
301,
324
],
"media_key": "3_2012770545582972928",
"media_results": {
"id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARvuzCWn16AACgACG+7MOugaoGcAAA==",
"result": {
"__typename": "ApiMedia",
"id": "QXBpTWVkaWE6DAABCgABG+7MJafXoAAKAAIb7sw66BqgZwAA",
"media_key": "3_2012770545582972928"
}
},
"media_url_https": "https://pbs.twimg.com/media/G-7MJafXoAABNib.jpg",
"original_info": {
"focus_rects": [
{
"h": 1434,
"w": 2560,
"x": 0,
"y": 0
},
{
"h": 1488,
"w": 1488,
"x": 0,
"y": 0
},
{
"h": 1488,
"w": 1305,
"x": 0,
"y": 0
},
{
"h": 1488,
"w": 744,
"x": 76,
"y": 0
},
{
"h": 1488,
"w": 2560,
"x": 0,
"y": 0
}
],
"height": 1488,
"width": 2560
},
"sizes": {
"large": {
"h": 1190,
"w": 2048
}
},
"type": "photo",
"url": "https://t.co/VPokzisOnY"
}
]
},
"card": null,
"place": {},
"entities": {
"urls": [
{
"display_url": "1x.tech/discover/world…",
"expanded_url": "https://www.1x.tech/discover/world-model-self-learning",
"indices": [
69,
92
],
"url": "https://t.co/1sPpUJBcrF"
}
]
},
"quoted_tweet": null,
"retweeted_tweet": null,
"isLimitedReply": false,
"article": null
},
"isLimitedReply": false,
"article": null
}