🐦 Twitter Post Details

Viewing enriched Twitter post

@XiaojieJin

Excited to share our new survey paper: the first comprehensive survey on Vision World Model (VWM), a joint effort by researchers from BJTU, ByteDance, Tencent, NUS, and more. 🌟 From Seeing to Knowing the World: A Survey of Vision World Models πŸš€ Our core message is a paradigm shift toward vision-centric world modeling: Vision should not be treated merely as an input modality. It should be the primary driver of how world models are represented, learned, and evaluated. 🌈 This is also the longstanding view behind our #VideoWorld series: learning directly from visual observation and interaction offers a scalable path for AI agents to acquire world knowledge, laying the foundation for higher machine intelligence. πŸ€” Why Vision World Models? From biological evolution to human intelligence, vision has been central to learning about the world through observation and interaction. AI should have this capability too. This motivates Vision World Models: models that learn world knowledge from visual data and simulate future world states conditioned on interaction. πŸ€– In this survey, we thoroughly review 400+ recent papers and provide a vision-centric roadmap for Vision World Models, covering architectures, functional roles, applications, evaluation protocols, datasets, benchmarks, and future outlook. Key takeaways: 1️⃣ Vision is a fundamental basis of intelligence and a rich source of world knowledge. We advocate vision-centric world modeling, where AI learns the physical and causal principles behind world evolution from visual data. 2️⃣ We propose a unified framework that decomposes Vision World Models into three core components: Vision Encoding β†’ Knowledge Learning β†’ Controllable Simulation and organize current methods into 4 major families and 7 representative architectures. 3️⃣ We review evaluation from three levels: Visual Quality, Physical Plausibility, and Task Performance, and group datasets/benchmarks into foundational world modeling and domain-specific world modeling. 4️⃣ We outline three directions for next-generation world models: Re-grounding in physical and causal knowledge, Re-evaluating beyond visual appearance, and Re-scaling toward generalist, reliable, and interaction-aware world models. Check out our paper and the continuously updated curated list of Vision World Model papers for more details! πŸ“„ Paper: https://t.co/Yq8hdSwJAl 🌐 Project Page: https://t.co/SJoapeLVUL πŸ“š Curated VWM Paper List: https://t.co/dBufH4pAVV #VisionWorldModel #WorldModel #Survey #VideoWorld #EmbodiedAI #Robotics #AI #CV

Media 1

πŸ“Š Media Metadata

{
  "media": [
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2055312702889939181/media_0.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2055312702889939181/media_0.jpg",
      "type": "photo",
      "filename": "media_0.jpg"
    }
  ],
  "processed_at": "2026-05-15T16:03:41.676781",
  "pipeline_version": "2.0"
}

πŸ”§ Raw API Response

{
  "type": "tweet",
  "id": "2055312702889939181",
  "url": "https://x.com/XiaojieJin/status/2055312702889939181",
  "twitterUrl": "https://twitter.com/XiaojieJin/status/2055312702889939181",
  "text": "Excited to share our new survey paper: the first comprehensive survey on Vision World Model (VWM), a joint effort by researchers from BJTU, ByteDance, Tencent, NUS, and more.  \n\n🌟 From Seeing to Knowing the World: A Survey of Vision World Models \n\nπŸš€ Our core message is a paradigm shift toward vision-centric world modeling:  \n\nVision should not be treated merely as an input modality. It should be the primary driver of how world models are represented, learned, and evaluated.  \n\n🌈 This is also the longstanding view behind our #VideoWorld series: learning directly from visual observation and interaction offers a scalable path for AI agents to acquire world knowledge, laying the foundation for higher machine intelligence.\n\nπŸ€” Why Vision World Models?\n       From biological evolution to human intelligence, vision has been central to learning about the world through observation and interaction. AI should have this capability too. This motivates Vision World Models: models that learn world knowledge from visual data and simulate future world states conditioned on interaction.\n\nπŸ€– In this survey, we thoroughly review 400+ recent papers and provide a vision-centric roadmap for Vision World Models, covering architectures, functional roles, applications, evaluation protocols, datasets, benchmarks, and future outlook.\n\nKey takeaways:\n1️⃣ Vision is a fundamental basis of intelligence and a rich source of world knowledge. We advocate vision-centric world modeling, where AI learns the physical and causal principles behind world evolution from visual data.\n\n2️⃣ We propose a unified framework that decomposes Vision World Models into three core components: Vision Encoding β†’ Knowledge Learning β†’ Controllable Simulation and organize current methods into 4 major families and 7 representative architectures.\n\n3️⃣ We review evaluation from three levels: Visual Quality, Physical Plausibility, and Task Performance, and group datasets/benchmarks into foundational world modeling and domain-specific world modeling.\n\n4️⃣ We outline three directions for next-generation world models: Re-grounding in physical and causal knowledge, Re-evaluating beyond visual appearance, and Re-scaling toward generalist, reliable, and interaction-aware world models.\n\nCheck out our paper and the continuously updated curated list of Vision World Model papers for more details!  \nπŸ“„ Paper: https://t.co/Yq8hdSwJAl \n🌐 Project Page: https://t.co/SJoapeLVUL \nπŸ“š Curated VWM Paper List: https://t.co/dBufH4pAVV\n\n#VisionWorldModel #WorldModel #Survey #VideoWorld #EmbodiedAI #Robotics #AI #CV",
  "source": "Twitter for iPhone",
  "retweetCount": 1,
  "replyCount": 2,
  "likeCount": 5,
  "quoteCount": 1,
  "viewCount": 1140,
  "createdAt": "Fri May 15 15:41:48 +0000 2026",
  "lang": "en",
  "bookmarkCount": 2,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2055312702889939181",
  "displayTextRange": [
    0,
    271
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "XiaojieJin",
    "url": "https://x.com/XiaojieJin",
    "twitterUrl": "https://twitter.com/XiaojieJin",
    "id": "795190899315609600",
    "name": "Xiaojie Jin",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1882107294739259392/IbVn_Eje_normal.jpg",
    "coverPicture": "",
    "description": "",
    "location": "USA",
    "followers": 387,
    "following": 141,
    "status": "",
    "canDm": false,
    "canMediaTag": true,
    "createdAt": "Sun Nov 06 09:07:39 +0000 2016",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 50,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 16,
    "statusesCount": 54,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2021469936363991162"
    ],
    "profile_bio": {
      "description": "Researcher in #artificialintelligence. Founding member, research scientist @Bytedance Research, US | previously @NUSingapore",
      "entities": {
        "description": {
          "hashtags": [
            {
              "indices": [
                14,
                37
              ],
              "text": "artificialintelligence"
            }
          ],
          "user_mentions": [
            {
              "id_str": "",
              "indices": [
                75,
                85
              ],
              "name": "",
              "screen_name": "Bytedance"
            },
            {
              "id_str": "",
              "indices": [
                112,
                124
              ],
              "name": "",
              "screen_name": "NUSingapore"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/oAJZLjJMt7",
        "expanded_url": "https://twitter.com/XiaojieJin/status/2055312702889939181/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {
            "faces": []
          },
          "orig": {
            "faces": []
          }
        },
        "id_str": "2055309700502257665",
        "indices": [
          272,
          295
        ],
        "media_key": "3_2055309700502257665",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARyF7Uh52gABCgACHIXwA4YaUO0AAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABHIXtSHnaAAEKAAIchfADhhpQ7QAA",
            "media_key": "3_2055309700502257665"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HIXtSHnaAAEKZMP.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 831,
              "w": 1484,
              "x": 0,
              "y": 0
            },
            {
              "h": 1280,
              "w": 1280,
              "x": 102,
              "y": 0
            },
            {
              "h": 1280,
              "w": 1123,
              "x": 181,
              "y": 0
            },
            {
              "h": 1280,
              "w": 640,
              "x": 422,
              "y": 0
            },
            {
              "h": 1280,
              "w": 1484,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1280,
          "width": 1484
        },
        "sizes": {
          "large": {
            "h": 1280,
            "w": 1484
          }
        },
        "type": "photo",
        "url": "https://t.co/oAJZLjJMt7"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [
      {
        "indices": [
          530,
          541
        ],
        "text": "VideoWorld"
      },
      {
        "indices": [
          2492,
          2509
        ],
        "text": "VisionWorldModel"
      },
      {
        "indices": [
          2510,
          2521
        ],
        "text": "WorldModel"
      },
      {
        "indices": [
          2522,
          2529
        ],
        "text": "Survey"
      },
      {
        "indices": [
          2530,
          2541
        ],
        "text": "VideoWorld"
      },
      {
        "indices": [
          2542,
          2553
        ],
        "text": "EmbodiedAI"
      },
      {
        "indices": [
          2554,
          2563
        ],
        "text": "Robotics"
      },
      {
        "indices": [
          2564,
          2567
        ],
        "text": "AI"
      },
      {
        "indices": [
          2568,
          2571
        ],
        "text": "CV"
      }
    ],
    "symbols": [],
    "urls": [
      {
        "display_url": "aiworldlab.github.io/survey/preprin…",
        "expanded_url": "https://aiworldlab.github.io/survey/preprint.pdf",
        "indices": [
          2375,
          2398
        ],
        "url": "https://t.co/Yq8hdSwJAl"
      },
      {
        "display_url": "aiworldlab.github.io/survey/",
        "expanded_url": "https://aiworldlab.github.io/survey/",
        "indices": [
          2416,
          2439
        ],
        "url": "https://t.co/SJoapeLVUL"
      },
      {
        "display_url": "github.com/AIWorldLab/Awe",
        "expanded_url": "https://github.com/AIWorldLab/Awe",
        "indices": [
          2467,
          2490
        ],
        "url": "https://t.co/dBufH4pAVV"
      }
    ],
    "user_mentions": []
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}