🐦 Twitter Post Details

Viewing enriched Twitter post

@omarsar0

UniversalRAG RAG systems retrieve knowledge to ground model responses. However, most existing approaches are limited to a single modality, typically text. For many real-world RAG systems, some queries need images. Others need videos. Many need combinations. This new research introduces UniversalRAG, a framework that retrieves and integrates knowledge from heterogeneous sources across diverse modalities and granularities. Real-world queries vary widely in what knowledge they need. A universal RAG framework that dynamically routes to the right modality and granularity serves diverse information needs that no single-corpus approach can address. Instead of forcing everything into one embedding space, UniversalRAG uses modality-aware routing. A router dynamically predicts which modality-specific corpus best matches the query, then performs targeted retrieval within it. This sidesteps the modality gap entirely by avoiding cross-modal comparisons. Beyond modality, the framework also handles granularity. Complex analytical questions may need full documents or complete videos. Simple factoid questions are better served with paragraphs or short clips. UniversalRAG organizes each modality into multiple granularity levels: paragraphs and documents for text, clips and full videos for video, plus tables and images. The router can be trained or training-free. The trained version uses inductive biases from existing benchmarks. The training-free version prompts frontier models like Gemini to predict the best modality-granularity pairs directly. Validation across 10 benchmarks spanning text, images, tables, and videos shows UniversalRAG outperforms both unimodal RAG baselines and unified embedding approaches by large margins on average. Paper: https://t.co/OAfR65bEm2 Learn to build effective Agentic RAG systems in our academy: https://t.co/JBU5beIoD0

Media 1

📊 Media Metadata

{
  "media": [
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2011442693134754243/media_0.jpg?",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2011442693134754243/media_0.jpg?",
      "type": "photo",
      "filename": "media_0.jpg"
    }
  ],
  "processed_at": "2026-01-18T17:31:31.650055",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2011442693134754243",
  "url": "https://x.com/omarsar0/status/2011442693134754243",
  "twitterUrl": "https://twitter.com/omarsar0/status/2011442693134754243",
  "text": "UniversalRAG\n\nRAG systems retrieve knowledge to ground model responses.\n\nHowever, most existing approaches are limited to a single modality, typically text.\n\nFor many real-world RAG systems, some queries need images. Others need videos. Many need combinations.\n\nThis new research introduces UniversalRAG, a framework that retrieves and integrates knowledge from heterogeneous sources across diverse modalities and granularities.\n\nReal-world queries vary widely in what knowledge they need. A universal RAG framework that dynamically routes to the right modality and granularity serves diverse information needs that no single-corpus approach can address.\n\nInstead of forcing everything into one embedding space, UniversalRAG uses modality-aware routing. A router dynamically predicts which modality-specific corpus best matches the query, then performs targeted retrieval within it. This sidesteps the modality gap entirely by avoiding cross-modal comparisons.\n\nBeyond modality, the framework also handles granularity. Complex analytical questions may need full documents or complete videos. Simple factoid questions are better served with paragraphs or short clips. UniversalRAG organizes each modality into multiple granularity levels: paragraphs and documents for text, clips and full videos for video, plus tables and images.\n\nThe router can be trained or training-free. The trained version uses inductive biases from existing benchmarks. The training-free version prompts frontier models like Gemini to predict the best modality-granularity pairs directly.\n\nValidation across 10 benchmarks spanning text, images, tables, and videos shows UniversalRAG outperforms both unimodal RAG baselines and unified embedding approaches by large margins on average.\n\nPaper: https://t.co/OAfR65bEm2\n\nLearn to build effective Agentic RAG systems in our academy: https://t.co/JBU5beIoD0",
  "source": "Twitter for iPhone",
  "retweetCount": 92,
  "replyCount": 26,
  "likeCount": 438,
  "quoteCount": 2,
  "viewCount": 34721,
  "createdAt": "Wed Jan 14 14:18:03 +0000 2026",
  "lang": "en",
  "bookmarkCount": 375,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2011442693134754243",
  "displayTextRange": [
    0,
    304
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "omarsar0",
    "url": "https://x.com/omarsar0",
    "twitterUrl": "https://twitter.com/omarsar0",
    "id": "3448284313",
    "name": "elvis",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/939313677647282181/vZjFWtAn_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/3448284313/1565974901",
    "description": "",
    "location": "DAIR.AI Academy",
    "followers": 285337,
    "following": 758,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Fri Sep 04 12:59:26 +0000 2015",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 34412,
    "hasCustomTimelines": true,
    "isTranslator": true,
    "mediaCount": 4448,
    "statusesCount": 17028,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2012877899360297408"
    ],
    "profile_bio": {
      "description": "Building @dair_ai • Prev: Meta AI, Elastic, PhD • New cohort: https://t.co/AqEvlVuYdM",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "dair-ai.thinkific.com/courses/claude…",
              "expanded_url": "https://dair-ai.thinkific.com/courses/claude-code-for-everyone-cohort-3",
              "indices": [
                62,
                85
              ],
              "url": "https://t.co/AqEvlVuYdM"
            }
          ],
          "user_mentions": [
            {
              "id_str": "0",
              "indices": [
                9,
                17
              ],
              "name": "",
              "screen_name": "dair_ai"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "dair.ai",
              "expanded_url": "https://www.dair.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/XQto5ypkSM"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/yDVvAjT3SI",
        "expanded_url": "https://twitter.com/omarsar0/status/2011442693134754243/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {},
          "orig": {}
        },
        "id_str": "2011442689389297664",
        "indices": [
          305,
          328
        ],
        "media_key": "3_2011442689389297664",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARvqFHgLGwAACgACG+oUeOpaIcMAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABG+oUeAsbAAAKAAIb6hR46lohwwAA",
            "media_key": "3_2011442689389297664"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/G-oUeAsbAAA1RAk.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 787,
              "w": 1406,
              "x": 0,
              "y": 0
            },
            {
              "h": 1406,
              "w": 1406,
              "x": 0,
              "y": 0
            },
            {
              "h": 1603,
              "w": 1406,
              "x": 0,
              "y": 0
            },
            {
              "h": 1798,
              "w": 899,
              "x": 0,
              "y": 0
            },
            {
              "h": 1798,
              "w": 1406,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1798,
          "width": 1406
        },
        "sizes": {
          "large": {
            "h": 1798,
            "w": 1406
          }
        },
        "type": "photo",
        "url": "https://t.co/yDVvAjT3SI"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "urls": [
      {
        "display_url": "arxiv.org/abs/2504.20734",
        "expanded_url": "https://arxiv.org/abs/2504.20734",
        "indices": [
          1766,
          1789
        ],
        "url": "https://t.co/OAfR65bEm2"
      },
      {
        "display_url": "dair-ai.thinkific.com",
        "expanded_url": "https://dair-ai.thinkific.com/",
        "indices": [
          1852,
          1875
        ],
        "url": "https://t.co/JBU5beIoD0"
      }
    ]
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "article": null
}