🐦 Twitter Post Details

Viewing enriched Twitter post

@jerryjliu0

Parsing PDFs is hard This past week I gave a few talks (at both AI Dev '26 by @DeepLearningAI and @Capgemini ) on why this is still such an open problem, and it’s even more important as agents become the consumers of documents, and need the OCR tools to read them properly. The fundamental issue is that PDFs are designed for print and display purposes, not to give back a linearized, semantically meaningful string of text. Text and tables are represented as a bunch of chars and lines, without any guaranteed order. This is what the community is solving with VLM-based approaches, including our own efforts around LlamaParse and ParseBench. If you’re interested in learning more about the problem, check out the blog post I wrote on this a while ago! https://t.co/740ZiAFyOk

Media 1
Media 2

📊 Media Metadata

{
  "media": [
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2050961097642086427/media_0.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2050961097642086427/media_0.jpg",
      "type": "photo",
      "filename": "media_0.jpg"
    },
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2050961097642086427/media_1.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2050961097642086427/media_1.jpg",
      "type": "photo",
      "filename": "media_1.jpg"
    }
  ],
  "processed_at": "2026-05-11T22:55:44.583752",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2050961097642086427",
  "url": "https://x.com/jerryjliu0/status/2050961097642086427",
  "twitterUrl": "https://twitter.com/jerryjliu0/status/2050961097642086427",
  "text": "Parsing PDFs is hard\n\nThis past week I gave a few talks (at both AI Dev '26 by @DeepLearningAI  and @Capgemini ) on why this is still such an open problem, and it’s even more important as agents become the consumers of documents, and need the OCR tools to read them properly. \n\nThe fundamental issue is that PDFs are designed for print and display purposes, not to give back a linearized, semantically meaningful string of text. Text and tables are represented as a bunch of chars and lines, without any guaranteed order.\n\nThis is what the community is solving with VLM-based approaches, including our own efforts around LlamaParse and ParseBench.\n\nIf you’re interested in learning more about the problem, check out the blog post I wrote on this a while ago!\n\nhttps://t.co/740ZiAFyOk",
  "source": "Twitter for iPhone",
  "retweetCount": 38,
  "replyCount": 21,
  "likeCount": 316,
  "quoteCount": 2,
  "viewCount": 25114,
  "createdAt": "Sun May 03 15:30:05 +0000 2026",
  "lang": "en",
  "bookmarkCount": 309,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2050961097642086427",
  "displayTextRange": [
    0,
    275
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "jerryjliu0",
    "url": "https://x.com/jerryjliu0",
    "twitterUrl": "https://twitter.com/jerryjliu0",
    "id": "369777416",
    "name": "Jerry Liu",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1283610285031460864/1Q4zYhtb_normal.jpg",
    "coverPicture": "",
    "description": "Parsing the world's hardest PDFs @llama_index. cofounder/CEO\n\nCareers: https://t.co/EUnMNmbCtx\nEnterprise: https://t.co/Ht5jwxSrQB",
    "location": "",
    "followers": 74145,
    "following": 1475,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Wed Sep 07 22:54:31 +0000 2011",
    "entities": {
      "description": {
        "urls": [
          {
            "display_url": "llamaindex.ai/careers",
            "expanded_url": "https://www.llamaindex.ai/careers",
            "indices": [
              71,
              94
            ],
            "url": "https://t.co/EUnMNmbCtx"
          },
          {
            "display_url": "llamaindex.ai/contact",
            "expanded_url": "https://www.llamaindex.ai/contact",
            "indices": [
              107,
              130
            ],
            "url": "https://t.co/Ht5jwxSrQB"
          }
        ]
      },
      "url": {
        "urls": [
          {
            "display_url": "llamaindex.ai",
            "expanded_url": "https://www.llamaindex.ai/",
            "indices": [
              0,
              23
            ],
            "url": "https://t.co/YiIfjVlzb6"
          }
        ]
      }
    },
    "fastFollowersCount": 0,
    "favouritesCount": 8726,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 1508,
    "statusesCount": 6942,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [],
    "profile_bio": {},
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.x.com/kwPYbA1ID6",
        "expanded_url": "https://x.com/jerryjliu0/status/2050961097642086427/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {
            "faces": []
          },
          "medium": {
            "faces": []
          },
          "orig": {
            "faces": []
          },
          "small": {
            "faces": []
          }
        },
        "id_str": "2050960989089304577",
        "indices": [
          276,
          299
        ],
        "media_key": "3_2050960989089304577",
        "media_results": {
          "result": {
            "media_key": "3_2050960989089304577"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HHZ6JzLaEAEKZb_.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 822,
              "w": 1468,
              "x": 0,
              "y": 0
            },
            {
              "h": 822,
              "w": 822,
              "x": 0,
              "y": 0
            },
            {
              "h": 822,
              "w": 721,
              "x": 43,
              "y": 0
            },
            {
              "h": 822,
              "w": 411,
              "x": 198,
              "y": 0
            },
            {
              "h": 822,
              "w": 1468,
              "x": 0,
              "y": 0
            }
          ],
          "height": 822,
          "width": 1468
        },
        "sizes": {
          "large": {
            "h": 822,
            "resize": "fit",
            "w": 1468
          },
          "medium": {
            "h": 672,
            "resize": "fit",
            "w": 1200
          },
          "small": {
            "h": 381,
            "resize": "fit",
            "w": 680
          },
          "thumb": {
            "h": 150,
            "resize": "crop",
            "w": 150
          }
        },
        "type": "photo",
        "url": "https://t.co/kwPYbA1ID6"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "llamaindex.ai/blog/why-readi…",
        "expanded_url": "https://www.llamaindex.ai/blog/why-reading-pdfs-is-hard",
        "indices": [
          760,
          783
        ],
        "url": "https://t.co/740ZiAFyOk"
      }
    ],
    "user_mentions": [
      {
        "id_str": "992153930095251456",
        "indices": [
          79,
          94
        ],
        "name": "DeepLearning.AI",
        "screen_name": "DeepLearningAI"
      },
      {
        "id_str": "14109159",
        "indices": [
          100,
          110
        ],
        "name": "Capgemini",
        "screen_name": "Capgemini"
      }
    ]
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "article": null
}