🐦 Twitter Post Details

Viewing enriched Twitter post

@jerryjliu0

PDF parsing is fun because there's an infinite variety of enteprise documents šŸ“‘. For each document category, there's a long tail of work to build more precise bounding boxes, confidence scores, and domain-specific annotations so that you provide any downstream agent rich metadata without it having to reinvent this from scratch. Take forms for example. Besides simply outputting it into markdown, we put in the work to detect every annotation, field, checkbox, and section. That way you immediately get structured information as to whether a form is filled without a separate LLM extraction step. You also get source citations for free! Doing this well is hard. āœ… There's a rabbit-hole of optimizations you can do, including extracting annotations for every document type. āœ… It's hard to properly render visual formats like charts, handwriting into digitalized information. The more you skip this step, the more work you're creating for any downstream agent. āœ… Precise bounding boxes are a necessity for precise citations on any type of document. You can aggressively tune the model + harness so that the accuracy/cost on any document subtype is much more competitive than the frontier models. Whether you're parsing forms (see the enriched forms option in "processing options") or any other type of doc, check out LlamaParse ! https://t.co/XYZmx5TFz8

Media 1

šŸ“Š Media Metadata

{
  "media": [
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2093200379983073354/media_0.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2093200379983073354/media_0.jpg",
      "type": "photo",
      "filename": "media_0.jpg"
    }
  ],
  "processed_at": "2026-08-28T18:01:23.263430",
  "pipeline_version": "2.0"
}

šŸ”§ Raw API Response

{
  "type": "tweet",
  "id": "2093200379983073354",
  "url": "https://x.com/jerryjliu0/status/2093200379983073354",
  "twitterUrl": "https://twitter.com/jerryjliu0/status/2093200379983073354",
  "text": "PDF parsing is fun because there's an infinite variety of enteprise documents šŸ“‘. For each document category, there's a long tail of work to build more precise bounding boxes, confidence scores, and domain-specific annotations so that you provide any downstream agent rich metadata without it having to reinvent this from scratch. \n\nTake forms for example. Besides simply outputting it into markdown, we put in the work to detect every annotation, field, checkbox, and section. That way you \n immediately get structured information as to whether a form is filled without a separate LLM extraction step. You also get source citations for free! \n\nDoing this well is hard. \nāœ… There's a rabbit-hole of optimizations you can do, including extracting annotations for every document type. \nāœ… It's hard to properly render visual formats like charts, handwriting into digitalized information. The more you skip this step, the more work you're creating for any downstream agent.\nāœ… Precise bounding boxes are a necessity for precise citations on any type of document. \n\nYou can aggressively tune the model + harness so that the accuracy/cost on any document subtype is much more competitive than the frontier models.\n\nWhether you're parsing forms (see the enriched forms option in \"processing options\") or any other type of doc, check out LlamaParse !\n\nhttps://t.co/XYZmx5TFz8",
  "source": "Twitter for iPhone",
  "retweetCount": 7,
  "replyCount": 19,
  "likeCount": 112,
  "quoteCount": 0,
  "viewCount": 8839,
  "createdAt": "Fri Aug 28 04:53:55 +0000 2026",
  "lang": "en",
  "bookmarkCount": 105,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2093200379983073354",
  "displayTextRange": [
    0,
    271
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "jerryjliu0",
    "url": "https://x.com/jerryjliu0",
    "twitterUrl": "https://twitter.com/jerryjliu0",
    "id": "369777416",
    "name": "Jerry Liu",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1283610285031460864/1Q4zYhtb_normal.jpg",
    "coverPicture": "",
    "description": "",
    "location": "",
    "followers": 82044,
    "following": 1573,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Wed Sep 07 22:54:31 +0000 2011",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 10218,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 1632,
    "statusesCount": 7525,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {
      "label": {
        "badge": {
          "url": "https://pbs.twimg.com/profile_images/1967920417760251904/0ytfduMQ_bigger.png"
        },
        "description": "LlamaIndex šŸ¦™",
        "url": {
          "url": "https://twitter.com/llama_index",
          "url_type": "DeepLink"
        },
        "user_label_display_type": "Badge",
        "user_label_type": "BusinessLabel"
      }
    },
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2092650965333950593"
    ],
    "profile_bio": {
      "description": "Parsing the world's hardest PDFs @llama_index. cofounder/CEO\n\nCareers: https://t.co/EUnMNmbCtx\nEnterprise: https://t.co/Ht5jwxSrQB",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "llamaindex.ai/careers",
              "expanded_url": "https://www.llamaindex.ai/careers",
              "indices": [
                71,
                94
              ],
              "url": "https://t.co/EUnMNmbCtx"
            },
            {
              "display_url": "llamaindex.ai/contact",
              "expanded_url": "https://www.llamaindex.ai/contact",
              "indices": [
                107,
                130
              ],
              "url": "https://t.co/Ht5jwxSrQB"
            }
          ],
          "user_mentions": [
            {
              "id_str": "",
              "indices": [
                33,
                45
              ],
              "name": "",
              "screen_name": "llama_index"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "llamaindex.ai",
              "expanded_url": "https://www.llamaindex.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/YiIfjVlzb6"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.x.com/FtIGd8luSf",
        "expanded_url": "https://x.com/jerryjliu0/status/2093200379983073354/photo/1",
        "ext_master_playlist_only": [],
        "ext_media_availability": {
          "status": "Available"
        },
        "ext_playlists": [],
        "features": {
          "large": {
            "faces": [
              {
                "h": 89,
                "w": 89,
                "x": 581,
                "y": 496
              }
            ]
          },
          "orig": {
            "faces": [
              {
                "h": 151,
                "w": 151,
                "x": 982,
                "y": 837
              }
            ]
          }
        },
        "id_str": "2093200136008839168",
        "indices": [
          272,
          295
        ],
        "media_key": "3_2093200136008839168",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAAR0Mim72mzAACgACHQyKp8SakEoAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABHQyKbvabMAAKAAIdDIqnxJqQSgAA",
            "media_key": "3_2093200136008839168"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HQyKbvabMAAiZwh.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 1935,
              "w": 3456,
              "x": 0,
              "y": 0
            },
            {
              "h": 1998,
              "w": 1998,
              "x": 1333,
              "y": 0
            },
            {
              "h": 1998,
              "w": 1753,
              "x": 1456,
              "y": 0
            },
            {
              "h": 1998,
              "w": 999,
              "x": 1833,
              "y": 0
            },
            {
              "h": 1998,
              "w": 3456,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1998,
          "width": 3456
        },
        "sizes": {
          "large": {
            "h": 1184,
            "w": 2048
          }
        },
        "type": "photo",
        "url": "https://t.co/FtIGd8luSf"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "cloud.llamaindex.ai",
        "expanded_url": "https://cloud.llamaindex.ai/",
        "indices": [
          1341,
          1364
        ],
        "url": "https://t.co/XYZmx5TFz8"
      }
    ],
    "user_mentions": []
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}