🐦 Twitter Post Details

Viewing enriched Twitter post

@llama_index

Word docs are one of the most common file formats people process in LlamaParse, and they've always been surprisingly frustrating to parse well. Here's the counterintuitive part: .docx actually has better structural information than most document formats. We just haven't been able to fully use it. Until now. A .docx file is a ZIP archive of XML files. That XML knows everything: cell boundaries, merged cells, column and row spans, nested tables, formatting tags, hyperlinks. A PDF of the same table has none of that. It's just text positioned at coordinates and line intersections that a parser has to reverse-engineer into structure. The hard part with Word XML isn't extracting the table content. It's knowing which page it's on. Word is a flow format — there are no page boundaries in the XML. Pagination depends on the renderer, fonts, margins, line-height. The same .docx renders differently in Word, LibreOffice, and Google Docs. We built a technique to resolve this, mapping Word XML table elements to their correct page positions in the rendered output. We now get the original document structure AND know exactly where each table appears. The quality improvement is most significant for: · Tables with rich cell formatting (bold, italic, strikethrough, superscript, lists inside cells) · Merged cells and column/row spans · Nested tables (tables inside table cells) If you're processing Word docs with table-heavy content, try it out. 📖 Full writeup: https://t.co/aAEFkvfycG

Media 1
Media 2

📊 Media Metadata

{
  "media": [
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2036836522536902801/media_0.jpg",
      "filename": "media_0.jpg"
    },
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2036836522536902801/media_1.png",
      "filename": "media_1.png"
    }
  ],
  "processed_at": "2026-03-25T16:17:03.721775",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2036836522536902801",
  "url": "https://x.com/llama_index/status/2036836522536902801",
  "twitterUrl": "https://twitter.com/llama_index/status/2036836522536902801",
  "text": "Word docs are one of the most common file formats people process in LlamaParse, and they've always been surprisingly frustrating to parse well.\nHere's the counterintuitive part: .docx actually has better structural information than most document formats. We just haven't been able to fully use it. Until now.\n\nA .docx file is a ZIP archive of XML files. That XML knows everything: cell boundaries, merged cells, column and row spans, nested tables, formatting tags, hyperlinks. A PDF of the same table has none of that. It's just text positioned at coordinates and line intersections that a parser has to reverse-engineer into structure.\n\nThe hard part with Word XML isn't extracting the table content. It's knowing which page it's on. Word is a flow format — there are no page boundaries in the XML. Pagination depends on the renderer, fonts, margins, line-height. The same .docx renders differently in Word, LibreOffice, and Google Docs.\n\nWe built a technique to resolve this, mapping Word XML table elements to their correct page positions in the rendered output. We now get the original document structure AND know exactly where each table appears.\nThe quality improvement is most significant for:\n\n· Tables with rich cell formatting (bold, italic, strikethrough, superscript, lists inside cells)\n· Merged cells and column/row spans\n· Nested tables (tables inside table cells)\n\nIf you're processing Word docs with table-heavy content, try it out.\n\n📖 Full writeup: https://t.co/aAEFkvfycG",
  "source": "Twitter for iPhone",
  "retweetCount": 0,
  "replyCount": 1,
  "likeCount": 3,
  "quoteCount": 0,
  "viewCount": 248,
  "createdAt": "Wed Mar 25 16:04:04 +0000 2026",
  "lang": "en",
  "bookmarkCount": 3,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2036836522536902801",
  "displayTextRange": [
    0,
    280
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "llama_index",
    "url": "https://x.com/llama_index",
    "twitterUrl": "https://twitter.com/llama_index",
    "id": "1604278358296055808",
    "name": "LlamaIndex 🦙",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": "Business",
    "profilePicture": "https://pbs.twimg.com/profile_images/1967920417760251904/0ytfduMQ_normal.png",
    "coverPicture": "https://pbs.twimg.com/profile_banners/1604278358296055808/1770092126",
    "description": "",
    "location": "",
    "followers": 111052,
    "following": 32,
    "status": "",
    "canDm": false,
    "canMediaTag": true,
    "createdAt": "Sun Dec 18 00:52:44 +0000 2022",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 1513,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 1846,
    "statusesCount": 3785,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2029767312195117278"
    ],
    "profile_bio": {
      "description": "AI Agents for document OCR + workflows\n\nLlamaParse: https://t.co/yQGTiRSfFL\nDocs: https://t.co/us6GCS14vD",
      "entities": {
        "description": {
          "hashtags": [],
          "symbols": [],
          "urls": [
            {
              "display_url": "cloud.llamaindex.ai",
              "expanded_url": "https://cloud.llamaindex.ai/",
              "indices": [
                52,
                75
              ],
              "url": "https://t.co/yQGTiRSfFL"
            },
            {
              "display_url": "developers.llamaindex.ai/python/cloud/",
              "expanded_url": "https://developers.llamaindex.ai/python/cloud/",
              "indices": [
                82,
                105
              ],
              "url": "https://t.co/us6GCS14vD"
            }
          ],
          "user_mentions": []
        },
        "url": {
          "urls": [
            {
              "display_url": "llamaindex.ai",
              "expanded_url": "https://www.llamaindex.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/epzefqPT9Z"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/JZ0IkE64jc",
        "expanded_url": "https://twitter.com/llama_index/status/2036836522536902801/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {
            "faces": [
              {
                "h": 51,
                "w": 51,
                "x": 250,
                "y": 356
              }
            ]
          },
          "orig": {
            "faces": [
              {
                "h": 51,
                "w": 51,
                "x": 250,
                "y": 356
              }
            ]
          }
        },
        "id_str": "2036836519760330752",
        "indices": [
          281,
          304
        ],
        "media_key": "3_2036836519760330752",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARxETAXp24AACgACHERMBo9aoJEAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABHERMBenbgAAKAAIcREwGj1qgkQAA",
            "media_key": "3_2036836519760330752"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HERMBenbgAARbja.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 672,
              "w": 1200,
              "x": 0,
              "y": 128
            },
            {
              "h": 800,
              "w": 800,
              "x": 0,
              "y": 0
            },
            {
              "h": 800,
              "w": 702,
              "x": 0,
              "y": 0
            },
            {
              "h": 800,
              "w": 400,
              "x": 10,
              "y": 0
            },
            {
              "h": 800,
              "w": 1200,
              "x": 0,
              "y": 0
            }
          ],
          "height": 800,
          "width": 1200
        },
        "sizes": {
          "large": {
            "h": 800,
            "w": 1200
          }
        },
        "type": "photo",
        "url": "https://t.co/JZ0IkE64jc"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "llamaindex.ai/blog/improving…",
        "expanded_url": "https://www.llamaindex.ai/blog/improving-table-parsing-for-word-docx-documents?utm_source=socials&utm_medium=li_social",
        "indices": [
          1468,
          1491
        ],
        "url": "https://t.co/aAEFkvfycG"
      }
    ],
    "user_mentions": []
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "article": null
}