🐦 Twitter Post Details

Viewing enriched Twitter post

@vllm_project

πŸŽ‰ Day-0 support for @deepseek_ai V4 Pro and Flash on vLLM β€” a new generation of DeepSeek model, purpose-built for tasks up to 1M tokens. Alongside the release, we're publishing a first-principles walkthrough of the new long-context attention and how we implemented it in vLLM. The new attention mechanism, in four moves: β€’ Shared K/V + inverse RoPE β†’ 2Γ— memory savings β€’ c4a / c128a KV compression β†’ 4×–128Γ— savings β€’ DeepSeek Sparse Attention over compressed tokens β€’ Short sliding window for locality across compression boundaries At 1M context, per-layer KV state is ~8.7Γ— smaller than a DeepSeek V3.2-style 61-layer stack (9.62 GiB vs 83.9 GiB, bf16). fp8 attention cache + fp4 indexer cache shrink it further. vLLM side: β€’ Unified hybrid KV cache β€” single logical block size (256 native positions) across all compression rates; compressor state folded into the SWA KV cache spec so prefix caching, disagg prefill, CUDA graphs and MTP reuse the same abstraction β€’ Three page-size buckets for the full 5-way cache stack β†’ no cross-kind fragmentation β€’ Fused kernels: compressor + RMSNorm + RoPE + cache insert (1.4–3Γ—), inverse RoPE + fp8 quant (2–3Γ—), Q-norm + KV RoPE + K insert (10–20Γ—) β€’ Multi-stream overlap of indexer vs main-KV compression vs SWA insertion Disaggregated serving is supported out of the box and strongly recommended for best performance. Follow our recipes site for verified commands for @nvidia Blackwell (B200, B300, GB200, GB300) and Hopper (H100/H200/H20) systems. Thanks to the @deepseek_ai team for open-sourcing DeepSeek V4, and to @inferact for landing day-0 support 🀝 πŸ“ Blog: https://t.co/Eh7vk6xVJy πŸ“– Recipes: https://t.co/jlWuzYyZeX πŸ€— https://t.co/IA9qAysqJk

Media 1
Media 2
Media 3
Media 4

πŸ“Š Media Metadata

{
  "media": [
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_0.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_0.jpg",
      "type": "photo",
      "filename": "media_0.jpg"
    },
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_1.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_1.jpg",
      "type": "photo",
      "filename": "media_1.jpg"
    },
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_2.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_2.jpg",
      "type": "photo",
      "filename": "media_2.jpg"
    },
    {
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_3.jpg",
      "media_url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2047520252851105796/media_3.jpg",
      "type": "photo",
      "filename": "media_3.jpg"
    }
  ],
  "processed_at": "2026-05-11T22:37:58.237873",
  "pipeline_version": "2.0"
}

πŸ”§ Raw API Response

{
  "type": "tweet",
  "id": "2047520252851105796",
  "url": "https://x.com/vllm_project/status/2047520252851105796",
  "twitterUrl": "https://twitter.com/vllm_project/status/2047520252851105796",
  "text": "πŸŽ‰ Day-0 support for @deepseek_ai V4 Pro and Flash on vLLM β€” a new generation of DeepSeek model, purpose-built for tasks up to 1M tokens. Alongside the release, we're publishing a first-principles walkthrough of the new long-context attention and how we implemented it in vLLM.\n\nThe new attention mechanism, in four moves:\nβ€’ Shared K/V + inverse RoPE β†’ 2Γ— memory savings\nβ€’ c4a / c128a KV compression β†’ 4×–128Γ— savings\nβ€’ DeepSeek Sparse Attention over compressed tokens\nβ€’ Short sliding window for locality across compression boundaries\n\nAt 1M context, per-layer KV state is ~8.7Γ— smaller than a DeepSeek V3.2-style 61-layer stack (9.62 GiB vs 83.9 GiB, bf16). fp8 attention cache + fp4 indexer cache shrink it further.\n\nvLLM side:\nβ€’ Unified hybrid KV cache β€” single logical block size (256 native positions) across all compression rates; compressor state folded into the SWA KV cache spec so prefix caching, disagg prefill, CUDA graphs and MTP reuse the same abstraction\nβ€’ Three page-size buckets for the full 5-way cache stack β†’ no cross-kind fragmentation\nβ€’ Fused kernels: compressor + RMSNorm + RoPE + cache insert (1.4–3Γ—), inverse RoPE + fp8 quant (2–3Γ—), Q-norm + KV RoPE + K insert (10–20Γ—)\nβ€’ Multi-stream overlap of indexer vs main-KV compression vs SWA insertion\n\nDisaggregated serving is supported out of the box and strongly recommended for best performance.\n\nFollow our recipes site for verified commands for @nvidia Blackwell (B200, B300, GB200, GB300) and Hopper (H100/H200/H20) systems.\n\nThanks to the @deepseek_ai team for open-sourcing DeepSeek V4, and to @inferact for landing day-0 support 🀝\n\nπŸ“ Blog: https://t.co/Eh7vk6xVJy\nπŸ“– Recipes: https://t.co/jlWuzYyZeX\nπŸ€— https://t.co/IA9qAysqJk",
  "source": "Twitter for iPhone",
  "retweetCount": 90,
  "replyCount": 17,
  "likeCount": 571,
  "quoteCount": 20,
  "viewCount": 122680,
  "createdAt": "Fri Apr 24 03:37:24 +0000 2026",
  "lang": "en",
  "bookmarkCount": 140,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2047520252851105796",
  "displayTextRange": [
    0,
    276
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "vllm_project",
    "url": "https://x.com/vllm_project",
    "twitterUrl": "https://twitter.com/vllm_project",
    "id": "1774187564276289536",
    "name": "vLLM",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/1774187681746182144/N_5NJ8B1_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/1774187564276289536/1733806062",
    "description": "",
    "location": "",
    "followers": 38270,
    "following": 36,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Sat Mar 30 21:31:01 +0000 2024",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 588,
    "hasCustomTimelines": true,
    "isTranslator": false,
    "mediaCount": 274,
    "statusesCount": 981,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [],
    "profile_bio": {
      "description": "A high-throughput and memory-efficient inference and serving engine for LLMs. Join https://t.co/lxJ0SfX5pJ to discuss together with the community!",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "slack.vllm.ai",
              "expanded_url": "http://slack.vllm.ai",
              "indices": [
                83,
                106
              ],
              "url": "https://t.co/lxJ0SfX5pJ"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "vllm.ai",
              "expanded_url": "https://vllm.ai",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/OSHR5C6pCV"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/ho9pSI4IXs",
        "expanded_url": "https://twitter.com/vllm_project/status/2047520252851105796/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {
            "faces": []
          },
          "orig": {
            "faces": []
          }
        },
        "id_str": "2047520157493522432",
        "indices": [
          277,
          300
        ],
        "media_key": "3_2047520157493522432",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARxqQLwp2nAACgACHGpA0l2bsAQAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABHGpAvCnacAAKAAIcakDSXZuwBAAA",
            "media_key": "3_2047520157493522432"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HGpAvCnacAAF0G5.jpg",
        "original_info": {
          "focus_rects": [
            {
              "h": 936,
              "w": 1672,
              "x": 0,
              "y": 0
            },
            {
              "h": 941,
              "w": 941,
              "x": 0,
              "y": 0
            },
            {
              "h": 941,
              "w": 825,
              "x": 0,
              "y": 0
            },
            {
              "h": 941,
              "w": 471,
              "x": 0,
              "y": 0
            },
            {
              "h": 941,
              "w": 1672,
              "x": 0,
              "y": 0
            }
          ],
          "height": 941,
          "width": 1672
        },
        "sizes": {
          "large": {
            "h": 941,
            "w": 1672
          }
        },
        "type": "photo",
        "url": "https://t.co/ho9pSI4IXs"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "vllm.ai/blog/deepseek-…",
        "expanded_url": "http://vllm.ai/blog/deepseek-v4",
        "indices": [
          1618,
          1641
        ],
        "url": "https://t.co/Eh7vk6xVJy"
      },
      {
        "display_url": "recipes.vllm.ai/deepseek-ai/De…",
        "expanded_url": "https://recipes.vllm.ai/deepseek-ai/DeepSeek-V4-Pro",
        "indices": [
          1653,
          1676
        ],
        "url": "https://t.co/jlWuzYyZeX"
      },
      {
        "display_url": "huggingface.co/deepseek-ai/De…",
        "expanded_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro",
        "indices": [
          1679,
          1702
        ],
        "url": "https://t.co/IA9qAysqJk"
      }
    ],
    "user_mentions": [
      {
        "id_str": "1714580962569588736",
        "indices": [
          20,
          32
        ],
        "name": "DeepSeek",
        "screen_name": "deepseek_ai"
      },
      {
        "id_str": "61559439",
        "indices": [
          1419,
          1426
        ],
        "name": "NVIDIA",
        "screen_name": "nvidia"
      },
      {
        "id_str": "1714580962569588736",
        "indices": [
          1515,
          1527
        ],
        "name": "DeepSeek",
        "screen_name": "deepseek_ai"
      },
      {
        "id_str": "1998160192903745536",
        "indices": [
          1571,
          1580
        ],
        "name": "Inferact",
        "screen_name": "inferact"
      }
    ]
  },
  "quoted_tweet": {
    "type": "tweet",
    "id": "2047516922263285776",
    "url": "https://x.com/deepseek_ai/status/2047516922263285776",
    "twitterUrl": "https://twitter.com/deepseek_ai/status/2047516922263285776",
    "text": "πŸš€ DeepSeek-V4 Preview is officially live & open-sourced! Welcome to the era of cost-effective 1M context length.\n\nπŸ”Ή DeepSeek-V4-Pro: 1.6T total / 49B active params. Performance rivaling the world's top closed-source models.\nπŸ”Ή DeepSeek-V4-Flash: 284B total / 13B active params. Your fast, efficient, and economical choice.\n\nTry it now at https://t.co/GCdiMzk1Dl via Expert Mode / Instant Mode. API is updated & available today!\n\nπŸ“„ Tech Report: https://t.co/drlDrxkYtp\nπŸ€— Open Weights: https://t.co/T13Y8i7SDM\n\n1/n",
    "source": "Twitter for iPhone",
    "retweetCount": 7679,
    "replyCount": 1605,
    "likeCount": 45214,
    "quoteCount": 2396,
    "viewCount": 9632950,
    "createdAt": "Fri Apr 24 03:24:09 +0000 2026",
    "lang": "en",
    "bookmarkCount": 9970,
    "isReply": false,
    "inReplyToId": null,
    "conversationId": "2047516922263285776",
    "displayTextRange": [
      0,
      280
    ],
    "inReplyToUserId": null,
    "inReplyToUsername": null,
    "author": {
      "type": "user",
      "userName": "deepseek_ai",
      "url": "https://x.com/deepseek_ai",
      "twitterUrl": "https://twitter.com/deepseek_ai",
      "id": "1714580962569588736",
      "name": "DeepSeek",
      "isVerified": false,
      "isBlueVerified": true,
      "verifiedType": null,
      "profilePicture": "https://pbs.twimg.com/profile_images/1717417613775757312/Uk1zNOj4_normal.jpg",
      "coverPicture": "https://pbs.twimg.com/profile_banners/1714580962569588736/1698208997",
      "description": "",
      "location": "",
      "followers": 1024618,
      "following": 0,
      "status": "",
      "canDm": false,
      "canMediaTag": true,
      "createdAt": "Wed Oct 18 09:55:45 +0000 2023",
      "entities": {
        "description": {
          "urls": []
        },
        "url": {}
      },
      "fastFollowersCount": 0,
      "favouritesCount": 32,
      "hasCustomTimelines": true,
      "isTranslator": false,
      "mediaCount": 105,
      "statusesCount": 167,
      "withheldInCountries": [],
      "affiliatesHighlightedLabel": {},
      "possiblySensitive": false,
      "pinnedTweetIds": [
        "2048440764368347611"
      ],
      "profile_bio": {
        "description": "Unravel the mystery of AGI with curiosity. Answer the essential question with long-termism.",
        "entities": {
          "description": {},
          "url": {
            "urls": [
              {
                "display_url": "deepseek.com",
                "expanded_url": "https://www.deepseek.com/",
                "indices": [
                  0,
                  23
                ],
                "url": "https://t.co/Un4k2rqn4o"
              }
            ]
          }
        }
      },
      "isAutomated": false,
      "automatedBy": null
    },
    "extendedEntities": {
      "media": [
        {
          "allow_download_status": {
            "allow_download": true
          },
          "display_url": "pic.twitter.com/n1AgwMIymu",
          "expanded_url": "https://twitter.com/deepseek_ai/status/2047516922263285776/photo/1",
          "ext_media_availability": {
            "status": "Available"
          },
          "features": {
            "large": {
              "faces": []
            },
            "orig": {
              "faces": []
            }
          },
          "id_str": "2047516800439767040",
          "indices": [
            281,
            304
          ],
          "media_key": "3_2047516800439767040",
          "media_results": {
            "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARxqPa6J21AACgACHGo9yucasBAAAA==",
            "result": {
              "__typename": "ApiMedia",
              "id": "QXBpTWVkaWE6DAABCgABHGo9ronbUAAKAAIcaj3K5xqwEAAA",
              "media_key": "3_2047516800439767040"
            }
          },
          "media_url_https": "https://pbs.twimg.com/media/HGo9ronbUAAKlRk.jpg",
          "original_info": {
            "focus_rects": [
              {
                "h": 554,
                "w": 989,
                "x": 910,
                "y": 0
              },
              {
                "h": 554,
                "w": 554,
                "x": 1127,
                "y": 0
              },
              {
                "h": 554,
                "w": 486,
                "x": 1161,
                "y": 0
              },
              {
                "h": 554,
                "w": 277,
                "x": 1266,
                "y": 0
              },
              {
                "h": 554,
                "w": 2809,
                "x": 0,
                "y": 0
              }
            ],
            "height": 554,
            "width": 2809
          },
          "sizes": {
            "large": {
              "h": 404,
              "w": 2048
            }
          },
          "type": "photo",
          "url": "https://t.co/n1AgwMIymu"
        }
      ]
    },
    "card": null,
    "place": {},
    "entities": {
      "hashtags": [],
      "symbols": [],
      "urls": [
        {
          "display_url": "chat.deepseek.com",
          "expanded_url": "http://chat.deepseek.com",
          "indices": [
            337,
            360
          ],
          "url": "https://t.co/GCdiMzk1Dl"
        },
        {
          "display_url": "huggingface.co/deepseek-ai/De…",
          "expanded_url": "https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf",
          "indices": [
            443,
            466
          ],
          "url": "https://t.co/drlDrxkYtp"
        },
        {
          "display_url": "huggingface.co/collections/de…",
          "expanded_url": "https://huggingface.co/collections/deepseek-ai/deepseek-v4",
          "indices": [
            483,
            506
          ],
          "url": "https://t.co/T13Y8i7SDM"
        }
      ],
      "user_mentions": []
    },
    "quoted_tweet": null,
    "retweeted_tweet": null,
    "isLimitedReply": false,
    "communityInfo": null,
    "article": null
  },
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}