🐦 Twitter Post Details

Viewing enriched Twitter post

@omarsar0

Cool idea from Nous Research. What if you could speed up long-context pretraining with a subquadratic wrapper that you remove before deployment? That is the idea behind Lighthouse Attention. The method wraps ordinary SDPA with a hierarchical, gradient-free selection layer that compresses and decompresses queries, keys, and values symmetrically, preserving left-to-right causality. Crucially, it can be removed near the end of training in a short recovery phase, so the deployed model still runs vanilla attention with no architectural cost at inference. Preliminary LLM experiments report faster total training time and lower final loss than full-attention baselines. Why does it matter? Most efficient-attention work either changes the deployment-time architecture or pays a quality tax to do so. A training-only wrapper that survives a clean recovery phase sidesteps both. If it scales, this becomes an important training-time speedup for long-context pretraining. Paper: https://t.co/9g5Ldnb1rV Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX

Media 1

📊 Media Metadata

{
  "media": [
    {
      "type": "photo",
      "url": "https://crmoxkoizveukayfjuyo.supabase.co/storage/v1/object/public/media/posts/2054224130103554359/media_0.png",
      "filename": "media_0.png"
    }
  ],
  "processed_at": "2026-05-12T15:49:06.822810",
  "pipeline_version": "2.0"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2054224130103554359",
  "url": "https://x.com/omarsar0/status/2054224130103554359",
  "twitterUrl": "https://twitter.com/omarsar0/status/2054224130103554359",
  "text": "Cool idea from Nous Research.\n\nWhat if you could speed up long-context pretraining with a subquadratic wrapper that you remove before deployment?\n\nThat is the idea behind Lighthouse Attention.\n\nThe method wraps ordinary SDPA with a hierarchical, gradient-free selection layer that compresses and decompresses queries, keys, and values symmetrically, preserving left-to-right causality.\n\nCrucially, it can be removed near the end of training in a short recovery phase, so the deployed model still runs vanilla attention with no architectural cost at inference.\n\nPreliminary LLM experiments report faster total training time and lower final loss than full-attention baselines.\n\nWhy does it matter?\n\nMost efficient-attention work either changes the deployment-time architecture or pays a quality tax to do so. A training-only wrapper that survives a clean recovery phase sidesteps both. If it scales, this becomes an important training-time speedup for long-context pretraining.\n\nPaper: https://t.co/9g5Ldnb1rV\n\nLearn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX",
  "source": "Twitter for iPhone",
  "retweetCount": 0,
  "replyCount": 3,
  "likeCount": 6,
  "quoteCount": 0,
  "viewCount": 489,
  "createdAt": "Tue May 12 15:36:12 +0000 2026",
  "lang": "en",
  "bookmarkCount": 10,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2054224130103554359",
  "displayTextRange": [
    0,
    280
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "omarsar0",
    "url": "https://x.com/omarsar0",
    "twitterUrl": "https://twitter.com/omarsar0",
    "id": "3448284313",
    "name": "elvis",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/939313677647282181/vZjFWtAn_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/3448284313/1565974901",
    "description": "",
    "location": "DAIR.AI Academy",
    "followers": 303431,
    "following": 831,
    "status": "",
    "canDm": true,
    "canMediaTag": true,
    "createdAt": "Fri Sep 04 12:59:26 +0000 2015",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 35846,
    "hasCustomTimelines": true,
    "isTranslator": true,
    "mediaCount": 4647,
    "statusesCount": 17839,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2053978221193130434"
    ],
    "profile_bio": {
      "description": "Building @dair_ai • Prev: Meta AI, Elastic, PhD • New AI learning portal: https://t.co/1e8RZKs4uX",
      "entities": {
        "description": {
          "urls": [
            {
              "display_url": "academy.dair.ai",
              "expanded_url": "https://academy.dair.ai/",
              "indices": [
                74,
                97
              ],
              "url": "https://t.co/1e8RZKs4uX"
            }
          ],
          "user_mentions": [
            {
              "id_str": "",
              "indices": [
                9,
                17
              ],
              "name": "",
              "screen_name": "dair_ai"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "dair.ai",
              "expanded_url": "https://www.dair.ai/",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/XQto5ypSIk"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {
    "media": [
      {
        "display_url": "pic.twitter.com/NFrdUbEoHy",
        "expanded_url": "https://twitter.com/omarsar0/status/2054224130103554359/photo/1",
        "ext_media_availability": {
          "status": "Available"
        },
        "features": {
          "large": {
            "faces": []
          },
          "orig": {
            "faces": []
          }
        },
        "id_str": "2054224126500589568",
        "indices": [
          281,
          304
        ],
        "media_key": "3_2054224126500589568",
        "media_results": {
          "id": "QXBpTWVkaWFSZXN1bHRzOgwAAQoAARyCEfWVGtAACgACHIIR9mvbsTcAAA==",
          "result": {
            "__typename": "ApiMedia",
            "id": "QXBpTWVkaWE6DAABCgABHIIR9ZUa0AAKAAIcghH2a9uxNwAA",
            "media_key": "3_2054224126500589568"
          }
        },
        "media_url_https": "https://pbs.twimg.com/media/HIIR9ZUa0AAsd57.png",
        "original_info": {
          "focus_rects": [
            {
              "h": 949,
              "w": 1694,
              "x": 0,
              "y": 0
            },
            {
              "h": 1694,
              "w": 1694,
              "x": 0,
              "y": 0
            },
            {
              "h": 1852,
              "w": 1625,
              "x": 0,
              "y": 0
            },
            {
              "h": 1852,
              "w": 926,
              "x": 45,
              "y": 0
            },
            {
              "h": 1852,
              "w": 1694,
              "x": 0,
              "y": 0
            }
          ],
          "height": 1852,
          "width": 1694
        },
        "sizes": {
          "large": {
            "h": 1852,
            "w": 1694
          }
        },
        "type": "photo",
        "url": "https://t.co/NFrdUbEoHy"
      }
    ]
  },
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [
      {
        "display_url": "arxiv.org/abs/2605.06554",
        "expanded_url": "https://arxiv.org/abs/2605.06554",
        "indices": [
          984,
          1007
        ],
        "url": "https://t.co/9g5Ldnb1rV"
      },
      {
        "display_url": "academy.dair.ai",
        "expanded_url": "https://academy.dair.ai/",
        "indices": [
          1060,
          1083
        ],
        "url": "https://t.co/1e8RZKs4uX"
      }
    ],
    "user_mentions": []
  },
  "quoted_tweet": null,
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}