🐦 Twitter Post Details

Viewing enriched Twitter post

@morganlinton

I am going to start to share the leaderboard at VulcanBench more regularly. VulcanBench is the only benchmark I know that benchmarks across effort levels, on an eval suite of 100% real engineering tasks. Getting ready to sunset eval suite 3 and eval suite 4 is almost ready to go. Here's the current leaderboard for Eval Suite 3, Grok 4.5 High is the highest scoring model still with Fable 5 Low, yes Low, right behind it. What I've been able to uncover with VulcanBench is that way too many people are running models at Max effort, thinking that buys them more accuracy, when really it just costs more and uses more tokens so takes more time. If you're still in the mode of, tell my agent to do something then go get coffee, you're probably still living in the past. You can move faster, with higher accuracy, the key is not thinking you need Max effort all the time.

📊 Media Metadata

{
  "score": 0.42,
  "score_components": {
    "author": 0.09,
    "engagement": 0.0,
    "quality": 0.12,
    "source": 0.135,
    "nlp": 0.05,
    "recency": 0.025
  },
  "scored_at": "2026-08-22T00:01:51.055576",
  "import_source": "api_import",
  "source_tagged_at": "2026-08-22T00:01:51.055588",
  "enriched": true,
  "enriched_at": "2026-08-22T00:01:51.055591"
}

🔧 Raw API Response

{
  "type": "tweet",
  "id": "2090837054536262007",
  "url": "https://x.com/morganlinton/status/2090837054536262007",
  "twitterUrl": "https://twitter.com/morganlinton/status/2090837054536262007",
  "text": "I am going to start to share the leaderboard at VulcanBench more regularly.\n\nVulcanBench is the only benchmark I know that benchmarks across effort levels, on an eval suite of 100% real engineering tasks.\n\nGetting ready to sunset eval suite 3 and eval suite 4 is almost ready to go.\n\nHere's the current leaderboard for Eval Suite 3, Grok 4.5 High is the highest scoring model still with Fable 5 Low, yes Low, right behind it.\n\nWhat I've been able to uncover with VulcanBench is that way too many people are running models at Max effort, thinking that buys them more accuracy, when really it just costs more and uses more tokens so takes more time.\n\nIf you're still in the mode of, tell my agent to do something then go get coffee, you're probably still living in the past. You can move faster, with higher accuracy, the key is not thinking you need Max effort all the time.",
  "source": "Twitter for iPhone",
  "retweetCount": 10,
  "replyCount": 15,
  "likeCount": 79,
  "quoteCount": 0,
  "viewCount": 71277,
  "createdAt": "Fri Aug 21 16:22:54 +0000 2026",
  "lang": "en",
  "bookmarkCount": 14,
  "isReply": false,
  "inReplyToId": null,
  "conversationId": "2090837054536262007",
  "displayTextRange": [
    0,
    278
  ],
  "inReplyToUserId": null,
  "inReplyToUsername": null,
  "author": {
    "type": "user",
    "userName": "morganlinton",
    "url": "https://x.com/morganlinton",
    "twitterUrl": "https://twitter.com/morganlinton",
    "id": "19016936",
    "name": "Morgan",
    "isVerified": false,
    "isBlueVerified": true,
    "verifiedType": null,
    "profilePicture": "https://pbs.twimg.com/profile_images/2058612692580184064/h0dOdYSc_normal.jpg",
    "coverPicture": "https://pbs.twimg.com/profile_banners/19016936/1787097440",
    "description": "",
    "location": "Incline Village, NV",
    "followers": 43451,
    "following": 796,
    "status": "",
    "canDm": true,
    "canMediaTag": false,
    "createdAt": "Thu Jan 15 09:54:24 +0000 2009",
    "entities": {
      "description": {
        "urls": []
      },
      "url": {}
    },
    "fastFollowersCount": 0,
    "favouritesCount": 85752,
    "hasCustomTimelines": true,
    "isTranslator": true,
    "mediaCount": 8531,
    "statusesCount": 79219,
    "withheldInCountries": [],
    "affiliatesHighlightedLabel": {},
    "possiblySensitive": false,
    "pinnedTweetIds": [
      "2090446623109476726"
    ],
    "profile_bio": {
      "description": "Cofounder @BoldMetrics: the AI body data engine. Building evals @VulcanBench: benchmarking models across effort levels on real coding tasks // Not an expert.",
      "entities": {
        "description": {
          "user_mentions": [
            {
              "id_str": "",
              "indices": [
                10,
                22
              ],
              "name": "",
              "screen_name": "BoldMetrics"
            },
            {
              "id_str": "",
              "indices": [
                64,
                76
              ],
              "name": "",
              "screen_name": "VulcanBench"
            }
          ]
        },
        "url": {
          "urls": [
            {
              "display_url": "modelrouting.substack.com",
              "expanded_url": "https://modelrouting.substack.com",
              "indices": [
                0,
                23
              ],
              "url": "https://t.co/xGsyc3Do8f"
            }
          ]
        }
      }
    },
    "isAutomated": false,
    "automatedBy": null
  },
  "extendedEntities": {},
  "card": null,
  "place": {},
  "entities": {
    "hashtags": [],
    "symbols": [],
    "urls": [],
    "user_mentions": []
  },
  "quoted_tweet": {
    "type": "tweet",
    "id": "2090836047391576365",
    "url": "",
    "twitterUrl": "",
    "text": "",
    "source": "Twitter for iPhone",
    "retweetCount": 0,
    "replyCount": 0,
    "likeCount": 0,
    "quoteCount": 0,
    "viewCount": 0,
    "createdAt": "",
    "lang": "",
    "bookmarkCount": 0,
    "isReply": false,
    "inReplyToId": null,
    "conversationId": "",
    "displayTextRange": [],
    "inReplyToUserId": null,
    "inReplyToUsername": null,
    "author": {},
    "extendedEntities": {},
    "card": null,
    "place": {},
    "entities": {},
    "quoted_tweet": null,
    "retweeted_tweet": null,
    "isLimitedReply": false,
    "communityInfo": null,
    "article": null
  },
  "retweeted_tweet": null,
  "isLimitedReply": false,
  "communityInfo": null,
  "article": null
}