@PatrickMoorhead
Enterprises do not buy tokens. They buy outcomes based on correct, completed work. Today @Signal_65, which I co-founded in 2023 with @danielnewmanUV and is led by President @ryanshrout who is a partner, launched PINNACLE, an enterprise agentic AI benchmark for enterprise CIOs, AI platform teams, AI developers and AI infrastructure buyers selecting the best models and systems for agents, and the vendors it compares. Capability leaderboards do well measuring the models. Infrastructure benchmarks do well measuring the box and clusters. Neither says whether a multi-step job came out right on the data you have, how many agents a platform sustains, or what one correct outcome costs. PINNACLE is the first and only benchmark that runs the same enterprise job on governed and as-found data generated from one ground truth, graded in code with no model judge, and publishes the gap per model. It is also the first and only benchmark that prices a correct task on a self-hosted GPU node as well as a hosted API, from the same graded runs, and can publish the utilization point where your own node beats buying tokens. The community has done a great job with certain elements, and we tried to pull it all together. PINNACLE runs generated multi-step jobs against a live filesystem, database, Python runtime and other tools, and grades the result in code against an answer key generated with the environment. Jobs run on governed data and on data as it accumulates, and quality, capacity and cost per correct task are measured in one program. Results are read through five enterprise personas: Knowledge Worker, Data Analyst, IT Professional, Customer Operations and Executive. Developer is coming later. Each reweights the same measurements by what that role fails on. Three of the more interesting, early findings, in Signal65 testing. More are on the website. On data as companies actually manage it: 43 of 44 model configurations lose ground and only one gains. Claude Opus 5 scores 92.5% on governed data and 98.3% on as-found data. The other 43 drop by a median of about 28 points, with Qwen3.5-397B-A17B losing 64 points. The top closed models invent answers more often than the best open models do. At 128K context, Claude Opus 5 answered 7.6% of questions the documents did not answer and GPT-5.6 Sol 10.3%. GLM-5.2 held to 1.5% and Qwen3.5-397B-A17B with reasoning on to 0.4%, and both retrieve at 96 or better. A correct task from a frontier API costs more than a dollar, not the dime the enterprise price sheet implies. The input an agent re-reads every round is 65% to 91% of the bill, which puts a correct task at $1.21 on GPT-5.6 Sol and $1.36 on Claude Opus 5. DeepSeek-V4-Flash on a leased B300 node delivers one for about 15 cents at full utilization and stays cheaper than GPT-5.6 Terra down to 30% utilization. The score still picks the model. The crossover prices the work. I want to be clear on where we are and where we aren’t. This is the first release, not the finished product. It covers only 44 model configurations from 30 base models, six hosted APIs and only three GPU platforms, with capacity measured on one node and one serving engine. We will be adding more capabilities weekly. I also want to point out that the methodology and governance are published so anyone can see how every number was produced, and how vendors get a factual review with no veto. I want to thank the teams at both @AMD and @NVIDIAAI for their feedback and support of the effort. I want to thank the Signal65 team under Ryan's leadership for their tireless work the past 6 months. This is just the beginning. Check out the details and other insights at https://t.co/WZzqmcTtwV