Your curated collection of saved posts and media

Showing 10 posts Β· last 14 days Β· by score
βž• Add New Post
A
Artificial Analysis
@ArtificialAnlys
πŸ“…
Aug 28, 2026
4d ago
πŸ†”68666138

Perplexity Search debuts on the Artificial Analysis Search Index, with all three context size variants taking top positions on the leaderboard The @perplexity_ai Search API comes with three context settings (low, medium, and high) that control how much extracted content each search result carries. We tested all three variants using our standardized methodology: the same model (GPT-5.6 Luna at medium reasoning), running inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web. Only the provider behind the search tool changes. Key results: ➀ Perplexity Search (medium) scores 80 on the Artificial Analysis Search Index, ahead of the previous leaders, Parallel (advanced) and Brave Search (LLM context), at 75. The high and low variants score 79 and 77 respectively. Its lead is concentrated in BrowseComp results, with AA-Omniscience and DeepSearchQA scoring comparably to other leading providers ➀ Efficient search payloads: smaller overall search results mean the model reads less per task, so Perplexity has the lowest model inference cost per task of providers we’ve tested so far, ranging from $0.028 to $0.034 across the three variants vs $0.036 for the next lowest provider ➀ Total cost per task is ~$0.091 for the medium and high context variants, at mid-pack latency. For comparison, Parallel (advanced) costs $0.084 per task and Brave (LLM context) costs $0.13 per task

Media 1
❀️204
likes
πŸ”16
retweets
πŸ–ΌοΈ Media
R
Sayak Paul
@RisingSayak
πŸ“…
Sep 01, 2026
14h ago
πŸ†”03605955

The only tutorial you should attend at #ECCV26 (kidding πŸ˜‚). But I am incredibly psyched to be doing this with the best bunch out there. I strongly feel this is a timely tutorial! We will present general approaches alongside our learnings from (post)-training impactful models like Flux2/3. Tutorial website: https://t.co/eig5dg3xwP Save your calendars!

Media 1
❀️203
likes
πŸ”25
retweets
πŸ–ΌοΈ Media
G
Gabriele Berton
@gabriberton
πŸ“…
Aug 24, 2026
8d ago
πŸ†”02237657

What is the simplest, cleanest technique for self-supervised learning (SSL)? RotNet (2018), without a doubt Take an image, rotate by either 0, 90, 180, 270 degrees, and predict its rotation. The model must learn useful features to solve the task. Read on... [1/2] https://t.co/jlqS7kXHeH

Media 1
❀️197
likes
πŸ”5
retweets
πŸ–ΌοΈ Media
G
GitHub
@github
πŸ“…
Aug 25, 2026
7d ago
πŸ†”60987573
⭐0.36

πŸ’‘ Learn the latest GitHub workflows by actually building with them. We’ve released 4 new GitHub Skills exercises designed to give developers practice with AI-powered development, agentic workflows, and code quality. πŸ§΅β¬‡οΈ

❀️193
likes
πŸ”12
retweets
C
clem πŸ€—
@ClementDelangue
πŸ“…
Aug 22, 2026
10d ago
πŸ†”15492806

NVIDIA built its own coding harness to optimize CUDA GPU kernels and achieved a 100% score on ARC-AGI-3’s 25 public games, solving all 183 levels. With agents, we'll move from a world where it's quite hard to run, optimize, post-train your own AI models and kernels to a world where virtually everybody can do it. 100 million AI builders when?

Media 1
❀️163
likes
πŸ”35
retweets
πŸ–ΌοΈ Media
Z
Zixuan Li
@ZixuanLi_
πŸ“…
Aug 28, 2026
4d ago
πŸ†”00487090
⭐0.32

GLM 5.3’s weights were just released, and it’s already available in Perplexity Computer.

@perplexity_ai β€’ Fri Aug 28 18:49

GLM 5.3 is now available in Perplexity Computer. Built for long-context, multimodal agent workloads, it beat GLM 5.2 on WANDR, our benchmark for large-scale, evidence-backed research. https://t.co/02FMedtoJb

❀️126
likes
πŸ”7
retweets
S
Spenser Skates
@spenserskates
πŸ“…
Sep 01, 2026
13h ago
πŸ†”72207384

Today we're launching Agent Analytics Every team shipping an AI agent has the same blind spot. Offline evals pass, you ship, and then you have no idea what's happening in production. AI fails silently. Users ask a question and get different answers. They all look 'engaged' in a classic dashboard. You don't know who got a great response and who got a terrible one. Agent Analytics solves it: - Every session scored out of the box on task completion, response quality, friction, safety, and negative feedback - Topic clustering across thousands of conversations, so you know if a failure hits 1 user or 10,000 - Eval agents that watch for regressions, and if you want will file a Linear ticket or the pull request themselves - Agent quality sits next to product data, so 'payment scheduling fails 31%' becomes 'which renewals did that cost us?' The Economist got their agent to a 96.9% task success rate and cut weekly failures 84%. Included on every plan. Free tier included. https://t.co/K5vqrlMLvd

Media 2
❀️122
likes
πŸ”41
retweets
πŸ–ΌοΈ Media
G
Google Gemma
@googlegemma
πŸ“…
Aug 28, 2026
4d ago
πŸ†”07963769

Gemma 4 E2B running fully locally on an 8GB NVIDIA Jetson Orin Nano. Using native image/audio input and function calling, this edge setup is capable of vision, voice interaction, and camera control, all within an 8GB memory footprint! https://t.co/ILdL9vnvdG

❀️119
likes
πŸ”8
retweets
πŸ–ΌοΈ Media
A
Artificial Analysis
@ArtificialAnlys
πŸ“…
Aug 21, 2026
11d ago
πŸ†”31994528

Announcing our new Speech Agent Arena, evaluating Speech to Speech models on real-world scenarios to analyze conversational preference and task success rate Existing Speech to Speech benchmarks cover reasoning, simulated agentic tasks, and conversational dynamics such as turn-taking and interruption handling. The Speech Agent Arena compares models and cascaded systems as humans complete real-world tasks, measuring conversational preference and successful tool use. This allows us to provide an evaluation which closer reflects real-world use, offering insight into which models users most prefer speaking with and how effectively those models support their requests. Overview of the Speech Agent Arena and Task Success Rate Human participants compare two hidden Speech to Speech models on the same assigned scenario, one of 15 agentic scenarios (tasks requiring tool calling, such as ordering takeout) or 20 non-agentic scenarios (tasks without tool calling, such as asking about opening hours). After separate live conversations with each model, participants select which they preferred, with these pairwise votes used to fit a Preference Elo score. For agentic scenarios, Task Success Rate is the share of eligible conversations (no participant deviations or unverifiable cases) where the model completed the requested action through the correct final tool call or calls. For all but a New Patient Dental Booking example, scenario model prompts, tool schemas and participant instructions are currently private to reduce overfitting. Additionally, the Speech Agent Arena currently uses a qualified pool of paid, screened third-party participants to conduct and evaluate agent interactions. Key results: ➀ Arena Preference Elo: @GoogleAI Gemini 3.1 Flash Live Preview - Minimal leads at 1,046 Elo, followed by Gemini 3.1 Flash Live Preview - High at 1,014, @OpenAI GPT-Realtime-1.5 at 1,000, GPT Realtime (Aug '25) at 944, and @ElevenLabs Agents (default cascaded system of Scribe v2 Realtime / GPT-4o Mini / Eleven v3, with pre-registered tool schema) at 937. In reviewed conversations, highly preferred models tended to respond quickly, sound more natural and produce fewer unnatural sounds or audio artifacts ➀ Task Success Rate: @SpaceXAI Grok Voice Think Fast 2.0 High leads at 94.7%, followed by @OpenAI GPT-Realtime-2.1 High at 91.5%, @ElevenLabs Agents (Default Cascaded System) at 90.5%, and GPT-Realtime-2 (High) at 89.8%, with GPT Realtime (Aug '25) and GPT-Realtime-2.1 Minimal tied at 89.4%. Gemini 3.1 Flash Live Preview - Minimal leads overall preference at 1,046 Elo but records a 74.6% Task Success Rate, showing that a preferred conversation does not always result in successful task completion - some conversations can sound as though the requested action was completed even when the required final tool call was unsuccessful We are continuing to expand our coverage of native and cascaded Speech to Speech systems, and welcome feedback as we add more models, providers and scenarios. See more details below ⬇️

Media 1
❀️119
likes
πŸ”11
retweets
πŸ–ΌοΈ Media
X
Florian Brand
@xeophon
πŸ“…
Aug 24, 2026
9d ago
πŸ†”81518646
⭐0.30

tired: hill climbing an eval by doing a synth env wired: hill climbing an eval by fixing the eval

❀️110
likes
πŸ”5
retweets