Your curated collection of saved posts and media
Perplexity Search debuts on the Artificial Analysis Search Index, with all three context size variants taking top positions on the leaderboard The @perplexity_ai Search API comes with three context settings (low, medium, and high) that control how much extracted content each search result carries. We tested all three variants using our standardized methodology: the same model (GPT-5.6 Luna at medium reasoning), running inside Stirrup, our open-source agent harness, with tools for searching and fetching pages from the web. Only the provider behind the search tool changes. Key results: β€ Perplexity Search (medium) scores 80 on the Artificial Analysis Search Index, ahead of the previous leaders, Parallel (advanced) and Brave Search (LLM context), at 75. The high and low variants score 79 and 77 respectively. Its lead is concentrated in BrowseComp results, with AA-Omniscience and DeepSearchQA scoring comparably to other leading providers β€ Efficient search payloads: smaller overall search results mean the model reads less per task, so Perplexity has the lowest model inference cost per task of providers weβve tested so far, ranging from $0.028 to $0.034 across the three variants vs $0.036 for the next lowest provider β€ Total cost per task is ~$0.091 for the medium and high context variants, at mid-pack latency. For comparison, Parallel (advanced) costs $0.084 per task and Brave (LLM context) costs $0.13 per task
The only tutorial you should attend at #ECCV26 (kidding π). But I am incredibly psyched to be doing this with the best bunch out there. I strongly feel this is a timely tutorial! We will present general approaches alongside our learnings from (post)-training impactful models like Flux2/3. Tutorial website: https://t.co/eig5dg3xwP Save your calendars!
What is the simplest, cleanest technique for self-supervised learning (SSL)? RotNet (2018), without a doubt Take an image, rotate by either 0, 90, 180, 270 degrees, and predict its rotation. The model must learn useful features to solve the task. Read on... [1/2] https://t.co/jlqS7kXHeH
π‘ Learn the latest GitHub workflows by actually building with them. Weβve released 4 new GitHub Skills exercises designed to give developers practice with AI-powered development, agentic workflows, and code quality. π§΅β¬οΈ
NVIDIA built its own coding harness to optimize CUDA GPU kernels and achieved a 100% score on ARC-AGI-3βs 25 public games, solving all 183 levels. With agents, we'll move from a world where it's quite hard to run, optimize, post-train your own AI models and kernels to a world where virtually everybody can do it. 100 million AI builders when?
GLM 5.3βs weights were just released, and itβs already available in Perplexity Computer.
GLM 5.3 is now available in Perplexity Computer. Built for long-context, multimodal agent workloads, it beat GLM 5.2 on WANDR, our benchmark for large-scale, evidence-backed research. https://t.co/02FMedtoJb
Today we're launching Agent Analytics Every team shipping an AI agent has the same blind spot. Offline evals pass, you ship, and then you have no idea what's happening in production. AI fails silently. Users ask a question and get different answers. They all look 'engaged' in a classic dashboard. You don't know who got a great response and who got a terrible one. Agent Analytics solves it: - Every session scored out of the box on task completion, response quality, friction, safety, and negative feedback - Topic clustering across thousands of conversations, so you know if a failure hits 1 user or 10,000 - Eval agents that watch for regressions, and if you want will file a Linear ticket or the pull request themselves - Agent quality sits next to product data, so 'payment scheduling fails 31%' becomes 'which renewals did that cost us?' The Economist got their agent to a 96.9% task success rate and cut weekly failures 84%. Included on every plan. Free tier included. https://t.co/K5vqrlMLvd
Gemma 4 E2B running fully locally on an 8GB NVIDIA Jetson Orin Nano. Using native image/audio input and function calling, this edge setup is capable of vision, voice interaction, and camera control, all within an 8GB memory footprint! https://t.co/ILdL9vnvdG
Announcing our new Speech Agent Arena, evaluating Speech to Speech models on real-world scenarios to analyze conversational preference and task success rate Existing Speech to Speech benchmarks cover reasoning, simulated agentic tasks, and conversational dynamics such as turn-taking and interruption handling. The Speech Agent Arena compares models and cascaded systems as humans complete real-world tasks, measuring conversational preference and successful tool use. This allows us to provide an evaluation which closer reflects real-world use, offering insight into which models users most prefer speaking with and how effectively those models support their requests. Overview of the Speech Agent Arena and Task Success Rate Human participants compare two hidden Speech to Speech models on the same assigned scenario, one of 15 agentic scenarios (tasks requiring tool calling, such as ordering takeout) or 20 non-agentic scenarios (tasks without tool calling, such as asking about opening hours). After separate live conversations with each model, participants select which they preferred, with these pairwise votes used to fit a Preference Elo score. For agentic scenarios, Task Success Rate is the share of eligible conversations (no participant deviations or unverifiable cases) where the model completed the requested action through the correct final tool call or calls. For all but a New Patient Dental Booking example, scenario model prompts, tool schemas and participant instructions are currently private to reduce overfitting. Additionally, the Speech Agent Arena currently uses a qualified pool of paid, screened third-party participants to conduct and evaluate agent interactions. Key results: β€ Arena Preference Elo: @GoogleAI Gemini 3.1 Flash Live Preview - Minimal leads at 1,046 Elo, followed by Gemini 3.1 Flash Live Preview - High at 1,014, @OpenAI GPT-Realtime-1.5 at 1,000, GPT Realtime (Aug '25) at 944, and @ElevenLabs Agents (default cascaded system of Scribe v2 Realtime / GPT-4o Mini / Eleven v3, with pre-registered tool schema) at 937. In reviewed conversations, highly preferred models tended to respond quickly, sound more natural and produce fewer unnatural sounds or audio artifacts β€ Task Success Rate: @SpaceXAI Grok Voice Think Fast 2.0 High leads at 94.7%, followed by @OpenAI GPT-Realtime-2.1 High at 91.5%, @ElevenLabs Agents (Default Cascaded System) at 90.5%, and GPT-Realtime-2 (High) at 89.8%, with GPT Realtime (Aug '25) and GPT-Realtime-2.1 Minimal tied at 89.4%. Gemini 3.1 Flash Live Preview - Minimal leads overall preference at 1,046 Elo but records a 74.6% Task Success Rate, showing that a preferred conversation does not always result in successful task completion - some conversations can sound as though the requested action was completed even when the required final tool call was unsuccessful We are continuing to expand our coverage of native and cascaded Speech to Speech systems, and welcome feedback as we add more models, providers and scenarios. See more details below β¬οΈ
tired: hill climbing an eval by doing a synth env wired: hill climbing an eval by fixing the eval