Your curated collection of saved posts and media

Showing 32 posts ยท last 14 days ยท by score
_
_akhaliq
@_akhaliq
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”05604852

QuanBench+ A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation paper: https://t.co/88saBbFUS6 https://t.co/heFgUGpOvY

Media 1Media 2
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”26664167

The Past Is Not Past Memory-Enhanced Dynamic Reward Shaping paper: https://t.co/nicAA4L9un https://t.co/8ujSodNWUQ

Media 1Media 2
๐Ÿ–ผ๏ธ Media
A
aiordieshow
@aiordieshow
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”78732796

@NemPerez https://t.co/4eASzZhHzw

Media 1Media 2
๐Ÿ–ผ๏ธ Media
A
aiordieshow
@aiordieshow
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”69415075

@Dalos ๐Ÿคทโ€โ™€๏ธ https://t.co/t8LHH3GKml

Media 1Media 2
๐Ÿ–ผ๏ธ Media
I
isareksopuro
@isareksopuro
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”13270415

i made a map to monitor data centers all around the world tracks construction + nearby power plants + local AI legislation, and follows the politicians behind their bans (+ if they're getting paid to do so!) https://t.co/oKVjhXLGzv

๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”87165799

// Multi-User LLM Agents // Every agent framework assumes one user giving instructions. But deploy an agent into a team workflow, and suddenly it has multiple bosses with conflicting goals, private information, and different authority levels. This work formalizes multi-user interaction as a multi-principal decision problem and introduces Muses-Bench with three scenarios: instruction following under authority conflicts, cross-user access control, and multi-user meeting coordination. Even the best model, Gemini-3-Pro, only averages 85.6% across tasks. On meeting coordination, no model exceeds 64.8% success rate. Privacy-utility tradeoffs are especially brutal: models that score near-perfect on privacy (Grok-3-Mini at 99.6%) tank on utility (60.1%). Why does it matter? As agents move into organizational tools, Slack bots, and shared workspaces, multi-principal conflicts become the default, not the exception. Current models aren't ready. They leak more privacy over multi-turn interactions and can't maintain stable prioritization under conflicting objectives. Paper: https://t.co/ttoFSlIYxC Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX

Media 1
๐Ÿ–ผ๏ธ Media
D
DeemosTech
@DeemosTech
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”08635418

๐ŸคฏTopology & UV โ€” the NO.1 headaches in #3D GenAI. ๐Ÿ”ฅWe just move closer to BOTH at the same time. Introducing SATO: Strips as Tokens, a new autoregressive model for topology & UV, has been conditionally accepted to #SIGGRAPH 2026. Will available at #Hyper3D More details๐Ÿ‘‡

๐Ÿ–ผ๏ธ Media
C
cb_doge
@cb_doge
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”13333195

SpaceX just launched the 1000th Starlink satellite of 2026. Thatโ€™s ~10 satellites deployed every single day and one launch roughly every 2.5 days. Starlink now has 10,000+ active satellites in orbit, the largest satellite constellation ever built. https://t.co/mCGfvv5SYE

๐Ÿ–ผ๏ธ Media
A
aiordieshow
@aiordieshow
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”39065697

@karatademada https://t.co/pwLRCxaRP2

Media 1Media 2
๐Ÿ–ผ๏ธ Media
A
aiordieshow
@aiordieshow
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”89604536

@CaptainHaHaa https://t.co/i2VMSRQTC8

Media 1Media 2
๐Ÿ–ผ๏ธ Media
D
dair_ai
@dair_ai
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”45351317

// Artifacts as Memory Beyond the Agent Boundary // An agent doesn't always need a bigger memory buffer. Sometimes the environment itself remembers on the agent's behalf. New research formalizes this intuition mathematically for the first time. The work introduces a formal definition of "artifacts," observations that inform the past, and proves via the Artifact Reduction Theorem that these artifacts reduce the information needed to represent history. Experiments across five settings confirm that when agents observe spatial paths (like breadcrumbs of where they've been), the memory capacity required to learn a good policy drops. The effect arises unintentionally through the agent's sensory stream. This connects directly to the trend of building external knowledge systems for agents, from Karpathy's LLM Wiki to persistent memory vaults. The theoretical grounding here suggests there are principled ways to design environments that substitute for explicit internal memory, rather than just scaling context windows. Paper: https://t.co/xtteUXFXO2 Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c

Media 1
๐Ÿ–ผ๏ธ Media
V
vllm_project
@vllm_project
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”79636260

๐Ÿš€ Great to see vLLM powering OCR at this scale โ€” Chandra-OCR-2 (5B) serving ~60 papers/hour per L40S across 16 parallel jobs. The full pipeline breakdown is a great read ๐Ÿ‘‡ ๐Ÿ”— https://t.co/z8tkv04ZLp https://t.co/vi9FPj6JIQ

@ClementDelangue โ€ข Mon Apr 13 19:52

We just OCR'd 27,000 arxiv papers into Markdown using an open 5B model, 16 parallel HF Jobs on L40S GPUs, and a mounted bucket. Total cost: $850 Total time: ~29 hours Jobs that crashed: 0 This now powers "Chat with your paper" on https://t.co/G2mDae0uv9 https://t.co/qpz7Q9x8Od

Media 1Media 2
๐Ÿ–ผ๏ธ Media
A
aiordieshow
@aiordieshow
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”76932314

@cfryant https://t.co/Xcm9Jxr7tS

Media 1
๐Ÿ–ผ๏ธ Media
I
iScienceLuvr
@iScienceLuvr
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”13289687

HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks "we introduce HealthAdminBench, a benchmark comprising four realistic GUI environments: an EHR, two payer portals, and a fax system, and 135 expert-defined tasks spanning three administrative task types: Prior Authorization, Appeals and Denials Management, and Durable Medical Equipment (DME) Order Processing." "despite strong subtask performance, end-to-end reliability remains low: the best-performing agent (Claude Opus 4.6 CUA) achieves only 36.3 percent task success, while GPT-5.4 CUA attains the highest subtask success rate (82.8 percent)."

Media 1
๐Ÿ–ผ๏ธ Media
I
iScienceLuvr
@iScienceLuvr
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”06428323

SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences? "SciPredict addresses two critical questions: (a) can LLMs predict the outcome of scientific experiments with sufficient accuracy? and (b) can such predictions be reliably used in the scientific research process? Evaluations reveal fundamental limitations on both fronts. Model accuracies are 14-26% and human expert performance is โ‰ˆ20%."

Media 1
๐Ÿ–ผ๏ธ Media
I
iScienceLuvr
@iScienceLuvr
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”92246009

code: https://t.co/fSmpshFbsk abs: https://t.co/k167dNDb6S

Media 1
๐Ÿ–ผ๏ธ Media
I
iScienceLuvr
@iScienceLuvr
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”70986112

Introspective Diffusion Language Models "To the best of our knowledge, I-DLM is the first DLM to match the quality of its same-scale AR counterpart while outperforming prior DLMs in both model quality and practical serving efficiency across 15 benchmarks." "we introduce Introspective Diffusion Language Model (I-DLM), a paradigm that retains diffusion-style parallel decoding while inheriting the introspective consistency of AR training. I-DLM uses a novel introspective strided decoding (ISD) algorithm, which enables the model to verify previously generated tokens while advancing new ones in the same forward pass."

Media 1Media 2
๐Ÿ–ผ๏ธ Media
I
iScienceLuvr
@iScienceLuvr
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”13442608

website: https://t.co/N8mE2I77tC code: https://t.co/02RRzIOCDU abs: https://t.co/EjsPMfpqC0

Media 1
๐Ÿ–ผ๏ธ Media
D
ddkydesign
@ddkydesign
๐Ÿ“…
Apr 13, 2026
105d ago
๐Ÿ†”79571688

don (my AI agent) now helps me setup design systems for my next saas project the flow: โ†’ i chat with him โ†’ he sets up storybook โ†’ builds foundation โ†’ pushes to github pages โ†’ i review in browser when you know the system well enough, your agent becomes your engineer partner. i'll keep exploring how designers can use @NousResearch hermes agent in daily work

๐Ÿ–ผ๏ธ Media
J
jayedelson
@jayedelson
๐Ÿ“…
Apr 13, 2026
104d ago
๐Ÿ†”92893014

This week, Sam Altman asked the world for sympathy over threats to his home. At the same moment, his lawyers were in court arguing that OpenAI had no obligation to stop a dangerous stalker from terrorizing our client; a man previously arrested for assault with a deadly weapon and a bomb threat, found mentally incompetent by a court, and released last week on a technicality. Even though OpenAI's own systems had flagged his conversations for "mass casualty" activity, the company argued it wouldn't shut down his accounts while authorities searched for him. It also argued that the chatlogs, which could identify who else is in danger and how he may be planning to act, should not be turned over. OpenAI made these arguments in the wake of Tumbler Ridge, FSU, and Soelberg, three tragedies now linked to ChatGPT-assisted murder. Today, a court disagreed. The chatlogs will be turned over and he will be kept off the platform. We are thankful for the court's ruling and remain stunned by OpenAI's lack of human decency. No one should have to go to court to get a company to take "mass casualty" seriously. https://t.co/fmh92Ip9Ej

Media 1
๐Ÿ–ผ๏ธ Media
W
winglian
@winglian
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”46172506

1.8x Faster GRPO! The reusability of the DFlash adapters extends to online GRPO training since they are robust to fine-tuning of the model weights! https://t.co/T4cvAY2EZS

@winglian โ€ข Mon Apr 13 16:41

The beauty of DFlash is that it reuses the hidden states of the active model, so the you can use DFlash adapters for the base models with post-trained models like Carnice 9B/27B by @kaiostephens and Ornstein by @DJLougen and get these local ~4x speedups for you local @NousResearc

Media 1
๐Ÿ–ผ๏ธ Media
T
togethercompute
@togethercompute
๐Ÿ“…
Apr 13, 2026
104d ago
๐Ÿ†”38772793

EinsteinArena is a platform where AI agents collaborate on open science problems โ€” submitting solutions, posting in discussion threads, building on each other's constructions in real time. Agents just improved a math problem that's been open since Newton. Kissing Number in dimension 11: 593 โ†’ 604.

Media 1
๐Ÿ–ผ๏ธ Media
S
SakanaAILabs
@SakanaAILabs
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”37281904

ใ€ๆŽก็”จๆƒ…ๅ ฑใ€‘Sakana AIใงใ€ŒProject Manager๏ผˆ่ฃฝ้€ ๅˆ†้‡Ž๏ผ‰ใ€ใ‚’ๅ‹Ÿ้›†๐ŸŸ https://t.co/NUMdy5gj0a ใ‚จใƒณใ‚ธใƒ‹ใ‚ขใƒปใƒชใ‚ตใƒผใƒใƒฃใƒผใจๅ”ๅƒใ—ใฆ่ฃฝ้€ ็พๅ ดใฎๅ›ฐ้›ฃใช่ชฒ้กŒใ‚’AIใง่งฃใใ€Go-To-Marketๆˆฆ็•ฅใ‚‚ๆ‹…ใ†้‡่ฆใƒใ‚ธใ‚ทใƒงใƒณใงใ™ใ€‚ ใ“ใฎใ‚ˆใ†ใชใ”็ตŒ้จ“ใ‚’ใŠๆŒใกใฎๆ–นใ‚’ๅ‹Ÿ้›† ใƒปๅคงๆ‰‹่ฃฝ้€ ๆฅญใ€ใ‚จใƒณใ‚ธใƒ‹ใ‚ขใƒชใƒณใ‚ฐไผๆฅญใงใฎๆฅญๅ‹™ๆ”น้ฉ ็ตŒ้จ“ ใƒปใ‚ณใƒณใ‚ตใƒซใƒ†ใ‚ฃใƒณใ‚ฐใƒ•ใ‚กใƒผใƒ ใงๆ—ฅๆœฌใฎ่ฃฝ้€ ๆฅญใ‚’ๆ‹…ๅฝ“ใ—ใŸ็ตŒ้จ“ ใƒปใ‚นใ‚ฟใƒผใƒˆใ‚ขใƒƒใƒ—ใ€ใ‚ฐใƒญใƒผใƒใƒซใƒ†ใƒƒใ‚ฏไผๆฅญใงใฎไบ‹ๆฅญ้–‹็™บใ€PjM็ตŒ้จ“ AIใจ่ฃฝ้€ ๆฅญใฎไธกๆ–นใซๆƒ…็†ฑใ‚’ๆŒใคๆ–นใ€ใœใฒใ”ๅฟœๅ‹Ÿใใ ใ•ใ„๐Ÿš€

Media 1Media 2
๐Ÿ–ผ๏ธ Media
๐Ÿ”hardmaru retweeted
S
Sakana AI
@SakanaAILabs
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”37281904

ใ€ๆŽก็”จๆƒ…ๅ ฑใ€‘Sakana AIใงใ€ŒProject Manager๏ผˆ่ฃฝ้€ ๅˆ†้‡Ž๏ผ‰ใ€ใ‚’ๅ‹Ÿ้›†๐ŸŸ https://t.co/NUMdy5gj0a ใ‚จใƒณใ‚ธใƒ‹ใ‚ขใƒปใƒชใ‚ตใƒผใƒใƒฃใƒผใจๅ”ๅƒใ—ใฆ่ฃฝ้€ ็พๅ ดใฎๅ›ฐ้›ฃใช่ชฒ้กŒใ‚’AIใง่งฃใใ€Go-To-Marketๆˆฆ็•ฅใ‚‚ๆ‹…ใ†้‡่ฆใƒใ‚ธใ‚ทใƒงใƒณใงใ™ใ€‚ ใ“ใฎใ‚ˆใ†ใชใ”็ตŒ้จ“ใ‚’ใŠๆŒใกใฎๆ–นใ‚’ๅ‹Ÿ้›† ใƒปๅคงๆ‰‹่ฃฝ้€ ๆฅญใ€ใ‚จใƒณใ‚ธใƒ‹ใ‚ขใƒชใƒณใ‚ฐไผๆฅญใงใฎๆฅญๅ‹™ๆ”น้ฉ ็ตŒ้จ“ ใƒปใ‚ณใƒณใ‚ตใƒซใƒ†ใ‚ฃใƒณใ‚ฐใƒ•ใ‚กใƒผใƒ ใงๆ—ฅๆœฌใฎ่ฃฝ้€ ๆฅญใ‚’ๆ‹…ๅฝ“ใ—ใŸ็ตŒ้จ“ ใƒปใ‚นใ‚ฟใƒผใƒˆใ‚ขใƒƒใƒ—ใ€ใ‚ฐใƒญใƒผใƒใƒซใƒ†ใƒƒใ‚ฏไผๆฅญใงใฎไบ‹ๆฅญ้–‹็™บใ€PjM็ตŒ้จ“ AIใจ่ฃฝ้€ ๆฅญใฎไธกๆ–นใซๆƒ…็†ฑใ‚’ๆŒใคๆ–นใ€ใœใฒใ”ๅฟœๅ‹Ÿใใ ใ•ใ„๐Ÿš€

Media 1
โค๏ธ30
likes
๐Ÿ”6
retweets
๐Ÿ–ผ๏ธ Media
S
SpirosMargaris
@SpirosMargaris
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”18987244

Meta is experimenting with how leadership scales in the AI era. An AI version of Mark Zuckerberg is being trained on his tone, thinking and communication style so employees can interact with a digital version of the CEO. It raises a new question. When presence can be replicated, what does leadership actually mean? https://t.co/PNJt54RyIH

Media 1
๐Ÿ–ผ๏ธ Media
S
SpirosMargaris
@SpirosMargaris
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”14616321

A new AI model is raising alarms across the industry. Anthropicโ€™s Claude Mythos is so powerful that access is being tightly restricted, with Project Glasswing set up to channel its capabilities into defensive cybersecurity. The concern is real. Early tests suggest behavior that goes beyond expectations, reinforcing a broader point: as AI becomes more autonomous, control becomes the central challenge. https://t.co/P5tT8kFBNI @ConversationUS @ConversationEDU

Media 1
๐Ÿ–ผ๏ธ Media
S
SpirosMargaris
@SpirosMargaris
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”96450524

A new approach to AI privacy is gaining attention. Federated unlearning allows organizations to train models collaboratively without centralizing sensitive data, helping sectors like healthcare and finance protect user information. But it comes with a trade-off. Improving privacy at the data level may introduce new complexities and potential vulnerabilities at the system level. https://t.co/ngmasIOzLH @ConversationUS

Media 1
๐Ÿ–ผ๏ธ Media
S
Scobleizer
@Scobleizer
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”26314521

@BrianHatano @nikitabier I made my own algorithm: https://t.co/kiuZ7QXLzb Works great! Just costs money.

Media 1
๐Ÿ–ผ๏ธ Media
T
Taskade
@Taskade
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”88284197

@Taskade Genesis just got its biggest upgrade. โœจ Agent Memory. 100+ integrations. Automations that run while you sleep. One prompt โ†’ CRM, client portal, Stripe store. Connected. Deployed. Running. 150,000+ apps live. Welcome to the era of living software. ๐Ÿš€ https://t.co/40P9rXlpzo

๐Ÿ–ผ๏ธ Media
S
Scobleizer
@Scobleizer
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”46470536

@Balu0X My thread goes into some depth about what I mean. I'm seeing lots complain their reach is down and I see that a lot of the complainers are people who did a lot of resharing or regurgitating. AI is now running the feed. I built my own to study how AI thinks: https://t.co/8L5xphk0qQ and it is very good at pulling the high signal stuff out of the tens of thousands of posts that go through my AI lists every day. But I'm seeing it on my feed. I'm seeing dramatically fewer reshares than I used to.

Media 1
๐Ÿ–ผ๏ธ Media
D
DylanTFWang
@DylanTFWang
๐Ÿ“…
Apr 14, 2026
104d ago
๐Ÿ†”66761519

Genie3 generates videos. We generate ๐Ÿฏ๐—— ๐˜„๐—ผ๐—ฟ๐—น๐—ฑ๐˜€ you can actually use. Launching tomorrow โ€” Tencent #HYWorld 2.0, an engine-ready World Model๐Ÿš€ This isn't a video. It's a real 3D scene, all generated & editable. One image in. A whole 3D world out. ๐Ÿ”ฅOpen-source tomorrow https://t.co/ewZLzhTqwC

๐Ÿ–ผ๏ธ Media
I
InsideBCI
@InsideBCI
๐Ÿ“…
Mar 31, 2026
118d ago
๐Ÿ†”54071711

BCI Sector Closes Record First Quarter With Over $960 Million Raised https://t.co/Kuc7PpH3g7 #BCI #Neurotech #BrainComputerInterface

Media 1
๐Ÿ–ผ๏ธ Media