Your curated collection of saved posts and media

Showing 10 posts Β· last 14 days Β· by score
βž• Add New Post
D
dex
@dexhorthy
πŸ“…
Aug 25, 2026
7d ago
πŸ†”01363518

You can’t claim β€œthe models are good enough that I don’t have to read the code”. Because if you’re not reading the code, then you can’t possibly know how much slop is getting in. We go live to @0xblacklight and @vaibcode for the scoop https://t.co/P0bVoM7L45

❀️29
likes
πŸ”4
retweets
πŸ–ΌοΈ Media
T
Jesse Zhang
@thejessezhang
πŸ“…
Aug 28, 2026
5d ago
πŸ†”31638922
⭐0.36

One of the trickiest problems you run into with voice AI in the real world is that people often call in messy environments. There might be other voices in the background, or the person might be talking to other people at the same time. Super cool work from Decagon Labs on how we've tackled this problem!

@DecagonAI β€’ Thu Aug 27 22:38

Conversational voice agents need to know when a different person is speaking, but not every background voice should count. We combined speaker embeddings with a post-trained audio-language model to determine when a speaker change matters. https://t.co/Cg5HKtyaUf

❀️27
likes
πŸ”3
retweets
H
Hugging Apps
@HuggingApps
πŸ“…
Aug 26, 2026
6d ago
πŸ†”98614848

Breeze TTS 2 by @BreezeBlueX dropped on Hugging Face as the #1 open weights TTS model on @ArtificialAnlys It does 🎨 Voice Design (prompt a voice) πŸŽ™οΈ Voice Reference (upload a voice) πŸŽ›οΈ Voice Direction (upload & prompt a voice) vibe it on Spaces ▢️ https://t.co/MCKL1l4y6R https://t.co/QTgLO50JtD

@ArtificialAnlys β€’ Tue Aug 25 23:51

Breeze TTS 2 is now the leading Open Weights TTS model in the Artificial Analysis Provider Voices Speech Arena, surpassing Fish Audio S2 Pro by 90 Elo points Breeze TTS 2 is the latest TTS model from @BreezeBlueX, supporting 50 languages, voice generation from text prompts, and

Media 2
❀️21
likes
πŸ”3
retweets
πŸ–ΌοΈ Media
D
Dawn Song
@dawnsongtweets
πŸ“…
Aug 31, 2026
1d ago
πŸ†”52099710

Introducing CUA-Lite 🧡 β€” an open platform for computer-use agents. Training and benchmarking CUAs (Computer-Use Agents) requires four core pieces: 1️⃣ Agents β€” the models and the scaffolding that drives them 2️⃣ Environments β€” runtime/sandboxes for agents to interact with, tasks & verifiers/graders 3️⃣ Traces β€” records of agent trajectories 4️⃣ Frameworks β€” to evaluate, SFT & RL-train agents Today, all four are fragmented. Every agent ships with its own implementation, often in a separate repo β€” there is no unified way to run them all. Every environment exposes its own interface and action space, often requiring an expensive VM sandbox for each verifiable task. Traces come in incompatible formats. And without common standards across the stack, every project ends up rebuilding its own tooling/framework for eval, SFT, and RL. CUA-Lite unifies the stack: β†’ One standardized interface & action space for agents and environments β†’ One standardized format for agent traces β†’ One framework for evaluation, SFT & RL β†’ Across desktop, browser & mobile And open resources plug straight in, creating the largest open collection of CUA agents, environments and traces, all in a unified format: πŸ€– 10+ CUAs, including GPT, Claude, Gemini, Qwen, Muse-Glimmer, UI-TARS 🌐 15+ benchmarks, including OSWorld, WebArena & AndroidWorld ⚑ Optional VM-free sandboxes with 30K+ verifiable tasks for training πŸ“š 10+ trace datasets, freely available on Hugging Face, including public datasets converted into the standardized format and fresh rollouts from frontier open-weight CUAs Led by @BerkeleyRDI , our goal is for CUA-Lite to become a community-driven, open-source ecosystem for computer-use agents. Join the community and contribute today: bring an environment (runtime/sandbox + tasks + verifier), traces, or an agent, and plug it into CUA-Lite!

Media 1
❀️21
likes
πŸ”7
retweets
πŸ–ΌοΈ Media
Y
Yunzhu Li
@YunzhuLiYZ
πŸ“…
Sep 01, 2026
9h ago
πŸ†”49790103
⭐0.34

Another glimpse of Real2Sim with Atlas: capture a real environment with just a number of casual photos from a phone, turn them into a sim, then simulate diverse robots navigating different trajectories. Atlas generates the RGB and depth observations they would see along the way.

@theworldlabs β€’ Tue Sep 01 17:28

For robotic simulation, Atlas reconstructs a space from just a few photos and generates the photorealistic RGB and depth data any robot's sensors would observe on any trajectory. Robots can now be trained and tested in far more spaces. Until now, scanning spaces like these requ

❀️21
likes
πŸ”1
retweets
πŸ”Nader Khalil🍊 retweeted
B
Ben Pouladian
@benitoz
πŸ“…
Aug 25, 2026
7d ago
πŸ†”57125342
⭐0.32

The Perplexity team and @NaderLikeLadder from NVIDIA demoed Computer for me during my lunch break at Hot Chips. This makes DGX Spark a genuinely easy-to-use, useful agentic workstation. Excited to plug it into my own workflows and see what it can do. (Excuse the etched hat it was sunny 🀣) Speed of light $NVDA

❀️20
likes
πŸ”3
retweets
B
Ben Pouladian
@benitoz
πŸ“…
Aug 25, 2026
7d ago
πŸ†”57125342

The Perplexity team and @NaderLikeLadder from NVIDIA demoed Computer for me during my lunch break at Hot Chips. This makes DGX Spark a genuinely easy-to-use, useful agentic workstation. Excited to plug it into my own workflows and see what it can do. (Excuse the etched hat it was sunny 🀣) Speed of light $NVDA

@NaderLikeLadder β€’ Tue Aug 25 15:16

2 months ago, we crashed Jensen's board meeting to show him Perplexity running locally on DGX Spark 🀣 (lol im holding architecture diagrams like he's gonna look at em during a board meeting) Local AI used to be for enthusiasts. Folks were running tiny quantized models on underp

Media 1
❀️20
likes
πŸ”3
retweets
πŸ–ΌοΈ Media
K
knut
@kmelve
πŸ“…
Aug 19, 2026
13d ago
πŸ†”97360578

sat down and actually read @rosmine's research paper on this. https://t.co/vxqTirXxrS first of all, i was guilty of having a "take" based on this post, pointing out how the copy in the screenshot is not a great example of what "good writing" is, despite scoring 100 on "human written" in @pangram. I recommend actually reading the work before commenting on it - because sloppy takes are as lazy as sloppy texts. (shame on me for joining the band wagon) but sleeping on it, i am grateful that @rosmine took the time to dive into this stuff. AI-slop fatigue is real, and we should support efforts to make agents produce communication that is clear and lucid, and doesn't feel overly synthetic but there are some assumptions in this, however, that is interesting to question when it comes to "what makes for good writing" as far as i understand, Deft is trained to have more diversity/variation in tokens, so less repetition of what we are recognizing as "AI-tells" (certain word and stylistic choices). And i totally agree with @rosmine that: "LLMs are not the cause of slop. Lack of effort/care is. If you spend days researching and planning a blog post, and put all the information into a detailed, well-structured outline, and ask ChatGPT to generate the post based on the outline, then the output will be interesting to read, even if the text has a lot of em-dashes." I argued the same in our eng blog announcement post yesterday: https://t.co/sV3cEP4S9G BUT! I still feel that this report (at least somewhat), but especially the various takes on it, conflates something sounding "human" with it being "good." Tricking @pangram doesn't make a text well written. Some reflections: - Making writing sound more "human" by means of adding more variation in word/style choices, doesn't make it better - What makes for a "good" text is highly contextual. If you are writing a recipe or instructions, repetition and stylistic stringency is highly desirable - AI-tells cuts deeper than just word choices, observant readers will start to be sensitive to the lack of certain devices, ways of arguing, structure, etc. - Interesting writing often comes from doing synthesis of unexpected things in a way that bring clarity. I have still to see that from LLMs as we usually interact with them (but it's probably possible to get them to do this). with that being said - i'm grateful that this research was shared with the wider community, and will applaud any effort to make the user experience of interacting with agents, and the stuff we make with agents, better. 🫑

@rosmine β€’ Tue Aug 18 18:13

Announcing Deft, a new AI lab for better writing, cofounded with @jmrphy See the picture for launch announcement the Deft model wrote for itself Currently, 86% of user queries are fully human according to pangram. This is still a small beta model and it might make mistakes. We

Media 1Media 2
❀️20
likes
πŸ”1
retweets
πŸ–ΌοΈ Media
M
Modular
@Modular
πŸ“…
Aug 24, 2026
8d ago
πŸ†”96199184

Modular's price-performance on @Zai_org's GLM-5.2 (Non-reasoning) lands right on the Pareto frontier in @ArtificialAnlys' latest benchmark: near-top speed without the near-top price tag. We're just getting started, and we're ready for GLM-5.3. Expect to see a lot more incredible results. πŸš€

Media 1
❀️19
likes
πŸ”2
retweets
πŸ–ΌοΈ Media
P
Peter Yang
@petergyang
πŸ“…
Aug 23, 2026
9d ago
πŸ†”41491820

β€œThe fundamentals for [AI evals] still apply. Start by looking at real data. What has changed is getting agents to help you look at it in a thoughtful way.” Here’s my new episode with @sh_reya and @HamelHusain, who have taught AI evals to 4,500+ engineers and PMs. I asked them to audit the evals I built for my creator skills live. They then demoed a free skill that you can use in Claude Code or Codex to build reusable evals from your feedback. Some quotes from Shreya and Hamel: β€œBottom-up evals come from looking at lots of sample outputs and turning that into eval criteria. AI is very bad at coming up with them. That’s all you.” β€œThe agent’s job is not to invent new feedback. But it can help you group and distill the feedback into actionable rubric criteria.” β€œAll your competitors can point Claude at their product and say, β€˜Find all the errors.’ What matters is how much taste you can infuse beyond that.” πŸ“Œ Watch now: https://t.co/BuTklgfbr4 Thanks to our sponsors: @WisprFlow: 4x faster than typing with your voice https://t.co/oqHJ8bN3ll @linear: The AI agent platform for modern teams https://t.co/tgWf9oL4bs

Media 2
❀️18
likes
πŸ”3
retweets
πŸ–ΌοΈ Media