Your curated collection of saved posts and media

Showing 32 posts ยท last 14 days ยท by score
Z
Zai_org
@Zai_org
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”04093443

GLM-5.1 Tool Calling Issue Fix & Chat Template Update If you are running GLM-5.1 with vLLM/SGLang and using tool calling, please update your chat template. https://t.co/XyyCucws82 Issue When using tool calling, frameworks including vLLM automatically convert plain-text tool message content into an array of content parts (`[{"type": "text", "text": "..."}]`) before passing it to the chat template. The original template only supported string-formatted tool content, causing array-formatted tool outputs to render empty. As a result, the model does not receive tool results and repeatedly triggers the same tool call in a loop. Affected Models All GLM-5.1 variants deployed with vLLM or SGLang. Fix Simply replace your existing `chat_template.jinja` with the updated version from the repository.

Media 1
๐Ÿ–ผ๏ธ Media
_
_NathanCalvin
@_NathanCalvin
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”26109585

This part of the 4.7 Opus system card is pretty neat and seems potentially worth emulating (Anthropic showed Mythos the private discussions/evidence underlying the system card and asked Mythos if the Opus system card accurately characterized that private evidence) https://t.co/4pf666ZB6m

@MaskedTorah โ€ข Thu Apr 16 14:46

@TheZvi Less juicy overall than last time, but I was happy we got to fit in section 6.1.3: https://t.co/8KYbb2DKgX

Media 1
๐Ÿ–ผ๏ธ Media
A
arankomatsuzaki
@arankomatsuzaki
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”28186936

Nearly 1/3 of surveyed people in Anthropic now think entry-level engineers and researchers are likely replaced by Mythos within 3 months https://t.co/QUozBxLUrR

Media 1
๐Ÿ–ผ๏ธ Media
Y
Yuchenj_UW
@Yuchenj_UW
๐Ÿ“…
Apr 12, 2026
105d ago
๐Ÿ†”71128202

This is really bad. The scary part in the US is that it doesnโ€™t matter whether you are the CEO of OpenAI or just a regular PhD. There are paid online websites that can find your address and phone number. I donโ€™t know how such personal info got out. https://t.co/jyDMhyXIZC

Media 1
๐Ÿ–ผ๏ธ Media
S
simonsarris
@simonsarris
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”01857015

why do the Japanese like their buns askew? https://t.co/6be3BqSJ7d

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”72637426

GameWorld Towards Standardized and Verifiable Evaluation of Multimodal Game Agents paper: https://t.co/IfbTgfNnSM https://t.co/gL3BURxzkV

๐Ÿ–ผ๏ธ Media
L
leftcurvedev_
@leftcurvedev_
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”17564814

New insane model from Jackrong on @huggingface ๐Ÿคฏ Qwen3.5-9B-GLM5.1-Distill-v1 ๐Ÿง  Distilled on GLM-5.1 reasoningโ€จโš™๏ธ Deeper thinking than base modelโ€จ๐Ÿงช Benchmarks coming soon โœ… Fits on 8GB VRAM โœ๏ธ New model after Qwopus/Gemopus After distilling Claude Opus 4.6, heโ€™s now back on the strongest open-source model! An MLX ๏ฃฟ version is also available on his huggingface page 27B model incoming? https://t.co/893FvH51jb

Media 1
๐Ÿ–ผ๏ธ Media
P
perplexity_ai
@perplexity_ai
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”85454518

Today we're releasing Personal Computer. Personal Computer integrates with the Perplexity Mac App for secure orchestration across your local files, native apps, and browser. Weโ€™re rolling this out to all Perplexity Max subscribers and everyone on the waitlist starting today. https://t.co/kxgFQFo7BB

๐Ÿ–ผ๏ธ Media
S
Scobleizer
@Scobleizer
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”34309517

@vicberggren Whew. I'm not gonna stop doing them anyway. Even if it costs me a few hundred bucks a month. My AI that builds https://t.co/kiuZ7QXLzb learns from my reshares what to look for.

Media 1
๐Ÿ–ผ๏ธ Media
H
HuggingPapers
@HuggingPapers
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”99577997

Qwen just released Qwen 3.6 on Hugging Face A 35B MoE vision-language model with 3B active parameters, featuring advanced agentic coding capabilities and thinking preservation. https://t.co/NmwjwxyN3m

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”77232166

Parcae Scaling Laws For Stable Looped Language Models paper: https://t.co/hUYU2x8STk https://t.co/ravMv2kbR3

Media 1
๐Ÿ–ผ๏ธ Media
_
_akhaliq
@_akhaliq
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”46286353

Geometric Context Transformer for Streaming 3D Reconstruction paper: https://t.co/3ad6iyi0cG https://t.co/Y0k4csiC11

Media 1
๐Ÿ–ผ๏ธ Media
E
emollick
@emollick
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”88130992

Claude remains irreducibly Claude. If you know, you know. (The fact that models have distinct personalities that are consistent across generations is technically interesting, it also makes it very easy to use new releases when they come along, because they feel very similar). https://t.co/imyGcPsYBI

Media 1
๐Ÿ–ผ๏ธ Media
A
Alibaba_Qwen
@Alibaba_Qwen
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”34243427

โšก Meet Qwen3.6-35B-A3B๏ผšNow Open-Source๏ผ๐Ÿš€๐Ÿš€ A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. ๐Ÿ”ฅ Agentic coding on par with models 10x its active size ๐Ÿ“ท Strong multimodal perception and reasoning ability ๐Ÿง  Multimodal thinking + non-thinking modes Efficient. Powerful. Versatile. Try it now๐Ÿ‘‡ Blog๏ผšhttps://t.co/EXx5y466su Qwen Studio๏ผšhttps://t.co/bg4tAU1p74 HuggingFace๏ผšhttps://t.co/w4pDX14DZS ModelScope๏ผšhttps://t.co/SuRyLzdQiO API๏ผˆโ€˜Qwen3.6-Flashโ€™ on Model Studio๏ผ‰๏ผšComing soon๏ฝž Stay tuned

Media 1Media 2
+2 more
๐Ÿ–ผ๏ธ Media
J
johnjnay
@johnjnay
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”74021826

today we launched the Legal AGI Lab. AI agents are beginning to operate in highly regulated environments like healthcare and financial services. but existing legal frameworks arenโ€™t ready for this. And this creates a bottleneck for the agentic economy. so we are conducting interdisciplinary legal & AI research on how agents should be governed, held liable, and measured and defining the legal architecture required for autonomous agents to operate safely in high-stakes environments. Norm sits at a unique intersection: we build AI agents, we deploy them with institutional clients, and we power Norm Law, an AI-native law firm operating on live legal work. that feedback loop between building, testing, and deploying is what makes this research different.

Media 1
๐Ÿ–ผ๏ธ Media
G
gerardsans
@gerardsans
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”74106489

@_winter_wonders A frequently cited example is OpenAIโ€™s Sora, which reportedly incurred extremely high compute costs, widely estimated in the range of up to around $1 million per day, while struggling to match that with sustainable revenue. The product has since been pulled back, often framed as a mismatch between cutting-edge capability and viable unit economics. This connects to the broader pattern in AI commercialization, including the earlier cybersecurity discussion: significant spending is increasingly justified under โ€œsafety,โ€ โ€œrisk,โ€ or โ€œcapabilityโ€ narratives, even when the underlying economic returns remain uncertain. Source: https://t.co/VTYh7hipAY

Media 1
๐Ÿ–ผ๏ธ Media
O
omarsar0
@omarsar0
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”71044536

What you need to know about Opus 4.7 * Takes instructions literally * Better vision means improved computer use and producing slides and other visual artifacts * Optimized for large-scale real-world analysis * Better at using file system-based memory https://t.co/tEywxsCxSV

@claudeai โ€ข Thu Apr 16 14:29

Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verifies its own outputs before reporting back. You can hand off your hardest work with less supervision. https://t.co/PtlRdpQcG

Media 1
๐Ÿ–ผ๏ธ Media
K
KenRoth
@KenRoth
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”06545014

Viktor Orbรกnโ€™s electoral loss in Hungary is as much a defeat for Trump and JD Vance. "Seldom have American leaders intervened so overtly in a foreign election, and seldom has their preferred candidate fared so badly." https://t.co/DbLhMz65ja

Media 1
๐Ÿ–ผ๏ธ Media
P
PyTorch
@PyTorch
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”33842758

๐ŸŽค Take the stage at #PyTorchCon North America! We are looking for technical deep dives & production stories for our return to San Jose this Oct 20-21. Check out our "Preparing to Submit" guide to help craft your proposal. ๐Ÿ—“๏ธ Deadline: June 7 Apply now: https://t.co/hLlKK7WxLD https://t.co/leYJj7nDfR

๐Ÿ–ผ๏ธ Media
G
github
@github
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”81161744

๐Ÿ†• @AnthropicAI's Claude Opus 4.7 is now generally available and rolling out in GitHub Copilot. Early testing shows โžก๏ธ It has stronger multi-step task performance and more reliable agentic execution โžก๏ธ Meaningful improvement in long-horizon reasoning and complex workflows Try it out in @code or Copilot CLI. https://t.co/8QFLkf0RqR

๐Ÿ–ผ๏ธ Media
T
tom_sachs
@tom_sachs
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”62987030

โ€œFurnitureโ€ opens at Salon 94 NYC Thursday, April 23, 2026 from 7-9pm EDT. See you there! On view from April 23 - June 20, 2026 Salon 94: 3 E 89th St, New York, NY 10128 Hours: Wednesday - Saturday 11:00am - 6:00pm EDT https://t.co/N3MIlW3bv3

Media 1Media 2
๐Ÿ–ผ๏ธ Media
M
MarioNawfal
@MarioNawfal
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”42060554

๐Ÿ‡ฟ๐Ÿ‡ฆ A senior South African politician just got caught lying about Starlink to protect mobile network operators. Parliamentary communications chair Khusela Diko claimed Starlink "doesn't move the needle" on school connectivity. Here's what she left out: 16,000 schools still have no internet after 12 years, a missed deadline, and mobile operators billions over budget. Starlink offered to connect 5,000 schools for free and was turned away. A rural mobile tower costs around $61,000 and can serve just a single school. The numbers don't lie. The politician did. Source: MyBroadband

@elonmusk โ€ข Thu Apr 16 07:35

Accurate

Media 1Media 2
๐Ÿ–ผ๏ธ Media
C
claudeai
@claudeai
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”93977612

Introducing Claude Opus 4.7, our most capable Opus model yet. It handles long-running tasks with more rigor, follows instructions more precisely, and verifies its own outputs before reporting back. You can hand off your hardest work with less supervision. https://t.co/PtlRdpQcG5

Media 1
๐Ÿ–ผ๏ธ Media
L
ltx_model
@ltx_model
๐Ÿ“…
Apr 10, 2026
107d ago
๐Ÿ†”14859378

T-7 days ๐Ÿ‡ซ๐Ÿ‡ท Open-source AI Art takes over Paris. 3 days. Hackathons. Art. Talks. 120 spots per day. https://t.co/f9RpcI3Cs5

Media 1
๐Ÿ–ผ๏ธ Media
๐Ÿ”ylecun retweeted
L
LTX
@ltx_model
๐Ÿ“…
Apr 10, 2026
107d ago
๐Ÿ†”14859378

T-7 days ๐Ÿ‡ซ๐Ÿ‡ท Open-source AI Art takes over Paris. 3 days. Hackathons. Art. Talks. 120 spots per day. https://t.co/f9RpcI3Cs5

Media 1
โค๏ธ228
likes
๐Ÿ”31
retweets
๐Ÿ–ผ๏ธ Media
W
whoiskatrin
@whoiskatrin
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”40225228

cloudflare just gave agents git this is one of those changes that will just quietly improve everything agents with proper version control @dillon_mulroy @elithrar @mattzcarey @thomas_ankcorn have done something incredible here https://t.co/4dFPie896A

Media 1
๐Ÿ–ผ๏ธ Media
D
dok2001
@dok2001
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”09510421

Email is older than the web. Itโ€™s also the best interface for agents. Everyone already has an address. No app to install. No SDK to integrate. Cloudflare Email Service is now in public beta. https://t.co/NSHGfuwM0o

Media 1
๐Ÿ–ผ๏ธ Media
M
MayukhBagchi4
@MayukhBagchi4
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”43508155

@Teknium @MiniMax_AI Just shipped a Hermes Agent skill with HuskyLens V2 on a Pi 5 gesture control, face recognition, emotion reading, all local. MiniMax M2.7 as the brain would be wild for this. What's the target hardware? https://t.co/UuBtK7YIyO

Media 1
๐Ÿ–ผ๏ธ Media
T
Teknium
@Teknium
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”18058279

@pepethehandsome @vmvarg4 @maxkolysh @browser_use Oh, it does look like it should lol, but I dont know If its another issue you can run /debug and dm me your logs https://t.co/W0gFhGdhCC

Media 1
๐Ÿ–ผ๏ธ Media
D
dair_ai
@dair_ai
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”14322393

Agent evals are drifting away from production reality. Most benchmarks use clean tasks, well-specified requirements, deterministic metrics, and retrospective curation. Production work is messier, with implicit constraints, fragmented multimodal inputs, undeclared domain knowledge, long-horizon deliverables, and expert judgment that evolves over time. This paper introduces AlphaEval, a production-grounded benchmark for evaluating agents as complete products. AlphaEval contains 94 tasks sourced from seven companies deploying AI agents in core business workflows, spanning six O*NET domains. It evaluates systems like Claude Code and Codex as commercial agent products, not just model APIs. The benchmark combines multiple evaluation paradigms: LLM-as-a-Judge, reference-driven metrics, formal verification, rubric-based assessment, automated UI testing, and domain-specific checks. Why it matters: organizations need benchmarks that start from real production requirements, then become executable evals with minimal friction. Paper: https://t.co/cbTGgTWoNl Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c

Media 1
๐Ÿ–ผ๏ธ Media
L
llama_index
@llama_index
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”91019571

LiteParse hit 4.3K+ GitHub stars in a few weeks. Today it officially joins the LlamaIndex ecosystem, with its own page at https://t.co/1tdQbEer9H. ~500 pages in 2 sec. 50+ formats. Zero cloud dependency. Already powering agents in Claude Code, Cursor, and production pipelines. In a few days out head of OSS, @LoganMarkewich, is hosting a live workshop: build a fintech due diligence agent with LiteParse โ†’ https://t.co/x0uMc0gR8O

๐Ÿ–ผ๏ธ Media
F
friesmakesfries
@friesmakesfries
๐Ÿ“…
Apr 16, 2026
101d ago
๐Ÿ†”41802481

hermes agent @NousResearch is fucking insane i know literally NOTHING about coding. ZERO. and i just built a fully functioning web app in minutes http://localhost:3000/ check it out @Teknium https://t.co/H0uvfhoNX5

Media 1
๐Ÿ–ผ๏ธ Media