@petergyang
“The fundamentals for [AI evals] still apply. Start by looking at real data. What has changed is getting agents to help you look at it in a thoughtful way.” Here’s my new episode with @sh_reya and @HamelHusain, who have taught AI evals to 4,500+ engineers and PMs. I asked them to audit the evals I built for my creator skills live. They then demoed a free skill that you can use in Claude Code or Codex to build reusable evals from your feedback. Some quotes from Shreya and Hamel: “Bottom-up evals come from looking at lots of sample outputs and turning that into eval criteria. AI is very bad at coming up with them. That’s all you.” “The agent’s job is not to invent new feedback. But it can help you group and distill the feedback into actionable rubric criteria.” “All your competitors can point Claude at their product and say, ‘Find all the errors.’ What matters is how much taste you can infuse beyond that.” 📌 Watch now: https://t.co/BuTklgfbr4 Thanks to our sponsors: @WisprFlow: 4x faster than typing with your voice https://t.co/oqHJ8bN3ll @linear: The AI agent platform for modern teams https://t.co/tgWf9oL4bs