Learn
Learn testing: by watching and doing.
Video walkthroughs, a podcast, and a path that picks up where the last lesson left off.
Video lessons are coming
First video is in the works, walking through the regression harness from the featured article. Videos will live on YouTube.
Follow the channelThe podcast hasn’t launched yet
Conversations on testing and AI, recorded, not scripted. Episode one is being planned.
What’s coming
A public roadmap, not a progress barCatching LLM Hallucinations With a Regression Harness
Published: read it.
Write Your First LLM Eval
Published: read it.
Building a Golden Dataset for LLM Evals
Published: read it.
The LLM Test Pyramid
Published: read it.
Testing LangGraph Agents
Published: read it.
Building an LLM-as-Judge Eval Pipeline
Published: read it.
Red-Teaming Your Own Guardrails
Published: read it.
Precision, Recall, and Vibes: Evaluating RAG Pipelines
Published: read it.
Can Your LLM Catch Its Own Bugs? Mutation Testing Meets AI
Published: read it.
Chaos Monkey for Agents
Published: read it.
The Tokens Are the Budget
Published: read it.
Schema or It Didn’t Happen: Testing MCP Tool Definitions
Published: read it.
AI Wrote the Test Cases: Now Who Tests the Test Cases?
Published: read it.
Turn It Loose: Building an Exploratory Testing Agent
Published: read it.
Teaching Your Test Suite to Heal Itself
Published: read it.
The Flakiness Score: Finding Your Worst Tests With Statistics
Published: read it.
Semantic Flaky-Test Clustering: When Text Similarity Isn’t Enough
Planned.
Timeouts and Zombies: Testing Agents Against Hung Tool Calls
Planned.