Your agent writes the code. It does not get to grade it.
Test Maze plugs into Claude Code, Cursor, Cline, Gemini CLI, Codex CLI or any MCP client. Your agent keeps writing code and running tests in your repo; Test Maze keeps the test cases, records every run against the exact commit and returns a verdict computed by fixed rules.
What goes wrong today
Self-grading drift
The agent edits the assertion until the test passes. The diff looks clean. The bug ships.
“Tests pass” on which commit?
A green summary from the agent says nothing about which working tree it ran against, or whether the tree was dirty.
Hidden flakiness
Retries hide real regressions. “Transient” failures pile up until the suite is unrunnable.
What changes
Deterministic verdicts
pdlc.verifyNo LLM in the pass/fail path. The same results always give the same verdict, with failure buckets and a reliability KPI your agent can act on.
Commit-pinned runs
testrun.create · testrun.record_resultsEvery verifier run carries gitSha, branch and workingTreeClean, so a green run means a specific commit was green.
A repair loop that points somewhere
nextStep: repair_code · repair_test · add_coverageOn failure the verdict says whether the code or the test is wrong and which case to look at first. No more guessing.
Code health on request
quality.smell_check · architecture.adviseA code-smell check with fixed rules, mapped to the refactorings that fix each smell, plus design-pattern advice ranked for your problem. Only the snippet you choose is sent.
Always on the right project
project.whoamiEach project folder keeps its own token in a git-ignored .env.testmaze, so two repos on one laptop never write into each other’s workspace.
How it works for developers
Best for: Developers
Hand over the ticket, the PRD or a one-liner.
You talk to your agent. It picks the tools.
Test Maze for the rest of your team
Ready to give your agent a verifier?
Create a free workspace, generate an MCP token and connect your coding agent in under five minutes.