regress
Regression testing for prompt changes

Before you ship a prompt change, know what breaks.

Save the real inputs and outputs you keep re-checking by hand. Paste a baseline and a candidate output side by side, mark pass or fail, and get an explicit improved / same / regressed verdict per case — instead of another spreadsheet tab or scrollback search through old chats.

01

Save a case

Real input + what a good answer looks like.

02

Compare outputs

Paste baseline vs. candidate, mark pass/fail.

03

See the verdict

Improved, same, or regressed — saved to history.

case — refund policy question
Baseline — v1 promptfail

Sorry, we can't help with that. Refunds aren't available at this point.

Candidate — v2 promptpass

I'm sorry for the trouble. Our refund window is 30 days, so this is outside it, but I can offer store credit.

Case verdictImproved

Illustration of the comparison workflow — not live data.

Built for the change you're about to make

Cases outlive the chat tab

Every input, expected behavior, and note is saved per project so the next prompt change starts from evidence, not memory.

No model connection required

Paste outputs from whatever you're using today. Mark each side pass or fail and the tool derives the verdict.

History you can return to

Every comparison run is kept per case, so you can see exactly when a case started failing.