Turn AI failures into rerunnable evals and targeted fixes.
LLM evaluation
Only the most recent Revisions are kept. Restoring one saves its content again as the newest.