Build and run evaluators across traces and experiments.
LLM evaluation
Only the most recent Revisions are kept. Restoring one saves its content again as the newest.