OpenEvalv0.5.3Documentation
GitHub72
START

Concepts & terminology

A metric measures what happened. A criterion defines what earns credit. A score is the credit awarded.

On this page

One vocabulary

A benchmark contains evals. Each eval defines a task and a rubric. Judges produce scores for the rubric's criteria. Runs also record metrics such as cost, tokens, and tool reliability.

TermMeaningExample
BenchmarkA collection of evals and their run configuration.A set of agent tasks evaluated across models.
SuiteAn optional display label for one benchmark among several that share a name. It never changes scoring.Frontier and Regressions, both named sql-bench.
EvalA task, its input and environment, and its grading specification.Produce a parameterized SQL query.
CriterionA named requirement being graded. Plural: criteria.safe_parameters
ScoreCredit awarded to a criterion, normalized from 0 to 1, or an aggregate of that credit.1 for full credit; 0 for no credit.
MetricAn observed or calculated measurement. It contributes to grading only when a criterion uses it.Cost in USD, token count, or tool error rate.
RubricThe criteria and rules for awarding scores.The conditions for accepting a parameterized query.
JudgeAn evaluator that applies grading rules using code, an LLM, or both.An LLM applying a written rubric.
JudgmentA judge's output, including any named criterion scores and supporting data.A safe_parameters score with evidence.
BenchmarkRunA recorded benchmark collection with its inputs and selected results.One retained comparison across models.
EvalRunOne candidate execution of an eval.A model's second repetition of the SQL task.
JudgeRunA grading execution against recorded evidence.A new judgment of a retained EvalRun.

Measurements become credit through a rule

RoleExampleMeaning
Metriccost_usd = 0.84An observed or calculated cost in USD.
Criterionwithin_budgetAward full credit when cost_usd is at most 1.00.
Criterion scorewithin_budget = 1Credit awarded by applying that rule.

A tool error rate of 0.2 is a measurement even though it falls between 0 and 1. Recording a metric does not automatically award or deduct credit. The author decides whether a criterion uses it.

Criterion IDs are author-defined. A name such as correct_answer has no built-in grading rule. A rubric explains how evidence becomes a score; a judgment records the result of applying it.

Definitions, executions, and evidence

An Eval is a reusable task definition; an EvalRun is one candidate execution. A JudgeRun applies grading rules to recorded evidence. A completed rejudge updates the active selection while retaining earlier judgments.

TermUse it for
RecordingRetained messages, events, native session archives, and workspace artifacts, subject to recorded capture coverage.
TraceAn ordered view of recorded model and tool activity.
EvidenceRecorded material that supports a judgment; citations identify the relevant parts.

One scoring contract

NameMeaning
## Criterion: id — LabelA criterion declaration in judge.md.
CriterionDefinition / rubricCriteriaCriterion definitions and their Markdown parser.
CriterionScore / judgment.scoresNormalized criterion scores, reasons, evidence, and source.
JudgeContextThe data supplied to a plain judge.ts function.
judge.md
## Criterion: safe_parameters — Uses bound parameters

Judge is the common name for code-based, LLM-based, and hybrid evaluators. judge.md and judge.ts can both contribute distinct criteria to the same eval. Their scores have equal weight regardless of which file produced them.

Scores accept booleans, finite numbers from 0 to 1, or null. Booleans become 0 or 1; null means unresolved. A code function can also return custom JSON. Only entries in scores contribute to grading; outputs without scores are unscored.

OpenEval 0.3.0 uses the canonical API and results schema 5. Preserve historical quotations, source titles, original recordings, and finalized judgments when carrying out a separate migration of an older store.

Find in documentation