# Concepts & terminology

A metric measures what happened. A criterion defines what earns credit. A score is the credit awarded.

<a id="terms"></a>

## One vocabulary

A benchmark contains evals. Each eval defines a task and a rubric. Judges produce scores for the rubric's criteria. Runs also record metrics such as cost, tokens, and tool reliability.

| Term | Meaning | Example |
| --- | --- | --- |
| Benchmark | A collection of evals and their run configuration. | A set of agent tasks evaluated across models. |
| Suite | An optional display label for one benchmark among several that share a name. It never changes scoring. | Frontier and Regressions, both named sql-bench. |
| Eval | A task, its input and environment, and its grading specification. | Produce a parameterized SQL query. |
| Criterion | A named requirement being graded. Plural: criteria. | safe_parameters |
| Score | Credit awarded to a criterion, normalized from 0 to 1, or an aggregate of that credit. | 1 for full credit; 0 for no credit. |
| Metric | An observed or calculated measurement. It contributes to grading only when a criterion uses it. | Cost in USD, token count, or tool error rate. |
| Rubric | The criteria and rules for awarding scores. | The conditions for accepting a parameterized query. |
| Judge | An evaluator that applies grading rules using code, an LLM, or both. | An LLM applying a written rubric. |
| Judgment | A judge's output, including any named criterion scores and supporting data. | A safe_parameters score with evidence. |
| BenchmarkRun | A recorded benchmark collection with its inputs and selected results. | One retained comparison across models. |
| EvalRun | One candidate execution of an eval. | A model's second repetition of the SQL task. |
| JudgeRun | A grading execution against recorded evidence. | A new judgment of a retained EvalRun. |

<a id="measurements-and-credit"></a>

## Measurements become credit through a rule

| Role | Example | Meaning |
| --- | --- | --- |
| Metric | cost_usd = 0.84 | An observed or calculated cost in USD. |
| Criterion | within_budget | Award full credit when cost_usd is at most 1.00. |
| Criterion score | within_budget = 1 | Credit awarded by applying that rule. |

A tool error rate of 0.2 is a measurement even though it falls between 0 and 1. Recording a metric does not automatically award or deduct credit. The author decides whether a criterion uses it.

Criterion IDs are author-defined. A name such as correct_answer has no built-in grading rule. A rubric explains how evidence becomes a score; a judgment records the result of applying it.

> **Aggregation**
>
> Average repetitions per criterion, average criteria within each eval, then average evals equally and multiply by 100. Required unresolved scores keep the final percentage unresolved. Category-specific views and unequal weights are future design work.

<a id="recordings"></a>

## Definitions, executions, and evidence

An Eval is a reusable task definition; an EvalRun is one candidate execution. A JudgeRun applies grading rules to recorded evidence. A completed rejudge updates the active selection while retaining earlier judgments.

| Term | Use it for |
| --- | --- |
| Recording | Retained messages, events, native session archives, and workspace artifacts, subject to recorded capture coverage. |
| Trace | An ordered view of recorded model and tool activity. |
| Evidence | Recorded material that supports a judgment; citations identify the relevant parts. |

<a id="contract"></a>

## One scoring contract

| Name | Meaning |
| --- | --- |
| ## Criterion: id — Label | A criterion declaration in judge.md. |
| CriterionDefinition / rubricCriteria | Criterion definitions and their Markdown parser. |
| CriterionScore / judgment.scores | Normalized criterion scores, reasons, evidence, and source. |
| JudgeContext | The data supplied to a plain judge.ts function. |

### judge.md

```markdown
## Criterion: safe_parameters — Uses bound parameters
```

Judge is the common name for code-based, LLM-based, and hybrid evaluators. judge.md and judge.ts can both contribute distinct criteria to the same eval. Their scores have equal weight regardless of which file produced them.

Scores accept booleans, finite numbers from 0 to 1, or null. Booleans become 0 or 1; null means unresolved. A code function can also return custom JSON. Only entries in scores contribute to grading; outputs without scores are unscored.

OpenEval 0.3.0 uses the canonical API and results schema 5. Preserve historical quotations, source titles, original recordings, and finalized judgments when carrying out a separate migration of an older store.

[Documentation index](https://openev.al/llms.txt)

