# Code & hybrid judges

Write an ordinary function. Read the recording. Return scores and your own data.

<a id="function"></a>

## A deterministic eval is a plain function

### evals/exact-answer/prompt.md

```markdown
Reply with exactly APPLE.
```

### evals/exact-answer/judge.ts

```typescript
import type { JudgeContext } from "@hona/openeval";

export default ({ response }: JudgeContext) => ({
  scores: { correct_answer: response.text === "APPLE" },
});
```

### benchmark.ts · no judge model needed

```typescript
import type { Benchmark } from "@hona/openeval";

export default {
  models: ["provider/candidate-model"],
  repetitions: 3,
} satisfies Benchmark;
```

The runner discovers judge.ts by convention. It receives JudgeContext and returns JSON, synchronously or asynchronously. response.text is always a string; missing text becomes an empty string. The execution outcome remains available on context.run.

| Returned value in scores | Normalized credit |
| --- | --- |
| true / false | 1 / 0 |
| 0.375 | 37.5% credit |
| null | Unresolved |
| Outside 0–1, NaN, or a string | A judging error |

The author chooses criterion IDs and all grading rules. No task-specific scorers or builder APIs are required. Custom JSON is retained as-is. An output without scores is unscored and cannot silently disappear from the benchmark denominator.

<a id="hybrid"></a>

## Both files contribute to one eval

### prompt.md

```markdown
Summarize the incident in ./incident.txt. Return a JSON object with a "summary" string.
```

### judge.md

```markdown
## Criterion: supported_summary — Source-supported summary

Pass when the summary covers the incident's impact and resolution
without contradicting the supplied report. Fail for missing or false
information. Accept equivalent concise wording.
```

### judge.ts

```typescript
import type { JudgeContext } from "@hona/openeval";

export default ({ response }: JudgeContext) => {
  try {
    const value = JSON.parse(response.text);
    return {
      scores: { json_shape: typeof value?.summary === "string" },
    };
  } catch {
    return { scores: { json_shape: false } };
  }
};
```

Supply incident.txt in the workspace. The LLM grades supported_summary, while code grades json_shape. Both read the same candidate recording. Duplicate criterion IDs are errors; neither source overwrites the other. A hybrid judgment becomes active only after both sources succeed.

> **Equal weights**
>
> The files themselves have no weight. Each named criterion contributes equally within the eval. A code score of 0.5 and an LLM score of 1 produce 75% for that repetition.

<a id="context"></a>

## All the recorded data, with lazy native access

| Context primitive | What it provides |
| --- | --- |
| response / prompt | The final root answer and exact task prompt. |
| run | The EvalRun, model reference, outcome, and execution inputs. Null for constructed controls. |
| metrics | Candidate-only usage, cost, tool reliability, compactions, and timing. |
| recording.events(filter?) | Native events with recorded sequence and time; filter by sessionID or type. |
| recording.tools(filter?) | Recorded invocations, inputs, outputs, errors, and durations. |
| recording.sessions() | All sessions belonging to the candidate's isolated native database. |
| recording.messages(sessionID?) | Complete paginated message history, including before compaction. |
| recording.export(sessionID?) | Native OpenCode session export. |
| workspace.files/read/text/diff | Verified initial and final file snapshots. |
| workspace.materialize(revision?) | A disposable workspace copy. This is not a sandbox for untrusted execution. |
| verification.run(request) | Run bounded argv commands over initial/final artifacts in a pinned OCI container. No credentials, host mounts, or network. Retain exit statuses, logs, and requested output files. |
| verification.read/text(result, path) | Read a retained verification output after checking its hash. This verifies reconstructed artifacts, not the original runtime's live state. |
| native.database() | Read-only SQLite access to a verified archive copy. |
| native.sdk() / native.schema() | The pinned OpenCode SDK/API and schema. SDK operations use a disposable copy. |

### independent inspection

```typescript
import { readRecording } from "@hona/openeval";

await using context = await readRecording("./results/RUN", "eval_ID");
const messages = await context.recording.messages();
const db = await context.native.database();
console.log(db.query("SELECT name FROM sqlite_master").all());
```

The runner owns reader lifetimes and initializes native services only when requested. Native reads are bound to the recorded database, not your live OpenCode service. The SDK opens supported native schemas on disposable copies; recorded and reader versions remain distinct metadata.

### Isolated artifact checks

```typescript
const result = await context.verification.run({
  revision: "final", timeoutMs: 60_000,
  commands: [["bun", "/verification/check.mjs"]],
  files: { "check.mjs": checkSource },
  artifacts: ["checks.json", "screenshot.png"],
});

return { scores: { correct: {
  value: result.state === "completed" && result.exitCode === 0,
  reason: "Explain the verified outcome.",
  evidence: [{ kind: "verification", id: result.id }],
} } };
```

Set judge.verification.image in benchmark.ts, then build the standard image with openeval image --verification. The image includes Bun, Python, and Playwright with Chromium. Check inputs remain outside the candidate. Artifact tests and thresholds belong to your judge, not the SDK. Boolean and numeric scores still work; structured scores optionally add reasons, evidence, and measurements.

<a id="measurements"></a>

## Common measurements are automatic

| Metric | Definition |
| --- | --- |
| cost / tokens | Deduplicated recorded OpenCode usage across candidate sessions, including recorded auxiliary usage such as compaction. |
| tools.errorRate | Failed / (succeeded + failed). Null when no terminal calls exist; unfinished calls are separate. |
| requests / compactions | Recorded starts, completions, failures, and retry events. |
| timing.modelActiveMs | Union of closed model-step and compaction intervals, including time within those steps. |
| timing.outputTokensPerSecond | Reported output tokens per model-active second; overlapping intervals count once. |

A successful shell tool returning a failed test is distinct from a native tool failure. Counts describe recorded native invocations, not inferred operations inside a batch. Missing usage is unavailable. Reported zero cost does not establish free service.

Token categories keep their native meanings. Avoid double-counting reasoning/output or cache categories. These observations affect a score only when the author's function or rubric explicitly uses them. Markdown judges can inspect the same primitives through candidate_evidence with action metrics.

<a id="execution"></a>

## Frozen inputs, bounded execution, inspectable results

```json
{
  "checks": [
    {
      "name": "Input schema",
      "pass": true
    },
    {
      "name": "Returned rows",
      "pass": false,
      "error": "Two rows missing"
    }
  ],
  "elapsedMs": 1400,
  "reportedCoverage": "3/5"
}
```

Illustrative author-owned check data. These observations do not create extra criterion scores; no candidate model was called.

![A code judgment with boolean and fractional criterion scores, returned JSON, and frozen source](https://openev.al/images/code-judgment.png)

Illustrative label-reading task. The code judge ran on constructed responses; no candidate model was called.

The host bundles local imports without executing the judge during planning. Its fingerprint covers the bundled code and the versions of imported packages, not the checkout location or lockfiles. Code and reference changes schedule rejudging; the candidate recording is reused.

judge.ts runs in a separate Bun process under judge.timeoutMs, so the host can stop synchronous loops. Code and hybrid evals grade finalized recordings; earlyStop is supported for Markdown-only evals. The viewer shows returned JSON, frozen source, process logs, and candidate metrics.

### code-only calibration

```typescript
import { recordEvidence, judgeEvidence } from "@hona/openeval";

const evidence = await recordEvidence({
  directory: "./controls/apple/evidence",
  prompt: "Reply with exactly APPLE.",
  response: "APPLE",
});
const result = await judgeEvidence({
  evidence,
  code: "./evals/exact-answer/judge.ts",
  directory: "./controls/apple/judge",
});
```

Code-only controls use no model calls unless the author writes one explicitly. The original JSON is retained, while normalized scores are stored with their source. Failed grading does not overwrite finalized judgments.

[Documentation index](https://openev.al/llms.txt)

