Code & hybrid judges
Write an ordinary function. Read the recording. Return scores and your own data.
On this page
A deterministic eval is a plain function
Reply with exactly APPLE.
import type { JudgeContext } from "@hona/openeval";
export default ({ response }: JudgeContext) => ({
scores: { correct_answer: response.text === "APPLE" },
});
import type { Benchmark } from "@hona/openeval";
export default {
models: ["provider/candidate-model"],
repetitions: 3,
} satisfies Benchmark;
The runner discovers judge.ts by convention. It receives JudgeContext and returns JSON, synchronously or asynchronously. response.text is always a string; missing text becomes an empty string. The execution outcome remains available on context.run.
| Returned value in scores | Normalized credit |
|---|---|
| true / false | 1 / 0 |
| 0.375 | 37.5% credit |
| null | Unresolved |
| Outside 0–1, NaN, or a string | A judging error |
The author chooses criterion IDs and all grading rules. No task-specific scorers or builder APIs are required. Custom JSON is retained as-is. An output without scores is unscored and cannot silently disappear from the benchmark denominator.
Both files contribute to one eval
Summarize the incident in ./incident.txt. Return a JSON object with a "summary" string.
## Criterion: supported_summary — Source-supported summary
Pass when the summary covers the incident's impact and resolution
without contradicting the supplied report. Fail for missing or false
information. Accept equivalent concise wording.
import type { JudgeContext } from "@hona/openeval";
export default ({ response }: JudgeContext) => {
try {
const value = JSON.parse(response.text);
return {
scores: { json_shape: typeof value?.summary === "string" },
};
} catch {
return { scores: { json_shape: false } };
}
};
Supply incident.txt in the workspace. The LLM grades supported_summary, while code grades json_shape. Both read the same candidate recording. Duplicate criterion IDs are errors; neither source overwrites the other. A hybrid judgment becomes active only after both sources succeed.
All the recorded data, with lazy native access
| Context primitive | What it provides |
|---|---|
| response / prompt | The final root answer and exact task prompt. |
| run | The EvalRun, model reference, outcome, and execution inputs. Null for constructed controls. |
| metrics | Candidate-only usage, cost, tool reliability, compactions, and timing. |
| recording.events(filter?) | Native events with recorded sequence and time; filter by sessionID or type. |
| recording.tools(filter?) | Recorded invocations, inputs, outputs, errors, and durations. |
| recording.sessions() | All sessions belonging to the candidate's isolated native database. |
| recording.messages(sessionID?) | Complete paginated message history, including before compaction. |
| recording.export(sessionID?) | Native OpenCode session export. |
| workspace.files/read/text/diff | Verified initial and final file snapshots. |
| workspace.materialize(revision?) | A disposable workspace copy. This is not a sandbox for untrusted execution. |
| verification.run(request) | Run bounded argv commands over initial/final artifacts in a pinned OCI container. No credentials, host mounts, or network. Retain exit statuses, logs, and requested output files. |
| verification.read/text(result, path) | Read a retained verification output after checking its hash. This verifies reconstructed artifacts, not the original runtime's live state. |
| native.database() | Read-only SQLite access to a verified archive copy. |
| native.sdk() / native.schema() | The pinned OpenCode SDK/API and schema. SDK operations use a disposable copy. |
import { readRecording } from "@hona/openeval";
await using context = await readRecording("./results/RUN", "eval_ID");
const messages = await context.recording.messages();
const db = await context.native.database();
console.log(db.query("SELECT name FROM sqlite_master").all());
The runner owns reader lifetimes and initializes native services only when requested. Native reads are bound to the recorded database, not your live OpenCode service. The SDK opens supported native schemas on disposable copies; recorded and reader versions remain distinct metadata.
const result = await context.verification.run({
revision: "final", timeoutMs: 60_000,
commands: [["bun", "/verification/check.mjs"]],
files: { "check.mjs": checkSource },
artifacts: ["checks.json", "screenshot.png"],
});
return { scores: { correct: {
value: result.state === "completed" && result.exitCode === 0,
reason: "Explain the verified outcome.",
evidence: [{ kind: "verification", id: result.id }],
} } };
Set judge.verification.image in benchmark.ts, then build the standard image with openeval image --verification. The image includes Bun, Python, and Playwright with Chromium. Check inputs remain outside the candidate. Artifact tests and thresholds belong to your judge, not the SDK. Boolean and numeric scores still work; structured scores optionally add reasons, evidence, and measurements.
Common measurements are automatic
| Metric | Definition |
|---|---|
| cost / tokens | Deduplicated recorded OpenCode usage across candidate sessions, including recorded auxiliary usage such as compaction. |
| tools.errorRate | Failed / (succeeded + failed). Null when no terminal calls exist; unfinished calls are separate. |
| requests / compactions | Recorded starts, completions, failures, and retry events. |
| timing.modelActiveMs | Union of closed model-step and compaction intervals, including time within those steps. |
| timing.outputTokensPerSecond | Reported output tokens per model-active second; overlapping intervals count once. |
A successful shell tool returning a failed test is distinct from a native tool failure. Counts describe recorded native invocations, not inferred operations inside a batch. Missing usage is unavailable. Reported zero cost does not establish free service.
Token categories keep their native meanings. Avoid double-counting reasoning/output or cache categories. These observations affect a score only when the author's function or rubric explicitly uses them. Markdown judges can inspect the same primitives through candidate_evidence with action metrics.
Frozen inputs, bounded execution, inspectable results
- Checks
Name Pass Error Input schema Pass — Returned rows Fail Two rows missing - Elapsed ms
- 1400
- Reported coverage
- 3/5
Illustrative author-owned check data. These observations do not create extra criterion scores; no candidate model was called.

The host bundles local imports without executing the judge during planning. Its fingerprint covers the bundled code and the versions of imported packages, not the checkout location or lockfiles. Code and reference changes schedule rejudging; the candidate recording is reused.
judge.ts runs in a separate Bun process under judge.timeoutMs, so the host can stop synchronous loops. Code and hybrid evals grade finalized recordings; earlyStop is supported for Markdown-only evals. The viewer shows returned JSON, frozen source, process logs, and candidate metrics.
import { recordEvidence, judgeEvidence } from "@hona/openeval";
const evidence = await recordEvidence({
directory: "./controls/apple/evidence",
prompt: "Reply with exactly APPLE.",
response: "APPLE",
});
const result = await judgeEvidence({
evidence,
code: "./evals/exact-answer/judge.ts",
directory: "./controls/apple/judge",
});
Code-only controls use no model calls unless the author writes one explicitly. The original JSON is retained, while normalized scores are stored with their source. Failed grading does not overwrite finalized judgments.