OpenEvalv0.5.3Documentation
GitHub72
AUTHOR

Code & hybrid judges

Write an ordinary function. Read the recording. Return scores and your own data.

On this page

A deterministic eval is a plain function

evals/exact-answer/prompt.md
Reply with exactly APPLE.
evals/exact-answer/judge.ts
import type { JudgeContext } from "@hona/openeval";

export default ({ response }: JudgeContext) => ({
  scores: { correct_answer: response.text === "APPLE" },
});
benchmark.ts · no judge model needed
import type { Benchmark } from "@hona/openeval";

export default {
  models: ["provider/candidate-model"],
  repetitions: 3,
} satisfies Benchmark;

The runner discovers judge.ts by convention. It receives JudgeContext and returns JSON, synchronously or asynchronously. response.text is always a string; missing text becomes an empty string. The execution outcome remains available on context.run.

Returned value in scoresNormalized credit
true / false1 / 0
0.37537.5% credit
nullUnresolved
Outside 0–1, NaN, or a stringA judging error

The author chooses criterion IDs and all grading rules. No task-specific scorers or builder APIs are required. Custom JSON is retained as-is. An output without scores is unscored and cannot silently disappear from the benchmark denominator.

Both files contribute to one eval

prompt.md
Summarize the incident in ./incident.txt. Return a JSON object with a "summary" string.
judge.md
## Criterion: supported_summary — Source-supported summary

Pass when the summary covers the incident's impact and resolution
without contradicting the supplied report. Fail for missing or false
information. Accept equivalent concise wording.
judge.ts
import type { JudgeContext } from "@hona/openeval";

export default ({ response }: JudgeContext) => {
  try {
    const value = JSON.parse(response.text);
    return {
      scores: { json_shape: typeof value?.summary === "string" },
    };
  } catch {
    return { scores: { json_shape: false } };
  }
};

Supply incident.txt in the workspace. The LLM grades supported_summary, while code grades json_shape. Both read the same candidate recording. Duplicate criterion IDs are errors; neither source overwrites the other. A hybrid judgment becomes active only after both sources succeed.

All the recorded data, with lazy native access

Context primitiveWhat it provides
response / promptThe final root answer and exact task prompt.
runThe EvalRun, model reference, outcome, and execution inputs. Null for constructed controls.
metricsCandidate-only usage, cost, tool reliability, compactions, and timing.
recording.events(filter?)Native events with recorded sequence and time; filter by sessionID or type.
recording.tools(filter?)Recorded invocations, inputs, outputs, errors, and durations.
recording.sessions()All sessions belonging to the candidate's isolated native database.
recording.messages(sessionID?)Complete paginated message history, including before compaction.
recording.export(sessionID?)Native OpenCode session export.
workspace.files/read/text/diffVerified initial and final file snapshots.
workspace.materialize(revision?)A disposable workspace copy. This is not a sandbox for untrusted execution.
verification.run(request)Run bounded argv commands over initial/final artifacts in a pinned OCI container. No credentials, host mounts, or network. Retain exit statuses, logs, and requested output files.
verification.read/text(result, path)Read a retained verification output after checking its hash. This verifies reconstructed artifacts, not the original runtime's live state.
native.database()Read-only SQLite access to a verified archive copy.
native.sdk() / native.schema()The pinned OpenCode SDK/API and schema. SDK operations use a disposable copy.
independent inspection
import { readRecording } from "@hona/openeval";

await using context = await readRecording("./results/RUN", "eval_ID");
const messages = await context.recording.messages();
const db = await context.native.database();
console.log(db.query("SELECT name FROM sqlite_master").all());

The runner owns reader lifetimes and initializes native services only when requested. Native reads are bound to the recorded database, not your live OpenCode service. The SDK opens supported native schemas on disposable copies; recorded and reader versions remain distinct metadata.

Isolated artifact checks
const result = await context.verification.run({
  revision: "final", timeoutMs: 60_000,
  commands: [["bun", "/verification/check.mjs"]],
  files: { "check.mjs": checkSource },
  artifacts: ["checks.json", "screenshot.png"],
});

return { scores: { correct: {
  value: result.state === "completed" && result.exitCode === 0,
  reason: "Explain the verified outcome.",
  evidence: [{ kind: "verification", id: result.id }],
} } };

Set judge.verification.image in benchmark.ts, then build the standard image with openeval image --verification. The image includes Bun, Python, and Playwright with Chromium. Check inputs remain outside the candidate. Artifact tests and thresholds belong to your judge, not the SDK. Boolean and numeric scores still work; structured scores optionally add reasons, evidence, and measurements.

Common measurements are automatic

MetricDefinition
cost / tokensDeduplicated recorded OpenCode usage across candidate sessions, including recorded auxiliary usage such as compaction.
tools.errorRateFailed / (succeeded + failed). Null when no terminal calls exist; unfinished calls are separate.
requests / compactionsRecorded starts, completions, failures, and retry events.
timing.modelActiveMsUnion of closed model-step and compaction intervals, including time within those steps.
timing.outputTokensPerSecondReported output tokens per model-active second; overlapping intervals count once.

A successful shell tool returning a failed test is distinct from a native tool failure. Counts describe recorded native invocations, not inferred operations inside a batch. Missing usage is unavailable. Reported zero cost does not establish free service.

Token categories keep their native meanings. Avoid double-counting reasoning/output or cache categories. These observations affect a score only when the author's function or rubric explicitly uses them. Markdown judges can inspect the same primitives through candidate_evidence with action metrics.

Frozen inputs, bounded execution, inspectable results

Checks
NamePassError
Input schemaPass—
Returned rowsFailTwo rows missing
Elapsed ms
1400
Reported coverage
3/5

Illustrative author-owned check data. These observations do not create extra criterion scores; no candidate model was called.

A code judgment with boolean and fractional criterion scores, returned JSON, and frozen source
Illustrative label-reading task. The code judge ran on constructed responses; no candidate model was called.Open full size

The host bundles local imports without executing the judge during planning. Its fingerprint covers the bundled code and the versions of imported packages, not the checkout location or lockfiles. Code and reference changes schedule rejudging; the candidate recording is reused.

judge.ts runs in a separate Bun process under judge.timeoutMs, so the host can stop synchronous loops. Code and hybrid evals grade finalized recordings; earlyStop is supported for Markdown-only evals. The viewer shows returned JSON, frozen source, process logs, and candidate metrics.

code-only calibration
import { recordEvidence, judgeEvidence } from "@hona/openeval";

const evidence = await recordEvidence({
  directory: "./controls/apple/evidence",
  prompt: "Reply with exactly APPLE.",
  response: "APPLE",
});
const result = await judgeEvidence({
  evidence,
  code: "./evals/exact-answer/judge.ts",
  directory: "./controls/apple/judge",
});

Code-only controls use no model calls unless the author writes one explicitly. The original JSON is retained, while normalized scores are stored with their source. Failed grading does not overwrite finalized judgments.

Find in documentation