# OpenEval > Write agent evals with task prompts, code or LLM judges, isolated OpenCode runs, and an evidence viewer. Documentation for @hona/openeval 0.5.3, using the 0.3.0 API and results schema 5. Pin an exact released SDK version. A benchmark contains evals. A criterion is a graded requirement; a score is awarded credit; a metric is a measurement. For setup, read the agent guide first. Read other pages as needed. Each linked Markdown page contains the same documentation as its HTML page. Existing page URLs also return Markdown when the Accept header prefers text/markdown or text/x-markdown over HTML. ## Start - [Agent setup guide](https://openev.al/agent-start.md): Interactive setup, prerequisites, local skill installation, one eval, model choices, and an optional first run. - [Overview](https://openev.al/index.md): Task, judge, run, inspect, and compare examples. - [Your first eval](https://openev.al/docs/quickstart/index.md): One SQL task. Two criteria. Choose a model, run it, and inspect the work. - [Concepts & terminology](https://openev.al/docs/terminology/index.md): A metric measures what happened. A criterion defines what earns credit. A score is the credit awarded. ## Author - [Write task prompts](https://openev.al/docs/prompts/index.md): Give the candidate a real task. Put the scoring rules in the rubric. - [Write judge rubrics](https://openev.al/docs/rubrics/index.md): Small, independent decisions with evidence you can inspect. - [Code & hybrid judges](https://openev.al/docs/code-judges/index.md): Write an ordinary function. Read the recording. Return scores and your own data. - [Prepare a workspace](https://openev.al/docs/workspaces/index.md): Give every candidate a controlled starting point. ## Run & inspect - [Run your benchmark](https://openev.al/docs/running/index.md): Plan a small batch, control cost, and keep one aggregate. - [Understand the scores](https://openev.al/docs/scoring/index.md): Equal eval weights. Visible unknowns. One final percentage per model. - [Read the evidence](https://openev.al/docs/evidence/index.md): Follow a score back to the work that earned it. ## Reference - [CLI & file reference](https://openev.al/docs/reference/index.md): The author-facing controls, in one place. ## Eval Writing skill - [Skill instructions](https://openev.al/skills/eval-writing/eval-writing.md): Public eval design and review workflow; includes links to its references. - [Skill catalog](https://openev.al/skills/index.json): Native OpenCode V2 HTTP catalog, with content-versioned files. - [Project skill download](https://openev.al/eval-writing.zip): Extract into a project to install .opencode/skills/eval-writing/ and all references. ## Optional - [SQL starter](https://openev.al/starter.zip): Benchmark files for the human quick start. - [Source repository](https://github.com/Hona/openeval): Public SDK, CLI, viewer, documentation, and the source-controlled agent-start.md. - [OpenCode V2 documentation](https://opencode.ai/v2/llms.txt): Installation, connections, models, variants, skills, and tools.