Skip to content

Evaluation suite

@kindgi/specs/eval-suite.schema.json, schema version 1.0.0.

Tenant-scoped evaluation suite definition — the registry shape (/v1/eval-suites). Suites define WHAT to evaluate; run execution + per-kind grader dispatch are separate concerns. kind selects the shape of spec; the registry treats spec as opaque JSON so new kinds can extend the catalog without breaking existing readers. Each kind's runtime consumer (eval-judge adapter for accuracy / pairwise / regression, HITL bridge for human-review, sandbox handler for custom) deserializes spec against its own contract. Case data lives in spec.

  • id (string, required): Stable suite id chosen by the caller (e.g. acme.drafting-accuracy).
  • tenantId (string, required): Tenant owning the suite. When published via HTTP, must match the caller tenant (server-derived from the token). Cross-tenant publish is rejected.
  • version (string, required): Semver — publishing a modified suite produces a new version. Eval runs are traceable back to the version in force at the time.
  • kind ("accuracy" | "pairwise" | "regression" | "human-review" | "benchmark" | "custom", required): Closed enum, extended additively — new kinds require a spec + validator update in tandem so the registry never accepts a kind no runtime consumer honors. accuracy — metric-based grading against ground truth (input → expectedOutput pairs; grader is an eval-judge adapter). pairwise — A/B comparisons between two agent/flow versions. regression — compare against a stored baseline run. human-review — delegates rubric grading to HITL. benchmark — reference to a standardized suite (name + version). custom — caller-supplied evaluator handler (sandboxed handler-as-code).
  • description (string)
  • spec (map of any, required): Kind-specific suite body. Registry validates only that this is an object; deeper validation is the runtime consumer's responsibility per kind. For accuracy, typically { cases: [{ input, expectedOutput }], grader?: { adapterId, config? } }. For pairwise, typically { prompts: [...], variantA: { agentId, version }, variantB: { agentId, version } }. For regression, typically { baseline: { runId } | { suiteId, version }, cases: [...] }. For human-review, typically { rubric: [...], reviewerRole }. For benchmark, typically { benchmark: { name, version } }. For custom, typically { handler: { modulePath, entrypointPath }, cases: [...] }.