Evaluation suite
@kindgi/specs/eval-suite.schema.json, schema version 1.0.0.
Tenant-scoped evaluation suite definition — the registry shape (/v1/eval-suites). Suites define WHAT to evaluate; run execution + per-kind grader dispatch are separate concerns. kind selects the shape of spec; the registry treats spec as opaque JSON so new kinds can extend the catalog without breaking existing readers. Each kind's runtime consumer (eval-judge adapter for accuracy / pairwise / regression, HITL bridge for human-review, sandbox handler for custom) deserializes spec against its own contract. Case data lives in spec.
id(string, required): Stable suite id chosen by the caller (e.g.acme.drafting-accuracy).tenantId(string, required): Tenant owning the suite. When published via HTTP, must match the caller tenant (server-derived from the token). Cross-tenant publish is rejected.version(string, required): Semver — publishing a modified suite produces a new version. Eval runs are traceable back to the version in force at the time.kind("accuracy"|"pairwise"|"regression"|"human-review"|"benchmark"|"custom", required): Closed enum, extended additively — new kinds require a spec + validator update in tandem so the registry never accepts a kind no runtime consumer honors.accuracy— metric-based grading against ground truth (input → expectedOutput pairs; grader is an eval-judge adapter).pairwise— A/B comparisons between two agent/flow versions.regression— compare against a stored baseline run.human-review— delegates rubric grading to HITL.benchmark— reference to a standardized suite (name + version).custom— caller-supplied evaluator handler (sandboxed handler-as-code).description(string)spec(map of any, required): Kind-specific suite body. Registry validates only that this is an object; deeper validation is the runtime consumer's responsibility per kind. Foraccuracy, typically{ cases: [{ input, expectedOutput }], grader?: { adapterId, config? } }. Forpairwise, typically{ prompts: [...], variantA: { agentId, version }, variantB: { agentId, version } }. Forregression, typically{ baseline: { runId } | { suiteId, version }, cases: [...] }. Forhuman-review, typically{ rubric: [...], reviewerRole }. Forbenchmark, typically{ benchmark: { name, version } }. Forcustom, typically{ handler: { modulePath, entrypointPath }, cases: [...] }.