All project evidence

Sentinel Eval Harness

Completed portfolio build

An evaluation harness for AI systems with repeatable datasets, deterministic and quality scorers, mutation testing, statistical comparison, persisted runs, an API, dashboard, and CI release gate.

Problem & approach

Problem

AI quality needs repeatable evaluation before a release can be justified.

Approach

Combine repeatable datasets, deterministic and quality scorers, statistical comparison, persisted runs, and a CI release gate.

Evidence

Inspect the evaluation interfaces, mutation tests, and release-gate implementation.

Product decision review

Who it serves

AI product and engineering teams deciding whether a prompt, model, or retrieval change is ready for release.

The decision this design supports

Make release criteria repeatable: version the evaluation cases and compare a candidate against explicit quality, latency, and cost limits. A good-looking answer from one manual test is not enough to justify a release.

Implemented workflow

  1. Define expected behavior in a versioned JSONL suite and preserve the dataset fingerprint.
  2. Run the fixture adapter first, then connect a real endpoint using the same adapter contract.
  3. Inspect case-level failures and robustness mutations alongside aggregate scores.
  4. Use the release gate and stored run evidence to explain a pass or block decision.

Design tradeoff

Transparent deterministic scorers are reproducible and inexpensive, but lexical matching cannot fully assess nuanced reasoning. They are useful as a baseline; richer judgments need their own validation rather than replacing every check with an opaque model score.

What to measure

Review pass rate, critical-case failures, baseline-to-candidate change, latency, and recorded cost. These are evaluation outputs, not claims of customer impact. Compare runs only when datasets and policies are controlled.

Current scope

A passing fixture run validates the configured test workflow. It does not establish general model safety or production readiness for an unseen use case.

Next evaluation

Expand the suite around observed user failures and compare scorer decisions against human review before raising release confidence.

Inspect the implementation

Release history

GitHub releases are tagged versions. Implementation code and commits can exist without a published release.

Inspect repository and README · GitHub, new tab