Twenty-two scorers, seven dimensions
Cognitive phase, perception focus, skill usage, behaviour triggers, empathy, length and a model-as-judge pass. The dimensions are about how a response reads to a person.
Evaluation for LLM output: 22 scorers across eight scenario types, saved baselines so you can tell a prompt improvement from a regression, and five red-team generators that throw prompt injection, jailbreaks, PII extraction, hallucination bait and bias probes at your own system before someone else does.
Across seven human-like dimensions: cognitive phase, perception focus, skill usage, behaviour triggers, empathy, length and judged quality.
Prompt injection, jailbreak, PII extraction, hallucination probes and bias detection, run as a suite.
Save an evaluation and later runs are compared to it, so a score that drops after a prompt change is reported.
Evaluation for agent output, built around the idea that a score is only useful next to the one before it.
Cognitive phase, perception focus, skill usage, behaviour triggers, empathy, length and a model-as-judge pass. The dimensions are about how a response reads to a person.
An evaluation can be saved as a baseline and later runs compared against it across prompt versions, models and configurations, so a change that costs quality shows up as a number.
Every suite, case and run carries its tenant on the context, so cross-tenant queries are structurally impossible.
SQLite or Postgres for production, in-memory for development, with every subsystem behind an interface.
Prompt injection, jailbreak, PII extraction, hallucination probes and bias detection, so resilience becomes a measurement.
Factual, creative, safety, summarisation, classification, extraction, conversation and reasoning, each able to run against a persona with dimension scoring.
22 built-in, across seven dimensions.
Eight types, from single turn to adversarial.
Comparison against a recorded run, not a vibe.
Adversarial suites run as part of the same harness.
Sentinel evaluates LLM output. Suites, cases, scorers, baselines and red-team generation, as a Go library with an optional Forge integration.
Running the red-team suite against Shield, the boundary layer of plain deny lists consistently catches more than the classifier layers. Semantic detection is the interesting part and the written-down list is the part that works. I keep expecting that to change and it has not yet.
Shipping something on Sentinel? Nobody is listed here yet. Tell me what you built and you will be the first.
Get listed →