← the whole session plugin/skills/governer-adaptive-score-consumer-guide/SKILL.md
How to consume the adaptive scoring response from the harness's local governer — every variable an agent can condition on, what each means, and how to build conditional agent behavior on top of it.
Consuming the Adaptive Score Response
When you run the governer's scorer, you get back a rich response with every variable you need to build conditional agent behavior. This guide explains what you get and how to use it.
How to call it
The scorer runs entirely locally — no server, no API key, no network call. Two ways to get it:
From the command line (what /governer itself does):
alignment-harness score --task "short description of the task" --category internal-tooling --ac 9 --hp 8 --ew 5 --json
--ac/--hp/--ew are the three risk answers (1-10, see below). --category is optional; add --touches-user-state --impact-tier revenue when it applies. --json prints the full structured response shown below in addition to the plain-language summary.
From code, if you're writing an adaptor that needs the raw object rather than shelling out:
const scorer = require('<path-to-harness-plugin>/lib/scorer/pipelineLeverageScorer');
const result = scorer.computeAdaptiveScore({ category: 'internal-tooling', agenticCompounding: 9, hallucinationPropagation: 8, existingWinsAtStake: 5 });
Either path runs the same arithmetic. If you've set up a category table of your own (see /alignment-harness:harness-setup → risk table), the scorer uses it automatically; until then it uses the shipped defaults and the score is reported as "uncalibrated."
Full response shape
{
"unifiedScore": 75,
"baseScore": 15,
"metaScore": 75,
"method": "meta-elevated",
"budget": "maximum",
"budgetDetails": { "label": "Maximum", "pipelineDepth": "L90", "compactionDepth": "deep", "reflectionTrigger": true, "confidenceFloor": 90 },
"base": {
"score": 15,
"method": "override",
"details": {
"category": "internal-tooling",
"overrideScore": 15,
"wasElevated": false,
"elevationReason": null,
"impactTier": null
}
},
"meta": {
"metaScore": 75,
"weakestDimension": { "name": "Existing Wins at Stake", "value": 5 },
"dimensions": {
"agenticCompounding": 9,
"hallucinationPropagation": 8,
"existingWinsAtStake": 5
}
},
"scoreThresholds": {
"skipPipelineBelow": 20,
"reflectAbove": 60,
"fullValidationAbove": 80,
"confidenceFloor": 90
},
"humanReadable": { "...": "the same numbers with long plain-language field names, meant to be dropped straight into agent context" }
}
There is no tier field on the response — don't key on one. There is no separate network round trip either: this is the literal return value of a pure function, so what you see above is exactly what's available, not a subset of a richer server response.
When two of the three risk answers (ac/hp/ew) are 8 or higher, metaScore is raised to a floor (80 by default, configurable via the criticalMassFloor switch — see alignment-harness config show) regardless of what the geometric mean alone would produce. This is deliberate: two seriously risky dimensions at once means the task gets full-depth treatment even if the third dimension looks mild.
If you call the scorer with none of ac, hp, ew set, meta comes back null. Any rule below that reads meta.dimensions… needs to check for that first.
Every variable you can condition on
Headline numbers (1-100, continuous)
| Field | What it means | When to use it |
|---|---|---|
unifiedScore |
The single number representing total reasoning depth needed. Max of baseScore and metaScore. | Default for any "how careful should I be?" decision. |
baseScore |
How important is this feature to users/business? Driven by category and impact tier (or your own risk table, once you've set one up). | When you need to know business importance independent of agent risk. |
metaScore |
How risky is this for an agent to execute? Geometric mean of ac/hp/ew × 10, with the critical-mass floor described above. | When you care specifically about agent execution risk, not business value. |
Method (what drove the score)
| Value | Meaning |
|---|---|
"base-driven" |
Business importance was higher than (or tied with) agent risk. The feature matters to users more than it's dangerous for agents. |
"meta-elevated" |
Agent risk exceeded business importance. The feature might be low-stakes business-wise but dangerous if the agent gets it wrong. |
A tie between baseScore and metaScore reports "base-driven". Use method when you need to know WHY the score is what it is — the same score can mean different things depending on whether base or meta drove it.
Individual meta-dimensions (1-10 each)
| Field path | What it measures |
|---|---|
meta.dimensions.agenticCompounding |
Does this change multiply across future agent decisions? 1 = one-off, 10 = changes how all future agents decide. |
meta.dimensions.hallucinationPropagation |
If the agent hallucinates here, how far does the error spread? 1 = stays in this file, 10 = every future decision inherits the mistake. |
meta.dimensions.existingWinsAtStake |
Could this break something already working? 1 = greenfield, 10 = touches auth/payment/shared middleware. |
These are the most powerful fields for conditional logic — but check meta isn't null first (see above). An agent building a code review step might only care about hallucinationPropagation. An agent deciding whether to run integration tests might only care about existingWinsAtStake.
Weakest dimension
| Field path | What it tells you |
|---|---|
meta.weakestDimension.name |
Which dimension scored lowest — the bottleneck holding the meta-score down. |
meta.weakestDimension.value |
The actual score of that dimension (1-10). |
Use this when you want to know what's pulling the score DOWN. If weakestDimension is "Existing Wins at Stake" at value 2, the feature is risky for agents but safe for existing code — you might skip regression tests.
Base scoring details
| Field path | What it tells you |
|---|---|
base.method |
How base score was determined: "manual", "override" (category), "sila", "sila-over-override" (if you've set up the IX-SILA nine-question model as your risk table). |
base.details.category |
The category used — from your own risk table if you've set one up (alignment-harness risk-table show), otherwise the shipped defaults. |
base.details.wasElevated |
Did the touchesUserState gate raise the score above the category default? |
base.details.elevationReason |
Human-readable explanation of why the score was elevated. |
base.details.impactTier |
The impact tier that drove elevation: "indirect", "config", "communication", "revenue". |
Score thresholds — use these instead of a bucket label
| Field | What it tells you |
|---|---|
scoreThresholds.skipPipelineBelow |
Below this unifiedScore, skip the governer pipeline entirely (default 20). |
scoreThresholds.reflectAbove |
At or above this, /reflect is required before acting (default 60). |
scoreThresholds.fullValidationAbove |
At or above this, every quality gate applies (default 80). |
scoreThresholds.confidenceFloor |
Minimum certainty (%) the agent should have before proceeding, scaled to unifiedScore. |
budget (minimal/standard/elevated/maximum) is a legacy bucket label kept for backward compatibility — prefer unifiedScore plus scoreThresholds directly over branching on the label.
Example conditional patterns
"Only run deep security review when hallucination could spread to users":
if meta && meta.dimensions.hallucinationPropagation >= 7
"Skip regression tests for greenfield features":
if meta && meta.dimensions.existingWinsAtStake <= 3
"Require human review when agent risk is high but business stakes are low":
if method === "meta-elevated" AND baseScore < 30 AND metaScore > 70
"Extra care with anything that changes agent decision-making":
if meta && meta.dimensions.agenticCompounding >= 8
"Scale test coverage continuously with the score":
coverageTarget = 60 + (unifiedScore * 0.35)
"Different behavior depending on what drove the score":
if method === "base-driven" → focus on UX validation
if method === "meta-elevated" → focus on code correctness and testing
"Only flag for staged release if it touches revenue AND agents could compound errors":
if base.details.impactTier === "revenue" AND meta && meta.dimensions.agenticCompounding >= 6
Building an adaptor
An adaptor is any skill or agent behavior that reads the adaptive score and adjusts its own depth based on the response. To build one:
- Call
alignment-harness score --json(or the local scorer function directly) at the start of your workflow, or read the already-written score for this session withalignment-harness status --jsonif the governer already ran. - Pick the fields you care about — you don't need all of them. A test depth adaptor might only use
unifiedScoreandexistingWinsAtStake. A compaction-depth adaptor might only useunifiedScoreandmethod. - Write your conditional logic using the fields directly — no buckets, no label translation. Branch on numbers and strings. Guard every
meta.*read with a null check. - Document your conditions in your skill file so other agents know what you're keying on.
There is no adaptor registry or framework. Your skill reads the response, makes decisions, and documents those decisions in its own SKILL.md. The scoring system provides the data. Your skill owns its logic.
Related
/governer runs this scorer as part of its own pipeline and writes the result to the session state file; alignment-harness status reads that file back. If you're adding a new step to the governer pipeline itself rather than just reacting to its score, see /governer-agent-pipeline-how-to-add-steps (if you have it).