← the whole session plugin/skills/governer-adaptive-score-consumer-guide/SKILL.md

How to consume the adaptive scoring response from the harness's local governer — every variable an agent can condition on, what each means, and how to build conditional agent behavior on top of it.

Consuming the Adaptive Score Response

When you run the governer's scorer, you get back a rich response with every variable you need to build conditional agent behavior. This guide explains what you get and how to use it.

How to call it

The scorer runs entirely locally — no server, no API key, no network call. Two ways to get it:

From the command line (what /governer itself does):

alignment-harness score --task "short description of the task" --category internal-tooling --ac 9 --hp 8 --ew 5 --json

--ac/--hp/--ew are the three risk answers (1-10, see below). --category is optional; add --touches-user-state --impact-tier revenue when it applies. --json prints the full structured response shown below in addition to the plain-language summary.

From code, if you're writing an adaptor that needs the raw object rather than shelling out:

const scorer = require('<path-to-harness-plugin>/lib/scorer/pipelineLeverageScorer');
const result = scorer.computeAdaptiveScore({ category: 'internal-tooling', agenticCompounding: 9, hallucinationPropagation: 8, existingWinsAtStake: 5 });

Either path runs the same arithmetic. If you've set up a category table of your own (see /alignment-harness:harness-setup → risk table), the scorer uses it automatically; until then it uses the shipped defaults and the score is reported as "uncalibrated."

Full response shape

{
  "unifiedScore": 75,
  "baseScore": 15,
  "metaScore": 75,
  "method": "meta-elevated",
  "budget": "maximum",
  "budgetDetails": { "label": "Maximum", "pipelineDepth": "L90", "compactionDepth": "deep", "reflectionTrigger": true, "confidenceFloor": 90 },
  "base": {
    "score": 15,
    "method": "override",
    "details": {
      "category": "internal-tooling",
      "overrideScore": 15,
      "wasElevated": false,
      "elevationReason": null,
      "impactTier": null
    }
  },
  "meta": {
    "metaScore": 75,
    "weakestDimension": { "name": "Existing Wins at Stake", "value": 5 },
    "dimensions": {
      "agenticCompounding": 9,
      "hallucinationPropagation": 8,
      "existingWinsAtStake": 5
    }
  },
  "scoreThresholds": {
    "skipPipelineBelow": 20,
    "reflectAbove": 60,
    "fullValidationAbove": 80,
    "confidenceFloor": 90
  },
  "humanReadable": { "...": "the same numbers with long plain-language field names, meant to be dropped straight into agent context" }
}

There is no tier field on the response — don't key on one. There is no separate network round trip either: this is the literal return value of a pure function, so what you see above is exactly what's available, not a subset of a richer server response.

When two of the three risk answers (ac/hp/ew) are 8 or higher, metaScore is raised to a floor (80 by default, configurable via the criticalMassFloor switch — see alignment-harness config show) regardless of what the geometric mean alone would produce. This is deliberate: two seriously risky dimensions at once means the task gets full-depth treatment even if the third dimension looks mild.

If you call the scorer with none of ac, hp, ew set, meta comes back null. Any rule below that reads meta.dimensions… needs to check for that first.

Every variable you can condition on

Headline numbers (1-100, continuous)

Field What it means When to use it
unifiedScore The single number representing total reasoning depth needed. Max of baseScore and metaScore. Default for any "how careful should I be?" decision.
baseScore How important is this feature to users/business? Driven by category and impact tier (or your own risk table, once you've set one up). When you need to know business importance independent of agent risk.
metaScore How risky is this for an agent to execute? Geometric mean of ac/hp/ew × 10, with the critical-mass floor described above. When you care specifically about agent execution risk, not business value.

Method (what drove the score)

Value Meaning
"base-driven" Business importance was higher than (or tied with) agent risk. The feature matters to users more than it's dangerous for agents.
"meta-elevated" Agent risk exceeded business importance. The feature might be low-stakes business-wise but dangerous if the agent gets it wrong.

A tie between baseScore and metaScore reports "base-driven". Use method when you need to know WHY the score is what it is — the same score can mean different things depending on whether base or meta drove it.

Individual meta-dimensions (1-10 each)

Field path What it measures
meta.dimensions.agenticCompounding Does this change multiply across future agent decisions? 1 = one-off, 10 = changes how all future agents decide.
meta.dimensions.hallucinationPropagation If the agent hallucinates here, how far does the error spread? 1 = stays in this file, 10 = every future decision inherits the mistake.
meta.dimensions.existingWinsAtStake Could this break something already working? 1 = greenfield, 10 = touches auth/payment/shared middleware.

These are the most powerful fields for conditional logic — but check meta isn't null first (see above). An agent building a code review step might only care about hallucinationPropagation. An agent deciding whether to run integration tests might only care about existingWinsAtStake.

Weakest dimension

Field path What it tells you
meta.weakestDimension.name Which dimension scored lowest — the bottleneck holding the meta-score down.
meta.weakestDimension.value The actual score of that dimension (1-10).

Use this when you want to know what's pulling the score DOWN. If weakestDimension is "Existing Wins at Stake" at value 2, the feature is risky for agents but safe for existing code — you might skip regression tests.

Base scoring details

Field path What it tells you
base.method How base score was determined: "manual", "override" (category), "sila", "sila-over-override" (if you've set up the IX-SILA nine-question model as your risk table).
base.details.category The category used — from your own risk table if you've set one up (alignment-harness risk-table show), otherwise the shipped defaults.
base.details.wasElevated Did the touchesUserState gate raise the score above the category default?
base.details.elevationReason Human-readable explanation of why the score was elevated.
base.details.impactTier The impact tier that drove elevation: "indirect", "config", "communication", "revenue".

Score thresholds — use these instead of a bucket label

Field What it tells you
scoreThresholds.skipPipelineBelow Below this unifiedScore, skip the governer pipeline entirely (default 20).
scoreThresholds.reflectAbove At or above this, /reflect is required before acting (default 60).
scoreThresholds.fullValidationAbove At or above this, every quality gate applies (default 80).
scoreThresholds.confidenceFloor Minimum certainty (%) the agent should have before proceeding, scaled to unifiedScore.

budget (minimal/standard/elevated/maximum) is a legacy bucket label kept for backward compatibility — prefer unifiedScore plus scoreThresholds directly over branching on the label.

Example conditional patterns

"Only run deep security review when hallucination could spread to users":

if meta && meta.dimensions.hallucinationPropagation >= 7

"Skip regression tests for greenfield features":

if meta && meta.dimensions.existingWinsAtStake <= 3

"Require human review when agent risk is high but business stakes are low":

if method === "meta-elevated" AND baseScore < 30 AND metaScore > 70

"Extra care with anything that changes agent decision-making":

if meta && meta.dimensions.agenticCompounding >= 8

"Scale test coverage continuously with the score":

coverageTarget = 60 + (unifiedScore * 0.35)

"Different behavior depending on what drove the score":

if method === "base-driven"  → focus on UX validation
if method === "meta-elevated" → focus on code correctness and testing

"Only flag for staged release if it touches revenue AND agents could compound errors":

if base.details.impactTier === "revenue" AND meta && meta.dimensions.agenticCompounding >= 6

Building an adaptor

An adaptor is any skill or agent behavior that reads the adaptive score and adjusts its own depth based on the response. To build one:

  1. Call alignment-harness score --json (or the local scorer function directly) at the start of your workflow, or read the already-written score for this session with alignment-harness status --json if the governer already ran.
  2. Pick the fields you care about — you don't need all of them. A test depth adaptor might only use unifiedScore and existingWinsAtStake. A compaction-depth adaptor might only use unifiedScore and method.
  3. Write your conditional logic using the fields directly — no buckets, no label translation. Branch on numbers and strings. Guard every meta.* read with a null check.
  4. Document your conditions in your skill file so other agents know what you're keying on.

There is no adaptor registry or framework. Your skill reads the response, makes decisions, and documents those decisions in its own SKILL.md. The scoring system provides the data. Your skill owns its logic.

/governer runs this scorer as part of its own pipeline and writes the result to the session state file; alignment-harness status reads that file back. If you're adding a new step to the governer pipeline itself rather than just reacting to its score, see /governer-agent-pipeline-how-to-add-steps (if you have it).