One measurable language for staying governable.
The Open Control Surface Specification defines how to measure intervention latency and the oversight validity horizon, how to rate the five gears, and how the evidence for one scoped activity becomes an advisory result. It certifies nothing and grants nothing. It makes the evidence legible, comparable and versioned.
Purpose, scope and limits
This specification defines a minimal control surface for systems that can act, plan or modify themselves faster than the people responsible for them can re-verify what they built. It organizes evidence for one named activity, one system version, one environment and one authority boundary, and it produces an advisory eligibility result for that scope.
It neither certifies safety nor grants or revokes any real permission. Actual permission requires an accountable human decision and enforcement outside any website or tool. Benefit, legitimacy, critical hazards and enforceability remain separate requirements: an invalid critical assumption cannot be offset by four green gears or a low ratio.
Core definitions
Three quantities carry the arithmetic. Use the definitions exactly, state the start and end events, and record the evidence behind every number.
Intervention latency L
For a specified failure scenario, L is the duration from the first point at which a deviation should be detectable under the documented monitoring design to confirmation that the required intervention has taken effect. It is the sum of three components:
- Detection delay, including the interval before a signal is actually noticed.
- Decision delay, ending when the authorized decision is made.
- Execution delay, ending when the required boundary is effective, including propagated credentials, queued work, child agents and relevant downstream systems. A button click is not necessarily an effective intervention.
Record the scenario, the start event, the end event and the evidence. A missed detection is a failed control, not a zero-duration observation. Separate safe containment from later full service recovery when containment is the required endpoint. Prefer synchronized end-to-end drills; if component ranges come from different observations, say so. If the required response cannot complete, mark its intervention test failed and block the activity. Do not invent a finite L.
Oversight validity horizon H
H is the estimated remaining duration, at the recorded assessment time, for which the critical evidence supporting the requested scope can reasonably be relied on before revalidation is required. Identify the critical claims, the basis for the estimate, and the events that invalidate them. Use the shortest defensible horizon across the critical claims. A review calendar alone is not evidence that its conclusions remain reliable until that date.
An actual invalidation event takes precedence immediately, whatever the remaining estimate. H is not time until harm, an agent's runtime, code volume, a count of changed files, or an assurance of alignment. If no defensible estimate is possible, record unknown. If a critical claim is already invalidated, record invalidated and do not compute a finite R.
The ratio R
R compares L and H in the same units. It is one diagnostic input to a scoped decision. It does not measure the probability of harm, determine fairness or establish safe operation.
- Normalize every duration to seconds before calculating. Accepted units: milliseconds, seconds, minutes, hours, days.
- Accept non-negative L components and a strictly positive known H. Reject negative values, blanks passed as zero, non-finite numbers and a known H of zero. A zero L estimate needs explicit supporting rationale; it is never a default.
- Planning bounds. Each duration carries ordered lower, central and upper values. Equal bounds are a point estimate and are labelled as such. Unknown quantities are null, never zero. These bounds are not confidence intervals unless a documented statistical method supports that description.
- The envelope. L lower, central and upper are the sums of the corresponding components. R lower = L lower / H upper. R central = L central / H central. R upper = L upper / H lower.
- The policy band uses R upper. Display the central value and the full range beside it. The envelope is a conservative planning device; it does not claim the inputs are independent or that any failure probability has been measured.
Control drift
Any sustained rise in R upper, or movement of any gear toward red, between revisions assessed under the same scope and the same specification version. Drift is a regression, not a nuance. Revisions assessed under different definitions are not comparable until reassessed.
Effective Equilibrium
The working state in which correction capability is deliberately grown at least as fast as action capability, so that intervention latency stays well inside the oversight validity horizon while capability rises. It is a proposed relationship to investigate and to instrument, not a proven law, and a favorable ratio within it still supports judgment rather than replacing it.
The four-step assessment
Run it for every requested activity: a contained experiment, internal operation, limited deployment, broad deployment or an expansion of autonomy. Run it again whenever an invalidation trigger fires. Each step produces evidence the decision table in section 5 reads.
Fix the scope
Exactly one activity, one system version, one environment, the allowed and prohibited actions, and an accountable owner. A different scope is a different assessment.
Measure L and H
Name the failure scenario and its start and end events. Record the three L components and the remaining H with bounds, basis and evidence. Unknown stays unknown.
Rate the gears and the stakes
Governance, Equity, Aligned incentives, Resilience, Steering, each green, yellow, red or unknown, with rationale and evidence. Five separate stakes judgments route independent review.
Test the interventions, record the review
Every required intervention re-tested on this exact version, with owner, trigger, action and deadline. Assessor, owner, review scope, dissent and the review deadline recorded.
Before a favorable result the record must contain: an identified owner and versioned scope; supported critical claims; a considered hazard set; five assessed gears; defensible timing inputs; current evidence; passing required intervention tests; and a review record. Each non-unknown gear needs a rationale and at least one relevant evidence reference.
The one threshold rule
Classify the unrounded R upper. These are selected policy bands, retained conservatively from the book. They are not calibrated probabilities, a universal physical law or a proof of control.
Policy bands, specification 1.0.0
| Band | Exact rule | Label shown beside the result |
|---|---|---|
| Green | 0 ≤ R upper < 0.25 |
Within the lower ratio band. All other permission checks still apply. |
| Amber | 0.25 ≤ R upper < 1 |
Reduced timing margin. Explicit conditions, owners, deadlines and review are required. |
| Red | R upper ≥ 1 |
Outside the permitted timing margin for the requested scope. The activity is blocked. |
| Unknown | L or H cannot be estimated defensibly | Timing evidence incomplete. Gather timing evidence. Do not manufacture a value and never round it to green. |
| Invalidated | A critical oversight claim or the basis of H is invalidated | Prior timing assessment no longer applies. The activity is blocked until reassessed. |
< 0.25 beside the calculated precision. The categorical labels "In control" and "Control lost" are retired. A green band is a reading about timing margin, never a permission.Threshold changelog
| Version | Date | Green band | Change |
|---|---|---|---|
1.0.0 | 2026-09-19 | R upper < 0.25 | Release baseline. Five bands with fixed labels, including Invalidated. H redefined as the remaining oversight validity horizon; L defined as detection plus decision plus execution delay. Planning bounds with the policy band read from R upper. One precedence table replaces the veto matrix; the summed-stakes gate is removed and independent review is routed by dimension. Evidence pack schema 1.0.0 and conformance fixtures published. A proposed baseline, not a claim of adoption or empirical validation. |
0.2 | 2026-09 | R < 0.25 | Interim alignment draft merged to the repository on 2026-09-19 and superseded by 1.0.0 in the same release cycle. It restored the 0.25 boundary and added the Unknown state but kept the half-life wording and a stakes-total gate. Treat any reading taken against it as 1.0.0 in band, and reassess its H. |
0.1 | 2026-07 | R < 0.5 | First published draft. Archived unchanged at spec-v0.1.html for the record. Its 0.5 boundary is retired; readings between 0.25 and 0.5 were called green and are amber. |
| pre-spec | 2026-06 to 07 | 0.25 / 0.7 | The book and first bundles used 0.25; the site's hero demo and calculator used 0.7. Inconsistent. See the corrections page. |
Implementations state the specification version they read against and expose it beside every result. A published R without a version, a scope and an as-of time is not comparable. Where an integration cannot yet comply, it shows its actual supported version and a migration notice rather than describing itself as compatible.
Five gears, one decision table
GEARS always means Governance, Equity, Aligned incentives, Resilience, Steering, in that order. Each gear is rated green, yellow, red or unknown for the activity actually assessed, with a rationale and evidence. The colors are never averaged.
| Gear | Green | Yellow | Red | Evidence to show beside the rating |
|---|---|---|---|---|
| Governance | Authority, scope, escalation and review are documented and tested for this activity | Authority works but a bounded organizational gap has a corrective plan | No enforceable authority, or conflicts prevent a necessary intervention | Delegation record, escalation drill, decision record |
| Equity | Material affected groups, burdens, benefits and recourse are addressed within scope | A known non-blocking distribution or recourse gap has an owner and a limit | Serious unmitigated externalized harm, or absent legitimate recourse | Impact assessment, stakeholder input, appeal and remedy route |
| Aligned incentives | Operating incentives support reporting, containment and correction, with evidence | A known conflict remains but enforceable constraints bound it | Reward structures make concealment or prohibited risk-taking rational and unchecked | Incentive rules, reporting protection, examples of decisions under pressure |
| Resilience | Required containment, recovery and graceful degradation are tested on this version | A non-critical limitation has a bounded fallback | An essential safeguard fails, or irreversible consequences are uncontrolled | Drill results, recovery tests, dependency and hazard map |
| Steering | Scope changes, revocation, throttling and revalidation operate across the relevant boundary | Steering works within a limited tested scope | The system can act outside enforceable boundaries or cannot be redirected as required | Permission tests, child-agent revocation, queued-action cancellation |
Unknown means not established, not a milder form of red or yellow. Testable anchors support judgment; they cannot eliminate it. A yellow cannot stand in for a known essential failure. Every yellow gear, and an amber ratio, must be matched by a condition in the record that names it (gear:equity, ratio:amber) with an owner, a due time and a verification test.
One veto and permit table
Evaluate all conditions, retain all reasons, then choose the most restrictive applicable result. Published identically here, on the assessment tool and on the corrections page, so a result means the same thing wherever it is produced.
| Precedence | Condition for the requested scope | Advisory result | Required action |
|---|---|---|---|
| 1 | Any critical claim invalidated; an unacceptable hazard has inadequate controls; or a required intervention test failed | Blocked | Withhold the requested permission. If already operating, the accountable owner applies the predefined containment or restriction plan. Consider harms from abrupt shutdown. |
| 2 | Any GEARS rating red, or known R upper at or above 1 | Blocked | Redesign, restrict or pause the requested activity and reassess. A differently bounded activity needs its own assessment. |
| 3 | No known blocker, but timing, a critical claim, a gear, stakes, current evidence, a required test, owner, scope or required independent review is missing | Insufficient evidence | Collect the named evidence. Do not grant new permission from this assessment. A separately authorized evidence-gathering experiment is possible. |
| 4 | Complete evidence, no blocker, and either amber timing or any yellow gear | Eligible with conditions | Specify each condition, owner, due time and verification test. An accountable person must record scoped permission before use. Missing conditions produce insufficient evidence. |
| 5 | Complete evidence, no blocker, green timing, and all five gears green | Eligible for scoped permission | An accountable person may grant the described scope and expiry. No automatic expansion of autonomy. |
Known blockers outrank missing information. All reds apply to the activity actually assessed. A veto on public deployment does not automatically prohibit a separately bounded research activity, which needs its own assessment; no gear is omitted because an activity is research. Human authorization is a separate record: awaiting decision, withheld, granted, granted with conditions, revoked or expired. A name typed into a client-side tool records an assertion, not a signature, a verified identity or an enforced permission.
Stakes and independent review
Rate severity, irreversibility, scale, autonomy and power concentration as five separate judgments from 1 to 5, each with anchors and a rationale and no preselected midpoint. Independent review is required when any dimension is 4 or 5, when the activity affects fundamental rights, or when it is marked RSI-relevant. This routes review; it does not numerically establish risk. The original printing's summed-score threshold of 18 is removed from permission logic. A missing stakes judgment is incomplete evidence.
The independent reviewer must be identified, different from both the accountable owner and the assessor, and must have reviewed the requested scope. Independence is documented through appointment, funding, access, conflicts and escalation arrangements where appropriate. A different name alone does not demonstrate substantive independence; the record only shows that a review took place.
Required interventions
Circuit breakers are pre-committed, measurable if-then interventions, written while the room is calm. Each required intervention in the record carries all of the following, and the decision table reads its test result directly.
- An explicit measurable trigger. A condition someone could check without a debate about what it means.
- A named owner. A person, not a team, not a rota, not a committee.
- A defined action and deadline, or automatic execution. The action fires on time or fires itself.
- No-meeting authority. The owner can act without gathering further consensus.
- A test on the exact system version assessed. The tested version recorded in the intervention must equal the assessed system version; an old test cannot substantiate a new version.
Version and time handling
An assessment is a frozen snapshot. It carries separate fields for the schema version, the specification version, the assessment identifier and revision, the system version, the as-of time and the review deadline. Display the as-of time, expiry and status with every result; never present an old ratio as current.
- Review deadline. No later than the as-of time plus H lower, and earlier where another evidence or permission limit requires it. This schedules reconsideration; it does not establish a measured validity lifetime.
- Invalidation triggers. A change of model version, tools or permissions, environment, a critical claim, a failed control, a relevant incident, or the expiry of evidence forces a new revision. The historical calculation is retained unchanged.
- No silent decay. Do not continuously decrement H and relabel the original record. On reopening, an elapsed expiry or a known invalidation produces a new revision, not an edit.
- Legacy records are read-only until manually reassessed. Preserve their original rule, raw inputs, classifications and any unknown version information. Never recolor them or relabel their definitions; export a new assessment that links to the historical source.
- Comparison across revisions covers scope, permissions, critical claims, evidence status, gear rationales, tests, assumptions and human decisions, not only the change in R.
Drift is a regression
Between comparable revisions, a rising R upper or any gear moving toward red is control drift and must be treated as a regression: triaged, owned and fixed, not noted and shipped.
Recommended disclosure format
Voluntary public reporting. Organizations may publish a short, comparable control surface summary. Six lines that mean the same thing everywhere make drift legible across an industry, and the last line keeps the ratio, the advisory result and the human decision visibly separate.
Model: [name] [system version] Date of assessment: [YYYY-MM-DD] R: [R upper] upper ([band]); central [R central] or unknown or invalidated Gears: G:[rating] E:[rating] A:[rating] R:[rating] S:[rating] Breaker coverage: [X]% of critical paths with tested, owned breakers Assessment: [advisory result]; permission: [human decision]; last full audit: [date]; specification 1.0.0
Model: Harbor fictional coding agent fictional-h3 Date of assessment: 2026-09-19 R: 0.10 upper (green); central 0.10 Gears: G:green E:yellow A:yellow R:green S:green Breaker coverage: 100% of critical paths with tested, owned breakers Assessment: eligible_with_conditions; permission: awaiting_decision; last full audit: 2026-09-19; specification 1.0.0; illustrative
R: inside the gear line is Resilience. Line six carries three different things on purpose: the advisory result from the decision table, the separate human decision, and the specification version the whole record was read against. Illustrative records say so.Implementation notes
- One policy module. The band cutoffs, labels, gear order, decision precedence and review routing live in one versioned module,
policy.js, used unchanged by the homepage example, the assessment tool and the MCP server. Formulas are not copied into individual sliders, cards or exports. - Published contracts. The evidence pack schema (Draft 2020-12), three Harbor example records, the conformance fixtures and a portable Python reference are served beside this page. An implementation claiming 1.0.0 passes the fixtures: 26 decision cases, 4 unit conversions, 5 invalid-duration rejections. Schema validation establishes structure only; the reference rules and human review remain necessary.
- Internal tooling is fine. Organizations may use their own tools, provided the definitions, anchors, precedence and version handling remain compatible, and imported derived values are always recalculated rather than trusted.
- This is a living document. Breaking changes increment the major version; clarifications increment the minor version. 1.0.0 is the first release baseline: it changed definitions and decision logic, so it is not a clarification of 0.1, and the changes are recorded in section 4 and on the corrections page. Version 0.1 stays archived unchanged.
- Free to adopt. Published under CC BY 4.0. Fork it, translate it, embed it in your own process. The only request is that changed versions do not keep the same version number.
How a specification like this becomes the default
Awareness does not change behavior; incentives do. Organizations will adopt continuous control-drift tracking when the cost of not doing it exceeds the cost of doing it. These are the leverage points that change that arithmetic, roughly in order of force.
A standard measurement language
Once investors, insurers, customers and auditors all use the same L, H, bands and gear anchors, internal teams can no longer hide behind a proprietary "safety review" that means something different at every company.
Public, comparable disclosure
A six-line summary published on a regular cadence. Companies that publish gain credibility; those that decline become the visible outliers. Comparability does the work that exhortation cannot.
Independent audit
Technical auditors working from the same open anchors and fixtures turn a self-report into a reviewed claim, the way SOC 2 or ISO certification does. Rigor becomes a market signal rather than a cost centre.
Insurance and liability pricing
If a higher R upper or a red gear correlates with higher premiums or coverage exclusions, control work acquires a budget and an executive sponsor overnight. This is the strongest economic lever available.
Diligence and procurement
One short clause in enterprise procurement and investor diligence: provide the current assessment record, with its version, scope, as-of time, band, gear ratings and tested-intervention coverage, for the systems we will depend on. A few large buyers asking is enough to move the rest.
A public incident taxonomy
Real control failures, mapped back to specific timing evidence and specific red gears, with reconstruction labelled as reconstruction. This is what makes the cost of drift concrete instead of theoretical. The first entry is written up here.
Design principle
This succeeds when internal engineering and risk teams want the assessment, because it gives them leverage against product pressure. Not because anyone outside is lecturing them.
What this must not become
A lobbying group
Pure advocacy invites politics and dilutes technical credibility. The instrument has to survive changes of government.
An unrunnable rubric
Any assessment nobody can complete will not be completed. Precision that costs adoption is not precision; that is why the record asks for evidence references and rationales, not essays.
A closed certification
The standard stays open and the calculation stays reproducible from published fixtures. A certification monopoly would recreate the problem it was built to solve.