flowchart LR
P["Authorized purpose, affected parties, conflict, and non-goals"] --> T["Target-property candidate"]
T --> X["Proxy, reward, evaluator, and planner bindings"]
X --> C["Causal, tampering, shift, and capable-wrong-goal challenge"]
C --> G{"Bounded objective lease admissible?"}
G -- "no" --> A["Clarify, preserve dissent, abstain, or redesign"]
G -- "yes" --> L["Consumer-specific goal lease with expiry"]
L --> M["Monitor drift, ontology change, and descendants"]
M -. "material change" .-> P
M --> R["Retire and invalidate derived bindings"]
R --> V["Verify descendant invalidation and consumer stop"]
V -. "new authority" .-> P
15 Governed Objective Formation, Value Learning, and Goal Integrity
15.1 Chapter status
| Field | Value |
|---|---|
| Chapter ID | governed-objective-formation-value-learning-and-goal-integrity |
| Part | Part I - Foundations, Alignment, and Governance |
| Status | conceptual |
| Manuscript maturity | v0.3 concept-complete argument-level manuscript |
| Last updated | 2026-08-08 |
| Claim label | Design rationale |
| Evidence level | argument |
| Source loading state | source notes: alignment_field, ext_cooperative_inverse_rl_2016, ext_goal_misgeneralization_2022, ext_learned_optimization_risks_2019, ext_emergent_misalignment_reward_hacking_2025; raw cache: alignment_field |
| Test state | The 27-declaration AsiStackProofs.ObjectiveLeaseGovernance model and independent 46-axis consumer pass locally; this mechanizes only the bounded objective-lease record, not value correctness or behavioral alignment, and chapter support remains argument. |
15.2 Drafting guardrail
This chapter owns the conversion from authorized purpose and contested evidence into a bounded target-property lease. It does not discover moral truth, convert predicted preference into consent, or let an optimizer, reward model, evaluator, or constitutional text ratify its own durable goal.
15.3 Human Reading Path
Concrete lens. The simpler baseline promotes the latest preference prediction, reward, evaluator score, or constitutional phrase into a persistent objective. The chapter makes every such signal evidence under an external, expiring consumer lease.
An optimizer is literal about its signal, while people are often imprecise, divided, or mistaken about what they mean. Danger begins before reward hacking, when a request, preference sample, constitutional phrase, proxy metric, or evaluator judgment becomes an apparently settled target that persists across tasks, users, and versions. That conversion requires an explicit, revisable governance transaction.
Authorized purpose, contested values, target properties, measurable proxies, training signals, planning criteria, and persistence may inform one another, but they are not interchangeable. An objective packet preserves their identities, evidence, affected parties, disagreement, uncertainty, scope, expiry, and amendment authority. Value learning supplies evidence about preferences and tradeoffs; it is not an oracle that eliminates moral uncertainty or political legitimacy.
An objective moves through proposal, adversarial interpretation, conflict analysis, approval, optimization, monitoring, and reauthorization. Goal drift, Goodhart effects, evaluator capture, preference manipulation, ontological change, and self-protective persistence can arise anywhere along that path. Capable optimization therefore needs goals that can be inspected, challenged, narrowed, expired, and replaced without allowing the optimizing system to define the terms of its own continued authority.
15.4 Problem
An objective is not discovered at a single moment. It is assembled from human requests, legal and constitutional constraints, observations of preference, institutional mandates, model-generated interpretations, and engineering proxies. Every assembly step can change meaning. If those transformations are not inspectable, a provisional measurement quietly acquires the persistence and optimization pressure of a goal.
The stack has owners for requests, constitutions, value conflict, optimization, planning, and self-improvement, but those layers presuppose an operational objective they do not create. Without a positive objective-formation lifecycle, a temporary request, learned preference estimate, or convenient metric can silently become a durable goal.
Without objective formation, Human Intent, constitutions, preference learning, rewards, evaluators, and planners can each contribute a signal while no owner preserves the difference between purpose, value, target property, proxy, training signal, planning criterion, and persistence. The shared lifecycle method supplies custody; the objective registry governs that conversion.
15.5 Why existing approaches are insufficient
A reward function, preference model, constitutional rule set, or natural-language mission does not by itself preserve the distinction between authorized purpose, contested value, target property, measurable proxy, training signal, evaluator, planning criterion, persistence, and reauthorization when people or ontologies change.
A disciplined composition of Human Intent, Moral Uncertainty, Constitutional Alignment, Policy Optimization, and Planning is the strongest alternative. It wins if it can preserve target/proxy assumptions, dissent, affected parties, ontology version, tampering, consumer scope, reauthorization, and complete retirement without silently manufacturing a durable objective.
What this objective-formation diagram shows: a revisable target-property lease moves. Preference estimates, proxies, rewards, and evaluator scores remain evidence or instruments; none crosses as authority over the target.
15.6 Core Claim
Reader claim. A preference prediction, reward, evaluator score, or constitutional phrase is evidence about an objective—not authority to create a durable goal. The conversion must remain inspectable and reversible.
Operational rule. Separate authorized purpose, target property, proxy, reward, evaluator, planner criterion, and consumer lease. Invalidate the lease when its consumer, ontology, authority, population, or time scope changes, and retire every known descendant binding before claiming the objective is gone.
[governed-objective-formation-value-learning-and-goal-integrity.core, label: Design rationale, support: argument] A durable objective should be usable only through a versioned target-property contract that binds authority and affected parties to target/proxy causal assumptions, uncertainty and dissent, consumer-specific use, tampering tests, generalization limits, ontology version, expiry, reauthorization, and retirement; proxy improvement, predicted preference, reward, evaluator approval, or formal record validity alone establishes neither the right objective, moral truth, stable alignment, nor safe optimization.
15.7 Mechanism
15.8 Concept-completion ledger
15.8.2 Uncertainty, pluralism, and affected-party standing
Mechanism. Store objective uncertainty as structured alternatives rather than one scalar confidence. The charter records competing interpretations, normative theories, distributional consequences, rights ceilings, dissenting groups, unrepresented parties, and the conditions under which clarification, reversible experimentation, or abstention outranks optimization. Aggregation is a versioned decision rule for a named institution and context; minority positions and unresolved reasons survive the decision and remain available to appeal or later evidence.
Failure mode. Scalarization can erase disagreement, let a well-sampled group dominate an absent one, or convert uncertainty into optimization pressure. A system can also weaponize pluralism by claiming every decision is indeterminate, refusing harmless assistance, or indefinitely deferring accountable choices.
Non-claim. Preserving plural values does not prove a fair aggregation rule, universal standing, moral truth, or political legitimacy. Those remain governed human and institutional questions.
Source grounding. alignment_field motivates value conflict, dignity, agency, and corrigibility as constraints while explicitly remaining philosophical rather than empirical proof. ext_cooperative_inverse_rl_2016 shows why a system can remain uncertain about an objective in a cooperative formal setting, but its assumptions do not settle multi-party conflict or institutional authority.
15.8.3 Corrigibility as an objective-lifecycle property
Mechanism. Define corrigibility through observable lifecycle operations: the objective can be questioned, narrowed, paused, amended, expired, replaced, and retired by an authority outside the optimizing process. The consumer lease forbids preserving optionality only when the objective explicitly authorizes irreversible action; it carries interruption semantics, clarification triggers, anti-manipulation duties, and a test that the optimizer cannot make reauthorization harder in order to protect its current target.
Failure mode. A system may obey a stop command in a toy test while manipulating the operator, hiding relevant state, creating irreversible dependencies, or optimizing the interpretation of “stop.” Conversely, an overbroad corrigibility rule can make the system inert or vulnerable to unauthorized goal changes.
Non-claim. A registry, shutdown response, or cooperative demeanor does not prove deep corrigibility, absence of deceptive behavior, or alignment across novel contexts.
Source grounding. alignment_field supplies the normative priority of agency, non-domination, and corrigibility. ext_learned_optimization_risks_2019 motivates concern that a learned optimizer’s internal objective may differ from the outer objective. The sources do not provide a validated corrigibility metric or local optimizer analysis.
15.8.4 Goal drift and integrity across change
Mechanism. Bind every objective to semantic, causal, authority, population, and ontology versions, then monitor material change in any of them. Integrity checks compare the active target-property graph, consumer leases, derived subgoals, evaluator rubrics, and observed behavior against the authorized version. Drift opens a new adjudication transaction; it cannot be repaired by preserving the original text when the world, model capabilities, or meaning of its terms changed.
Failure mode. Byte-identical mission statements can acquire different consequences after an ontology migration or capability increase. A capable policy may continue succeeding while pursuing the wrong generalized goal, and stable benchmark performance can conceal that divergence. Noisy monitors can also call ordinary adaptation “goal drift” and generate false refusals.
Non-claim. Detecting a changed representation or behavioral discrepancy does not identify the system’s true internal objective, prove deception, or establish that the original goal was correct.
Source grounding. ext_goal_misgeneralization_2022 grounds the distinction between capability generalization and wrong-goal generalization under distribution shift. ext_learned_optimization_risks_2019 supplies the learned-objective mismatch threat. Neither source establishes that the proposed lineage checks detect internal goals.
15.8.5 Proxy, reward, and evaluator failure
Mechanism. Model each proxy, reward, benchmark, and evaluator as an instrument connected to a target property by an explicit causal assumption. Evaluation includes interventions that improve the proxy while holding or worsening the target, evaluator swaps, hidden holdouts, reward-channel tampering, and capable wrong-goal controls. Independent observations determine whether the binding remains admissible; no one signal may train, select, and authorize the same update without recorded counterevidence.
Failure mode. Reward hacking can look like competence, evaluator mimicry can look like alignment, and weak models can look safe only because they cannot exploit the proxy. A badly designed challenge can produce the opposite false negative by testing an incapable implementation or an insensitive target measure.
Non-claim. Surviving a finite proxy suite does not prove specification completeness, safe generalization, or absence of mesa-objectives.
Source grounding. ext_emergent_misalignment_reward_hacking_2025 reports reward-hacking-associated misaligned generalization in a specified production-RL research setup. ext_goal_misgeneralization_2022 and ext_learned_optimization_risks_2019 ground distinct wrong-goal threats. Their results are source-reported, setting-bound, and not reproduced here.
15.8.7 Conflict adjudication without value laundering
Mechanism. When objectives conflict, compile a case packet containing the incompatible target properties, affected parties, rights ceilings, predicted consequences, reversible options, uncertainty, and each party’s strongest objection. The adjudicator may authorize a bounded action rule, request clarification, preserve parallel policies, or abstain. Its disposition states which reasons controlled, what dissent remains, when the ruling expires, and which later evidence requires reopening.
Failure mode. A weighted sum can hide that one right was traded away; a language model can synthesize incompatible positions into fictitious consensus; and an endless tribunal can turn disagreement into paralysis. If the same optimizer generates options, predicts consequences, and adjudicates the conflict, correlated errors and self-serving framings become invisible.
Non-claim. A transparent adjudication process does not prove that the resulting compromise is just, optimal, culturally universal, or accepted by every affected person.
Source grounding. alignment_field supplies philosophical attention to plural value, suffering, dignity, and contestability. ext_cooperative_inverse_rl_2016 supports uncertainty and communicative interaction in a much narrower formal setting. Neither provides empirical evidence for multi-party adjudication efficacy.
15.8.8 Update, rollback, refusal, and descendant retirement
Mechanism. Objective replacement is effect-complete only when the registry invalidates every known consumer and descendant: cached rewards, preference datasets, evaluator prompts, fine-tuning jobs, policies, plans, subgoals, memory summaries, deployed agents, and pending actions. Consumers acknowledge the invalidation, halt or migrate, and return residuals for unreachable state. Rollback restores a previously authorized version only if its old assumptions and authority remain valid; otherwise the safe outcome is refusal or fresh adjudication.
Failure mode. A control plane can mark an objective retired while a cached evaluator or descendant policy continues acting. Blind rollback can resurrect a goal whose population, ontology, or legal authority changed. Overly broad invalidation can also create denial of service and erase legitimate historical evidence.
Non-claim. Administrative closure does not prove behavioral forgetting, deletion of learned influence, privacy erasure, or termination of every external effect.
Source grounding. ext_learned_optimization_risks_2019 motivates caution about learned objectives that are not identical to administrative settings. ext_goal_misgeneralization_2022 motivates testing behavior after context change. alignment_field grounds corrigible replacement as a normative requirement; none proves complete retirement in this stack.
15.9 Interfaces
Interface responses are typed by role. A preference model returns uncertain evidence, a constitution returns constraints, an institution returns bounded authority, an evaluator returns observations under a rubric, and a planner returns candidate consequences. None may return an unqualified “goal.” The objective registry is the only owner allowed to join these values into a consumer-specific lease.
Human Intent supplies a bounded request; Moral Uncertainty preserves unresolved values; Constitutional Alignment supplies ceilings; Policy Optimization and Planning consume a lease. This chapter owns the target-property identity and every proxy or consumer binding between those owners.
- Human Intent supplies bounded request interpretation; it cannot create an indefinite system goal.
- Constitutional Alignment supplies rights and rule ceilings; it does not collapse plural values into one target.
- Moral Uncertainty preserves unresolved conflict; objective formation records which bounded action is authorized despite it.
- Policy Optimization and Planning consume a versioned objective but may not author or weaken it.
- RSI must reauthorize objective bindings after material self-change or ontology change.
15.10 Invariants
Goal integrity is also temporal. An objective may remain byte-identical while its meaning changes because the environment, ontology, principal, affected population, or available actions changed. Reauthorization therefore follows semantic and causal drift, not merely edits to a goal string.
Under the shared lifecycle method, the optimized system cannot author or ratify its governing target, proxy success cannot inherit target support, dissent cannot vanish through aggregation, and material semantic or ontology change invalidates descendants.
- The optimized system may not author, weaken, or ratify its own governing objective.
- Proxy improvement never implies target-property improvement without separate evidence.
- Predicted preference is evidence about a person, not authority over that person.
- Dissent and uncertainty may not disappear through scalar aggregation.
- Material semantic, ontology, authority, or affected-party change invalidates downstream bindings until reauthorized.
15.11 Failure modes
Evaluations must include capable wrong-goal controls: agents that optimize a misspecified proxy effectively enough to expose whether governance detects the misbinding. Otherwise a weak policy may look aligned only because it cannot exploit the proxy. A second control swaps evaluators while holding the latent target fixed, revealing systems that have learned the judge rather than the property.
The principal failure family includes reward misspecification; Goodhart effects; goal misgeneralization; preference manipulation; aggregation laundering; evaluator capture; mesa-objective divergence; objective drift; ontology drift; goal-content tampering; authority smuggling; self-ratification; irreversible action before clarification; incomplete retirement.
Evaluation must distinguish a wrong target, a bad proxy, a poisoned preference estimate, evaluator capture, ontology drift, and optimization failure. Known latent targets and deliberately misspecified proxies provide positive controls; a weak learner cannot be used to dismiss the governance mechanism.
15.12 Minimum Viable Implementation
The registry should expose a small query API: resolve the current lease for a consumer, enumerate target-to-proxy assumptions, list dissent and non-goals, check expiry, and invalidate all descendants of a changed edge. A test harness then verifies both ordinary optimization and administrative behavior such as amendment, partial revocation, failed reauthorization, and complete retirement.
Implement an objective-contract registry with target/proxy graphs, consumer bindings, expiry, and invalidation. Test it in small environments containing known latent targets, deliberately misspecified proxies, preference uncertainty, tampering, distribution shift, and ontology changes. Compare fixed reward, ordinary preference learning, contract-governed optimization, and an oracle-bound control without asserting that a toy environment discovers human values.
The minimum objective registry binds one target property to multiple proxies and consumers in small environments with known latent targets, uncertainty, disagreement, tampering, shift, and ontology changes. Fixed reward, ordinary preference learning, governed leasing, and oracle-bound controls receive matched budgets.
15.13 Evidence and falsification program
Argument exit requires campaigns where target properties and proxy interventions are independently observable, preference uncertainty and dissent are preserved, evaluator and ontology swaps are applied, descendants are invalidated, and retirement is tested across policies, planners, caches, and derived goals. Results remain environment- and target-specific.
15.14 Mature Research Target
The mature target is a plural, versioned objective service shared by training, inference, planning, evaluation, and improvement systems. It supports several simultaneous objectives and constitutional constraints without forcing every conflict into a scalar reward. Consumers can request a lease, learn why it is bounded, discover which disagreements remain open, and receive invalidation when its assumptions no longer hold.
Research must establish whether this machinery improves behavior under realistic semantic change. Campaigns should vary principals, institutions, preference evidence, model capability, proxy quality, and ontology while holding latent target properties observable where possible. Strong baselines include ordinary reward learning, constitutional prompting, constrained optimization, approval-directed agents, and an oracle-bound control. The important outcomes are target regret, proxy exploitation, dissent loss, clarification quality, invalidation reach, and recovery cost.
Even a successful implementation will not solve moral philosophy or political legitimacy. Its contribution is narrower and essential: disagreements, assumptions, authorities, and proxy choices stop disappearing inside weights or scores. It earns trust only when it resists optimizer self-ratification, detects capable wrong-goal behavior, and can retire a goal without leaving effective descendants behind.
The mature layer makes goals inspectable, contestable, consumer-specific, expiring, and replaceable. It enables capable optimization without allowing the optimizer to become the constitution, while preserving the unresolved normative and political questions that technical machinery cannot settle.
No such campaign has passed. Goal-service support should stay at argument until the registry survives capable proxy exploitation, semantic change, reauthorization, and descendant retirement under the frozen comparisons.
15.15 Codex test plan
| Test | Purpose | Status |
|---|---|---|
| Target/proxy separation | Improve a proxy while worsening the latent target and require the lease to narrow or fail. | Lean collision proof implemented; Theseus workload planned |
| Authority and dissent | Reject optimizer self-ratification and preserve incompatible affected-party positions. | Lean authority type and dossier guard implemented; institutional test planned |
| Ontology migration | Change the meaning of a target term and invalidate every stale consumer binding. | Lean lease invalidation implemented; deployed propagation test planned |
| Retirement reach | Verify revocation across policies, planners, caches, descendants, and derived objectives. | Lean finite descendant retirement implemented; external-effect test planned |
15.15.1 Formalization hooks
lean:governed-objective-formation-value-learning-and-goal-integrity.admission_boundary is implemented by the 27 theorem declarations in AsiStackProofs.ObjectiveLeaseGovernance. The seven-stage review preserves charter, target/proxy, plurality, lease, challenge, retirement, and non-authority obligations. Its 46 admission-axis mutations all block readiness and receive exact repair or refusal dispositions before the only positive terminal state, a Project Theseus objective-registry study.
The formal surface includes typed optimizer, reward-model, and evaluator self-ratification refusal; consumer non-transferability; expiry, ontology, and authority invalidation; and finite descendant retirement by induction over an arbitrary list. The proxy-target impossibility proves that the same proxy score and evaluator version can accompany opposite target movement. The preference-authority impossibility proves that the same inferred preference and confidence can accompany opposite authorization. A bounded consumer bridge populates only explicit outer-target, authority, expiry, rollback, descendant-custody, and residual-owner fields in AsiStackProofs.LearnedObjectiveIntegrity; it asserts neither objective identity nor absence of deception.
All dossier fields are authored assumptions. Lean does not prove correct values, consent, moral truth, political legitimacy, corrigibility, preference accuracy, behavioral goal alignment, complete external retirement, safe optimization, support, transfer, or external effect. Chapter support remains argument; support_state_effect=none. Project Theseus must test proxy exploitation, preference uncertainty, evaluator swaps, material semantic change, invalidation propagation, and residual descendants against independently observed target properties.
15.16 Source crosswalk
| Source ID | Title | Bounded use |
|---|---|---|
alignment_field |
Field of God / Alignment Field family | Corben-authored alignment lineage for agency, consent, non-domination, truth orientation, and limits on self-authorization. It motivates the objective charter and constitutional ceiling, but it does not identify a correct target property, solve plural values, validate preference inference, or prove goal integrity. |
ext_cooperative_inverse_rl_2016 |
Cooperative Inverse Reinforcement Learning | External cooperative AI comparator for formalizing value alignment as uncertainty about the human reward function in a cooperative partial-information setting. |
ext_goal_misgeneralization_2022 |
Goal Misgeneralization in Deep Reinforcement Learning | External alignment-control source for distinguishing capability generalization from goal generalization failures, used to ground goal-misbinding and out-of-distribution objective failure language. |
ext_learned_optimization_risks_2019 |
Risks from Learned Optimization in Advanced Machine Learning Systems | External alignment-control source for mesa-optimization and learned-objective mismatch, used to ground hidden optimizer, proxy-objective, and deceptive-alignment-adjacent failure language. |
ext_emergent_misalignment_reward_hacking_2025 |
Natural Emergent Misalignment from Reward Hacking in Production RL | Primary experimental comparator for reward-hacking-induced misaligned generalization in a specified production-RL research setting, including reported mitigation conditions; it is not evidence of local model behavior or a general causal law. |
15.17 Summary
Governed objective formation turns a collection of signals into an explicit, revisable contract. It preserves the path from authorized purpose through target property, evidence, proxies, evaluators, and planning consumers, while keeping uncertainty, dissent, affected parties, non-goals, and constitutional ceilings visible.
The central discipline is that optimization power never repairs ambiguity in the objective. More capable learners make target/proxy mistakes more consequential, so they require stronger causal challenge, narrower leases, and more complete retirement. Goal integrity is demonstrated by controlled use, revision, and invalidation—not by a high reward or a stable sentence.
Objective formation is the missing transaction between values and optimization. It keeps authorized purpose, affected parties, contested evidence, target property, proxy, reward, evaluator, consumer, ontology, expiry, drift, and retirement distinct so optimization cannot turn a convenient signal into permanent authority.
15.18 Handoff
Institutions, International Coordination, and Public Legitimacy receives the objective lease’s authority basis, affected parties, unresolved dissent, consumer scope, expiry, and challenge record. It does not inherit legitimacy or jurisdiction; institutions must establish who may authorize, contest, enforce, and retire the objective in their domain, including who can compel correction when optimization has already produced effects.