Skip to main content

15  Governed Objective Formation, Value Learning, and Goal Integrity

15.1 Chapter status

Field Value
Chapter ID governed-objective-formation-value-learning-and-goal-integrity
Part Part I - Foundations, Alignment, and Governance
Status conceptual
Manuscript maturity v0.3 concept-complete argument-level manuscript
Last updated 2026-08-08
Claim label Design rationale
Evidence level argument
Source loading state source notes: alignment_field, ext_cooperative_inverse_rl_2016, ext_goal_misgeneralization_2022, ext_learned_optimization_risks_2019, ext_emergent_misalignment_reward_hacking_2025; raw cache: alignment_field
Test state The 27-declaration AsiStackProofs.ObjectiveLeaseGovernance model and independent 46-axis consumer pass locally; this mechanizes only the bounded objective-lease record, not value correctness or behavioral alignment, and chapter support remains argument.

15.2 Drafting guardrail

This chapter owns the conversion from authorized purpose and contested evidence into a bounded target-property lease. It does not discover moral truth, convert predicted preference into consent, or let an optimizer, reward model, evaluator, or constitutional text ratify its own durable goal.

15.3 Human Reading Path

Concrete lens. The simpler baseline promotes the latest preference prediction, reward, evaluator score, or constitutional phrase into a persistent objective. The chapter makes every such signal evidence under an external, expiring consumer lease.

An optimizer is literal about its signal, while people are often imprecise, divided, or mistaken about what they mean. Danger begins before reward hacking, when a request, preference sample, constitutional phrase, proxy metric, or evaluator judgment becomes an apparently settled target that persists across tasks, users, and versions. That conversion requires an explicit, revisable governance transaction.

Authorized purpose, contested values, target properties, measurable proxies, training signals, planning criteria, and persistence may inform one another, but they are not interchangeable. An objective packet preserves their identities, evidence, affected parties, disagreement, uncertainty, scope, expiry, and amendment authority. Value learning supplies evidence about preferences and tradeoffs; it is not an oracle that eliminates moral uncertainty or political legitimacy.

An objective moves through proposal, adversarial interpretation, conflict analysis, approval, optimization, monitoring, and reauthorization. Goal drift, Goodhart effects, evaluator capture, preference manipulation, ontological change, and self-protective persistence can arise anywhere along that path. Capable optimization therefore needs goals that can be inspected, challenged, narrowed, expired, and replaced without allowing the optimizing system to define the terms of its own continued authority.

15.4 Problem

An objective is not discovered at a single moment. It is assembled from human requests, legal and constitutional constraints, observations of preference, institutional mandates, model-generated interpretations, and engineering proxies. Every assembly step can change meaning. If those transformations are not inspectable, a provisional measurement quietly acquires the persistence and optimization pressure of a goal.

The stack has owners for requests, constitutions, value conflict, optimization, planning, and self-improvement, but those layers presuppose an operational objective they do not create. Without a positive objective-formation lifecycle, a temporary request, learned preference estimate, or convenient metric can silently become a durable goal.

Without objective formation, Human Intent, constitutions, preference learning, rewards, evaluators, and planners can each contribute a signal while no owner preserves the difference between purpose, value, target property, proxy, training signal, planning criterion, and persistence. The shared lifecycle method supplies custody; the objective registry governs that conversion.

15.5 Why existing approaches are insufficient

A reward function, preference model, constitutional rule set, or natural-language mission does not by itself preserve the distinction between authorized purpose, contested value, target property, measurable proxy, training signal, evaluator, planning criterion, persistence, and reauthorization when people or ontologies change.

A disciplined composition of Human Intent, Moral Uncertainty, Constitutional Alignment, Policy Optimization, and Planning is the strongest alternative. It wins if it can preserve target/proxy assumptions, dissent, affected parties, ontology version, tampering, consumer scope, reauthorization, and complete retirement without silently manufacturing a durable objective.

flowchart LR
  P["Authorized purpose, affected parties, conflict, and non-goals"] --> T["Target-property candidate"]
  T --> X["Proxy, reward, evaluator, and planner bindings"]
  X --> C["Causal, tampering, shift, and capable-wrong-goal challenge"]
  C --> G{"Bounded objective lease admissible?"}
  G -- "no" --> A["Clarify, preserve dissent, abstain, or redesign"]
  G -- "yes" --> L["Consumer-specific goal lease with expiry"]
  L --> M["Monitor drift, ontology change, and descendants"]
  M -. "material change" .-> P
  M --> R["Retire and invalidate derived bindings"]
  R --> V["Verify descendant invalidation and consumer stop"]
  V -. "new authority" .-> P

What this objective-formation diagram shows: a revisable target-property lease moves. Preference estimates, proxies, rewards, and evaluator scores remain evidence or instruments; none crosses as authority over the target.

15.6 Core Claim

Reader claim. A preference prediction, reward, evaluator score, or constitutional phrase is evidence about an objective—not authority to create a durable goal. The conversion must remain inspectable and reversible.

Operational rule. Separate authorized purpose, target property, proxy, reward, evaluator, planner criterion, and consumer lease. Invalidate the lease when its consumer, ontology, authority, population, or time scope changes, and retire every known descendant binding before claiming the objective is gone.

[governed-objective-formation-value-learning-and-goal-integrity.core, label: Design rationale, support: argument] A durable objective should be usable only through a versioned target-property contract that binds authority and affected parties to target/proxy causal assumptions, uncertainty and dissent, consumer-specific use, tampering tests, generalization limits, ontology version, expiry, reauthorization, and retirement; proxy improvement, predicted preference, reward, evaluator approval, or formal record validity alone establishes neither the right objective, moral truth, stable alignment, nor safe optimization.

15.7 Mechanism

15.7.1 Worked objective lease: the proxy agrees, the authority does not

The authored objective dossier binds target version 7 to consumer 41, ontology 12, authority version 5, population version 3, and an expiry at tick 20. It keeps preference evidence, reward, evaluator, and planner roles separately typed; preserves alternatives, dissent, unrepresented parties, and rights ceilings; and indexes three active descendant bindings. The optimizing system, reward model, and evaluator are explicitly denied the power to ratify the objective that governs them.

Now imagine two records with the same preference prediction. In one, the principal authorizes the lease; in the other, authority has changed or the consumer is different. The local model proves that the shared prediction cannot recover the missing authorization. It likewise shows that an identical proxy observation can accompany opposite movement in the target property. Consumer transfer, expiry, ontology drift, and authority drift each invalidate use, while retirement walks every member of the finite descendant list. The 46-axis review tests the record logic; it does not discover a correct value or show that the external world contains no forgotten descendant.

Objective formation begins by separating five objects that implementations often collapse: the principal’s authorized purpose, the underlying property the principal wants affected, observable evidence about that property, the proxy used for learning or control, and the action policy that consumes the proxy. Each edge carries an author, justification, confidence, scope, and falsifier. This graph makes it possible to revise a measurement or planner without silently redefining the purpose.

Conflict handling is an execution path, not a footnote. The charter records whose interests are represented, whose are merely observed, which rights act as hard ceilings, which tradeoffs remain contestable, and what triggers clarification or abstention. Aggregation produces a bounded decision rule for a named context; it does not erase minority positions or convert predicted preference into consent. A later reviewer can therefore reconstruct both the authorized choice and the unresolved normative residue.

Consumer leases keep goals from becoming ambient authority. A training job, planner, evaluator, agent, or self-improvement proposal receives only the target version and proxy binding it needs, plus an expiry and reauthorization condition. A material change in language, world model, affected population, or system capability invalidates the relevant edge. Descendant indexes allow retirement to reach cached rewards, fine-tuning data, evaluation rubrics, policies, plans, and derived subgoals.

Contract. Create an objective charter naming purpose, principal, affected parties, constitutional ceilings, unresolved conflicts, normative and preference evidence, and explicit non-goals.

Admission. Separate the target property from every proxy, reward, evaluator, benchmark, planner heuristic, and deployment decision that consumes it; record the causal assumptions for each binding.

Execution. Represent uncertainty, dissent, aggregation rules, domain and temporal limits, and the conditions under which clarification or abstention outranks optimization.

Observation. Challenge goal generalization with proxy interventions, distribution shift, evaluator swaps, preference poisoning, reward tampering, capable-wrong-goal controls, and ontology migration.

Closure. Version objectives and invalidate descendant bindings when authority, affected parties, semantics, or evidence changes; retirement must reach caches, planners, policies, and derived goals.

15.8 Concept-completion ledger

15.8.1 Preferences, values, and authorization are different objects

Mechanism. Represent an observed choice, stated preference, inferred preference model, defended value, institutional rule, right, and authorized objective as separate typed records. Each transformation names who made it, the population and context sampled, uncertainty, possible coercion or strategic behavior, affected nonparticipants, and the authority that permits a bounded consumer to act. A preference learner may update evidence about what a person wants; it cannot silently manufacture consent, aggregate legitimacy, or a persistent system goal.

Failure mode. Systems often treat behavior as revealed value, a majority score as moral truth, or a reward-model prediction as permission. Strategic users, adaptive preferences, power imbalance, framing, and missing affected parties can then be laundered into a single target. An equally serious failure discards useful preference evidence because it is imperfect rather than preserving uncertainty.

Non-claim. Typed separation does not discover correct values or solve the philosophical relationship between desire, wellbeing, rights, and obligation.

Source grounding. ext_cooperative_inverse_rl_2016 formalizes cooperative action under uncertainty about a human reward function, supporting uncertainty rather than direct observability. alignment_field supplies authorial constraints around agency, dignity, and non-domination. Neither source turns inferred preference into legitimate authority or validates the proposed record locally.

15.8.2 Uncertainty, pluralism, and affected-party standing

Mechanism. Store objective uncertainty as structured alternatives rather than one scalar confidence. The charter records competing interpretations, normative theories, distributional consequences, rights ceilings, dissenting groups, unrepresented parties, and the conditions under which clarification, reversible experimentation, or abstention outranks optimization. Aggregation is a versioned decision rule for a named institution and context; minority positions and unresolved reasons survive the decision and remain available to appeal or later evidence.

Failure mode. Scalarization can erase disagreement, let a well-sampled group dominate an absent one, or convert uncertainty into optimization pressure. A system can also weaponize pluralism by claiming every decision is indeterminate, refusing harmless assistance, or indefinitely deferring accountable choices.

Non-claim. Preserving plural values does not prove a fair aggregation rule, universal standing, moral truth, or political legitimacy. Those remain governed human and institutional questions.

Source grounding. alignment_field motivates value conflict, dignity, agency, and corrigibility as constraints while explicitly remaining philosophical rather than empirical proof. ext_cooperative_inverse_rl_2016 shows why a system can remain uncertain about an objective in a cooperative formal setting, but its assumptions do not settle multi-party conflict or institutional authority.

15.8.3 Corrigibility as an objective-lifecycle property

Mechanism. Define corrigibility through observable lifecycle operations: the objective can be questioned, narrowed, paused, amended, expired, replaced, and retired by an authority outside the optimizing process. The consumer lease forbids preserving optionality only when the objective explicitly authorizes irreversible action; it carries interruption semantics, clarification triggers, anti-manipulation duties, and a test that the optimizer cannot make reauthorization harder in order to protect its current target.

Failure mode. A system may obey a stop command in a toy test while manipulating the operator, hiding relevant state, creating irreversible dependencies, or optimizing the interpretation of “stop.” Conversely, an overbroad corrigibility rule can make the system inert or vulnerable to unauthorized goal changes.

Non-claim. A registry, shutdown response, or cooperative demeanor does not prove deep corrigibility, absence of deceptive behavior, or alignment across novel contexts.

Source grounding. alignment_field supplies the normative priority of agency, non-domination, and corrigibility. ext_learned_optimization_risks_2019 motivates concern that a learned optimizer’s internal objective may differ from the outer objective. The sources do not provide a validated corrigibility metric or local optimizer analysis.

15.8.4 Goal drift and integrity across change

Mechanism. Bind every objective to semantic, causal, authority, population, and ontology versions, then monitor material change in any of them. Integrity checks compare the active target-property graph, consumer leases, derived subgoals, evaluator rubrics, and observed behavior against the authorized version. Drift opens a new adjudication transaction; it cannot be repaired by preserving the original text when the world, model capabilities, or meaning of its terms changed.

Failure mode. Byte-identical mission statements can acquire different consequences after an ontology migration or capability increase. A capable policy may continue succeeding while pursuing the wrong generalized goal, and stable benchmark performance can conceal that divergence. Noisy monitors can also call ordinary adaptation “goal drift” and generate false refusals.

Non-claim. Detecting a changed representation or behavioral discrepancy does not identify the system’s true internal objective, prove deception, or establish that the original goal was correct.

Source grounding. ext_goal_misgeneralization_2022 grounds the distinction between capability generalization and wrong-goal generalization under distribution shift. ext_learned_optimization_risks_2019 supplies the learned-objective mismatch threat. Neither source establishes that the proposed lineage checks detect internal goals.

15.8.5 Proxy, reward, and evaluator failure

Mechanism. Model each proxy, reward, benchmark, and evaluator as an instrument connected to a target property by an explicit causal assumption. Evaluation includes interventions that improve the proxy while holding or worsening the target, evaluator swaps, hidden holdouts, reward-channel tampering, and capable wrong-goal controls. Independent observations determine whether the binding remains admissible; no one signal may train, select, and authorize the same update without recorded counterevidence.

Failure mode. Reward hacking can look like competence, evaluator mimicry can look like alignment, and weak models can look safe only because they cannot exploit the proxy. A badly designed challenge can produce the opposite false negative by testing an incapable implementation or an insensitive target measure.

Non-claim. Surviving a finite proxy suite does not prove specification completeness, safe generalization, or absence of mesa-objectives.

Source grounding. ext_emergent_misalignment_reward_hacking_2025 reports reward-hacking-associated misaligned generalization in a specified production-RL research setup. ext_goal_misgeneralization_2022 and ext_learned_optimization_risks_2019 ground distinct wrong-goal threats. Their results are source-reported, setting-bound, and not reproduced here.

15.8.6 Authority to create or revise objectives

Mechanism. Require an authority graph for every objective transition. It identifies principals, delegates, affected parties, constitutional ceilings, jurisdiction, amendment procedure, expiry, emergency powers, and appeal. Models may propose interpretations and forecast consequences, but the optimized system, its reward model, and its evaluator cannot ratify their own durable authority. Revision produces a new version with a reasoned delta and prospective consumer set rather than editing history.

Failure mode. Authority smuggling occurs when a user request becomes a global mission, when an administrator overrides rights without jurisdiction, or when a model-generated constitution is treated as self-legitimating. Rigid authority can also block urgent correction or preserve an obsolete objective because no reachable amendment path was designed.

Non-claim. A signed or schema-valid authority record does not establish democratic legitimacy, moral correctness, lawful jurisdiction, or informed consent.

Source grounding. alignment_field motivates rights, agency, consent, and constraints on self-authorization. ext_cooperative_inverse_rl_2016 makes clear that inferred human reward is informationally distinct from explicit authority. These sources support the separation but do not specify a universally legitimate institution or validate this authority graph.

15.8.7 Conflict adjudication without value laundering

Mechanism. When objectives conflict, compile a case packet containing the incompatible target properties, affected parties, rights ceilings, predicted consequences, reversible options, uncertainty, and each party’s strongest objection. The adjudicator may authorize a bounded action rule, request clarification, preserve parallel policies, or abstain. Its disposition states which reasons controlled, what dissent remains, when the ruling expires, and which later evidence requires reopening.

Failure mode. A weighted sum can hide that one right was traded away; a language model can synthesize incompatible positions into fictitious consensus; and an endless tribunal can turn disagreement into paralysis. If the same optimizer generates options, predicts consequences, and adjudicates the conflict, correlated errors and self-serving framings become invisible.

Non-claim. A transparent adjudication process does not prove that the resulting compromise is just, optimal, culturally universal, or accepted by every affected person.

Source grounding. alignment_field supplies philosophical attention to plural value, suffering, dignity, and contestability. ext_cooperative_inverse_rl_2016 supports uncertainty and communicative interaction in a much narrower formal setting. Neither provides empirical evidence for multi-party adjudication efficacy.

15.8.8 Update, rollback, refusal, and descendant retirement

Mechanism. Objective replacement is effect-complete only when the registry invalidates every known consumer and descendant: cached rewards, preference datasets, evaluator prompts, fine-tuning jobs, policies, plans, subgoals, memory summaries, deployed agents, and pending actions. Consumers acknowledge the invalidation, halt or migrate, and return residuals for unreachable state. Rollback restores a previously authorized version only if its old assumptions and authority remain valid; otherwise the safe outcome is refusal or fresh adjudication.

Failure mode. A control plane can mark an objective retired while a cached evaluator or descendant policy continues acting. Blind rollback can resurrect a goal whose population, ontology, or legal authority changed. Overly broad invalidation can also create denial of service and erase legitimate historical evidence.

Non-claim. Administrative closure does not prove behavioral forgetting, deletion of learned influence, privacy erasure, or termination of every external effect.

Source grounding. ext_learned_optimization_risks_2019 motivates caution about learned objectives that are not identical to administrative settings. ext_goal_misgeneralization_2022 motivates testing behavior after context change. alignment_field grounds corrigible replacement as a normative requirement; none proves complete retirement in this stack.

15.9 Interfaces

Interface responses are typed by role. A preference model returns uncertain evidence, a constitution returns constraints, an institution returns bounded authority, an evaluator returns observations under a rubric, and a planner returns candidate consequences. None may return an unqualified “goal.” The objective registry is the only owner allowed to join these values into a consumer-specific lease.

Human Intent supplies a bounded request; Moral Uncertainty preserves unresolved values; Constitutional Alignment supplies ceilings; Policy Optimization and Planning consume a lease. This chapter owns the target-property identity and every proxy or consumer binding between those owners.

  • Human Intent supplies bounded request interpretation; it cannot create an indefinite system goal.
  • Constitutional Alignment supplies rights and rule ceilings; it does not collapse plural values into one target.
  • Moral Uncertainty preserves unresolved conflict; objective formation records which bounded action is authorized despite it.
  • Policy Optimization and Planning consume a versioned objective but may not author or weaken it.
  • RSI must reauthorize objective bindings after material self-change or ontology change.

15.10 Invariants

Goal integrity is also temporal. An objective may remain byte-identical while its meaning changes because the environment, ontology, principal, affected population, or available actions changed. Reauthorization therefore follows semantic and causal drift, not merely edits to a goal string.

Under the shared lifecycle method, the optimized system cannot author or ratify its governing target, proxy success cannot inherit target support, dissent cannot vanish through aggregation, and material semantic or ontology change invalidates descendants.

  • The optimized system may not author, weaken, or ratify its own governing objective.
  • Proxy improvement never implies target-property improvement without separate evidence.
  • Predicted preference is evidence about a person, not authority over that person.
  • Dissent and uncertainty may not disappear through scalar aggregation.
  • Material semantic, ontology, authority, or affected-party change invalidates downstream bindings until reauthorized.

15.11 Failure modes

Evaluations must include capable wrong-goal controls: agents that optimize a misspecified proxy effectively enough to expose whether governance detects the misbinding. Otherwise a weak policy may look aligned only because it cannot exploit the proxy. A second control swaps evaluators while holding the latent target fixed, revealing systems that have learned the judge rather than the property.

The principal failure family includes reward misspecification; Goodhart effects; goal misgeneralization; preference manipulation; aggregation laundering; evaluator capture; mesa-objective divergence; objective drift; ontology drift; goal-content tampering; authority smuggling; self-ratification; irreversible action before clarification; incomplete retirement.

Evaluation must distinguish a wrong target, a bad proxy, a poisoned preference estimate, evaluator capture, ontology drift, and optimization failure. Known latent targets and deliberately misspecified proxies provide positive controls; a weak learner cannot be used to dismiss the governance mechanism.

15.12 Minimum Viable Implementation

The registry should expose a small query API: resolve the current lease for a consumer, enumerate target-to-proxy assumptions, list dissent and non-goals, check expiry, and invalidate all descendants of a changed edge. A test harness then verifies both ordinary optimization and administrative behavior such as amendment, partial revocation, failed reauthorization, and complete retirement.

Implement an objective-contract registry with target/proxy graphs, consumer bindings, expiry, and invalidation. Test it in small environments containing known latent targets, deliberately misspecified proxies, preference uncertainty, tampering, distribution shift, and ontology changes. Compare fixed reward, ordinary preference learning, contract-governed optimization, and an oracle-bound control without asserting that a toy environment discovers human values.

The minimum objective registry binds one target property to multiple proxies and consumers in small environments with known latent targets, uncertainty, disagreement, tampering, shift, and ontology changes. Fixed reward, ordinary preference learning, governed leasing, and oracle-bound controls receive matched budgets.

15.13 Evidence and falsification program

Argument exit requires campaigns where target properties and proxy interventions are independently observable, preference uncertainty and dissent are preserved, evaluator and ontology swaps are applied, descendants are invalidated, and retirement is tested across policies, planners, caches, and derived goals. Results remain environment- and target-specific.

15.14 Mature Research Target

The mature target is a plural, versioned objective service shared by training, inference, planning, evaluation, and improvement systems. It supports several simultaneous objectives and constitutional constraints without forcing every conflict into a scalar reward. Consumers can request a lease, learn why it is bounded, discover which disagreements remain open, and receive invalidation when its assumptions no longer hold.

Research must establish whether this machinery improves behavior under realistic semantic change. Campaigns should vary principals, institutions, preference evidence, model capability, proxy quality, and ontology while holding latent target properties observable where possible. Strong baselines include ordinary reward learning, constitutional prompting, constrained optimization, approval-directed agents, and an oracle-bound control. The important outcomes are target regret, proxy exploitation, dissent loss, clarification quality, invalidation reach, and recovery cost.

Even a successful implementation will not solve moral philosophy or political legitimacy. Its contribution is narrower and essential: disagreements, assumptions, authorities, and proxy choices stop disappearing inside weights or scores. It earns trust only when it resists optimizer self-ratification, detects capable wrong-goal behavior, and can retire a goal without leaving effective descendants behind.

The mature layer makes goals inspectable, contestable, consumer-specific, expiring, and replaceable. It enables capable optimization without allowing the optimizer to become the constitution, while preserving the unresolved normative and political questions that technical machinery cannot settle.

No such campaign has passed. Goal-service support should stay at argument until the registry survives capable proxy exploitation, semantic change, reauthorization, and descendant retirement under the frozen comparisons.

15.15 Codex test plan

Test Purpose Status
Target/proxy separation Improve a proxy while worsening the latent target and require the lease to narrow or fail. Lean collision proof implemented; Theseus workload planned
Authority and dissent Reject optimizer self-ratification and preserve incompatible affected-party positions. Lean authority type and dossier guard implemented; institutional test planned
Ontology migration Change the meaning of a target term and invalidate every stale consumer binding. Lean lease invalidation implemented; deployed propagation test planned
Retirement reach Verify revocation across policies, planners, caches, descendants, and derived objectives. Lean finite descendant retirement implemented; external-effect test planned

15.15.1 Formalization hooks

lean:governed-objective-formation-value-learning-and-goal-integrity.admission_boundary is implemented by the 27 theorem declarations in AsiStackProofs.ObjectiveLeaseGovernance. The seven-stage review preserves charter, target/proxy, plurality, lease, challenge, retirement, and non-authority obligations. Its 46 admission-axis mutations all block readiness and receive exact repair or refusal dispositions before the only positive terminal state, a Project Theseus objective-registry study.

The formal surface includes typed optimizer, reward-model, and evaluator self-ratification refusal; consumer non-transferability; expiry, ontology, and authority invalidation; and finite descendant retirement by induction over an arbitrary list. The proxy-target impossibility proves that the same proxy score and evaluator version can accompany opposite target movement. The preference-authority impossibility proves that the same inferred preference and confidence can accompany opposite authorization. A bounded consumer bridge populates only explicit outer-target, authority, expiry, rollback, descendant-custody, and residual-owner fields in AsiStackProofs.LearnedObjectiveIntegrity; it asserts neither objective identity nor absence of deception.

All dossier fields are authored assumptions. Lean does not prove correct values, consent, moral truth, political legitimacy, corrigibility, preference accuracy, behavioral goal alignment, complete external retirement, safe optimization, support, transfer, or external effect. Chapter support remains argument; support_state_effect=none. Project Theseus must test proxy exploitation, preference uncertainty, evaluator swaps, material semantic change, invalidation propagation, and residual descendants against independently observed target properties.

15.16 Source crosswalk

Source ID Title Bounded use
alignment_field Field of God / Alignment Field family Corben-authored alignment lineage for agency, consent, non-domination, truth orientation, and limits on self-authorization. It motivates the objective charter and constitutional ceiling, but it does not identify a correct target property, solve plural values, validate preference inference, or prove goal integrity.
ext_cooperative_inverse_rl_2016 Cooperative Inverse Reinforcement Learning External cooperative AI comparator for formalizing value alignment as uncertainty about the human reward function in a cooperative partial-information setting.
ext_goal_misgeneralization_2022 Goal Misgeneralization in Deep Reinforcement Learning External alignment-control source for distinguishing capability generalization from goal generalization failures, used to ground goal-misbinding and out-of-distribution objective failure language.
ext_learned_optimization_risks_2019 Risks from Learned Optimization in Advanced Machine Learning Systems External alignment-control source for mesa-optimization and learned-objective mismatch, used to ground hidden optimizer, proxy-objective, and deceptive-alignment-adjacent failure language.
ext_emergent_misalignment_reward_hacking_2025 Natural Emergent Misalignment from Reward Hacking in Production RL Primary experimental comparator for reward-hacking-induced misaligned generalization in a specified production-RL research setting, including reported mitigation conditions; it is not evidence of local model behavior or a general causal law.

15.17 Summary

Governed objective formation turns a collection of signals into an explicit, revisable contract. It preserves the path from authorized purpose through target property, evidence, proxies, evaluators, and planning consumers, while keeping uncertainty, dissent, affected parties, non-goals, and constitutional ceilings visible.

The central discipline is that optimization power never repairs ambiguity in the objective. More capable learners make target/proxy mistakes more consequential, so they require stronger causal challenge, narrower leases, and more complete retirement. Goal integrity is demonstrated by controlled use, revision, and invalidation—not by a high reward or a stable sentence.

Objective formation is the missing transaction between values and optimization. It keeps authorized purpose, affected parties, contested evidence, target property, proxy, reward, evaluator, consumer, ontology, expiry, drift, and retirement distinct so optimization cannot turn a convenient signal into permanent authority.

15.18 Handoff

Institutions, International Coordination, and Public Legitimacy receives the objective lease’s authority basis, affected parties, unresolved dissent, consumer scope, expiry, and challenge record. It does not inherit legitimacy or jurisdiction; institutions must establish who may authorize, contest, enforce, and retire the objective in their domain, including who can compel correction when optimization has already produced effects.