flowchart LR
Q["Claim and exposed quantifiers"] --> A["Population, sample, hypothesis, algorithm, and metric assumptions"]
A --> M["Bound, fit, MDL lens, or emergence hypothesis"]
M --> F["Prospective forecast over scale or distribution"]
F --> H["Held-out model sizes, tasks, compute regimes, and shifts"]
H --> C{"Calibration and alternatives survive?"}
C -- "no" --> N["Narrow, reparameterize, split regimes, or retire"]
C -- "yes" --> R["Bounded generalization or scaling record"]
N --> E["Record failed forecast and alternative explanation"]
R --> D["Bind decision use, uncertainty, and transfer limit"]
D -. "forecast miss or new substrate" .-> Q
R -. "new regime" .-> Q
57 Learning Theory, Generalization, and Scaling Science
57.1 Chapter status
| Field | Value |
|---|---|
| Chapter ID | learning-theory-generalization-and-scaling-science |
| Part | Part III - Routing, Compression, Representation, and Substrates |
| Status | conceptual |
| Manuscript maturity | v0.2 complete argument-level manuscript |
| Last updated | 2026-07-24 |
| Claim label | Design rationale |
| Evidence level | argument |
| Source loading state | source notes: simulation_scaling, ext_scaling_laws_neural_language_models_2020, ext_mdl_tutorial_2004, ext_weakness_generalization_2023, ext_weak_to_strong_generalization_2023, ext_information_bottleneck_2000, ext_valiant_theory_learnable_1984, ext_deep_double_descent_2020, ext_emergent_abilities_llms_2022, ext_emergent_abilities_mirage_2023, ext_no_free_lunch_inductive_bias_2024, ext_eggroll_hyperscale_es_2026, ext_forward_forward_2022; raw cache: simulation_scaling |
| Test state | The chapter defines a minimum implementation and falsification plan; no chapter-core promotion follows from prose or source synthesis. |
57.2 Drafting guardrail
This chapter owns the assumptions and predictive burden behind generalization, transfer, emergence, and scaling claims. It does not turn a fitted curve, asymptotic bound, interpolation result, compression ratio, benchmark jump, or larger model into a universal forecast.
57.3 Human Reading Path
Concrete lens. The curve-fitting baseline reports only the winning family after seeing the holdout. The governed forecast freezes both alternatives and preserves all attempts before results open.
Learning theory is most useful when it states exactly what must be true for a prediction to travel. Modern systems may fit training data, interpolate, grok late, improve smoothly in loss while jumping on thresholded metrics, or break an old scaling curve after an architecture change. Every theorem, empirical regularity, and forecast has a domain whose assumptions must remain attached.
Optimization success, in-distribution generalization, distribution shift, transfer, robustness, calibration, emergence, and extrapolation are different claims. A learning packet records the model family, data-generating assumptions, sampling and dependence structure, objective, compute regime, measured variables, uncertainty, observed range, extrapolation target, and known violations. Scaling laws remain useful conditional models rather than universal promises about intelligence or safety.
The lifecycle runs from regime definition through experimental design, curve fitting or derivation, held-out prediction, intervention, architecture change, and prospective validation. Benchmark contamination, multiple-comparison selection, threshold artifacts, changing data quality, hidden compute, survivor bias, and distant extrapolation remain visible. Theory should narrow expectations and identify the next experiment, while transfer beyond the measured regime remains a hypothesis until prospective evidence earns a broader claim.
57.4 Problem
Mechanism. Bind every generalization, transfer, emergence, or scaling statement to a dated claim contract naming population, sampling process, data support, hypothesis family, algorithm, optimization path, metric, compute regime, uncertainty, alternatives, consumer, and expiry. Failure mode. An attractive result can silently travel to a different distribution, architecture, scale, or consequence. Non-claim. Contract completeness does not make the forecast true. Source grounding. PAC, scaling, MDL, information, and emergence sources motivate distinct fields; none supplies a universal contract.
The central risk is not that every forecast is wrong, but that a useful local regularity is promoted beyond the assumptions that made it useful. Scaling can change the optimizer’s reachable solutions, the data mixture, the effective hypothesis class, and the evaluation itself. A curve fitted across nominal model sizes may therefore confound several scientific interventions at once.
The stack repeatedly relies on claims that learning will generalize, capabilities will transfer, losses will scale, or phase changes will appear outside observed training support. Those claims require an owner that connects assumptions about data, hypothesis class, optimization, inductive bias, compute, and evaluation to an exact prediction and failure envelope.
Without this owner, empirical curves, learning-theory bounds, benchmark thresholds, and scaling narratives can each look authoritative while their populations, algorithms, metrics, compute regimes, alternatives, and prediction horizons differ. The shared lifecycle method supplies custody; scaling science supplies the assumption ledger and forecast registry.
57.5 Why existing approaches are insufficient
Mechanism. Treat a bound as a conditional implication whose quantifiers, hypothesis class, loss, confidence, sample process, computational assumptions, and tail behavior remain visible beside the result. Failure mode. Asymptotic or average guarantees can become practical, adaptive, rare-event, or safety claims after their assumptions disappear. Non-claim. A correct theorem does not certify a trained foundation model. Source grounding. Valiant supplies foundational conditional structure; modern-system applicability remains an open empirical and formal question.
PAC bounds, capacity measures, stability, compression, information theory, scaling-law fits, emergence plots, and benchmark curves each illuminate part of learning. None is a universal theory of deep networks, and none automatically transports from a loss metric, architecture family, data regime, optimizer, or scale range to downstream capability or safety.
Benchmark Ratchets, Resource Economics, optimizer studies, and ordinary statistical review are the strongest alternative composition. It wins if it can preserve quantifiers, data support, hypothesis class, algorithm and optimization bias, metric transformations, breakpoint alternatives, held-out forecasts, and transfer boundaries without this owner.
What this scaling-science diagram shows: an assumption-bound prediction moves. A retrospective fit or threshold crossing cannot cross as evidence for future scale, broad transfer, or an underlying phase transition.
57.6 Core Claim
[learning-theory-generalization-and-scaling-science.core, label: Design rationale, support: argument] A generalization, transfer, emergence, or scaling assertion should be accepted only through a dated claim contract that binds population and sampling assumptions, data support, hypothesis and algorithm, optimization and inductive bias, complexity or explanatory lens, metric, compute regime, uncertainty, breakpoint tests, held-out prediction, alternatives, and transfer boundary; a bound, fit, interpolation result, compression ratio, benchmark jump, or larger model alone establishes neither broad generalization, capability emergence, safety, nor future scale behavior.
Reader claim. A smooth curve through observed scales is a retrospective description until it predicts a frozen, genuinely held-out scale and survives plausible breakpoint alternatives.
Operational rule. Freeze the population, metric, compute regime, candidate families, fitting rule, uncertainty interval, breakpoint alternatives, and forecast horizon before opening held-out results. Keep failed runs and every preregistered alternative in the denominator; regime changes invalidate the receipt.
57.6.1 Worked scaling forecast: two models enter, the holdout stays closed
A scaling study observes one metric over a bounded compute range and proposes two candidate families, model 1 and model 2. Both are preregistered and both must be scored on the held-out scale; attempts 11, 12, and 13 remain in the denominator even if one run failed. The population, sample process, algorithm, architecture, metric 11, compute regime 13, and horizon 17 are frozen before the held-out result is opened.
If the final report shows only the winning curve, changes the metric after inspection, or applies the receipt to a new architecture or compute regime, the forecast is not prospective evidence. The finite review gives exact repair or refusal dispositions to all 45 admission-axis mutations and invalidates receipts under seven population or regime changes. It also shows that retrospective fit does not identify prospective coverage and that a threshold jump does not identify a mechanism change. The result governs claim custody; it does not establish generalization, emergence, transfer, future scaling accuracy, or safety.
57.7 Mechanism
Mechanism. Maintain an assumption and identifiability ledger that names the inductive biases introduced by data, architecture, optimizer, parameterization, augmentation, objective, tools, and evaluation. Compare alternative explanations that agree on observed data but diverge under shift. Failure mode. “No free lunch” can become empty skepticism, while one observed preference can be mistaken for robust transfer. Non-claim. Naming bias does not identify its causal contribution. Source grounding. The no-free-lunch note frames explicit assumptions and limits, not a solved theory.
Contract. State the population, sampling process, shift model, hypothesis family, learning algorithm, optimization path, inductive bias, metric, and consumer before making a generalization claim.
Admission. Use multiple explanatory lenses—capacity and complexity, stability, compression and MDL, information, margins, implicit bias, interpolation, and benign-overfitting hypotheses—without laundering one into another.
Execution. Fit scaling relations with uncertainty, held-out scales, architecture and data identifiers, failed runs, compute accounting, breakpoint diagnostics, and prospective forecasts.
Observation. Distinguish smooth underlying performance from thresholded metrics and test whether apparent emergence survives metric, prompting, sampling, and denominator changes.
Closure. Challenge transfer and compositionality under natural shift, targeted shift, new task structure, optimizer change, architecture replacement, and data contamination.
57.7.1 Sample complexity: expose the quantifiers before quoting the bound
Mechanism. Use compression, MDL, information bottlenecks, stability, margins, and capacity as separate explanatory lenses, each with its coding scheme, target variable, residual burden, and downstream verification cost. Failure mode. Complexity can be moved into prompts, tools, human labor, or an evaluator, and the chosen relevance variable can omit safety-critical information. Non-claim. Short descriptions or compressed representations are not understanding, truth, or generalization certificates. Source grounding. MDL and information-bottleneck sources provide bounded vocabulary; no local scorer or objective has been validated.
In a probably-approximately-correct framing, a learning statement has the shape “with probability at least (1-), the learned hypothesis has error at most ()” under a stated sampling process, target or comparator class, hypothesis family, loss, and computational model. Sample complexity describes how the required number of examples scales with those choices. Valiant’s foundational contribution is not a magical formula for modern networks; it is a discipline for making learnability conditional and resource-aware.
For a foundation model, the practical contract must additionally identify pretraining and evaluation populations, dependence and duplication, token and example units, contamination, curriculum, augmentation, post-training, in-context demonstrations, tool access, and the downstream task. A bound for an i.i.d. concept class cannot be presented as a guarantee under open-world shift.
Mechanism. State the target accuracy and confidence, effective sample unit, assumption set, complexity quantity, learner, and consumer; then check whether the real data process satisfies the premises closely enough for the bound to have decision value.
Failure mode. A formally true but enormous or assumption-mismatched bound is cited as practical assurance.
Explicit non-claim. A PAC-style bound does not establish semantic correctness, robustness, safety, or transfer outside its distribution and hypothesis assumptions.
57.7.2 Interpolation and double descent: capacity is not a monotone risk dial
Mechanism. Sweep effective complexity, sample size, and training duration densely enough to locate interpolation and non-monotone regions, retaining all failed runs and matched train/test curves. Failure mode. Sparse endpoints, nominal parameter counts, or censored failures can manufacture a smooth monotone story. Non-claim. Observing double descent in one configuration does not make it universal or causal. Source grounding. Deep Double Descent supplies configuration-bound empirical pressure; this chapter has not reproduced it.
Classical bias-variance intuition often predicts a U-shaped test-error curve as capacity grows. Modern overparameterized systems can interpolate training data and sometimes improve beyond the interpolation threshold. Nakkiran et al. report model-wise, sample-wise, and epoch-wise double-descent behavior in their studied settings and use effective model complexity to connect the regimes.
The result does not mean “bigger is always better,” “more data can generally hurt,” or “overfitting is solved.” It means that model size, optimization time, data amount, label noise, regularization, and the learned solution interact. Scaling experiments should therefore sweep across the suspected interpolation region, retain failed runs, and report training error, test error, calibration, subgroup behavior, and robustness together.
Mechanism. Locate interpolation and optimization transitions empirically, then compare monotone, U-shaped, double-descent, and breakpoint models on held-out scale points.
Failure mode. Sparse scale points skip the peak and produce a confident but wrong monotone forecast.
Explicit non-claim. Observed double descent in one architecture, dataset, noise regime, or metric does not transfer automatically to another.
57.7.3 MDL and Kolmogorov-style simplicity: an explanatory lens, not an oracle
Mechanism. Register prospective scaling forecasts with exact model, data, optimizer, metric, hardware accounting, fit range, candidate curve families, uncertainty, breakpoints, held-out scales, and scoring rules. Failure mode. Retrospective fitting and excluded failed runs can make nearly any curve appear predictive. Non-claim. A loss forecast does not automatically predict capability, risk, cost, or safety. Source grounding. Neural scaling laws supply source-reported power-law behavior within their regime; the registry is an untested governance design.
Minimum Description Length trades the cost of describing a model against the cost of describing residual data under that model. Kolmogorov complexity supplies an idealized notion of shortest description but is not generally computable. Practical codecs, priors, parameterizations, and training procedures therefore define which “simplicity” is actually measured.
The lens is valuable because it forces a model to pay for complexity and unexplained residuals. It can also mislead. Two functionally equivalent networks may have different byte encodings; a small description can encode a harmful shortcut; compression may reflect regularity in the benchmark rather than causal structure.
Mechanism. Name the coding scheme, model class, data order, residual model, side information, and decoder. Compare predictive performance and description length on held-out data instead of using parameter count as a proxy.
Failure mode. “Shorter” becomes “truer,” “understands,” or “safer.”
Explicit non-claim. Compression is evidence about a representation under a code, not proof of semantic understanding or generalization.
Bennett (ext_weakness_generalization_2023) sharpens this boundary by separating description length from weakness, defined in the paper as the cardinality of a hypothesis’s extension: a weaker hypothesis is less specific because it is compatible with more statements. In the paper’s finite enactive-cognition formalism, under a uniform distribution over the defined task space, maximizing weakness is proved necessary and sufficient for maximizing the probability of generalizing from a child task to an unknown parent. A constructed counterexample then shows that a shortest valid hypothesis need not be the weakest valid hypothesis.
That result is a conditional comparator, not a universal replacement for MDL. The vocabulary fixes which statements and extensions exist, and the uniform task prior does substantial work; under a nonuniform deployment distribution, another inductive bias can perform better on the tasks that matter. The source-reported experiments compare weakness and description length only on toy 8-bit binary addition and multiplication. They do not test modern neural networks, natural task distributions, transfer, safety, or the paper’s later speculations about hallucination and grokking.
Mechanism. For every proposed simplicity or generalization proxy, record the representation language, task prior, extension semantics, tie-breaking rule, and target distribution; compare shortest, least-specific, and task-weighted hypotheses rather than treating them as synonyms.
Failure mode. A theorem under a uniform task distribution is reported as a universal learning law, or the chosen vocabulary silently makes the preferred hypothesis look weak.
Explicit non-claim. The paper does not show that maximizing weakness trains better neural networks or that description length is useless on structured, nonuniform real tasks.
57.7.4 Transfer claims are distribution, search, verifier, and budget relative
A reusable abstraction does not have one context-free transfer value. Its effect is indexed by the future task distribution, lineage split, representation language, search procedure, verifier, and resource budget. The same abstraction can help when retrieval is accurate and verification is cheap, do nothing when the baseline already finds the solution, and harm when vocabulary growth or branching consumes a bounded search budget.
The evaluation unit is therefore a frozen tuple rather than an artifact name. Within-family transfer and cross-family transfer receive different labels; lineage-related tasks cannot be counted as independent evidence merely because their surface forms differ. Positive average transfer also retains heterogeneity, failures, unknown outcomes, timeouts, and cost rather than becoming a universal generalization claim.
From Compression to Forward Transfer turns these constraints into proposed program-synthesis interventions. It offers no measured learning curve or scaling law. Its contribution here is to expose the quantifiers that a future transfer claim must carry.
57.7.5 Scaling laws: forecast registries instead of retrospective straight lines
Mechanism. Test apparent emergence with continuous and probabilistic metrics, denser checkpoints, uncertainty intervals, prompt and sampling variants, contamination controls, and prospective breakpoint predictions. Failure mode. Exact-match thresholds, floors, small samples, and retrospective task choice can create cliffs or hide consequential crossings. Non-claim. Showing a metric artifact in one task does not refute every phase transition. Source grounding. Emergent Abilities and the Mirage critique provide competing source-bounded interpretations, not a settled general law.
Kaplan et al. report power-law relations between cross-entropy loss, model size, data, and compute for the studied model family and regime. Such fits can support resource planning when their domains, uncertainty, and held-out predictions are preserved. They do not automatically predict downstream capability, new architectures, post-training, tool use, or safety.
A serious scaling record freezes candidate functional forms, compute accounting, run-selection rules, data quality, architecture, optimizer, metric, and forecast horizon before the held-out scales are observed. It retains failed and undertrained runs. Breakpoint, saturation, mixture, and regime-change alternatives compete with one power law.
Failure mode. a visually clean retrospective fit inherits a narrow error bar while failed runs and hyperparameter changes disappear.
Explicit non-claim. predictable loss scaling is not predictable deployment behavior.
57.7.6 Changing the learner changes the claim
A learning-theory statement is scoped not only to a model class and data distribution but also to the credit estimator. Reverse-mode empirical-risk minimization, score-function policy gradients, evolution strategies, and local forward objectives expose different noise, bias, dependence, and sample-cost structures. A result cannot silently replace one with another because both eventually change weights.
EGGROLL’s analysis is a useful example of disciplined transport. It studies when aggregated low-rank perturbations approach a Gaussian evolution-strategy update in a high-dimensional linearized regime, with explicit noise scaling, rank dependence, and smoothness conditions. That supports a bounded consistency story for its estimator. It does not prove that the estimator finds a globally good solution, beats backpropagation, or retains the approximation after large nonlinear updates. Forward-Forward provides an even sharper evidence boundary: it proposes a different local credit architecture, but its preliminary small-scale experiments are not a generalization theory for foundation models.
Every theoretical learning claim therefore records at least: hypothesis and parameter class; objective and evaluator; perturbation or gradient estimator; population/query dependence; data reuse; initialization; update regime; asymptotic variable; constants and scaling assumptions; and the observable to which the conclusion applies. When the architecture, quantization, learning rule, or evaluator changes, the claim returns to assumption review instead of inheriting authority by vocabulary.
57.7.7 Emergence: distinguish a system transition from a measurement threshold
Mechanism. Challenge transfer under natural and targeted distribution shift, new task structure, weak supervision, architecture replacement, optimizer change, data contamination, and resource constraints, with source-domain and target-domain denominators kept separate. Failure mode. Shared pretraining, leaked labels, easy target metrics, or one model family can simulate transfer. Non-claim. Weak-to-strong improvement does not establish honest elicitation, alignment, or superhuman oversight. Source grounding. The weak-to-strong paper is a bounded proof of concept; the broader transfer contract remains untested.
Wei et al. organize abilities reported as absent in smaller language models and present in larger ones. Schaeffer et al. provide a direct competing explanation: discontinuous metrics and limited test samples can turn smooth changes in underlying performance into sharp-looking thresholds.
Both sources belong in the same subsection because the scientific question is not settled by naming “emergence.” A claim should be tested under continuous and probabilistic metrics, larger samples, repeated seeds, prompt and scoring variants, calibrated uncertainty, and a prospectively stated breakpoint. Then ask whether an internal or behavioral mechanism changed, whether only the decision threshold changed, or whether the available data cannot distinguish them.
Mechanism. preregister a smooth model, thresholded-observation model, and genuine-breakpoint model; score all three on held-out scales and tasks.
Failure mode. exact-match accuracy jumps because a smoothly improving token distribution finally crosses an argmax boundary.
Explicit non-claim. metric artifacts do not prove that every capability is smooth; an apparent breakpoint does not prove a new internal algorithm.
57.8 Interfaces
Every interface is a claim translation. Training records expose what was run; benchmarks expose measurements under a protocol; resource accounting exposes the consumed compute and data; architecture owners expose changes in inductive bias. The forecast registry joins these inputs but returns only a dated prediction with an interval and domain, never permission to allocate resources or deploy.
Optimizers and training systems supply algorithms and run records; benchmarks supply measured outcomes; Resource Economics supplies compute and physical cost; capability thresholds consume bounded forecasts. This chapter owns the assumptions and prediction tests that connect them.
- Governed Model Training owns faithful run execution; this chapter interprets what the resulting learning may generalize.
- Efficient ASI owns resource hypotheses; scaling science supplies bounded predictions, not resource authority.
- Benchmark Ratchets owns measurement renewal and held-out integrity.
- Replaceable Substrates and optimizer sections expose architecture and algorithm changes that can break old laws.
57.9 Invariants
Failed and censored runs remain part of the scientific denominator. Changing a prompt, metric, data filter, optimizer, stopping rule, or model family creates a new regime unless the forecast contract explicitly modeled that change. Uncertainty expands rather than disappears when the extrapolation crosses an unsupported regime boundary.
Under the shared lifecycle method, quantifiers remain explicit, interpolation is not transfer, metric transformation cannot manufacture emergence, retrospective fit is not prospective prediction, and each new scale or distribution can invalidate the record.
- Every claim names its distribution, metric, algorithm, architecture, scale range, and uncertainty.
- Training fit, in-distribution generalization, transfer, compositionality, and safety remain distinct.
- Loss scaling does not automatically predict thresholded capability or risk.
- Apparent emergence is retested against metric and sampling artifacts.
- Extrapolation beyond observed support remains a forecast until prospectively checked.
57.10 Failure modes
Curve fitting can also leak the future through hyperparameter selection, architecture iteration, benchmark reuse, or informal knowledge of larger runs. A nominally held-out point is not prospective if those decisions used its neighborhood. The registry records what was known when the forecast was frozen and distinguishes a clean forecast from retrospective reconstruction.
The principal failure family includes vacuous bounds; distribution laundering; test contamination; post-hoc curve fitting; breakpoint omission; metric-threshold emergence; architecture-regime shift; optimizer confounding; double-descent surprise; grokking misread as magic; compression-as-understanding; loss-to-capability substitution; credit-assignment ambiguity; failed-run censoring.
Evaluation must compare multiple plausible functional forms and simpler baselines, reserve future scales or tasks, expose selection and correction, and measure calibration as well as point error. Positive controls use synthetic regimes with known breakpoints; an underpowered scale range cannot refute discontinuity or generalization mechanisms.
57.11 Minimum Viable Implementation
The notebook should emit an immutable forecast card before opening each held-out point. The card names candidate models, priors or fitting rules, training runs eligible for the fit, uncertainty method, metric transformations, prediction intervals, break criteria, and how abstention is scored. A replay mode verifies that the frozen inputs reproduce the forecast and that later corrections do not overwrite the original.
Create a claim-contract and forecasting notebook over several public small-model runs. Freeze candidate curve families and explanatory lenses, hold out scale points and task families, record failed runs and tuning, test breakpoint alternatives, and compare prediction intervals for loss, calibration, and downstream tasks. The result can evaluate local forecast discipline, not establish a universal law.
The minimum registry freezes competing curve families, compute and data accounting, metric transformations, breakpoint tests, and held-out model sizes or tasks before final outcomes. It reports forecast error, interval coverage, rank stability, correction, and total experimentation cost.
57.12 Evidence and falsification program
Argument exit requires repeated prospective forecasts across model families, task families, compute regimes, and distribution shifts, with preregistered alternatives, held-out scales, calibration, breakpoint sensitivity, and independent reanalysis. The result supports only the stated population, algorithm, metric, and horizon.
57.13 Mature Research Target
A mature scaling science is a portfolio of competing, prospectively scored models rather than one canonical curve. It connects mechanistic hypotheses, learning-theory assumptions, empirical fits, resource models, and natural shift tests to a common forecast ledger. Models gain or lose weight through calibration across new scales, tasks, data regimes, optimizers, and architectures, while failed predictions remain visible.
The research target makes interventions informative. Instead of changing model size, data quality, optimizer, architecture, and post-training together, campaigns isolate or factorially vary the suspected causes. Continuous and thresholded metrics, in-distribution and transfer tasks, average and subgroup performance, and loss and capability outcomes are scored separately. Known synthetic breakpoints and smooth regimes verify that the analysis can recover both kinds of behavior.
Its decision value is judged against simpler heuristics: latest-run extrapolation, constant elasticity, nearest-neighbor scale, or no forecast. Useful science should improve interval coverage, resource planning, experiment selection, and early detection of regime change without overstating future capability or safety. When alternatives remain indistinguishable, the mature answer is a wider interval or a targeted next experiment.
A mature scaling science says not merely that performance rose, but which assumptions generated a forecast, where it failed, what changed at the boundary, and how much confidence survives transfer. It guides resource and safety decisions while remaining easy to falsify and retire.
57.14 Codex test plan
| Test | Purpose | Status |
|---|---|---|
| Quantifier exposure | Reject a bound or generalization claim whose population, sample, hypothesis, algorithm, or probability statement is missing. | planned |
| Metric-emergence control | Apply monotone and nonlinear metric transforms and distinguish threshold artifacts from behavioral discontinuities. | planned |
| Prospective scale holdout | Freeze competing curves before opening larger models, tasks, or compute regimes. | planned |
| Alternative and correction ledger | Preserve failed fits, selected forms, breakpoint searches, and forecast revisions in the denominator. | planned |
57.15 Formalization hooks
Implemented formalization: lean:learning-theory-generalization-and-scaling-science.admission_boundary binds this chapter to AsiStackProofs.LearningTheoryForecastReview. Its 38 theorem declarations define a reachable six-transition forecast review and test 45 admission-axis mutations. A complete authored dossier reaches only eligibility for a Project Theseus prospective forecast campaign; every one-axis omission or forbidden broad claim reaches repair with an exact disposition.
The mechanized core proves forecast custody rather than forecast truth. Attempt-identity collection composes over arbitrary finite lists and retains every member. A complete denominator must count each attempt, while an omitted attempt or unscored preregistered alternative rejects completeness. Expiry, extrapolation beyond observed support, and an alternative-scoring shortfall remain rejecting under adverse monotone changes. Population, sample process, algorithm, architecture, metric, compute regime, and forecast horizon changes invalidate receipts.
Two information-loss constructions show why equal retrospective-fit signals cannot determine prospective held-out coverage and why equal threshold-metric signals cannot determine whether an underlying mechanism changed. The consumer bridge maps a missing prospective holdout to rejection by Benchmark Ratchets. These are theorems about encoded records and transitions; the model does not prove generalization, transfer, emergence, scaling accuracy, calibration, safety, deployment readiness, support, or release. Chapter support remains argument and support_state_effect remains none. Those claims require the Project Theseus prospective forecast campaign across model, task, data, optimizer, architecture, metric, scale, and distribution-shift families.
57.16 Source crosswalk
| Source ID | Title | Bounded use |
|---|---|---|
simulation_scaling |
Simulation Scaling Law | Corben-authored conceptual lineage for resource-constrained scope, clock speed, and fidelity. It motivates explicit scale and fidelity denominators, but it does not provide a validated neural scaling law, learning-theory bound, prospective forecast, emergence result, or transfer guarantee. |
ext_scaling_laws_neural_language_models_2020 |
Scaling Laws for Neural Language Models | Empirical study reporting power-law relationships between cross-entropy loss, model size, data, and compute in its model family. These fitted relations are source-reported, metric- and regime-bound, and do not automatically forecast downstream capabilities, safety, or other architectures. |
ext_mdl_tutorial_2004 |
A tutorial introduction to the minimum description length principle | External description-length source for model/data tradeoffs, compression as inductive discipline, and residual/error-accounting vocabulary. |
ext_weakness_generalization_2023 |
The Optimal Choice of Hypothesis Is the Weakest, Not the Shortest | Conditional formal comparator that distinguishes extension-based weakness from description length under a finite language and uniform task distribution; its toy arithmetic trials and speculative neural-network discussion do not establish broad neural generalization. |
ext_weak_to_strong_generalization_2023 |
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision | Primary weak-to-strong-supervision comparator for a capability-gap envelope, held-out outcome audit, ceiling comparison, and explicit disanalogies between current weak-model studies and superhuman oversight; it does not establish local supervision quality, reliable elicitation, alignment, safety, or an ASI Stack result. |
ext_information_bottleneck_2000 |
The information bottleneck method | External representation-compression source for relevance-preserving compression, bottleneck variables, mutual-information tradeoffs, and compression/utility separation. |
ext_valiant_theory_learnable_1984 |
A Theory of the Learnable | Foundational source for making learnability conditional on accuracy, confidence, resource, sampling, and concept-class assumptions; not a foundation-model guarantee. |
ext_deep_double_descent_2020 |
Deep Double Descent | Configuration-bound evidence that generalization can be non-monotone across model, sample, and training-time regimes; not a universal larger-is-better or larger-is-worse law. |
ext_emergent_abilities_llms_2022 |
Emergent Abilities of Large Language Models | Frames reported sharp task changes under scaling as a forecast and measurement problem; not proof of a mechanistic phase transition. |
ext_emergent_abilities_mirage_2023 |
Are Emergent Abilities of Large Language Models a Mirage? | Supplies the competing metric-artifact explanation and continuous-measurement requirement; does not prove that all emergence is illusory. |
ext_eggroll_hyperscale_es_2026 |
Evolution Strategies at the Hyperscale | Supplies a bounded high-dimensional consistency analysis under its declared linearization, smoothness, noise-scaling, and rank assumptions; not a global convergence or superiority theorem. |
ext_forward_forward_2022 |
The Forward-Forward Algorithm | Supplies a preliminary local-credit architecture whose evidence boundary prevents small-scale feasibility from becoming a large-model generalization claim. |
57.16.1 Manifest source assignment reconciliation
These rows keep Learning Theory, Generalization, and Scaling Science’s manifest assignments visible at their recorded review boundary. Passage review does not establish local reproduction, performance, safety, deployment, or support-state movement.
| Source | Intake role | Boundary |
|---|---|---|
ext_no_free_lunch_inductive_bias_2024 |
Metadata-first comparator: The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning. ICML 2024 treatment connecting no-free-lunch limits, Kolmogorov complexity, and inductive bias. It supports explicit assumption accounting; it does not show that all learning problems are equally hard or identify the right bias for a deployment. | No passage-level source claim, local implementation, reproduction, safety, performance, deployment, support-state, or ASI result is established by this reconciliation row. |
forward_transfer_program_synthesis |
Passage-reviewed comparator: From Compression to Forward Transfer: Evaluating Reusable Knowledge in Program Synthesis. Makes reusable-knowledge transfer explicitly relative to a task lineage, task distribution, search procedure, verifier, and resource budget, and separates within-family positive transfer from cross-family transfer. | No universal generalization, scaling law, distribution-independent benefit, or empirical transfer result is established. No local implementation, reproduction, performance, safety, deployment, support-state, or ASI result is established by this reconciliation row. |
57.17 No-free-lunch, identifiability, and the assumption ledger
Generalization is never assumption-free. No-free-lunch results do not say that learning is futile; they say that averaged over sufficiently unconstrained problem classes, no learner has a universal advantage. Success comes from inductive bias: architecture, objective, data distribution, invariance, regularization, optimizer, retrieval, tools, or an environmental structure that makes some hypotheses more plausible than others. The ICML treatment recorded as ext_no_free_lunch_inductive_bias_2024 reinforces why the book must state those biases instead of treating benchmark success as a property of learning in general.
The same discipline applies to identifiability. Observational agreement may leave multiple predictors, causal structures, reward functions, latent states, or world models equally compatible with the data. More scale can narrow estimation error inside an assumed family while leaving this ambiguity intact. The chapter therefore adds an assumption ledger to every learning claim:
| Ledger field | Required disclosure |
|---|---|
| Task and distribution | sampling process, support, shift model, and excluded cases |
| Hypothesis bias | architecture, representation, invariance, objective, and regularization |
| Supervision | label or feedback generator, noise, missingness, selection, and dependence |
| Identifiability | which target is uniquely recoverable, equivalent, partial, or unresolved |
| Compute and search | optimizer, initialization, budget, stopping rule, and selection denominator |
| Scaling scope | observed range, functional form, uncertainty, breaks, and extrapolation horizon |
| Intervention boundary | observational, interventional, simulated, or policy-induced evidence |
| Failure witness | a constructed environment in which the claim should fail |
A scaling law is conditional on this ledger. A representation that generalizes under one symmetry can fail when the symmetry is broken. An apparently emergent ability can change with metric resolution or prompting. A causal target can be unidentified even with infinite observational data. None of these limits negates useful empirical regularities; they prevent the regularity from being promoted past its assumptions.
The minimum falsification test pairs every positive learning result with an assumption-breaking control. If performance survives, the ledger narrows the suspected dependency. If it fails, the result remains valuable at the bounded scope. “We do not know which assumption carried the gain” is an admissible outcome; “the model learned generally” is not.
57.18 Summary
Learning claims become decision-relevant only when their quantifiers, population, data support, learner, optimizer, architecture, metric, resource regime, uncertainty, alternatives, and prediction horizon remain attached. Bounds, compression, interpolation behavior, scaling fits, and apparent emergence are complementary lenses with different assumptions and failure modes.
The forecast registry changes the culture from explaining yesterday’s curve to risking a prediction about tomorrow’s held-out point. It preserves failed runs, selection decisions, breakpoint alternatives, metric transformations, and corrections, allowing transfer to broaden only through prospective evidence across materially different regimes.
Learning theory and scaling science are valuable when they make assumptions and predictions easier to challenge. The forecast discipline keeps bounds, fits, simplicity lenses, interpolation, emergence, compute, metrics, uncertainty, alternatives, and transfer limits visible so retrospective regularity cannot become destiny or acquire unwarranted authority over future architectures and safety decisions.
57.19 Handoff
Readiness Gates, Residual Escrow, and Quarantine receives only the forecast’s exact population, algorithm, metric, compute regime, interval, calibration, expiry, and unresolved alternatives. Readiness must not treat the forecast as a measured capability, a threshold crossing, authority to deploy, or a substitute for a direct evaluation when that evaluation is feasible.