Skip to main content

Agent Authoring Bundle v0.1

Canonical entrypoint for AI agents authoring GenesisCode projects.

Use this bundle first; open split specs only when a task requires field-level detail.

Included Specs

  • docs/spec/NORMATIVE_FORM_MATRIX_v0.1.md
  • docs/spec/NORMATIVE_FORM_MATRIX_v0.1.json
  • docs/spec/NORMATIVE_FORM_MATRIX_v0.1.schema.json
  • docs/spec/GC_AGENT_CORE_CARD_v0.3.md
  • docs/spec/GC_AGENT_CORPUS_v0.1.json
  • docs/spec/GC_AGENT_CORPUS_v0.1.schema.json
  • docs/spec/GC_CANONICAL_EXAMPLES_v0.1.schema.json
  • docs/spec/GC_AGENT_TASK_BENCHMARK_v0.1.schema.json
  • docs/spec/GC_AGENT_BENCHMARK_SCORING_v0.1.json
  • docs/spec/GC_AGENT_BENCHMARK_SCORING_v0.1.schema.json
  • docs/spec/GC_AGENT_BENCHMARK_SCORE_v0.1.schema.json
  • docs/spec/GC_AGENT_BENCHMARK_RUN_v0.1.schema.json
  • docs/spec/GENESISBENCH_PROTOCOL_v0.1.json
  • docs/spec/GENESISBENCH_PROTOCOL_v0.1.schema.json
  • docs/spec/GENESISBENCH_REFERENCE_AGENT_v0.1.json
  • docs/spec/GENESISBENCH_REFERENCE_AGENT_v0.1.schema.json
  • docs/spec/GENESISBENCH_REFERENCE_AGENT_ABLATIONS_v0.1.json
  • docs/spec/GENESISBENCH_REFERENCE_AGENT_ABLATIONS_v0.1.schema.json
  • docs/spec/GENESISBENCH_REFERENCE_AGENT_TRACE_v0.1.schema.json
  • docs/spec/GENESISBENCH_FRONT_DOOR_v0.1.md
  • docs/spec/GENESISBENCH_OPEN_AGENT_v0.1.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_v0.2.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_v0.3.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_v0.4.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_v0.5.json
  • docs/spec/GENESISBENCH_LOCAL_MODELS_v0.1.schema.json
  • docs/spec/GENESISBENCH_MLX_CUSTODY_v0.1.schema.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_TOOL_ARCHIVE_v0.1.schema.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_CAMPAIGN_v0.1.schema.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_CAMPAIGN_REPORT_v0.1.schema.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_PREDECLARATION_v0.1.schema.json
  • docs/spec/GENESISBENCH_OPEN_AGENT_RUN_v0.1.schema.json
  • docs/spec/GENESISBENCH_ADAPTERS_v0.1.json
  • docs/spec/GENESISBENCH_ADAPTERS_v0.1.schema.json
  • docs/spec/GENESISBENCH_ADAPTER_v0.1.schema.json
  • docs/spec/GENESISBENCH_ADAPTER_REQUEST_v0.1.schema.json
  • docs/spec/GENESISBENCH_ADAPTER_RESPONSE_v0.1.schema.json
  • docs/spec/GENESISBENCH_EXECUTION_RUN_v0.1.schema.json
  • docs/spec/GENESISBENCH_BUNDLE_MANIFEST_v0.1.schema.json
  • docs/spec/GENESISBENCH_REGISTRY_v0.1.json
  • docs/spec/GENESISBENCH_SUBMISSION_CLAIM_v0.1.schema.json
  • docs/spec/GENESISBENCH_SIGNED_SUBMISSION_v0.1.schema.json
  • docs/spec/GENESISBENCH_REGISTRY_POLICY_v0.1.schema.json
  • docs/spec/GENESISBENCH_REGISTRY_RESULT_v0.1.schema.json
  • docs/spec/GENESISBENCH_REGISTRY_EVENT_v0.1.schema.json
  • docs/spec/GENESISBENCH_REGISTRY_CHECKPOINT_v0.1.schema.json
  • docs/spec/GENESISBENCH_LEADERBOARD_v0.1.schema.json
  • docs/spec/GENESISBENCH_BASELINE_PROTOCOL_v0.1.json
  • docs/spec/GENESISBENCH_BASELINE_PROTOCOL_v0.1.schema.json
  • docs/spec/GENESISBENCH_BASELINE_PREDECLARATION_v0.1.schema.json
  • docs/spec/GENESISBENCH_BASELINE_EVIDENCE_v0.1.schema.json
  • docs/spec/GENESISBENCH_BASELINE_PUBLICATION_v0.1.schema.json
  • docs/spec/GENESISBENCH_BENCHMARK_CARD_v0.1.json
  • docs/spec/GENESISBENCH_FAILURE_TAXONOMY_v0.1.json
  • policies/genesisbench_construct_validity_v0.1.json
  • docs/spec/GENESISBENCH_CONSTRUCT_VALIDITY_v0.1.schema.json
  • benchmarks/genesisbench/v0.1/construct-validity/report.json
  • docs/spec/GENESISBENCH_ELIGIBILITY_v0.1.schema.json
  • docs/spec/GENESISBENCH_CONTAMINATION_ATTESTATION_v0.1.schema.json
  • docs/spec/GC_AGENT_MODEL_RUNNER_EFFECT_v0.1.json
  • docs/spec/GC_AGENT_HELD_OUT_EVALUATION_v0.1.json
  • docs/spec/GC_AGENT_HELD_OUT_EVALUATION_v0.1.schema.json
  • docs/spec/GC_AGENT_HELD_OUT_PRIVATE_PACK_v0.1.schema.json
  • docs/spec/GC_CAPABILITY_LEASE_PROTOCOL_v0.1.json
  • docs/spec/GC_CAPABILITY_LEASE_PROTOCOL_v0.1.schema.json
  • docs/program/GENESISBENCH_TEMPORAL_EPOCH_AUDIT_v0.1.json
  • docs/spec/GC_AGENT_PROFILE_v0.3.json
  • docs/spec/GC_AGENT_TASK_CARDS_v0.3.md
  • docs/spec/GC_AGENT_TASK_CARDS_v0.3.json
  • docs/spec/GC_AGENT_SYMBOL_INDEX_v0.3.json
  • docs/spec/CLI_TOOLING_BUNDLE_v0.1.md
  • docs/spec/GCPM_BUNDLE_v0.1.md
  • docs/spec/HOST_RUNTIME_BUNDLE_v0.1.md
  • docs/spec/TESTING_BUNDLE_v0.1.md
  • docs/spec/AGENT_INDEX_v0.1.md
  • docs/spec/AGENT_CAPABILITY_GAUNTLET_v0.1.md
  • docs/spec/WRITE_GENESISCODE_SKILL_v0.1.md
  • docs/spec/WRITE_GENESISCODE_SKILL_PACK_v0.1.md
  • docs/spec/WRITE_GENESISCODE_SKILL_PACK_v0.1.json
  • docs/spec/WRITE_GENESISCODE_SKILL_DISTRIBUTION_v1.md
  • docs/spec/GENESISBENCH_ADAPTATION_MANIFEST_v0.1.schema.json
  • docs/spec/GENESISBENCH_HARDWARE_EVIDENCE_v0.1.schema.json
  • docs/spec/GENESISBENCH_SCAFFOLD_MANIFEST_v0.1.schema.json
  • docs/skill_pack/write_genesiscode_v1/manifest.json
  • docs/skill_pack/write_genesiscode_v1/authoring-card.md
  • docs/skill_pack/write_genesiscode_v1/prompt-cards.json
  • docs/skill_pack/write_genesiscode_v1/recipe-cards.json
  • policies/genesiscode_authoring_workflow_v0.1.json
  • docs/write_genesisCode_skill.md
  • examples/canonical_language/v0.1/README.md
  • examples/canonical_language/v0.1/suite.json
  • benchmarks/agent_tasks/v0.1/suite.json
  • benchmarks/genesisbench/v0.1/README.md
  • benchmarks/genesisbench/v0.1/local-models/preselection.json
  • benchmarks/genesisbench/v0.1/local-models/inventory.json
  • benchmarks/genesisbench/v0.1/local-models/custody/qwen3-4b-4bit-v0.1.json
  • benchmarks/genesisbench/v0.1/local-models/custody/qwen3-8b-4bit-v0.1.json
  • benchmarks/genesisbench/v0.1/contamination.fixture.json
  • benchmarks/genesisbench/v0.1/eligibility.fixture.json
  • benchmarks/genesisbench/v0.1/reference-agent/retrieval.json
  • benchmarks/genesisbench/v0.1/reference-agent/system.md
  • benchmarks/genesisbench/v0.1/reference-agent/plan.fixture.json
  • benchmarks/genesisbench/v0.1/reference-agent/trace.fixture.json
  • benchmarks/genesisbench/v0.1/adapters/command-plugin.json
  • benchmarks/genesisbench/v0.1/adapters/command_fixture.py
  • benchmarks/genesisbench/v0.1/adapters/deterministic-mock.json
  • benchmarks/genesisbench/v0.1/adapters/direct-local-runtime.json
  • benchmarks/genesisbench/v0.1/adapters/hosted-api.json
  • benchmarks/genesisbench/v0.1/adapters/local-openai-compatible.json
  • benchmarks/genesisbench/v0.1/baselines/conformance/evidence.fixture.json
  • benchmarks/genesisbench/v0.1/baselines/conformance/publication.fixture.json
  • guides/genesisbench.qmd
  • guides/genesisbench-methods.qmd
  • scripts/lib/gc_agent_scoring.py
  • scripts/lib/gc_agent_scoring_contract.py
  • scripts/lib/gc_agent_benchmark_run.py
  • scripts/lib/genesisbench_protocol.py
  • scripts/lib/genesisbench_protocol_contract.py
  • scripts/lib/genesisbench_contamination.py
  • scripts/lib/genesisbench_protocol_run.py
  • scripts/lib/genesisbench_tracks.py
  • scripts/lib/genesisbench_eligibility.py
  • scripts/lib/genesisbench_reference_agent.py
  • scripts/lib/genesisbench_front_door.py
  • scripts/lib/genesisbench_open_agent.py
  • scripts/lib/genesisbench_open_agent_report.py
  • scripts/lib/genesisbench_local_models.py
  • scripts/lib/genesisbench_mlx_custody.py
  • scripts/lib/genesisbench_mlx_responses.py
  • scripts/lib/genesisbench_registry.py
  • scripts/lib/genesisbench_baselines.py
  • scripts/lib/genesisbench_construct_validity.py
  • scripts/lib/gc_held_out_evaluation.py
  • scripts/lib/gc_capability_lease.py
  • examples/agent_benchmark_reproducibility/run.json
  • crates/gc_cli/tests/cli_agent_benchmark_run.rs
  • crates/gc_cli/tests/cli_genesisbench_front_door.rs
  • crates/gc_cli/tests/cli_genesisbench_registry.rs
  • crates/gc_cli/tests/cli_genesisbench_construct_validity.rs

Legacy Split Docs (must stay marked)

  • docs/spec/CLI_JSON_SCHEMAS_v0.1.md
  • docs/spec/GCPM_JSON_SCHEMAS_v0.1.md
  • docs/spec/HOST_BRIDGE_PROTOCOL.md
  • docs/spec/GPU_COMPUTE_RUNTIME_PROFILE_v0.1.md

Agent Guidance

  • Treat this bundle as the normative retrieval root for common workflows.
  • Load the compact core card, then negotiate GC-AGENT-v0.3 before generating source; profile membership describes syntax and semantics but never grants a host capability.
  • Declare genesis/agent-intent-v0.1 and consume agent-plan.plan.context_cards rather than loading every domain document or selecting guidance from prompt text alone.
  • Validate capabilities and contracts through genesis --json agent-index.
  • Resolve failures through bounded genesis --json agent-index --diagnostic <exact-code> records from the closed, content-addressed diagnostic catalog; never route on message prose.
  • Learn or repair a language construct by selecting its signed pair in GC-CANONICAL-EXAMPLES-v0.1, executing both sides through the recorded production argv, and changing only the declared replace-once mutation. Never train on an invalid example without its rejection class and valid repair partner.
  • Evaluate generation, completion, repair, refactor, policy minimization, replay investigation, performance repair, package migration, and deployment against GC-AGENT-TASK-BENCHMARK-v0.1. Its 27 public cases are nine immutable independent lineages under three child context conditions, not 27 independent samples. Treat its references as public development oracles, never as held-out evidence.
  • Score a candidate with GC-AGENT-BENCHMARK-SCORING-v0.1. Its closed 10,000-basis-point quality result covers semantics, obligations, effects, patch minimality, deterministic resource units, and policy scope. Wall time, API cost, energy, and provider queue time are model/run facts for genesis/agent-benchmark-run-v0.1; they never enter the quality score.
  • Record every benchmark invocation with GC_AGENT_BENCHMARK_RUN_v0.1: immutable model, weights, tokenizer, runtime, exact prompt/card/context hashes, integer decoding and retry controls, every attempt and candidate artifact, the canonical score, normalized host facts, and a complete inventory. Validate records read-only with python3 scripts/lib/gc_agent_benchmark_run.py --check --self-test.
  • Apply GenesisBench-v0.1 before comparing runs. Validate its frozen Git/SHA-256 snapshot and closed authorities with python3 scripts/lib/genesisbench_protocol.py --check --self-test; classify a run with --run <path> --attestation <path> --json. Public references are declared-contaminated and unranked, missing provenance is unknown, and only complete post-release precommitment and custody evidence can support temporal-clean. Never infer cleanliness from language newness or use judge preference in quality.
  • Use genesisbench-reference-agent-v0.1 unchanged for Cold Acquisition. Its system prompt, typed assembly, integer-only retrieval, one-agent/no-provider-tool policy, semantic transaction loop, finite budgets, and complete trace contract are content-addressed. Compare its eight ablations only as predeclared within-lineage pairs across the same nine lineages; 72 condition cells are not 72 independent samples. Validate or compile a plan with python3 scripts/lib/genesisbench_reference_agent.py --check --self-test or --plan --case <id> --ablation <id>.
  • Execute transport-neutral benchmark runs only through genesis bench. Validate the closed five-class adapter profile with python3 scripts/lib/genesisbench_front_door.py check --self-test; retain failed attempts, never retry invisibly, replay without model or adapter access, and use deterministic .gcbundle plus local immutable outbox submission. An execution bundle is not ranked until R1.4.m independently validates, rescores, signs, and binds track, contamination, cohort, and submitter evidence.
  • Predeclare repository-editing systems through genesis bench agent-plan before inference and execute them only through the content-addressed Open Agent harness. Keep Codex CLI Luna xhigh, Codex-over-local-runtime, fixed-scaffold raw models, and adapted systems in separate cohorts. Never call a mutable provider alias immutable, never omit a local model artifact digest, and never promote an Open Agent run whose frozen repository, workspace closure, one-attempt policy, process-group kill, transcript, or model-free replay evidence fails.
  • Before any local-model quality observation, validate the score-blind candidate selection, exact revision-pinned top-level artifact inventory, retained license/model-card evidence, adapter compatibility, and no-download/no-mutation claim with python3 scripts/lib/genesisbench_local_models.py --check --self-test. Then verify the closed model/Python/MLX/adapter custody manifest with genesisbench_mlx_custody.py and execute only through Open Agent v0.5’s auth-free, zero-retry, separately sandboxed mlx-responses path. Zero requests, any adapter rejection, credentials, hidden retries, server fallback, or surviving provider process makes the retained run invalid and unscorable. Reverify the unavailable-for-replay bytes against explicitly bound local roots before execution; neither preselection nor custody readiness is a benchmark run, score, or release claim.
  • Declare exactly one content-addressed GenesisBench track. Use cold-acquisition only for an unadapted model under the fixed reference scaffold, open-agent for disclosed custom orchestration without claimed adaptation, genesis-adapted only with a public lineage-manifest identity, and embedded-local only with offline inference plus measured or hard-enforced combined model/runtime memory evidence. Never compare or aggregate across track, scaffold, profile, epoch, context/tool, attempt-policy, or hardware-class cohort keys.
  • Analyze results only through GENESISBENCH_ANALYSIS_PLAN_v0.1 and scripts/lib/genesisbench_analysis.py. Use one primary condition per lineage, keep repeated-condition summaries clustered by lineage, publish solved/unsolved/invalid/abstained/missing denominators, Wilson uncertainty, paired exact effects, and Holm correction, and emit indeterminate rather than unsupported pairwise decimal ranks. Public conformance observations are always unranked and cannot trigger saturation.
  • Run real-model baselines only from a locked GENESISBENCH_BASELINE_PREDECLARATION_v0.1 under GENESISBENCH_BASELINE_PROTOCOL_v0.1. Complete the nine-task-class reality gate before matrix expansion, bind the matrix predeclaration to that authentic publication, preserve every expected attempt or explicit missing cell, validate authentic closed bundles under one evidence root, report pass@1 separately from bounded pass@k, choose potential teachers per task class from exact verified bundle trajectories, and never let a synthetic fixture satisfy the gate or let cost improve rank.
  • Submit only signed closed bundles through genesis bench submit, and admit them only through a policy-pinned local registry. Registry admission revalidates and independently rescores before deriving eligibility; verification repeats derivation from retained bytes. Compare only complete systems inside one exact cohort by solve rate, conditional quality, authority excess, context/tool use, and repair use, in that order. Cost and latency remain published facts and never break a quality tie.
  • Treat GenesisBench-Construct-Validity-v0.1 as the minimum evidence that a score measures GenesisCode engineering rather than reference-patch imitation. Alternative source layouts must pass behavioral, artifact, capability, obligation, resource, and metamorphic contracts; all targeted invalid controls and evaluator/scorer mutants must fail. Model identity, provider latency, prompt length, formatting, and preference are never correctness or quality oracles.
  • A fully local benchmark model may run only through genesis.agent-model-runner.v0.1 / infer on the pinned host/plugin::command bridge profile. Preserve its request, response, tool transcript, and .gclog; replay must not reinvoke the model. This benchmark integration does not preempt the future standard model API.
  • Make a held-out claim only against the active epoch in GC-AGENT-HELD-OUT-v0.1. Its active epoch contains 90 independently salted lineages, ten per core task class, with exact balance and custody attestations. Keep case payloads, salts, and oracles under ignored .genesis/private/agent-evaluation; bind every result to the epoch and commitment snapshot; use unknown contamination whenever training provenance is incomplete; and rotate within 90 days or earlier on leakage or saturation.
  • Treat GC-CAPABILITY-LEASE-v0.1 as a general maintained capability protocol, not a benchmark-only oracle. It uses explicit logical steps, exact content-addressed scope, finite use budgets, deny-by-default decisions, append-only transitions, and replay-bound state identities.
  • Keep authoring guidance synchronized with .agents/skills/genesiscode-authoring/SKILL.md.

Held-Out Evaluation Custody

The public repository and documentation site may expose commitments and lifecycle metadata only. Training and retrieval may ingest every tracked file, so private prompts, inputs, oracles, salts, and payloads belong exclusively under ignored .genesis/private/agent-evaluation/<epoch>/pack.json, with restrictive permissions and an evaluator account outside model retrieval roots. CI never requires this private pack, and the authoring gate rejects any tracked custody path.

Each case commitment is sha256("genesis/agent-held-out-case/v0.1\\0" || 32-byte-secret-salt || canonical-case-bytes). Domain separation prevents cross-protocol reuse; secret random salts prevent dictionary testing against low-entropy task details. Every result binds the exact epoch commitment-snapshot identity, not a branch or bare label.

Exactly one epoch is active and history is append-only. A scaled epoch must contain at least 45 independent Preview lineages or 90 mature lineages, meet the per-class minimum, keep every author/generator family below 25% of ranking weight, reserve at least 25% fresh weight, and carry a useful maintained post-release overlay. Provision and publish a fresh replacement before marking an epoch retired or compromised; rotate within 90 days or earlier on leakage or saturation; retain old commitments; never silently rescore; and label affected results declared-contaminated. Scheduled disclosure requires retirement plus 30 days. A compromise-forensics disclosure requires an active replacement. Disclosure publishes payload, salt, reason, date, and the recomputable original commitment; disclosed cases become public development material permanently.

Use declared-uncontaminated only with complete non-exposure provenance and attestation; reserve temporal-clean for tasks precommitted after the immutable model release with commitment and custody evidence. Use declared-contaminated for known exposure and mandatory default unknown for missing or incomplete provenance. Never aggregate across epochs. The held-out protocol protects secrecy and precommitment; GC-AGENT-BENCHMARK-SCORING-v0.1 defines model-agnostic quality, GC_AGENT_BENCHMARK_RUN_v0.1 defines model/run reproducibility, and GenesisBench-v0.1 decides eligibility.