Agent Authoring Bundle v0.1
Canonical entrypoint for AI agents authoring GenesisCode projects.
Use this bundle first; open split specs only when a task requires field-level detail.
Included Specs
docs/spec/NORMATIVE_FORM_MATRIX_v0.1.mddocs/spec/NORMATIVE_FORM_MATRIX_v0.1.jsondocs/spec/NORMATIVE_FORM_MATRIX_v0.1.schema.jsondocs/spec/GC_AGENT_CORE_CARD_v0.3.mddocs/spec/GC_AGENT_CORPUS_v0.1.jsondocs/spec/GC_AGENT_CORPUS_v0.1.schema.jsondocs/spec/GC_CANONICAL_EXAMPLES_v0.1.schema.jsondocs/spec/GC_AGENT_TASK_BENCHMARK_v0.1.schema.jsondocs/spec/GC_AGENT_BENCHMARK_SCORING_v0.1.jsondocs/spec/GC_AGENT_BENCHMARK_SCORING_v0.1.schema.jsondocs/spec/GC_AGENT_BENCHMARK_SCORE_v0.1.schema.jsondocs/spec/GC_AGENT_BENCHMARK_RUN_v0.1.schema.jsondocs/spec/GENESISBENCH_PROTOCOL_v0.1.jsondocs/spec/GENESISBENCH_PROTOCOL_v0.1.schema.jsondocs/spec/GENESISBENCH_REFERENCE_AGENT_v0.1.jsondocs/spec/GENESISBENCH_REFERENCE_AGENT_v0.1.schema.jsondocs/spec/GENESISBENCH_REFERENCE_AGENT_ABLATIONS_v0.1.jsondocs/spec/GENESISBENCH_REFERENCE_AGENT_ABLATIONS_v0.1.schema.jsondocs/spec/GENESISBENCH_REFERENCE_AGENT_TRACE_v0.1.schema.jsondocs/spec/GENESISBENCH_FRONT_DOOR_v0.1.mddocs/spec/GENESISBENCH_OPEN_AGENT_v0.1.jsondocs/spec/GENESISBENCH_OPEN_AGENT_v0.2.jsondocs/spec/GENESISBENCH_OPEN_AGENT_v0.3.jsondocs/spec/GENESISBENCH_OPEN_AGENT_v0.4.jsondocs/spec/GENESISBENCH_OPEN_AGENT_v0.5.jsondocs/spec/GENESISBENCH_LOCAL_MODELS_v0.1.schema.jsondocs/spec/GENESISBENCH_MLX_CUSTODY_v0.1.schema.jsondocs/spec/GENESISBENCH_OPEN_AGENT_TOOL_ARCHIVE_v0.1.schema.jsondocs/spec/GENESISBENCH_OPEN_AGENT_CAMPAIGN_v0.1.schema.jsondocs/spec/GENESISBENCH_OPEN_AGENT_CAMPAIGN_REPORT_v0.1.schema.jsondocs/spec/GENESISBENCH_OPEN_AGENT_PREDECLARATION_v0.1.schema.jsondocs/spec/GENESISBENCH_OPEN_AGENT_RUN_v0.1.schema.jsondocs/spec/GENESISBENCH_ADAPTERS_v0.1.jsondocs/spec/GENESISBENCH_ADAPTERS_v0.1.schema.jsondocs/spec/GENESISBENCH_ADAPTER_v0.1.schema.jsondocs/spec/GENESISBENCH_ADAPTER_REQUEST_v0.1.schema.jsondocs/spec/GENESISBENCH_ADAPTER_RESPONSE_v0.1.schema.jsondocs/spec/GENESISBENCH_EXECUTION_RUN_v0.1.schema.jsondocs/spec/GENESISBENCH_BUNDLE_MANIFEST_v0.1.schema.jsondocs/spec/GENESISBENCH_REGISTRY_v0.1.jsondocs/spec/GENESISBENCH_SUBMISSION_CLAIM_v0.1.schema.jsondocs/spec/GENESISBENCH_SIGNED_SUBMISSION_v0.1.schema.jsondocs/spec/GENESISBENCH_REGISTRY_POLICY_v0.1.schema.jsondocs/spec/GENESISBENCH_REGISTRY_RESULT_v0.1.schema.jsondocs/spec/GENESISBENCH_REGISTRY_EVENT_v0.1.schema.jsondocs/spec/GENESISBENCH_REGISTRY_CHECKPOINT_v0.1.schema.jsondocs/spec/GENESISBENCH_LEADERBOARD_v0.1.schema.jsondocs/spec/GENESISBENCH_BASELINE_PROTOCOL_v0.1.jsondocs/spec/GENESISBENCH_BASELINE_PROTOCOL_v0.1.schema.jsondocs/spec/GENESISBENCH_BASELINE_PREDECLARATION_v0.1.schema.jsondocs/spec/GENESISBENCH_BASELINE_EVIDENCE_v0.1.schema.jsondocs/spec/GENESISBENCH_BASELINE_PUBLICATION_v0.1.schema.jsondocs/spec/GENESISBENCH_BENCHMARK_CARD_v0.1.jsondocs/spec/GENESISBENCH_FAILURE_TAXONOMY_v0.1.jsonpolicies/genesisbench_construct_validity_v0.1.jsondocs/spec/GENESISBENCH_CONSTRUCT_VALIDITY_v0.1.schema.jsonbenchmarks/genesisbench/v0.1/construct-validity/report.jsondocs/spec/GENESISBENCH_ELIGIBILITY_v0.1.schema.jsondocs/spec/GENESISBENCH_CONTAMINATION_ATTESTATION_v0.1.schema.jsondocs/spec/GC_AGENT_MODEL_RUNNER_EFFECT_v0.1.jsondocs/spec/GC_AGENT_HELD_OUT_EVALUATION_v0.1.jsondocs/spec/GC_AGENT_HELD_OUT_EVALUATION_v0.1.schema.jsondocs/spec/GC_AGENT_HELD_OUT_PRIVATE_PACK_v0.1.schema.jsondocs/spec/GC_CAPABILITY_LEASE_PROTOCOL_v0.1.jsondocs/spec/GC_CAPABILITY_LEASE_PROTOCOL_v0.1.schema.jsondocs/program/GENESISBENCH_TEMPORAL_EPOCH_AUDIT_v0.1.jsondocs/spec/GC_AGENT_PROFILE_v0.3.jsondocs/spec/GC_AGENT_TASK_CARDS_v0.3.mddocs/spec/GC_AGENT_TASK_CARDS_v0.3.jsondocs/spec/GC_AGENT_SYMBOL_INDEX_v0.3.jsondocs/spec/CLI_TOOLING_BUNDLE_v0.1.mddocs/spec/GCPM_BUNDLE_v0.1.mddocs/spec/HOST_RUNTIME_BUNDLE_v0.1.mddocs/spec/TESTING_BUNDLE_v0.1.mddocs/spec/AGENT_INDEX_v0.1.mddocs/spec/AGENT_CAPABILITY_GAUNTLET_v0.1.mddocs/spec/WRITE_GENESISCODE_SKILL_v0.1.mddocs/spec/WRITE_GENESISCODE_SKILL_PACK_v0.1.mddocs/spec/WRITE_GENESISCODE_SKILL_PACK_v0.1.jsondocs/spec/WRITE_GENESISCODE_SKILL_DISTRIBUTION_v1.mddocs/spec/GENESISBENCH_ADAPTATION_MANIFEST_v0.1.schema.jsondocs/spec/GENESISBENCH_HARDWARE_EVIDENCE_v0.1.schema.jsondocs/spec/GENESISBENCH_SCAFFOLD_MANIFEST_v0.1.schema.jsondocs/skill_pack/write_genesiscode_v1/manifest.jsondocs/skill_pack/write_genesiscode_v1/authoring-card.mddocs/skill_pack/write_genesiscode_v1/prompt-cards.jsondocs/skill_pack/write_genesiscode_v1/recipe-cards.jsonpolicies/genesiscode_authoring_workflow_v0.1.jsondocs/write_genesisCode_skill.mdexamples/canonical_language/v0.1/README.mdexamples/canonical_language/v0.1/suite.jsonbenchmarks/agent_tasks/v0.1/suite.jsonbenchmarks/genesisbench/v0.1/README.mdbenchmarks/genesisbench/v0.1/local-models/preselection.jsonbenchmarks/genesisbench/v0.1/local-models/inventory.jsonbenchmarks/genesisbench/v0.1/local-models/custody/qwen3-4b-4bit-v0.1.jsonbenchmarks/genesisbench/v0.1/local-models/custody/qwen3-8b-4bit-v0.1.jsonbenchmarks/genesisbench/v0.1/contamination.fixture.jsonbenchmarks/genesisbench/v0.1/eligibility.fixture.jsonbenchmarks/genesisbench/v0.1/reference-agent/retrieval.jsonbenchmarks/genesisbench/v0.1/reference-agent/system.mdbenchmarks/genesisbench/v0.1/reference-agent/plan.fixture.jsonbenchmarks/genesisbench/v0.1/reference-agent/trace.fixture.jsonbenchmarks/genesisbench/v0.1/adapters/command-plugin.jsonbenchmarks/genesisbench/v0.1/adapters/command_fixture.pybenchmarks/genesisbench/v0.1/adapters/deterministic-mock.jsonbenchmarks/genesisbench/v0.1/adapters/direct-local-runtime.jsonbenchmarks/genesisbench/v0.1/adapters/hosted-api.jsonbenchmarks/genesisbench/v0.1/adapters/local-openai-compatible.jsonbenchmarks/genesisbench/v0.1/baselines/conformance/evidence.fixture.jsonbenchmarks/genesisbench/v0.1/baselines/conformance/publication.fixture.jsonguides/genesisbench.qmdguides/genesisbench-methods.qmdscripts/lib/gc_agent_scoring.pyscripts/lib/gc_agent_scoring_contract.pyscripts/lib/gc_agent_benchmark_run.pyscripts/lib/genesisbench_protocol.pyscripts/lib/genesisbench_protocol_contract.pyscripts/lib/genesisbench_contamination.pyscripts/lib/genesisbench_protocol_run.pyscripts/lib/genesisbench_tracks.pyscripts/lib/genesisbench_eligibility.pyscripts/lib/genesisbench_reference_agent.pyscripts/lib/genesisbench_front_door.pyscripts/lib/genesisbench_open_agent.pyscripts/lib/genesisbench_open_agent_report.pyscripts/lib/genesisbench_local_models.pyscripts/lib/genesisbench_mlx_custody.pyscripts/lib/genesisbench_mlx_responses.pyscripts/lib/genesisbench_registry.pyscripts/lib/genesisbench_baselines.pyscripts/lib/genesisbench_construct_validity.pyscripts/lib/gc_held_out_evaluation.pyscripts/lib/gc_capability_lease.pyexamples/agent_benchmark_reproducibility/run.jsoncrates/gc_cli/tests/cli_agent_benchmark_run.rscrates/gc_cli/tests/cli_genesisbench_front_door.rscrates/gc_cli/tests/cli_genesisbench_registry.rscrates/gc_cli/tests/cli_genesisbench_construct_validity.rs
Legacy Split Docs (must stay marked)
docs/spec/CLI_JSON_SCHEMAS_v0.1.mddocs/spec/GCPM_JSON_SCHEMAS_v0.1.mddocs/spec/HOST_BRIDGE_PROTOCOL.mddocs/spec/GPU_COMPUTE_RUNTIME_PROFILE_v0.1.md
Agent Guidance
- Treat this bundle as the normative retrieval root for common workflows.
- Load the compact core card, then negotiate
GC-AGENT-v0.3before generating source; profile membership describes syntax and semantics but never grants a host capability. - Declare
genesis/agent-intent-v0.1and consumeagent-plan.plan.context_cardsrather than loading every domain document or selecting guidance from prompt text alone. - Validate capabilities and contracts through
genesis --json agent-index. - Resolve failures through bounded
genesis --json agent-index --diagnostic <exact-code>records from the closed, content-addressed diagnostic catalog; never route on message prose. - Learn or repair a language construct by selecting its signed pair in
GC-CANONICAL-EXAMPLES-v0.1, executing both sides through the recorded production argv, and changing only the declaredreplace-oncemutation. Never train on an invalid example without its rejection class and valid repair partner. - Evaluate generation, completion, repair, refactor, policy minimization, replay investigation, performance repair, package migration, and deployment against
GC-AGENT-TASK-BENCHMARK-v0.1. Its 27 public cases are nine immutable independent lineages under three child context conditions, not 27 independent samples. Treat its references as public development oracles, never as held-out evidence. - Score a candidate with
GC-AGENT-BENCHMARK-SCORING-v0.1. Its closed 10,000-basis-point quality result covers semantics, obligations, effects, patch minimality, deterministic resource units, and policy scope. Wall time, API cost, energy, and provider queue time are model/run facts forgenesis/agent-benchmark-run-v0.1; they never enter the quality score. - Record every benchmark invocation with
GC_AGENT_BENCHMARK_RUN_v0.1: immutable model, weights, tokenizer, runtime, exact prompt/card/context hashes, integer decoding and retry controls, every attempt and candidate artifact, the canonical score, normalized host facts, and a complete inventory. Validate records read-only withpython3 scripts/lib/gc_agent_benchmark_run.py --check --self-test. - Apply
GenesisBench-v0.1before comparing runs. Validate its frozen Git/SHA-256 snapshot and closed authorities withpython3 scripts/lib/genesisbench_protocol.py --check --self-test; classify a run with--run <path> --attestation <path> --json. Public references aredeclared-contaminatedand unranked, missing provenance isunknown, and only complete post-release precommitment and custody evidence can supporttemporal-clean. Never infer cleanliness from language newness or use judge preference in quality. - Use
genesisbench-reference-agent-v0.1unchanged for Cold Acquisition. Its system prompt, typed assembly, integer-only retrieval, one-agent/no-provider-tool policy, semantic transaction loop, finite budgets, and complete trace contract are content-addressed. Compare its eight ablations only as predeclared within-lineage pairs across the same nine lineages; 72 condition cells are not 72 independent samples. Validate or compile a plan withpython3 scripts/lib/genesisbench_reference_agent.py --check --self-testor--plan --case <id> --ablation <id>. - Execute transport-neutral benchmark runs only through
genesis bench. Validate the closed five-class adapter profile withpython3 scripts/lib/genesisbench_front_door.py check --self-test; retain failed attempts, never retry invisibly, replay without model or adapter access, and use deterministic.gcbundleplus local immutable outbox submission. An execution bundle is not ranked until R1.4.m independently validates, rescores, signs, and binds track, contamination, cohort, and submitter evidence. - Predeclare repository-editing systems through
genesis bench agent-planbefore inference and execute them only through the content-addressed Open Agent harness. Keep Codex CLI Lunaxhigh, Codex-over-local-runtime, fixed-scaffold raw models, and adapted systems in separate cohorts. Never call a mutable provider alias immutable, never omit a local model artifact digest, and never promote an Open Agent run whose frozen repository, workspace closure, one-attempt policy, process-group kill, transcript, or model-free replay evidence fails. - Before any local-model quality observation, validate the score-blind candidate selection, exact revision-pinned top-level artifact inventory, retained license/model-card evidence, adapter compatibility, and no-download/no-mutation claim with
python3 scripts/lib/genesisbench_local_models.py --check --self-test. Then verify the closed model/Python/MLX/adapter custody manifest withgenesisbench_mlx_custody.pyand execute only through Open Agent v0.5’s auth-free, zero-retry, separately sandboxedmlx-responsespath. Zero requests, any adapter rejection, credentials, hidden retries, server fallback, or surviving provider process makes the retained run invalid and unscorable. Reverify the unavailable-for-replay bytes against explicitly bound local roots before execution; neither preselection nor custody readiness is a benchmark run, score, or release claim. - Declare exactly one content-addressed GenesisBench track. Use
cold-acquisitiononly for an unadapted model under the fixed reference scaffold,open-agentfor disclosed custom orchestration without claimed adaptation,genesis-adaptedonly with a public lineage-manifest identity, andembedded-localonly with offline inference plus measured or hard-enforced combined model/runtime memory evidence. Never compare or aggregate across track, scaffold, profile, epoch, context/tool, attempt-policy, or hardware-class cohort keys. - Analyze results only through
GENESISBENCH_ANALYSIS_PLAN_v0.1andscripts/lib/genesisbench_analysis.py. Use one primary condition per lineage, keep repeated-condition summaries clustered by lineage, publish solved/unsolved/invalid/abstained/missing denominators, Wilson uncertainty, paired exact effects, and Holm correction, and emitindeterminaterather than unsupported pairwise decimal ranks. Public conformance observations are always unranked and cannot trigger saturation. - Run real-model baselines only from a locked
GENESISBENCH_BASELINE_PREDECLARATION_v0.1underGENESISBENCH_BASELINE_PROTOCOL_v0.1. Complete the nine-task-class reality gate before matrix expansion, bind the matrix predeclaration to that authentic publication, preserve every expected attempt or explicit missing cell, validate authentic closed bundles under one evidence root, report pass@1 separately from bounded pass@k, choose potential teachers per task class from exact verified bundle trajectories, and never let a synthetic fixture satisfy the gate or let cost improve rank. - Submit only signed closed bundles through
genesis bench submit, and admit them only through a policy-pinned local registry. Registry admission revalidates and independently rescores before deriving eligibility; verification repeats derivation from retained bytes. Compare only complete systems inside one exact cohort by solve rate, conditional quality, authority excess, context/tool use, and repair use, in that order. Cost and latency remain published facts and never break a quality tie. - Treat
GenesisBench-Construct-Validity-v0.1as the minimum evidence that a score measures GenesisCode engineering rather than reference-patch imitation. Alternative source layouts must pass behavioral, artifact, capability, obligation, resource, and metamorphic contracts; all targeted invalid controls and evaluator/scorer mutants must fail. Model identity, provider latency, prompt length, formatting, and preference are never correctness or quality oracles. - A fully local benchmark model may run only through
genesis.agent-model-runner.v0.1/inferon the pinnedhost/plugin::commandbridge profile. Preserve its request, response, tool transcript, and.gclog; replay must not reinvoke the model. This benchmark integration does not preempt the future standard model API. - Make a held-out claim only against the active epoch in
GC-AGENT-HELD-OUT-v0.1. Its active epoch contains 90 independently salted lineages, ten per core task class, with exact balance and custody attestations. Keep case payloads, salts, and oracles under ignored.genesis/private/agent-evaluation; bind every result to the epoch and commitment snapshot; useunknowncontamination whenever training provenance is incomplete; and rotate within 90 days or earlier on leakage or saturation. - Treat
GC-CAPABILITY-LEASE-v0.1as a general maintained capability protocol, not a benchmark-only oracle. It uses explicit logical steps, exact content-addressed scope, finite use budgets, deny-by-default decisions, append-only transitions, and replay-bound state identities. - Keep authoring guidance synchronized with
.agents/skills/genesiscode-authoring/SKILL.md.
Held-Out Evaluation Custody
The public repository and documentation site may expose commitments and lifecycle metadata only. Training and retrieval may ingest every tracked file, so private prompts, inputs, oracles, salts, and payloads belong exclusively under ignored .genesis/private/agent-evaluation/<epoch>/pack.json, with restrictive permissions and an evaluator account outside model retrieval roots. CI never requires this private pack, and the authoring gate rejects any tracked custody path.
Each case commitment is sha256("genesis/agent-held-out-case/v0.1\\0" || 32-byte-secret-salt || canonical-case-bytes). Domain separation prevents cross-protocol reuse; secret random salts prevent dictionary testing against low-entropy task details. Every result binds the exact epoch commitment-snapshot identity, not a branch or bare label.
Exactly one epoch is active and history is append-only. A scaled epoch must contain at least 45 independent Preview lineages or 90 mature lineages, meet the per-class minimum, keep every author/generator family below 25% of ranking weight, reserve at least 25% fresh weight, and carry a useful maintained post-release overlay. Provision and publish a fresh replacement before marking an epoch retired or compromised; rotate within 90 days or earlier on leakage or saturation; retain old commitments; never silently rescore; and label affected results declared-contaminated. Scheduled disclosure requires retirement plus 30 days. A compromise-forensics disclosure requires an active replacement. Disclosure publishes payload, salt, reason, date, and the recomputable original commitment; disclosed cases become public development material permanently.
Use declared-uncontaminated only with complete non-exposure provenance and attestation; reserve temporal-clean for tasks precommitted after the immutable model release with commitment and custody evidence. Use declared-contaminated for known exposure and mandatory default unknown for missing or incomplete provenance. Never aggregate across epochs. The held-out protocol protects secrecy and precommitment; GC-AGENT-BENCHMARK-SCORING-v0.1 defines model-agnostic quality, GC_AGENT_BENCHMARK_RUN_v0.1 defines model/run reproducibility, and GenesisBench-v0.1 decides eligibility.