Skip to content

[Paper] Run cold-shape cross-run learning ablation v2 #112

Description

@Darkroom4364

Context

The current paper/evaluation text is honest that the initial cross-run learning ablation is negative at 5 iterations with strong curated priors. That weakens the central system thesis that the persistent shape-indexed database compounds across runs.

Scope

Design and run a second-generation ablation that gives cross-run learning a fair test:

  • held-out cold shape buckets, not only shapes close to curated priors;
  • cold random/stateless baseline vs database-seeded baseline vs cost-model ranking;
  • longer budget than 5 iterations;
  • multiple seeds with iterations-to-90-percent-best, final regret, and final throughput;
  • clear handling of curated starter configs so they do not dominate both conditions.

Acceptance criteria

  • Reproducible command and checked-in JSON/MD artifacts under docs/results/.
  • Paper/docs update that either validates the compounding-learning claim or explicitly reframes it if the result is still negative.
  • Artifact includes raw per-iteration histories, not just final summaries.

Related historical issues: #2, #12, #75.

Metadata

Metadata

Assignees

No one assigned

    Labels

    evaluationKernelBench, benchmarks, reproducibilitypaperPaper figure, table, or narrativeresearchNovel contributions, ablation, paper work

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions