Context
The current paper/evaluation text is honest that the initial cross-run learning ablation is negative at 5 iterations with strong curated priors. That weakens the central system thesis that the persistent shape-indexed database compounds across runs.
Scope
Design and run a second-generation ablation that gives cross-run learning a fair test:
- held-out cold shape buckets, not only shapes close to curated priors;
- cold random/stateless baseline vs database-seeded baseline vs cost-model ranking;
- longer budget than 5 iterations;
- multiple seeds with iterations-to-90-percent-best, final regret, and final throughput;
- clear handling of curated starter configs so they do not dominate both conditions.
Acceptance criteria
- Reproducible command and checked-in JSON/MD artifacts under
docs/results/.
- Paper/docs update that either validates the compounding-learning claim or explicitly reframes it if the result is still negative.
- Artifact includes raw per-iteration histories, not just final summaries.
Related historical issues: #2, #12, #75.
Context
The current paper/evaluation text is honest that the initial cross-run learning ablation is negative at 5 iterations with strong curated priors. That weakens the central system thesis that the persistent shape-indexed database compounds across runs.
Scope
Design and run a second-generation ablation that gives cross-run learning a fair test:
Acceptance criteria
docs/results/.Related historical issues: #2, #12, #75.