Note: This roadmap describes intended directions, not commitments. Priorities will shift based on user feedback and real-world testing.
For the underlying readiness framework (6 dimensions, gate criteria,
branching forms), see docs/PLATFORM_READINESS.md.
Status: Released (2026, initial alpha)
- Hybrid Search (Vector + BM25 + keyword)
- Graph-RAG with ontology (12 relation types)
- 3-stage security model (RBAC + ABAC + Instruction Isolation)
- JWT auth + rate limiting + audit log
- Self-evolution scaffolding (Patch Pipeline 4-Gate)
- Multi-LLM routing (Ollama-based)
- Multimodal hooks (image / video / audio)
- Web search (Tavily + DuckDuckGo)
- 11-trait personality system
- Knowledge tracker (8 abilities + 6 domains)
- Internationalization (English + Korean UI)
- Real-data validation pending
- Multimodal pipeline integration limited
- Self-evolution untested at scale
Theme: Make the v0.1 capabilities trustworthy enough to recommend to a second user. Six axes, all of them P0/P1. All six axes complete. Axis 6 (second-user gate) closed on 2026-05-13 with the user secured β v0.2 β v0.3 gate clear, project formally entering v0.3 Platform Skeleton.
Goal: no single file owns more than one responsibility.
- Split
core/reasoning_engine.pyintocore/reasoning/package:engine.py(orchestration),pipeline.py(loop),modes.py(mode dispatch). PRs #37 / #38 / #39. - Consolidate
memory_*intocore/memory/package with documented public API. PR #35. - Public typed interfaces for:
Retriever,GraphEngine,PolicyEngine,Reasoner,OutputFilter. PR #50. -
tools/self/sub-process boundary β deferred to v0.3 (no observed need at v0.2 scale; in-proc access is contained).
Done when: β
import graph is acyclic and each module has < 20 KB.
Goal: every change is measured against the same yardsticks.
- STEP 7 13-query suite locked as committed regression baseline
(
eval/regression/step7_*.json,scripts/bench.py). PR #52. - RAGAS integrated for retrieval / faithfulness / answer
relevance. Live
/query/integration. PRs #51 / #64 / #66. -
scripts/bench.pywith--check/--update-baseline. PR #52. - PR-contract:
core/{retrieval,graph,reasoning}PRs paste bench numbers. CLAUDE.md rule 2 + CONTRIBUTING. PR #43. - LegalBench subset β intentionally deferred (domain-coupled; contradicts the "no parallel domains" mother-platform rule until v1.0).
Done when: β a PR cannot land without bench numbers attached.
Goal: any answer can be debugged without re-running it.
-
trace_idContextVar end-to-end + structured stage logs (auth β retrieve β rerank β graph β tool β answer β complete). PR #67. -
JAMES_TRACE_STDOUTconsole mirror (default ON for the single-user operator workflow). PRs #71 / #75. -
GET /admin/trace/{trace_id}full pipeline replay. PR #82. - Per-stage
p50 / p90 / p99 / maxlatency histograms inGET /admin/metrics?window_hours=24. PR #83. - 7-day auto-prune (
JAMES_TRACE_RETENTION_DAYSenv, default 7). PR #84.
Done when: β a hallucination report can be diagnosed by trace_id alone.
Goal: policy is a layer, not a sprinkle.
-
core/policy_engine.pyβ single point of role/sensitivity decisions. Wired into retrieval / graph / output / tools. 10 PRs (#50, #53, #54, #56, #57, #58, #59, #60, #61, #63). - Capability tokens for tool access (no direct fs path strings). PRs #57 / #58 / #59.
- Multimodal inputs flagged + quarantined before joining LLM context. PRs #60 / #61 / #63.
- Risky-coding hard-refuse policy at
pre_check. PR #70. - External red-team pass on prompt injection β deferred to v0.4 (per ROADMAP β needs an external partner, not internal work).
Done when: β removing the policy engine breaks 6+ modules.
Goal: self-evolution cannot deploy without a human.
- Opt-in env flag
JAMES_ENABLE_EVOLUTION=0(default off) +JAMES_AUTO_APPROVEsafety check (refuses to start withoutJAMES_DEV_MODE). PR #69. - feedback β candidate β eval β human approval β deploy β rollback pipeline.
- Eval gate via
scripts/bench.py --checksubprocess on every/admin/patch/approvedeploy. PR #77. - Audit log records
approver_username/approver_role/approved_at/approval_method/before_metrics/after_metrics. Lifecycle JSONLjames_patch_log.jsonl. PRs #69 / #77. - Auto-rollback on bench regression + lifecycle log records the
ROLLED_BACKevent. Tested for byte-identical recovery under simulated mid-write crash. PR #78. -
GET /admin/patch/audit?since=&approver=&outcome=&limit=operator-facing query endpoint. PR #79.
Done when: β
any deployed patch has an approver_username field
in the audit DB, and deploy without it is rejected.
Goal: numbers from real data, not just synthetic.
- Wiki corpus to 161 entities (concept 62 / org 57 / person 11 / document 31, hard-deduped via PR #28).
- STEP 7 13-query suite includes negative / dedup / lang-mix / security / meta categories.
- Multimodal pipeline integration (image / video / audio, OCR-poison quarantine). PRs #60 / #61 / #63.
- Edge case discovery: #5 / #6 / #7 / #8 / #11 / #14 / #20 all closed via real-data feedback loops.
- Second-user end-to-end bench run: secured (2026-05-13). v0.2 β v0.3 gate clear.
The following moved to v0.3 to keep v0.2 focused:
- Self-evolution end-to-end demonstration β folded into Axis 5
- Performance profiling β after Axis 1 (premature otherwise)
- Tutorial documentation β after Axis 1 stabilizes
-
Change Request primitive β generalise the approver_username pattern. β v0.2.x landed (CR-A~D); CR-E moved to v0.3. v0.1 hard-coded approver tracking for self-evolution alone (Axis 5). Every other write β wiki edits, workspace job runs, ontology patches, config saves β landed with no proposal, no review, no diff, no rollback. v0.2.x shipped
core/change_request.pycore/change_request_apply.py+ 6 endpoints + workspace UI + two target types (wiki_entityPR #243,run_jobsPR #240). Spans Axis 4 (security boundary generalisation) + Axis 5 (controlled write authorisation generalised beyond self-evolution).
- Trust zone documented in
docs/ARCHITECTURE.md Β§5.6. - Cycle plan + frozen schema + invariants in
docs/handovers/v0.2.x-cr-track.md. - Deliberately closed enum for
target_typeβ the registration API surface is the v0.3 plugin contract. - Multi-approver, team / project / department, external notification routing β all deferred to v0.3.
- Self-evolution gate wrap (CR-E) descoped from v0.2.x (4 locked JSONL-shape test files + eval-gate + rollback chain too large for the cycle without regression risk). Moved to v0.3 deliverables β see v0.3 "Change Request β finish the primitive".
- Done-when (full): every audit row carrying an
approver_usernametoday goes throughcore/change_request.pyβ achieved forwiki_entity/run_jobs; self-evolution approvers still bypass until CR-E.
-
Audit Phase 4b-2 β remove 16 JSONL writer sites. The JSONL β SQLite migration completed every reader path (PRs #206 / #207 / #208 / #210 / #211): the legacy
/admin/auditendpoint,/admin/dashboard,/code/surface/all queryaudit_logdirectly, and the new/admin/audit/listcategories (tools/attack/system) surface the mirrored rows in the admin UI. JSONL writers remain in place as a belt-and-suspenders safety net becausecore/audit_bridge.mirror_to_audit_dbis best-effort (try/except), so a silent SQLite write failure would otherwise be unrecoverable.- Re-entry criteria: 2β4 weeks of production use with no
SQLite mirror gaps observed (compare
audit_logrow count vs JSONL line count weekly), then drop the writers. - Alternative if mirror reliability is uncertain: gate JSONL
writers behind an env var (
JAMES_AUDIT_JSONL=1, default off) before deletion. - Touch points:
core/security_layer.log_attack+log_system_event; 9 module-local_log_systemcopies (llm/router, gemma_client, graph_engine, query_expander, orchestrator, retrieval_engine, memory/loom, memory/trust, reasoning/engine); 5 tool writers (tools/router, tools/code/{sandbox,code_reader,code_analyzer, code_editor}).
- Re-entry criteria: 2β4 weeks of production use with no
SQLite mirror gaps observed (compare
Closure summary (2026-05-25): Platform Skeleton theme completed across three Zenodo-archived releases:
- v0.3.0 β V3' Protocol v1 methodology spec (
docs/research/v3prime-protocol-v1.md) - v0.3.1 (DOI
10.5281/zenodo.20363998) β D1 Adaptive Budgeting closure with 7-tier natural-stop gradient - v0.3.2 (DOI
10.5281/zenodo.20372649) β D5 Auto-routing on Provider Contract (10-PR sequence #474β#484) - v0.3.3 (DOI
10.5281/zenodo.20374227) β D6 retry-wiring follow-up cycle (PR #486 wiring + #487 audit + #488 native Ollama done_reason)
ROADMAP Β§Plugin contract / Change Request / Knowledge cascade / Governance / Carryover / v0.3-only follow-ups / Measurement framework Direction 1+4+5+6(J) β all checked. Done-when: 3/4 fully satisfied + 1/4 partial (first external pack author trial awaits Ali mid-June Gemini PR). v0.4 entry triggered.
Theme: define and freeze the extension contract that all future domain packs will be built against.
Required for: any domain pack work (forbidden until this gate passes).
-
core/plugins/base.pyβ typed interfaces for 4 plugin types (β PR #344, Track C PR-C2):OntologyPack(entity types, relations, hierarchies)PromptPack(system prompts, few-shot examples per task)UIPanel(server-rendered admin/user widgets)Scorer(custom retrieval/answer scoring overrides)
-
core/plugins/loader.pyβJAMES_PACKS=general,referenceenv-driven dynamic loader; SemVer enforcement (β PR-C3 #409). Coexists with the pre-existing reasoning-backend plugin loader atcore/reasoning/backends/_load_plugins(PR #326); the two use separate env vars (JAMES_PACKSvsJAMES_PLUGINS) per design memo Option A -
core/plugins/manifest.pyβpack.yamlschema with closedlicense:enum (β PR-C3 #409) -
core/plugins/registry.pyβ slot registry per Protocol type (β PR-C3 #409) -
packs/general/β first first-party pack as no-op overlay dogfood (β PR-C5a #413). Skeleton only:GeneralOntology+GeneralPromptssatisfy the Protocols with empty values; existing JAMES defaults incore/relations_schema.py+core/reasoning/modes/remain authoritative. Server startup wiring shipped in PR-C5b #418 (2026-05-23) β the loader now runs at FastAPI startup and the pack registers into the process-wide registry. -
JAMES_WORKSPACE=env var for multi-instance hosting (β PR-C6 #410 β resolver only; β PR-C6.b #421, 2026-05-23 βconfig.pyconsumes the resolver forRAW_DIR/WIKI_DIR/UPLOAD_DIR/CHROMA_DIR; default unset behavior is byte-identical to pre-v0.3) -
docs/PLUGIN_AUTHORING.mdβ author guide (β PR-C7 #412) -
docs/VERSIONING.mdβ SemVer + 12-month deprecation policy (β PR-C8 #411) - Eval contract β every pack passes static manifest + slot-import
+ ruff via
scripts/eval_pack.py, enforced by the.github/workflows/packs-eval.ymlCI gate (β PR-C9 #419, 2026-05-23). Heavyweight RAGAS layer reserved as a follow-up step in the same workflow. -
scripts/dogfood_check.py+ CI hook β runtime check of the 4 end-to-end contract invariants of the dogfood gate (default loadspacks/general/;JAMES_PACKS=''refused; missing pack refused; path-traversal refused) (β PR-C10 #420, 2026-05-23).
β Plugin API status: 8/8 PRs landed (100%) as of 2026-05-24
post-PR-C10/C6.b. The eight-PR sequence (PR-C2 / C3 / C5a / C5b / C6 /
C6.b / C7 / C8 / C9 / C10) closed in the 2026-05-22~24 window. Full
breakdown: docs/handovers/v0.3.x-audit-2026-05-23.md.
β All other v0.3 gate items closed in the 2026-05-24 marathon session: CR-E shadow rows via Stage B 4-PR sequence (#433/#434/ #435/#436), module size violations via Stage C 5-PR sequence (#427/#428/#429/#431/#432), Audit Phase 4b-2 JSONL-writer removal via Stage D.1 (#430).
- CR-A~D completed v0.2.x β scoping (CR-A), state machine + apply dispatcher + wiki_entity (CR-B, PRs #237/#243), workspace UI panel (CR-C, PR #239), run_jobs apply path (CR-D, PR #240).
- CR-E: route self-evolution approvals
(
/admin/patch/approve,/admin/proposals/{id}/approve|reject) throughcore/change_request.pyas a shadow row so the unified audit shape becomes part of the platform contract. β Stage B 4-PR sequence complete 2026-05-24: - CR-E.1 (#433) β target_type enum + no-op apply dispatcher (self_evo_patch/self_evo_proposal) - CR-E.2 (#434) β/admin/patch/approveshadow write + close on bench-pass / apply_failed / regression - CR-E.3 (#435) β/admin/proposals/{id}/approve|rejectshadow write + close - CR-E.4 (#436) β/admin/patch/auditunifies legacy JSONL + CR-shadow read viainclude_shadow=True(default ON in the endpoint, default OFF in the library function for back-compat) Dual-write permanent. Legacyjames_patch_log.jsonl+james_evo_log.jsonlremain authoritative for the deploy timeline; CR row is additive shadow with a no-op apply handler. The 4 locked JSONL-shape tests (test_self_evolution_gate,test_evolution_bench_gate,test_evolution_rollback,test_evolution_audit_query) all stay green. - (Stretch) Open the
target_typeregistration API to plugins β today it's a closed enum on purpose; this is the surface every plugin pack will hook through.
- Replaced the v0.2 single-
confidencefield withsources: [{doc_id, weight, role, ts}]. 5-phase plan completed: Phase A schema (#266), Phase B ingestion (#269), Phase C delete (#270), Phase D modify (#274), Phase E graph editor (#271, #273). Two hotfixes for cross-doc aggregation + noisy-OR derivation (#349, #350). 12 invariants locked intests/test_relations_schema.py+tests/test_phase_b_ingestion_sources.py. Postmortem:docs/postmortems/2026-05-20-knowledge-cascade-defects.md.
- License decision for v0.3: MIT held. Trigger conditions
(T1βT5), conversion procedure, and pre-built infrastructure
(CLA Β§4-bis relicensing grant, plugin
license:field, trademark + patent tracks) committed todocs/LICENSE_PLAN.md(2026-05-11). - CLA Assistant install +
docs/legal/CLA.md+.github/workflows/cla.ymlβ operator window closed 2026-05-20 (B-3/B-4/B-5/B-6 all verified end-to-end with a dry-run from a separate GitHub account). External contributors can sign before opening their first PR with the relicensing grant in place. -
THIRD_PARTY_LICENSES.md(dependency inventory; license-strength independent) β β PR #414 (2026-05-23,idna 3.13 β 3.15security pin refresh; full pip-licenses snapshot at repo root). Re-run only on dependency churn. - Quarterly trigger monitoring β first measurement recorded at the
v0.3 release in
docs/LICENSE_PLAN.md Β§8β β PR #415 (third quarterly measurement, 2026-05-23, stars 5β13, triggers 0/5). - Trademark + patent tracks opened (lawyer consult scheduled,
progress logged in
docs/LICENSE_PLAN.md Β§6 / Β§7) β operator track
-
core/memory/store.pysplit β was 21 KB at v0.2 close; currently ~17 KB on origin/main (9a756b7). Reached compliance without an explicit split (incidental shrinkage from later refactors). No action needed. - Audit Phase 4b-2 β remove 16 JSONL writer sites β β
Stage D.1 (#430, 2026-05-24). 14 files touched, ~158 lines
removed / 47 added (entry dicts + mirror_* calls retained; only
the
try: open(JSONL_PATH).write(json.dumps(entry))blocks + orphaned path constants +import jsonremoved). SQLiteaudit_logtable viacore/audit_bridge.pyis now the sole sink.tests/test_router_capability.pywas migrated to themirror_to_audit_dbcapture-via-mock pattern (9/9 pass).
These items were identified during the v0.3 cycle but are not in the original design memos:
- i18n contract invariant β backend
label_keypattern across 7 modules (~316 entries). Convention test in 5test_*_i18n.pyfiles. Established by the 2026-05-22 sweep series (PRs #393/#394/#395/#396/#397/#398/#400/#404). - Reasoning cap defaults β
DEFAULT_MAX_TOKENSfor query_rewriter / planner / reflect / verify bumped from 200/400/400/400 to 4096 each, above the ~500-token reasoning floor measured ongemma4:e4b(PR #399 + V3'.a/.b/.c/.d 4-stage validation set, PR #407). - CASCADE vs EVENT separation axiom β established by PR #401 (architecture memo). v0.4 first ship bundle rescoped to T1+T7+T2 (PR #403).
- Module size gate violations discovered by 2026-05-23 audit
β β
Stage C 5-PR sequence complete 2026-05-24. All five files
now split into mixin/composition packages with every sub-module
β€17 KB:
- C.1
core/wiki_generator.py51 KB β 6-file mixin package (max 17 KB) via PR #427 - C.2core/cascade.py24 KB β 4-file package (max 12 KB, Phase C delete + Phase D modify + shared helpers) via PR #428 - C.3core/character_profile.py24 KB β 4-file mixin package (max 11 KB) via PR #429 - C.4core/security_layer.py23 KB β 5-file package (max 7.2 KB; landed after Stage D.1's JSONL removal shrunk the file ~50 lines first) via PR #431 - C.5core/graph_editor.py20 KB β 3-file package (max 11 KB, edge mutation writes / helpers / facade) via PR #432 Pure refactor β no signatures, return shapes, or behaviour changed. PLATFORM_READINESS Gate v0.3 dimension A fully clear. - admin.html β 3-page split β inventory completed (PR #385); actual split (Option-A Phase 3) deferred. Operator decision whether to land before v0.4 entry or run as v0.3.x parallel.
The V3'.a/.b/.c/.d (4-stage cap-budget validation, PR #407) + V3'.e (substitution/synthesis mode split, PR #440) sweeps together established a reproducible 3-axis measurement framework for LLM cost behavior. Robin Converse's 2026-05-23 same-day cross-stack replication (issue #448, repo triavalabs/gemma4-26b-mode-split) adopted the JAMES JSON schema as the cross-stack analysis template β graduating the framework from "data we publish" to "schema another lab analyses against".
Three confirmed axes:
- Mode split (substitution = deterministic / synthesis = variable). Robin original + JAMES e4b + Robin 26b β confirmed on both stacks.
- Workload gradient (heavy / light / no synthesis cost scales with task weight). JAMES V3'.e + V3'.a~d β confirmed on e4b.
- Model-scale efficiency (synthesis cost β 1/param-count; substitution invariant). Robin 26b 2026-05-23 β new axis, single point so far.
Six follow-up directions queued for v0.3.x cycle (none are domain features; all are platform-skeleton strengthening or research):
- Direction 4 β Substitution bypass verification on e4b
(1-day, JAMES standalone). Patch
v3prime_e_mode_split.pyto recordunique_response_countper cell; verify Robin's Finding 1 ("mode bypasses sampling layer") on our e4b data. Ifunique=1confirmed β publishable mechanism axis 1 strengthened. - Direction 1 β Adaptive Budgeting in Cognitive Middleware
(closed 2026-05-25). PR #461 (TaskBudget module +
experiment driver + cognitive-stages extension), PR #463
(v2 heuristic CAP_LIGHT 800β1200 + closure result docs +
7-tier monotonic natural-stop gradient, 62β1681 tokens,
27Γ dynamic range, cross-sweep stable <5% per tier), PR
#469 (entity_extract cap 1500β4096 alignment). 4 cognitive
stages all zero truncation + zero quality regression. Ships
as safe / latency-positive / memory-positive /
defensive-bound. Sub-finding:
verifyis a high-clustering cognitive stage (~12.5% unique across 40 baseline calls) β adds task-type axis to Mechanism 2. - Direction 2 β Task-weight metric formalization β absorbed into Direction 5 as a dependent fragment, not run as an independent 1-2 week cycle. The 7-tier natural-stop gradient from Direction 1 closure provides the ground-truth dataset. Per Build-don't-broadcast principle: research-only directions are run on collaborator cadence; D2 in isolation is collaborator-coupled but the piece D5 needs is product work and lands inside D5.
- Direction 6(J) β Methodology spec standalone (closed
2026-05-24, PR #457).
docs/research/v3prime-protocol-v1.mdshipped: 441 lines, 12 sections + KO, JSON schema frozen v1, 12-month grace policy, BibTeX (author=Hashevolution), Β§11 external adopter table (JAMES + Triava Labs). Other labs can cite the protocol without depending on JAMES code. - Direction 5 β Auto-routing on Provider Contract
(closed 2026-05-25, 10-PR sequence PR #474-#483). New
core/reasoning/router.py(D5.A skeleton + D5.B capability tags + D5.C.1 policy decision tree) consumes Direction 1's budget to select backend (small/medium/largetier Γlocal/sovereign/cloudprovider). 5-stage wiring complete (query_rewriter/planner/reflect/verify/ synth viatrace_helpers) β every production LLM call path consults the router, gated byJAMES_AUTO_ROUTERflag (default OFF β byte-identical to pre-D5). Audit rowreason:routeper resolve call. Verify stage is grounding-critical β policy escalates tolargetier when registered (e.g.JAMES_ENABLE_CLAUDE_BACKEND=1). Provider Contract surface unchanged (router sits above L1). Cross-lingual RAG option 3 (core/entity_alias_pack.py+graph_engine.build_entity_map_snapshotaugmentation, PR #483) bundled β KOβEN surface forms now resolve to the same wiki entity_id without per-install frontmatter edits. Per Build-don't-broadcast principle (memory:feedback_build_dont_broadcast) this is a product cycle β no public broadcast, no Robin coupling. Operator-run STEP 7 sweep numbers (D1 7-tier ground truth) integrate into the closure result doc when operator runs the measurement; the wiring is bench-neutral at flag-OFF default. - Direction 3 β Cross-family generalization (~2-3 weeks, research). V3'.e on Llama 3.1 / Qwen 2.5 / DeepSeek v2 via Ollama. Output: mode-split universality vs Gemma-only judgment. Result first-shared with Robin (axis-1 owner) per collaborator informational-notice pattern.
- Direction 6(I) β Joint paper consolidation (Track 5 essential). Three-author byline (Ali / Robin / Hashevolution) after Directions 1-5 produce input. Track 4 scope-lock note applied.
Collaborator interaction pattern: Track 1 success pattern (land first + "no action needed" DM) applies to all directions except 6(I) joint paper. Direction 3 result first-shared to Robin; Direction 5 design preview shared to Ali; otherwise informational notice on PR landing.
v0.4 / v0.5 alignment: Direction 1 (adaptive budgeting) +
Direction 5 (auto-routing) integrate naturally with v0.4 Layer 4
GOVERNANCE/EVENT tracks (budget decisions = audit trail; routing
decisions = events). Direction 5 also strengthens v0.5 Domain
Pilot value proposition (domain pack-aware model routing).
Detailed handover: docs/handovers/v0.3.x-measurement-framework-track.md.
| Criterion | Status |
|---|---|
A new contributor can build a no-op pack from docs/PLUGIN_AUTHORING.md alone in < 1 day, load it, and observe its effect |
|
The dogfood test passes: packs/general/ produces byte-identical STEP 7 results to v0.2 main; deleting the pack breaks the server cleanly |
β
β packs/general/ registers at startup via PR-C5b (#418); scripts/dogfood_check.py (PR-C10 #420) locks the four runtime invariants (default loads general; empty JAMES_PACKS refused; missing pack refused; path-traversal refused) on every PR via .github/workflows/packs-eval.yml. Byte-identity is preserved because the overlay is empty. |
| Every self-evolution approval row has a paired Change Request row (CR-E acceptance) | β
β Stage B 4-PR sequence complete. PR #433 (target enum + no-op apply), PR #434 (/admin/patch/approve shadow write), PR #435 (`/admin/proposals/{id}/approve |
| CLA Assistant blocks any unsigned external PR at the workflow gate | β β verified end-to-end 2026-05-20 |
β 3/4 fully satisfied + 1/4 partially satisfied as of 2026-05-24
(post-Stage-B/C/D marathon session). Only remaining partial:
first external author trial against PLUGIN_AUTHORING.md
(infrastructure complete; awaits Ali's mid-June Gemini backend PR
as the first external trial). See
docs/handovers/v0.3.x-audit-2026-05-23.md for the staged plan and
the post-marathon reconciliation table.
- Any domain-specific pack (legal, food, retail)
- External plugin marketplace
- Plugin signing infrastructure beyond manifest hash
- Multi-approver workflows / team-project-department scoping on CR (still v0.3 plugin contract surface; ships only if a real second user needs it)
v0.4.0 β Layer 4 Lifecycle Semantics (entered 2026-05-25, T1+T7+T2 first bundle shipped 2026-05-27 at Sprint 5)
Status: T1 (Temporal Validity) + T7 (Supersede Chain) + T2 (Contradiction Arbitration) first bundle landed via Sprint 5 (8 PRs, #523~#543). v0.4.0 is release-ready β the CASCADE / EVENT separation invariant is provable end-to-end via the release-gating tests in tests/test_t7_release_gating_invariants.py.
Entry handover: docs/handovers/v0.4.0-entry-track.md β 6-sprint plan covering data correctness, UI consistency, plumbing, retrieval quality, Layer 4 main theme (T1+T2+T7), and long-term backlog. Sprint 0 (this entry) + Sprint 1 (graph entity-event relation diagnostic + language detection&matching) ran first.
Sprint 5 first-bundle entry: docs/handovers/v0.4.0-sprint5-layer4-first-bundle-entry.md β the locked-decision 7-PR sequence (PR-0 schema validators β PR-T1.A migration β PR-T1.B expiration cascade β PR-T7.A supersede chain β PR-T7.B release-gating invariants β PR-T2.A classifier β PR-T2.B A-path routing β PR-T2.C B-path routing β PR-T7.C closure) that shipped here.
Theme: deepen the memory lifecycle beyond the v0.3 Layer 3 cascade (Memory OS) into Layer 4 β two complementary tracks EVENT/TEMPORAL (semantic evolution) and GOVERNANCE (write-time control), plus a Layer 3 extension for causality. Seven areas (T1βT7) chosen from the 2026-05-21 user critique series.
Core architectural axiom β CASCADE vs EVENT separation
(docs/architecture/memory-lifecycle-architecture.md Β§1.5):
| Class | Trigger | Handling | Example |
|---|---|---|---|
| Invalidated β CASCADE (Layer 3) | Explicit invalidation signal (doc delete, ingestion error, source revoke) | Source removal + propagation + orphan sweep | Mis-uploaded doc / withdrawn report / typo correction |
| Superseded β EVENT/TEMPORAL (Layer 4-A) | New doc / time progression / policy change / state transition | Validity window close + supersede chain (past data preserved) | CEO change / policy revision / contract termination |
Many naive RAG/KG systems collapse both into one mechanism
(overwrite/append) β contradiction explosion, stale retrieval,
hallucination increase. v0.4 enforces the separation as a
test-level invariant (T7 invariants:
supersede_does_not_trigger_cascade /
cascade_preserves_supersede_chain).
Why v0.4 was retargeted from "Domain Pilot": the v0.3 cascade (Phase AβE + 2 hotfixes) revealed that "Layer 3 alone" doesn't cover real operational governance β stale facts accumulate, contradicting sources reject silently (Gate 5), manual entries have no reviewer hierarchy, and event sourcing for arbitrary-time graph replay is absent. The Domain Pilot moves to v0.5 (it requires Layer 4 governance to be production-credible anyway).
Required for: any v0.5 domain pilot (forbidden until Layer 4 T1+T7+T2 minimum lands and the CASCADE/EVENT separation invariant is provable).
| Stage | Area | Track | Goal |
|---|---|---|---|
| T1 | Temporal Validity & Expiration | EVENT | Fact valid window definition + auto-expiration (mark, not delete) |
| T2 | Deterministic Contradiction Arbitration | GOVERNANCE β routing | LLM-free deterministic resolution + A/B classification (CASCADE vs EVENT) |
| T3 | Evidence Aging & Trust Decay | EVENT | Per-source weight decay by domain function |
| T4 | Reviewer Authority Hierarchy | GOVERNANCE | Multi-level manual source governance (analyst / manager / admin) |
| T5 | Replayable Audit Graph | EVENT | Event-sourced reconstruction, arbitrary-time graph replay |
| T6 | Causality Chain Tracking | CASCADE extension | Derived fact base-fact tracking + auto-invalidation propagation |
| T7 β | Supersede Chain | EVENT | Edge-level status: {active, superseded_by}. World change β past preserved + new edge created + chain walk |
Full design: docs/design/v0.4-lifecycle-semantics-roadmap.md.
Reference architecture: docs/architecture/memory-lifecycle-architecture.md.
- v0.4.0 (2026 Q4 ~ 2027 Q1): T1 + T7 + T2 (EVENT core + routing). T1 (validity window) and T7 (supersede chain) ship together β separation is meaningless without both. T2's A/B routing enforces the separation at write time (misclassification becomes immediately testable).
- v0.4.1 (2027 Q1 ~ Q2): T6 β Causality Chain (CASCADE extension)
- v0.4.2 (2027 Q2): T3 β Evidence Aging (EVENT)
- v0.4.3 / v0.5 prep (2027 Q3 ~ Q4): T4 + T5 β Reviewer Authority + Replayable Audit
- Existing 12 locked in
tests/test_relations_schema.py+tests/test_phase_b_ingestion_sources.py(Layer 3 cascade). - New T1βT7 invariants: ~27 (cumulative 39). Per-stage
contract test files following the 2026-05-22 i18n sweep
pattern (
tests/test_*_i18n.pyΓ 5 already pin label_key conventions β same_contract-style files for T1βT7). - Critical separation invariants (T7):
test_supersede_does_not_trigger_cascadeβ EVENT mutation must not call Layer 3cascade_removetest_cascade_preserves_supersede_chainβ CASCADE invalidation must not breakstatus.superseded_bylinks (active history preserved)test_historical_replay_via_chainβreconstruct_view_at(t)accurately reconstructs the graph state at any past time
- i18n
label_keyβ every UI-exposed label from new Layer 4 modules (reviewer_rank, approval_state, contradiction reason, aging policy names) must follow the backend label_key contract established by the 2026-05-22 sweep across 7 modules. Convention test in the correspondingtest_*_i18n.pyblocks silent regression. - Bench gate (CLAUDE.md rule 2) β T1 expiration cascade,
T6 invalidation cascade, and any other change touching
core/retrieval/core/graph/core/reasoningmust paste STEP 7 bench numbers in PR body. - Module size gate (CLAUDE.md rule 5) β new
core/lifecycle/*.pyfiles each < 20 KB. Split first if a change would push over.
- β
T1 + T7 + T2 (minimum) shipped, with new invariants green
(
tests/test_lifecycle_*.py+tests/test_t7_*.py+tests/test_t2_*.py, 123 tests). - β
Separation invariant provable β
test_t7_release_gating_invariants.pypins three release-gating invariants against the actual wiki fixture (not mocks):test_supersede_does_not_trigger_cascade,test_cascade_preserves_supersede_chain,test_historical_replay_via_chain. The B-class CEO-change scenario (Joby CEO Alice β Bob) is end-to-end exercised intest_supersede_chain_replayable_after_dispatch. - β
Operator scenario A (CASCADE):
cascade_remove_doc_from_sourcesunchanged from v0.3 + now reachable viacontradiction_router.route_a_invalidatewith audit row carryingmutation_type=invalidated. - β
Operator scenario B (EVENT):
expiration_cascade(T1) +supersede_edge(T7) +route_b_supersede(T2.C) land. Edges are markedstatus.active=Falsenot deleted; historical replay works viareconstruct_view_at. - β
Reference architecture memo + 7-area design memo published
(
docs/architecture/memory-lifecycle-architecture.md+docs/design/v0.4-lifecycle-semantics-roadmap.md). - βοΈ Carried into v0.4.1: ingestion-path caller that invokes
dispatch_contradictionat a real contradiction-detection point. Sprint 5 ships the primitive surface + the routing wire; the call site is the next operator integration. End-to-end STEP 7 supersede bench (Alice β Bob historical replay) deferred to the ingestion-wire PR.
- Any domain pack (legal / food / retail) β moved to v0.5
- External customer onboarding playbook
- Public eval results in
eval/RESULTS.md - 6-month production track record
Status: β CLOSED. T6 Causality Chain landed as the v0.4.1 cycle's main theme (5-PR sequence T6.A β T6.B β T6.C β T6.C.b refinement β T6.D + the v0.4.0 carry-over T2.D-1/2/2.b/3 + QVT Ξ± track full closure). 21 PRs in a single multi-day session (#548 ~ #566).
Entry handover: docs/handovers/v0.4.1-t6-causality-chain-entry.md β 4-LOCK decisions (eager trigger / operator-tagged + LLM flag / strict cycle reject / C.b foundational-vs-corroborative semantics β Decision 4 refined in T6.C.b vs the original memo text).
- PR-T6.A (#562) β
derived_fromschema field +validate_edge_t6_derived_fromcycle validator (Decision 3 LOCK) +apply_t6_edge_defaultsidempotent helper.scripts/migrate_v041_lifecycle.py(--dry-run/--apply/--verify/ pre-write snapshot). 23 contract tests. - PR-T6.B (#563) β
core/lifecycle/derivation.pyextract_derivation_chain(operator-tagged path +JAMES_T6_LLM_DERIVATIONflag-gated LLM-inferred path). 14 contract tests. - PR-T6.C (#564) β
core/lifecycle/causality.pyinvalidate_derived_facts(base_fact_id, entity_root, *, additional_empty_bases, audit_emit). Soft-invalidate (status.active=False + mutation_type=invalidated, sources preserved for T7 replay). Atomic per-file writes. 19 contract tests. - PR-T6.C.b (#565) β foundational-vs-corroborative refinement.
transitive/inferredare hard deps (any base empty β invalidate);operatoris corroborative (only invalidates when no hard deps AND all operator bases empty). 22 tests (19 original + 3 new C.b cases). - PR-T6.D (#566) β cascade integration.
cascade_remove_doc_from_sourcesnow callsinvalidate_derived_factsfor every relation whose sources became empty (single walk batched viaadditional_empty_bases).tests/test_t6_release_gating_invariants.py(5 tests):test_derived_invalidated_when_base_removed,test_partial_base_loss_preserves_derived(T6.C.b),test_self_reference_rejected_at_write+test_two_hop_cycle_rejected_at_write(Decision 3),test_cascade_invalidate_emits_audit_row. Run against tmpdir wiki fixtures + realcascade_remove_doc_from_sourcesβ no mocks of the cascade itself.
- PR-T2.D-1 (#558) β
core/lifecycle/contradiction_ingest_detector.py(P1 different_tail + P2 divergent_validity patterns). 19 tests. - PR-T2.D-2 (#559) β
core/lifecycle/ingest_contradiction.py+_merge.pypre-merge hook. Flag-gated byJAMES_T2D_INGEST_DISPATCH(default OFF). 10 tests. - PR-T2.D-2.b (#561) β A_invalidate cascade race fix via
PendingCascadedeferred-execution. 15 tests. - PR-T2.D-3 (#560) β step7 v6 q17 "Anthropicμ CEOλ λꡬμΌ?" + acceptance integration test. 6 tests.
- PR-Ξ±-1 (#550) β
docs/design/v0.4-qvt-alpha-non-saturating-oracle.md. 3-axis non-saturating oracle + per-PR Quality Delta Card pattern + 5 exemption labels. - PR-Ξ±-2 (#551) β step7 fixture v4 β v5 (
gold_signals+abstention_truth). - PR-Ξ±-3 (#552) β
eval/qvt/oracle.py+scripts/qvt_capture_baseline.py(operator wrapper). - PR-Ξ±-4 (#553) β
.github/PULL_REQUEST_TEMPLATE.md+ CLAUDE.md rule 2 extension +docs/ARCHITECTURE.mdΒ§5.7.10 QVT subsystem. - PR-Ξ±-3 baseline capture (#555) β
eval/qvt/baseline_2a31b20.json(N=3 paired reruns, canonical reference). - PR-Ξ±-3 oracle calibration (#556) β Korean security-block phrase additions +
blocked=Trueshort-circuit.abstention_f10.29 β 0.67 median.
- Replayable RAG positioning (#548) β README + ARCHITECTURE category framing.
- F9 cycle full closure (#549) β q15 zero-recall closure result doc + audit-trail bench JSON.
- v0.4.0 post-mint DOI (#554) β
10.5281/zenodo.20411354. - v0.4.1 entry memo (#557).
| Flag | Default | Verification |
|---|---|---|
JAMES_T2D_INGEST_DISPATCH (T2.D-2) |
OFF | _merge.py pre-merge hook only fires when =1 |
JAMES_T6_LLM_DERIVATION (T6.B) |
OFF | extract_derivation_chain LLM-inferred path needs provider AND flag |
| T6.D cascade integration | ON (default-on per cycle scope) | byte-identical retrieval because migration adds derived_from: []; no actual derivations yet |
- β T6 5-PR sequence merged with all 83+ contract tests green.
- β
T6.D 4 release-gating invariants run against real wiki fixtures (
tests/test_t6_release_gating_invariants.py). - β T2.D ingestion wiring (carry-over from v0.4.0) shipped with race-free pending_cascades pattern.
- β QVT Ξ± track complete: oracle module + canonical baseline JSON + PR-gate template + CLAUDE.md rule 2 extension + ARCHITECTURE.md Β§5.7.10.
- β
Migration script (
scripts/migrate_v041_lifecycle.py) is idempotent + byte-stable;--verifymode confirms. - β 69 pre-T6 cascade tests still green (no regression).
- β
Closure docs published:
docs/release_notes_v0.4.1.md+ CHANGELOG[0.4.1]+.zenodo.jsonv0.4.1.
- T3 Evidence Aging β confidence decay over time (EVENT track)
- T4 Reviewer Authority Hierarchy β multi-level governance (GOVERNANCE track)
- T5 Replayable Audit Graph β full event-sourced reconstruction (partial in v0.4.0 via
reconstruct_view_at) - ~~v0.4-end QVT ablation matrix capture (18 cells Γ N=3 reruns, ~20 hours operator-run)~~ β re-shaped 2026-05-30~31 into the Ξ±-5 cycle below.
JAMES_T2D_INGEST_DISPATCHdefault flip to ON (waits for fixture coverage)- Production seeding of
derived_from(waits for operator workflow or v0.4.2+ LLM path)
Theme: the v0.4-end ablation matrix that v0.4.1's "out of scope" line deferred β re-shaped end-to-end during execution. The single bullet "18 cells Γ N=3, ~20 h" turned out to be too thin a frame for what the matrix actually needs to deliver. Recording the corrected sizing, the external-benchmark adoption, and the dual purpose so the next operator entering this cycle reads the right shape on the way in.
- Routing decision β flag-ON / tier-gating for each Layer 4 cognitive
routing layer (
AUTO_ROUTERD5,ADAPTIVE_BUDGETD1,SCOPE_ROUTINGLEO). The original framing. - Reasoning-capability evidence β publishable proof that JAMES's
layer stack outperforms peers on external benchmarks. Ties to
eval/RESULTS.md'sexternal benchmarkline that v0.3 deferred β this cycle is the unblock.
Both purposes must be served by every cell verdict; bucket-(d)
measurement artifacts must be removed before either narrative ships
externally (per feedback_oracle_phrase_artifacts 4-step rule).
- 19 cells (18 standard + 1 sanity
M_M/L1 think=ON) β extra cell carries A2 default-flip evidence inside the matrix. - 5-axis oracle β 3 quality (path / graded / abstention) + 2 cost (token_cost, latency_cost) per cell, Pareto verdict (strong-adopt / adopt / efficiency-adopt / tier-gated / reject / zero). Cost integration was added to serve user requirement #3 (token / time efficiency).
- Workspace runtime (MultiHop-RAG balanced-100, 931-entity
workspace, A2 think=OFF):
- Ingest: 264 min for 183 articles (latency grew with workspace size β first 25 s, last 197 s)
- Baseline N=1: 107 min (100 queries Γ ~64 s)
- T0 smoke (N=1, 3 cells): ~5.3 h estimate
- Full N=3 sanity-included T0+T1: ~16 h
- Full N=3 T0+T1+T2 across all 3 tiers: ~30 h (vs original 20 h)
- Adaptive tiering β
--t0-smoke(~2-5 h) gates whether T1 (M_M, ~5.5 h) is worth running; T1 gates whether T2 (M_S + M_L, ~13 h) adds tier-gating signal.
eval/RESULTS.md originally said "BEIR / MS MARCO β too large for
laptop, defer to v0.3" and "KLUE-RC β no clean public option yet,
watch for v0.4." This cycle finally lands an external benchmark via
MultiHop-RAG (Tang & Yang, EMNLP 2024) β 2,556 multi-hop QA over
609 news articles, CC-BY-4.0, with 4 question types (comparison /
inference / temporal / null) that map directly onto the routing /
abstention measurement need.
Workspace isolation via the existing JAMES_WORKSPACE env
(config.py:74, core/plugins/workspace.py) β zero code change to
the data-dir resolver. Production wiki is untouched throughout the
cycle.
| PR | Layer | Note |
|---|---|---|
| #615 | reset | MultiHop-RAG external benchmark + 5-axis + per-question-type matrix |
| #616 | exec | Ingest wrapper + bench timeout fix + path-axis finding |
| #617 | docs | Β§7.4 first per-question-type baseline signal |
| #618 | fix (d) | source-recall β bench was dropping response.sources |
| #619 | fix (d) | abstention phrases β gemma4:e4b English refusals |
| #620 | tool | re-score tool + Β§7.4 76% β 36% correction |
| #621 | docs | 4-bucket diagnostic taxonomy + dual purpose |
| #622 | fix | render_report defensive path + bucket retroactive |
| #623 | fix (d) | session_id suite-aware + 3 narrow abstention phrases |
| #624 | docs | ROADMAP Β§v0.4.x Ξ±-5 cycle entry (this section, initial form) |
| #625 | fix (a) | matrix runner β bench subprocess respects --suite (was hardcoded step7) |
| #626 | docs | T0 smoke result analysis TEMPLATE (qvt-ablation-T0-smoke-analysis-TEMPLATE.md) |
| #627 | docs | Pareto verdict walk-through + CLAUDE.md rule 2 fix label exempt + diagnostic post-mortem |
| #628 | feat | T0 analysis fill script β auto-populate template from cell JSONs |
| #629 | docs | oracle.py β package split design (deferred to post-matrix) |
| #630 | docs | Ξ±-5 cycle summary outline β reviewer-ready closure read |
| #631 | docs | publishable narrative draft β "Don't Build a Layer for the Bug" |
| #632 | docs | bucket-(c) LLM-judge abstention detector design memo |
| #633 | feat | qvt_promote_findings.py β auto-draft memory entries from findings.md |
All 19 cycle PRs above land on main between f7762a3 and ba50c47.
Two landmark findings of the cycle β path_recall = 0 and
null_query hallucination = 76% β were both bucket-(d) oracle
artifacts, fixed in #618 / #619 / #623. Without the 4-step
verification rule (memory feedback_oracle_phrase_artifacts) they
would have generated wrong-bucket follow-ups (new citation layer,
grounding architecture change) and the matrix verdicts would have
read as "JAMES fails the benchmark" when the failure was on the
measurement side.
Wrong-fix avoided count: 3 (path_recall=0 β new citation layer prevented #618; 76% hallucination β grounding rewrite prevented #619+#623; matrix near-zero verdict β AUTO_ROUTER decommission prevented #625). Cumulative system code change attributable to Ξ±-5 measurement debt: 0 lines.
Mid-cycle infra additions (#626β#633): 1 template + 1 fill script + 2 design memos + 1 closure outline + 1 publishable narrative + 1 finding-promotion script. Locked deferred-bucket-(c) follow-up + locked the methodology lesson in a publishable form before the cycle closes.
- 5-axis matrix report at
reports/promo-assets/v0.4-qvt-ablation-matrix-20260531T065209.md. Both L1/M_M (baseline) and L5/M_M (full stack) classified reject vs corrected baseline; sanity cell L1/M_M-thinkON also reject. - T0 analysis at
reports/research-runs/qvt-ablation-T0-smoke-result-20260531-1552.md. - Rescore audit at
reports/research-runs/qvt-ablation-rescore-summary.mdβ 3/3 cells QQ-bugged at write time, all rescored in one pass. - Findings log at
reports/research-runs/qvt-ablation-findings.mdβ all entriesbucket:-tagged (a/d). - ROADMAP / backlog sync β this section locked, status moved from "in flight" to "T0 closed". A1 in backlog Β§2 marked closed.
- Routing-flag default-flip PRs per layer β not applicable at production tier (Branch B verdict). Tier-gated assessment requires T1 (M_S + M_L) which is post-T0 operator decision.
- T1 / T2 operator decision β see Β§"T0 verdict + post-closure" below.
| Read | Number | Compared to |
|---|---|---|
| L1/M_M baseline (production) | path 0.419 / graded 0.327 / abst_f1 0.591 | corrected baseline 3a961a3_rescored (path 0.404 / graded 0.343 / abst_f1 0.609) |
| L5/M_M full stack (all routing ON) | path 0.412 / graded 0.317 / abst_f1 0.500 | abst_f1 -0.091 vs L1 β real regression |
| Sanity L1/M_M think=ON | path 0.404 / graded 0.370 / abst_f1 0.533 / token -109 chars / latency -1.7 s | think=ON not unambiguously winning; A2 default-flip gate (real-query QDC) unchanged |
Verdict (post-closure-corrected 2026-05-31 PM): routing-layer
stack as exercised in Ξ±-5 reduces to
ADAPTIVE_BUDGET + SCOPE_ROUTING + (AUTO_ROUTER no-op) because the
matrix env registered only the always-on ollama_local backend
(core/reasoning/backends/__init__.py:329); without multi-tier
backend registration the routing policy in core/reasoning/router.py
collapses to legacy on every branch. The combined stack regresses
abstention_f1 by 0.091 without meaningful cost benefit
(token -1.5%, latency +3.6%) at the M_M production tier. Both cells
classify reject under the 5-axis Pareto rule.
AUTO_ROUTER verdict is not in evidence at this cycle; proper
measurement requires multi-tier backend registration (engineering
preconditions documented in Ξ±-6 design memo). ADAPTIVE_BUDGET was
under-instrumented β judged by quality axes when its design intent
(per core/reasoning/budget.py docstring) is cost-optimization at
quality-neutral. Per-layer isolation (L2 / L3 / L4 cells) defers to
T1. Branch B of the publishable narrative Β§6 applies as
"conditional Branch B" β routing-layer stack inert as measured by
quality axes, with explicit caveats. Opt-in flags remain available;
no default-flip PRs file at this verdict.
Cycle totals: 36 PRs (#608 β #645+), 5 measurement-side fixes
(3 bucket-d + 2 bucket-a), 4 wrong-fix-averted, 0 lines of JAMES
code changed against Ξ±-5 measurement debt. The publishable claim
(see Β§6.1 of reports/research-runs/alpha-5-publishable-narrative-DRAFT.md)
is the cycle's load-bearing artifact.
T1 (M_M + more rows) and T2 (M_S + M_L) are operator-budget calls. Recommended:
- T1 (~5.5 h on M_M) β measures whether individual layers (L2/L3/L4) show effects that all-on L5 smears together. Could surface ADAPTIVE_BUDGET (L3) wins inference-query graded while AUTO_ROUTER (L2) hurts abstention.
- T2 (~13 h on M_S + M_L) β tests the tier-gated routing hypothesis: smaller models might benefit MORE from routing (compensating for base-capability gap); larger models might benefit LESS. T2 closes the cycle's "model is universally tested" caveat.
Both deferable. T0 closure is itself shippable as "production-tier evidence: routing layers inert" + Branch B publication. T1/T2 fold into an Ξ±-5.1 or Ξ±-6 cycle.
- bucket-(c) LLM-judge abstention detector β 6 of 9 remaining FN on baseline_f7762a3 use "not possible to" / "is not available" phrasings that broader phrase additions would FP-flood. An LLM classifier ("does the answer refuse to answer the question?") would resolve them without phrase-overlap risk. Deferred candidate.
- Korean-corpus version of MultiHop-RAG β current cycle is English
only (gemma4:e4b handles both per
core/i18n.py, but the matrix measures English routing only). No public Korean multi-hop RAG benchmark exists; translation cost vs. value is a v0.5+ question. - Production-real-query Quality Delta Card β bridging the
benchmark-vs-production gap for the A2 default-flip
(
JAMES_GEMMA4_E4B_THINK_OFF) decision. Plan is documented but the card itself is a separate cycle.
Theme: successor to Ξ±-5. Ξ±-5 toggled 5 routing flags at one model tier and found the routing stack inert at production tier (Branch B). Ξ±-6 reshapes to:
- Sector-level ablation (10 sectors of JAMES infrastructure instead of 5 routing flags). Cells are sector combinations, not flag combinations.
- Multi-LLM extension β gemma3 1b/4b/12b/27b + gemma4:e4b + (later) cross-family (qwen2.5 / llama3.1 / deepseek-v2).
- JAMES-vs-vanilla comparison β Ξ±-5's L1 baseline already had every other JAMES sector on; Ξ±-6's C_minus cell strips them off and measures bare gemma against benchmark.
Cycle sizing (vs original Ξ±-6 design memo Β§3):
- 4 phases (Phase 0 sector flags / Phase 1 M_M / Phase 2 M_S / Phase 3a scale ladder / Phase 3b cross-family deferred).
- 75+ PRs through Phase 3a entry, 8 wrong-fix-averted, 0 lines of JAMES code changed against measurement debt.
- Per-tier wall-clock: M_XS ~15 min / M_S ~30 min / M_M ~107 min / M_L ~65 min / M_XL ~6-8h (GPU/CPU split, see below).
| Phase | Tier | Status |
|---|---|---|
| Phase 1 | M_M (gemma4:e4b) | β closed 2026-06-01 AM |
| Phase 2 | M_S (gemma3:4b) | β closed 2026-06-01 AM β tier-gated hypothesis REVERSED |
| Phase 3a step 1 | M_XS (gemma3:1b) | β closed 2026-06-01 PM |
| Phase 3a step 2 | M_L (gemma3:12b) | β closed 2026-06-01 PM β gemma4-only hypothesis REVERSED |
| Phase 3a step 3 | M_XL (gemma3:27b) | π‘ in flight (bench timeout overrides applied, see PR-pending) |
| Phase 3a closure | recovery curve + closure PR | β³ post-27b |
| Phase 3b | cross-family (qwen / llama / deepseek) | deferred (gated on Phase 3a verdict) |
| Finding | Tier | Status |
|---|---|---|
| S4 citation tier-invariant β path Ξ +0.397~+0.420 across 1b/4b/12b/e4b (4-point series, 27b pending) | βββ candidate | universal-law promotion gated on 27b confirmation + cross-fixture sanity |
| JAMES S5 effect non-monotonic in pure-model abstention (1b 0 / 4b -0.074 / 12b +0.375 / e4b +0.033) | ββ partial | mechanism candidate = instruction-following capacity threshold |
| Graph layer regresses graded (-0.054 at M_M C_rag-graph step) | ββ measurement debt | resolves on Ξ±-7 graph top-K fix; logged in qvt-ablation-findings.md |
| Withdrawn framings β "JAMES amplifier", "inverted-U capability floor", "gemma3 vs gemma4 family", "REVERSES tier-gated hypothesis" | β self-deception | superseded by 12b / 27b data; honest framing memory applied |
reports/research-runs/alpha-6-cycle-pr-index.md.
- Phase 3a 1b analysis (
alpha-6-phase-3a-gemma3-1b-analysis-20260601.md) with post-12b reconciliation header - Phase 3a 12b analysis (
alpha-6-phase-3a-gemma3-12b-analysis-20260601.md) with honest framing tier table - Phase 3a 27b analysis (post-bench)
- Recovery curve doc (
alpha-6-phase-3a-recovery-curve-20260601.md) with Β§5 withdrawn-claims registry - Closure PR consolidating all 3 analyses + recovery curve + 2 timeout fix patches + CLAUDE.md sync + Ξ±-7/Ξ±-8 design memo seeds
Theme: first JAMES code change against Ξ±-5 + Ξ±-6 measurement debt.
Per Ξ±-6 Phase 1 Β§3 finding, the graph layer (expand_dynamic in
core/graph_engine.py:314) is the net regressor on graded at M_M
(-0.054). 41-161 entities surface per query at the 931-entity workspace,
swamping the LLM context window with low-signal nodes.
Status: design memo landed
(docs/design/v0.4-alpha-7-graph-topk.md); implementation gated on
Ξ±-6 closure + re-baseline.
Scope (per design memo Β§2):
- New module
core/graph_topk.py(~3 KB target) β post-DFS top-K filter, default K=10 - Tighten
DFS_SCORE_THRESHOLD0.05 β 0.08 - Wire into
expand_dynamicreturn - Avoids forcing
core/graph_engine.pysplit (already 20.4 KB, marginally over the 20 KB gate β new module routes around it) - Re-baseline
multihop_ragpost-PR (~107 min operator action) - 5-axis Quality Delta Card with per-question-type cross-tab
Acceptance band (per design memo Β§4): graded_answer Ξ β₯ +0.030
for adopt; +0.010-0.030 for tier-gated; β€ +0.010 β reject + sub-finding
investigation.
Honest framing (per memory feedback_finding_size_honest_framing):
- β operational β top-K filtering is a standard GraphRAG technique (MS GraphRAG, Neo4j LLM KG Builder both do variants). PR's value is the measured Ξ numbers, not the mechanism.
- Don't frame as "JAMES discovered top-K filtering" β false claim.
Cycle dependencies:
- Predecessor: Ξ±-6 closure (the re-baseline depends on Phase 3a closure being on main)
- Successor: Ξ±-8 ontology typed-filter (A/B compares against Ξ±-7's K-bound baseline)
- Parallel: T3 Evidence Aging (orthogonal axis)
Status: cycle closed as REJECT. PR #680 closed without merge. Ξ±-7 graph top-K is a research artifact, not a production change.
5-tier remeasurement verdict (1000 queries, 0 timeouts, 0 errors, 10/10 cells written):
| Tier | Ξ±-6 contribution | Ξ±-7 contribution | change | mode change |
|---|---|---|---|---|
| M_XS (1b) | 0.000 | 0.000 | 0 | unchanged (inert) |
| M_S (4b) | -0.074 | -0.067 | +0.007 | unchanged (disrupt) |
| M_M (e4b) | +0.033 | -0.283 | -0.316 | CRASH amplify β disrupt |
| M_L (12b) | +0.375 | -0.076 | -0.451 | CRASH create β disrupt |
| M_XL (27b) | -0.181 | -0.359 | -0.178 | worse disrupt |
Mechanism (CONFIRMED): top-K=10 removes "evidence-of-absence" signal at every tier where pure-LLM had nonzero abstention capability. Universal regression. K=50/K=25 tuning would not fix β wrong-knob mechanism, not wrong-value problem.
S4 universal-law candidate (path Ξ) survives the context reshape: spread Β±0.013 across 5 tiers (vs Ξ±-6 Β±0.022). βββ candidate strengthens β citation pipeline is graph-entity-count independent.
Wrong-fix-averted: 10th cumulative (7 Ξ±-5 + 2 Ξ±-6 + 1 Ξ±-7). Ξ±-7 cycle's value = preventing the wrong fix from shipping, plus mechanism finding informing Ξ±-8.
Carry-forward to Ξ±-8:
- Bucket-(d) phrase additions (7 total: 5 Ξ±-7 PR + 2 follow-up) β eligible for separate docs-only PR for oracle improvement.
- Ξ±-7 closure analysis preserved at
reports/research-runs/alpha-7-closure-analysis-20260602.md. - Ξ±-7 5-tier 10 cell JSONs preserved at
workspaces/hotpot_eval/reports/research-runs/qvt-ablation-cells/.
Memory entries added 2026-06-02:
feedback_alpha7_top_k_destroys_positive_contributionsfeedback_s4_citation_survives_context_reshapefeedback_12b_pure_capability_misclassified_as_plateauproject_alpha_7_closure_state
Theme: measures whether typed filtering (only surface entities
whose type matches query intent) beats Ξ±-7's type-agnostic K
filter on multihop_rag. A/B isolation: Ξ = C_rag-ontology β C_rag-graph_post-Ξ±-7.
Status: design memo landed
(docs/design/v0.4-alpha-8-ontology-typed-filter.md); implementation
gated on Ξ±-7 closure.
Scope (per design memo Β§2 + Β§3):
- Extend
core/ontology.py(10 KB β ~14-15 KB; under gate):- Abstract root
Entity+ENTITY_TYPESregistry - 5 new horizontal types:
event,date,location,quantity,project(additive; existingperson / org / concept / documentunchanged) - 6 new relations:
OCCURRED_AT,HAPPENED_ON,LOCATED_IN,INVOLVES,MEASURED_AS,WORKED_ON
- Abstract root
- Migration cost: 0 β additive only; old types and 931 wiki entities unchanged.
- New sector flag
JAMES_DISABLE_TYPED_FILTER(disable-polarity, default OFF = sector ON = production byte-identical) - New matrix cell
C_rag-ontology(betweenC_rag-graphandC_rag-full) - Query intent β expected-type classifier (keyword heuristic; LLM classifier deferred to v0.5+)
- 5-axis QDC + per-question-type cross-tab against post-Ξ±-7 C_rag-graph
Acceptance band (per design memo Β§4):
- ββ adopt β
graded Ξ β₯ +0.030 - tier-gated β
+0.010 β€ graded Ξ < +0.030(enable per question_type only) - reject β
graded Ξ < +0.010(heuristic top-K is sufficient; ontology layer stays in codebase but disabled by default)
Rule #1 boundary test (per design memo Β§2.3): each proposed type
checked horizontally. The 5 accepted types pass; regulation,
transaction, recipe flagged for deferral or rejection.
Forward compat hook (per design memo Β§3.4): since: field anchor
allows v0.5 domain packs to register pack-specific types via the same
mechanism without absorbing them into the mother schema.
Honest framing (per memory feedback_finding_size_honest_framing):
- ββ partial if Ξ exceeds noise β typed filtering is a known technique; JAMES-specific value is the empirical position, not the mechanism
- β operational if Ξ β€ noise β confirms heuristic top-K is sufficient; finding logged either way
- β Don't claim "JAMES discovered typed semantics matter" β decades- old territory
Cycle dependencies:
- Predecessor: Ξ±-7 closure (baseline = post-Ξ±-7 C_rag-graph)
- Successor: cross-fixture sanity (deferred to v0.5 entry)
- Parallel: T3 Evidence Aging (orthogonal)
Graph-touching cycles serial to avoid measurement baseline drift:
Ξ±-6 closure β Ξ±-7 β Ξ±-8. Confidence/time-axis cycles parallel to
graph track: T3 Evidence Aging is orthogonal to entity surfacing/typing
and can land in any cycle alongside the graph track.
The v0.5 first-domain candidate is enterprise internal knowledge ontology (horizontal, audit/ownership/correction moat) per 2026-06-01 strategy handover Β§6. This framing applies to v0.5 domain selection criteria, not to v0.4 cycle scope. Mother-platform rule #1 holds β v0.4.x cycles add only horizontal infrastructure (top-K, abstract types, forward-compat hooks), no domain-specific types.
v0.5 entry declared 2026-06-12 after v0.4.4 closure (LRB v0.2.3
S3 publication-scale + cycle Ξ³ 4-bench infrastructure closure; Zenodo
DOI 10.5281/zenodo.20652679).
Operator decision at entry: "μ΄μ λ μ§μ§ μν°νλΌμ΄μ¦ μ¨ν¨λ‘μ§ μ₯μ°©μΌλ‘ κ°λ€" (entry handover).
v0.5 closed 2026-06-12 PM with 21 PRs (#841 β #861) per the v0.5 close handover. 23 additional PRs (#863 β #886) landed between cycle close and 2026-06-13 implementing most of the close handover Β§5 work queue β see the v0.6 entry skeleton for the post-close work-queue status sweep.
| Stream | Scope | Status at 2026-06-13 |
|---|---|---|
| A β Pre-LOI materials + Hashevolution dogfooding | Customer-facing μλ£ v0.4.4 sync (7 docs) + Hashevolution own-company scenario (μλλ¦¬μ€ C: internal dogfooding + external outreach λ³ν) + G1.a/G1.b/G1.c tenant-id contract + G2.a/G2.b/G2.c approval-evidence contract | A.1 / A.2 β ; A.3 / A.4 operator-pending; G1+G2 SaaS-readiness trio β (#860 / #861 / #869 / #870 / #882 / #883) |
| B β Enterprise ontology framework (mother-level) | B.1 audit + B.2 multi-tenant + B.3 plugin API + B.5 enterprise document ontology + G8.a-c mount + SDK.a-c trio | B.5.a-d β ; B.1 audit + G3/G4/G5/G7 β ; B.2/B.3 design memos β ; G8.a-c β ; SDK.a-c trio β (#875 / #876 / #881); G8.d LOI-blocked |
| C β Measurement infra carry-over | v0.2.3b S3 cross-model + D-alce + D-2wiki + HR full sweep + arXiv submission + graph-RAG synthesis | Graph-RAG Step 1 β (#864 / #877, +0.41 path_coverage βββ n=3); Step 2 driver scaffold β (#885); all others operator-attended |
| D β LOI-gated (blocked) | Customer-specific NDA/DPA/MSA + ingestion + pilot kickoff + vertical pack (legal) + G8.d capability grant workflow + F.2 CR.e customer theming | BLOCKED until LOI signed (Fork A of v0.6 entry contract) |
- Track F.1 β Time-Travel Dashboard quartet (#865 / #878 / #879 / #880): TT.a timestamp picker + TT.b corpus-state diff renderer (audit replay overlay) + TT.c reasoning trail replay at time T + TT.d now-vs-T diff view modal. Surfaces v0.4.2 T5
reconstruct_graph_at+ v0.5 G3reconstruct_corpus_view_at+replay_auditprimitives as a single operator dashboard. - Track F.2 β Change Review Workspace quartet (#866 / #867 / #873 / #874): CR.a list page + CR.b detail modal with diff renderer + CR.c contradiction-arbiter visualisation + CR.d approve/reject buttons with G2.a evidence capture wire.
- Track C β CSP nonce middleware (#884):
core/security/csp_nonce.pyper-request nonce primitive +build_security_headerscomposition seam +request.state.csp_noncemiddleware wire-in.script-srcflag safe today;style-srcflag reserved for UI #6 inline-style migration.
| Rule | Status at 2026-06-13 |
|---|---|
| CLAUDE.md rule #1 (no domain features until v1.0) | PRESERVED across all 44 cycle + post-close PRs β 4-layer protection contract (code-level capability gate + doc-level "Out of scope" + naming-level domain-agnostic + trigger-level LOI tagging) held throughout |
| #2 (bench numbers + Quality Delta Card on core/ PRs) | Preserved (all 44 PRs touch UI / lifecycle primitives / security / SDK packaging / docs β none touch core/retrieval / core/graph traversal / core/reasoning) |
| #3 (self-evolution opt-in) | Preserved |
#4 (architecture changes require architecture label) |
Preserved |
#5 (core/ 20 KB module size) |
Preserved for NEW files; five legacy modules grandfathered (largest core/reasoning/reflect.py at 29.2 KB β split planned, tracked in v0.6 entry skeleton Β§5 NEW solo-doable items) |
| NEW #6 Dogfooding evidence β Dim F evidence | Hashevolution own-use = iteration feedback; Dim F still requires external customer pilot |
| NEW #7 Enterprise ontology = mother-level | Vertical-pack (legal etc.) waits for LOI or v1.0; B.5 horizontal subtypes β landed |
| NEW #8 External-facing claims gated on measurement evidence | Zenodo DOI + result.json refs mandatory; Step 1 graph-RAG finding cites n=3 paired evidence per feedback_n1_verdict_inflation_n3_caught |
- v0.5 entry handover doc (
docs/handovers/v0.5-entry-2026-06-12.md) - Pre-LOI material refresh (7 docs, PR #840)
- Hashevolution own-company scenario doc (PR #840)
- B.5.a-d enterprise document ontology (4 PRs, #841 / #842 / #843 / #844)
- B.1 audit + 4 gap closures (G3 / G4 / G5 / G7) β 5 PRs (#845 β #849)
- B.2 / B.3 design memos (G1+G2 multi-tenant contract / G8 plugin-API stability β 2 PRs, #850 / #851)
- UI improvement stream (6 PRs, #852 β #858)
- External evaluation disclosure (PR #857)
- Server-side hardening (3 PRs, #859 β #861 β security headers + G1.a tenant-id + G2.a approval-evidence)
- v0.5 close handover (PR #862, 2026-06-12 PM)
- Track A G1.b replay-side tenant filter + G2.b CR merge wire-in + G1.c deployment guide + G2.c OIDC + asyncio variant (4 PRs, #869 / #870 / #882 / #883)
- Track B G8.a-c ontology pack mount + SDK.a-c trio (6 PRs, #868 / #871 / #872 / #875 / #876 / #881)
- Track C CSP nonce middleware (PR #884)
- Track F.1 Time-Travel Dashboard quartet (4 PRs, #865 / #878 / #879 / #880)
- Track F.2 Change Review Workspace quartet (4 PRs, #866 / #867 / #873 / #874)
- Graph-RAG synthesis Step 1 + Step 2 scaffold (3 PRs, #864 / #877 / #885)
- v0.6 entry preparation skeleton (PR #886) β bridges v0.5 close β v0.6 entry under 2-fork contract
- Stream A.3 operator: Hashevolution dogfooding start (operator action)
- Stream A.4 operator: Tier S contact list + outreach (operator action)
- Graph-RAG synthesis Step 2 cross-model measurement (~14 h wall, operator-launchable via
scripts/research/graph_rag_synth_step2_cross_model.py) - v0.2.3b LLM-grounded S3 cross-model run (operator-attended)
- D-alce real NLI verifier + paper v1.4 (operator-attended)
- HR full sweep n=100 (operator-attended)
- arXiv preprint submission (Path A defer; pending operator outreach outcome)
Dim F gate (β₯6 month external customer pilot + measured success
metrics per docs/PLATFORM_READINESS.md Β§3) is NOT cleared.
The 2-fork v0.6 entry contract from the
v0.6 entry skeleton Β§4 governs:
- Fork A β LOI signed β Track D vertical pack scoping begins. G8.d capability grant workflow + F.2 CR.e customer theming unblock.
- Fork B β 6-month no-LOI β reassess strategy. v0.5 mother-platform work stays the baseline; v0.6 cycle re-scopes around the next strategic direction.
Until one resolves, mother-platform hardening continues; vertical content stays BLOCKED per CLAUDE.md rule #1.
Theme: prove the platform contract by running ONE real domain in production for 6 months with one external customer. Moved here from v0.4 so Layer 4 governance lands first β see v0.4 retarget rationale above.
Status note (2026-06-13): this section describes Fork A of the v0.6 entry contract (LOI signed β vertical pack scoping begins). v0.5 cycle close (above) shipped the mother-platform infrastructure that this fork builds on β G8.a-c ontology pack mount mechanism, G1/G2 SaaS-readiness primitives, SDK.a-c trio for third-party pack authors, Time-Travel Dashboard surface for audit operators. The Dim F gate clock starts only when a customer LOI is signed; until then, v0.6 cycle entry waits on the operator's decision per the v0.6 entry skeleton Β§4.
Required for: a second domain pack (forbidden until this gate passes).
- One domain pack chosen from informal candidates
(likely
packs/legal/orpacks/food/) β selection criteria: signed PoC interest + clear legal liability boundary - Customer onboarding playbook (
docs/CUSTOMER_ONBOARDING.md) - External red-team pass on prompt injection (replaces pattern-only defense with ML guard + patterns)
- Public eval results in
eval/RESULTS.md(mother + first pack) - 6-month no-core-regression production track record
- One paying or formal-PoC customer has run the deployment for 6 months with no core code change attributable to their domain needs (only pack-level changes).
- A second domain pack
- Vertical Product packaging
- Public marketplace
Theme: make domain branching safe for outsiders. After this gate, external developers can publish their own packs.
- HTTPS / SSO / SAML / LDAP β production defaults
- Multi-tenancy (per-tenant data isolation, per-tenant pack selection, quota management)
- SOC 2 or ISO 27001 readiness assessment
- Backup / restore / rollback CLI tested under simulated failure
- Prometheus + OpenTelemetry exporters
- Public SDK and plugin author guide finalized
- Bus factor β₯ 2 (one non-maintainer with full commit/review history)
- Annual external red-team schedule established
- A third party (not a customer) builds and publishes a pack against the v1.0 SDK without contacting maintainers.
- The platform survives a single-maintainer 30-day absence with no customer-visible regressions.
- Vertical Products (separate business decision per domain after v1.0)
- Federation across multiple JAMES instances (Beyond v1.0 section)
Entry trigger (none of which is currently met): a real deployed corpus whose float32 vector index exceeds available RAM β concretely > ~1M documents at the default 384-dim, or a pilot that measures a vector-store memory bottleneck. Until a trigger fires this stays a note, not work: at 384-dim even 1M docs β 1.5 GB fits the reference 32 GB machine, so adopting compression now would be optimization against an unproven premise (cf. the cloud-tier S6/S7 shelving lesson).
Priority ladder when the trigger fires β ranked by
(memory saved Γ recall kept) Γ· replay/audit cost, not memory alone,
because byte-identical replay (RAB / reconstruct_graph_at(t)) is the
crown-jewel constraint:
- 1. Int8 scalar quantization (~4Γ) β replay-safe: a fixed, audit-logged scale makes it deterministic; backend-agnostic (quantize at the embedding level before ChromaDB). First and lowest-risk. Entry PR must paste a 5-axis Quality Delta Card (rule #2) β the "Recall within 1%" claim is re-measured on JAMES's own QVT/LRB queries, not imported from the literature.
- 2. Dimensionality reduction β requires a model swap (the
default
paraphrase-multilingual-MiniLM-L12-v2is multilingual and not Matryoshka, so head-truncation is unsafe). Forces full re-ingestion + KO/EN bilingual re-validation (bilingual regression has bitten before β seefeedback_d2_v2_softener_bilingual_regression). Separate measurement cycle. - 3. FAISS IVFPQ (~10β100Γ) β strongest memory win but the
trained codebook is a replay liability: historical replays must
pin their era's codebook. Also a backend change (ChromaDB β
FAISS = trust-boundary /
docs/ARCHITECTURE.mdPR, rule #4). Architecture decision precedes any build. - 4. Binary + rescoring β only at very large scale; keeps float32 for rescore (so the "32Γ" is index-only), and rescore ordering must be made deterministic for audit reproducibility.
Rule #1 is not implicated (horizontal infra, not a domain feature) β this is a priority/trigger question, not a permission one.
After v1.0, growth is by domain accumulation, not core feature
addition. See docs/PLATFORM_READINESS.md Β§4 for the three branching
forms (Pack / Distribution / Vertical Product) and selection criteria.
Long-considered, no commitment:
- Optional Neo4j backend β migrate from markdown wiki to graph DB, Cypher query support, backward compatibility with markdown (was tentatively v0.3; reframed as post-v1.0 optimization)
- Multi-agent system β specialist agents (researcher, coder, security), agent-to-agent communication, task decomposition (was tentatively v0.3; reframed as post-v1.0 capability)
- OpenAI-compatible API for drop-in replacement
- Streaming responses + Webhook support
- Federation: connect multiple JAMES instances
- On-device fine-tuning: LoRA adapters per user
- Edge deployment: smaller models for embedded use
- Plugin marketplace: community-contributed packs
- Visual graph editor: web UI for ontology editing
- Voice interface: ASR + TTS pipeline
- GitHub Issues: feature requests, prioritized by upvotes
- Discussions: longer-form proposals
- Pull Requests: implement what you need
We prioritize based on:
- Security-critical fixes (immediate)
- Real-data feedback from users
- Strategic alignment with the project's direction
- Community contribution (volunteer-friendly tasks first)
Domain-specific feature requests during v0.2 β v0.4 are out of scope.
See docs/handovers/v0.2.1-business-track.md Β§3 for the rationale
and the "no parallel domains" rule.
We follow Semantic Versioning:
MAJOR.MINOR.PATCH-PRERELEASE0.x.yversions may contain breaking changes1.0.0and beyond will follow strict semver- After v1.0 ships, plugin API gets its own SemVer track with a
12-month deprecation guarantee (see
docs/PLATFORM_READINESS.mdGate v0.3 criteria)
Last updated: 2026-05-22 β v0.4 retarget to Layer 4
Lifecycle Semantics + CASCADE/EVENT λΆλ¦¬ axiom μ±ν. v0.3 cycle
μ Knowledge Cascade (Layer 3 Memory OS) κ° μμ νλλ©°, 2026-05-21
μ¬μ©μ λΉν μ리μ¦μμ "Layer 3 alone" μ governance gap (stale facts
/ Gate 5 silent reject / manual hierarchy μμ / event sourcing λΆμ¬)
λλ¬λ¨. νμ architectural insight λ‘ CASCADE (invalidation
propagation) μ EVENT/TEMPORAL (semantic evolution) μ λͺ
ν λΆλ¦¬
axiom μ±ν β λ μ’
λ₯ mutation μ κ°μ λ©μ»€λμ¦μΌλ‘ μ²λ¦¬νλ©΄
contradiction νλ°. 7 μμ (T1βT7, T7 supersede chain μ κ·) λμμΈ
λ©λͺ¨ (docs/architecture/memory-lifecycle-architecture.md Β§1.5 +
docs/design/v0.4-lifecycle-semantics-roadmap.md Β§9.5) μμ±. v0.4
첫 ship bundle = T1+T7+T2 (EVENT ν΅μ¬ + A/B routing). κΈ°μ‘΄ v0.4
(First Domain Pilot) β v0.5 λ‘ shift. 2026-05-22 i18n sweep μ리μ¦
(7 PR, ~304 entries, 7 backend λͺ¨λ) κ° label_key ν¨ν΄μ platform
invariant λ‘ lock-in β v0.4 μ κ· λΌλ²¨λ λμΌ contract νμ. Open
issues: 0.
Prior update (2026-05-13): v0.3 μ§μ
μ μ. Axis 6 λ λ²μ§Έ
μ¬μ©μ κ²μ΄νΈλ λͺ¨μ§ μλ£ β 6 axes λͺ¨λ ν΅κ³Ό β v0.2 β v0.3 gate clear.
v0.2.x κ° λν Change Request κΈ°λ° (core/change_request.py +
wiki_entity + run_jobs) + dependabot 6 high-severity λ§κ° +
μ κ³΅λ² 3 PR (#252/253/254: μΉ νμ΅ LLM-triple μμ / UNRESOLVED sweep /
λ
Έμ΄μ¦ cleanup μ€ν¬λ¦½νΈ). v0.3 Platform Skeleton νΈλ νμ± β License /
CLA / Plugin API / Knowledge Cascade / CR-E κ° λ³Έ μ¬μ΄ν΄λΆν° deliverable.