ATOCore

Author	SHA1	Message	Date
Anto01	7650c339a2	audit: verify batch2 claims and findings	2026-04-12 19:06:51 +00:00
Anto01	abc8af5f7e	audit: record extraction pipeline findings	2026-04-12 16:20:42 +00:00
Anto01	b790e7eb30	audit: record final 2026-04-12 findings	2026-04-12 13:03:10 +00:00
Anto01	2b79680167	chore(ledger): Wave 2 ingestion + codex audit response session log Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-12 07:57:32 -04:00
Anto01	39d73e91b4	fix(R6): fall back to interaction.project when LLM returns empty Codex R6: the LLM extractor accepted the model's project field verbatim. When the model returned empty string, clearly p06 memories got promoted as project='', making them invisible to the p06 project-memory band and explaining the p06-offline-design harness failure. Fix: if model returns empty project but interaction.project is set, inherit the interaction's project. Model-supplied project still takes precedence when non-empty. Two new tests lock the fallback and precedence behaviors. R5 acknowledged (LLM extractor not yet wired into API — next task). Test count: 278 -> 280. Harness re-run pending after deploy. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-12 07:37:14 -04:00
Anto01	89c7964237	audit: record 2026-04-12 review findings	2026-04-12 11:31:32 +00:00
Anto01	146f2e4a5e	chore: Day 8 — close mini-phase with before/after metrics Mini-phase complete. Before/after deltas: Metric Before After ───────────────────────────────────────── Rule extractor recall 0% 0% (unchanged, deprioritized) LLM extractor recall n/a 100% (new, claude -p haiku) LLM candidate yield n/a 2.55/interaction First triage accept rate n/a 31% (16/51) Active memories 20 36 (+16) p06-polisher memories 2 16 (+14) atocore memories 0 5 (+5) Retrieval harness 6/6 15/18 (expanded to 18 fixtures) Test count 264 278 (+14) 3 remaining harness failures are budget-contention on the p06 memory band: the specific memory a fixture targets ranks 4th+ and the 25% budget only holds 2-3 entries. Not a ranking bug — the per-entry 250-char cap was the one justified tweak; a second budget change risks regressing other fixtures per Codex's Day 7 hard gate. Ledger updated: Orientation, Session Log, main_tip, harness line. Next on the roadmap (from DEV-LEDGER Active Plan / docs/next-steps): - Wave 2 trusted operational ingestion (p04/p05/p06 dashboards) - Finish OpenClaw integration (Phase 8) - Auto-triage (multi-model second pass to reduce human review) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-12 06:41:42 -04:00
Anto01	b98a658831	chore(ledger): Day 4 complete + first triage done Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-12 06:06:38 -04:00
Anto01	330ecfb6a6	chore(ledger): Day 2 baseline escalated to Day 4 gate early Day 2 extractor eval baseline on a 20-interaction labeled set shows 0% yield / 0% recall / 0% precision. The 5 false negatives span 5 distinct miss classes, matching the pattern Codex's Day 4 hard gate was designed to catch but arriving two days early. No extractor code change on main. Day 1+2 artifacts committed on working branch 'claude/extractor-eval-loop' at `7d8d599`. Day 4 decision (keep rule-expanding vs prototype LLM-assisted mode) is escalated to Antoine for ratification before Day 3 work touches any extractor. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-11 15:12:58 -04:00
Anto01	d9dc55f841	docs: formalize DEV-LEDGER review protocol	2026-04-11 15:03:33 -04:00
Anto01	81307cec47	chore: ledger session log — wire protocol commit	2026-04-11 14:46:50 -04:00
Anto01	59331e522d	feat: DEV-LEDGER.md as shared operating memory + session protocol The ledger is the one-file source of truth for "what is currently true" across Claude/Codex/human sessions: - Orientation (live SHA, main tip, test count, harness state) - Active Plan (currently Codex's 8-day extractor + harness plan with hard gates and fail-early thresholds) - Open Review Findings (P1/P2, status) - Recent Decisions (bounded to last 20) - Session Log (bounded to last 20) - Working Rules (no parallel work, branching rule, P1 block) Narrative docs under docs/ sometimes lag reality; the ledger does not. Every session MUST read it at start and append a Session Log line before ending. AGENTS.md: added a new "Session protocol" section at the top that points at the ledger. Applies to any agent (Claude, Codex, future). CLAUDE.md (new, project-local): project instructions for Claude Code in this repo. Points at DEV-LEDGER.md and AGENTS.md, spells out the deploy workflow and the Claude/Codex working model. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>	2026-04-11 14:46:21 -04:00

12 Commits