Commit Graph

102 Commits

Author SHA1 Message Date
Kral
8ef85c2713 Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 09:04:09 +02:00
Kral
a465c33e3f D: own-test mutation scores of the accepted trajectories (metadata only), stage 2 set rebuilt with the score
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 07:36:39 +02:00
Kral
4002d889d5 own-test report and chain script; work list 2026-10-06 06:16:29 +02:00
Kral
c5f1a36df0 stage 2 builder: scrub of other runs' leftover objects, drop of foreign reads; rebuild (36 train + 4 valid), doc generator, HF dataset updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:16:02 +02:00
Kral
c2c3803b1b G: EPOD acceptance tests; harness: hide and clean other runs' mid-name objects (proxy, teardown, sweep); lock leak cause (Eclipse restart during a write)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:15:08 +02:00
Kral
0c0e07fe96 F: eval slots for INTF/TABL/STRU/MSAG/exception (+K), Kral spot-check sheet, step 0 in the restart plan; D: own-test mutation scoring (running)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:06:41 +02:00
Kral
7f03849a86 C+E: stage 2 builder (mask, 48k, CLAS cap, family split, hook), private HF dataset, bf16 mixed training script with memory test, memory table, ratio proposal
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 06:01:49 +02:00
Kral
f5611cbf56 data analysis doc: corrected two counts 2026-10-06 05:56:31 +02:00
Kral
542fc8ec31 B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 05:56:23 +02:00
Kral
1ddb9a7c65 Series A result: local Qwen 3 of 20 accepted, failures are loops and search loops; dashboard public copy
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 05:45:24 +02:00
Kral
01f3372e2c STATE: 21:43 outage was an Eclipse restart (confirmed)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 22:29:44 +02:00
Kral
9c753a37ff Infrastructure outage handling (MCP/A4H down: wait, clean up, rerun), dashboard fix, budget limit 129
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 22:27:48 +02:00
Kral
5c07026bbb serve_remote.sh: pick ~/qwen-venv, refuse an old mlx-lm with a clear message; docs for mlx-lm 0.32.0 on the MacBook
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 22:16:56 +02:00
Kral
6984fc7917 Snapshot: new training tasks (balanced generation), docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 22:08:59 +02:00
Kral
eb13453097 Series A: local Qwen on the MacBook (remote server), 12 h window with hard stop, clean pause on server loss, dashboard card, docs/remote-model.md
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 21:57:01 +02:00
Kral
dc8d99a913 Dashboard: ignore an old STOPPED.txt while a controller runs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 21:32:23 +02:00
Kral
301c221d8e Budget guard from panel 50.00: limit 128
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 21:31:15 +02:00
Kral
59496cb08b Lock leak: controlled reproduction attempts and method (not reproduced), enqueue reader; restart plan for 12 October
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 19:47:16 +02:00
Kral
4432ac896c Only kinds below their target share are generated and run; one-controller flock guard
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 19:39:37 +02:00
Kral
976274d48a Balanced generation under the pipeline, per-kind brake; STATE: type mix fix
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 19:27:43 +02:00
Kral
55f4068330 STRU and exception tasks accepted; G6 cascade fix, exception class mutants, RTTI unit notes
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 19:25:09 +02:00
Kral
b119f1afac Object type mix: INTF, TABL, STRU, MSAG, exception tasks (harness G2, mutants, generator notes), balanced generator, kind-deficit job order, dashboard mix card
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 18:52:29 +02:00
Kral
454db6626c Review notes: token length for bf16 test (p95 48k), DDLS reject reasons, G1034 lock analysis (not reproduced)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 18:21:44 +02:00
Kral
224ba1a5f8 Summary: correct the guard line; pipeline summary reads the limit from .env
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 18:05:49 +02:00
Kral
5bcd1bf9cd Stage 2 summary at 50 accepted trajectories
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 18:02:52 +02:00
Kral
332fb4601d Live status page (SAP colors), rewritten every 5 minutes
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 17:56:30 +02:00
Kral
bc6827b744 Budget guard from panel 41.07: ratio 1.2, limit 121
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 17:47:39 +02:00
Kral
704cfa3ede Budget guard from the panel: ratio 1.35 for trajectory runs, limit 123
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 15:54:10 +02:00
Kral
ae37617a0d Budget guard: ratio 2.07 measured, limit 142; limit read from .env at each check
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 14:47:54 +02:00
Kral
c062df2ec5 Trajectory runs: tool-call budget 100 for tasks with a CDS contract object (eval keeps 60)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 14:08:33 +02:00
Kral
13360292dd Pipeline controller: plan 2 (hard +, error +50 %), backlog throttle, 2 trajectory workers, stop rules, summaries every 50; budget reserve
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 13:35:54 +02:00
Kral
1c4667a0fc MCP load test (read, heavy read, write modes)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 13:26:13 +02:00
Kral
9de3911578 K variants for training (free text, EPOD tool names), second attempt only for failed tasks, docs/epod-syntax-hint.md
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:50:43 +02:00
Kral
35f0eb0e0c Overlap limit for whole spec 0.75, slot logs; STATE and roadmap: stage 2 data progress
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:04:11 +02:00
Kral
61b903f7b3 Acceptance filter: score, end reason, harness errors, loop trimming, repair marker
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:03:54 +02:00
Kral
c2d4997566 Converter to the Qwen 3.8 chat template (thinking off) with tokenizer round-trip check
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:03:54 +02:00
Kral
d5e43e1a83 Trajectory record (messages, raw tool results, metadata; reasoning apart) and trajectory runner
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:03:54 +02:00
Kral
a4eb567e4c Proxy: local abaplint messages on a bare 'save failed' (EPOD syntax check cannot see the rejected source)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:03:54 +02:00
Kral
229f862600 Generator training mode: train pool, eval overlap check, category mix, error-targeted tasks
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 11:53:07 +02:00
Kral
3a9f9e75fc Step 400 check: Mac -13.4 % vs GPU -35.0 %; next GPU run bf16 mixed with stage 2; findings in STATE and roadmap
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 11:46:21 +02:00
Kral
39dcdd9846 Step 200 check: GPU -26.0 % vs Mac -10.7 %, full run cancelled at step ~418
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 20:48:49 +02:00
Kral
0c0982f666 Overfit conversion test passed (Mac -98.5 %, GPU -99.97 %); full run started, alpha 32
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 19:34:24 +02:00
Kral
20ef7b2331 Stage 1 pipeline test: 30.5 s/step, converter checked, valid loss 0.848 with test adapter
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 19:24:25 +02:00
Kral
5447874fd3 Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 18:51:40 +02:00
Kral
a2ba9e7b44 serve.sh back to Qwen 3.8 27B 4-bit
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:02:59 +02:00
Kral
9f85a71ca5 Roadmap: Qwen baseline result
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:02:02 +02:00
Kral
a368233d26 Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:01:52 +02:00
Kral
dbe6077a48 Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 13:11:53 +02:00
Kral
5532a5b226 baseline: setup failure is not recorded as a result; repair_stats tolerates a missing trajectory
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 12:07:33 +02:00
Kral
d15e190f72 baseline.py: --model, --base-url, --enable-thinking-false (Qwen run on MacBook)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 11:39:33 +02:00