Commit Graph

  • 4261054eac Restart plan step 0b: rerun of the four empirical filter runs (G0183 G0180 G0143 G0185) without touching the eval set; empirical --rerun main Kral 2026-10-06 09:07:10 +02:00
  • 8ef85c2713 Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list Kral 2026-10-06 09:04:09 +02:00
  • a465c33e3f D: own-test mutation scores of the accepted trajectories (metadata only), stage 2 set rebuilt with the score Kral 2026-10-06 07:36:39 +02:00
  • 4002d889d5 own-test report and chain script; work list Kral 2026-10-06 06:16:29 +02:00
  • c5f1a36df0 stage 2 builder: scrub of other runs' leftover objects, drop of foreign reads; rebuild (36 train + 4 valid), doc generator, HF dataset updated Kral 2026-10-06 06:16:02 +02:00
  • c2c3803b1b G: EPOD acceptance tests; harness: hide and clean other runs' mid-name objects (proxy, teardown, sweep); lock leak cause (Eclipse restart during a write) Kral 2026-10-06 06:15:08 +02:00
  • 0c0e07fe96 F: eval slots for INTF/TABL/STRU/MSAG/exception (+K), Kral spot-check sheet, step 0 in the restart plan; D: own-test mutation scoring (running) Kral 2026-10-06 06:06:41 +02:00
  • 7f03849a86 C+E: stage 2 builder (mask, 48k, CLAS cap, family split, hook), private HF dataset, bf16 mixed training script with memory test, memory table, ratio proposal Kral 2026-10-06 06:01:49 +02:00
  • f5611cbf56 data analysis doc: corrected two counts Kral 2026-10-06 05:56:31 +02:00
  • 542fc8ec31 B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard) Kral 2026-10-06 05:56:23 +02:00
  • 1ddb9a7c65 Series A result: local Qwen 3 of 20 accepted, failures are loops and search loops; dashboard public copy Kral 2026-10-06 05:45:24 +02:00
  • 01f3372e2c STATE: 21:43 outage was an Eclipse restart (confirmed) Kral 2026-10-05 22:29:44 +02:00
  • 9c753a37ff Infrastructure outage handling (MCP/A4H down: wait, clean up, rerun), dashboard fix, budget limit 129 Kral 2026-10-05 22:27:48 +02:00
  • 5c07026bbb serve_remote.sh: pick ~/qwen-venv, refuse an old mlx-lm with a clear message; docs for mlx-lm 0.32.0 on the MacBook Kral 2026-10-05 22:16:56 +02:00
  • 6984fc7917 Snapshot: new training tasks (balanced generation), docs Kral 2026-10-05 22:08:59 +02:00
  • eb13453097 Series A: local Qwen on the MacBook (remote server), 12 h window with hard stop, clean pause on server loss, dashboard card, docs/remote-model.md Kral 2026-10-05 21:57:01 +02:00
  • dc8d99a913 Dashboard: ignore an old STOPPED.txt while a controller runs Kral 2026-10-05 21:32:23 +02:00
  • 301c221d8e Budget guard from panel 50.00: limit 128 Kral 2026-10-05 21:31:15 +02:00
  • 59496cb08b Lock leak: controlled reproduction attempts and method (not reproduced), enqueue reader; restart plan for 12 October Kral 2026-10-05 19:47:16 +02:00
  • 4432ac896c Only kinds below their target share are generated and run; one-controller flock guard Kral 2026-10-05 19:39:37 +02:00
  • 976274d48a Balanced generation under the pipeline, per-kind brake; STATE: type mix fix Kral 2026-10-05 19:27:43 +02:00
  • 55f4068330 STRU and exception tasks accepted; G6 cascade fix, exception class mutants, RTTI unit notes Kral 2026-10-05 19:25:09 +02:00
  • b119f1afac Object type mix: INTF, TABL, STRU, MSAG, exception tasks (harness G2, mutants, generator notes), balanced generator, kind-deficit job order, dashboard mix card Kral 2026-10-05 18:52:29 +02:00
  • 454db6626c Review notes: token length for bf16 test (p95 48k), DDLS reject reasons, G1034 lock analysis (not reproduced) Kral 2026-10-05 18:21:44 +02:00
  • 224ba1a5f8 Summary: correct the guard line; pipeline summary reads the limit from .env Kral 2026-10-05 18:05:49 +02:00
  • 5bcd1bf9cd Stage 2 summary at 50 accepted trajectories Kral 2026-10-05 18:02:52 +02:00
  • 332fb4601d Live status page (SAP colors), rewritten every 5 minutes Kral 2026-10-05 17:56:30 +02:00
  • bc6827b744 Budget guard from panel 41.07: ratio 1.2, limit 121 Kral 2026-10-05 17:47:39 +02:00
  • 704cfa3ede Budget guard from the panel: ratio 1.35 for trajectory runs, limit 123 Kral 2026-10-05 15:54:10 +02:00
  • ae37617a0d Budget guard: ratio 2.07 measured, limit 142; limit read from .env at each check Kral 2026-10-05 14:47:54 +02:00
  • c062df2ec5 Trajectory runs: tool-call budget 100 for tasks with a CDS contract object (eval keeps 60) Kral 2026-10-05 14:08:33 +02:00
  • 13360292dd Pipeline controller: plan 2 (hard +, error +50 %), backlog throttle, 2 trajectory workers, stop rules, summaries every 50; budget reserve Kral 2026-10-05 13:35:54 +02:00
  • 1c4667a0fc MCP load test (read, heavy read, write modes) Kral 2026-10-05 13:26:13 +02:00
  • 9de3911578 K variants for training (free text, EPOD tool names), second attempt only for failed tasks, docs/epod-syntax-hint.md Kral 2026-10-05 12:50:43 +02:00
  • 35f0eb0e0c Overlap limit for whole spec 0.75, slot logs; STATE and roadmap: stage 2 data progress Kral 2026-10-05 12:04:11 +02:00
  • 61b903f7b3 Acceptance filter: score, end reason, harness errors, loop trimming, repair marker Kral 2026-10-05 12:03:54 +02:00
  • c2d4997566 Converter to the Qwen 3.8 chat template (thinking off) with tokenizer round-trip check Kral 2026-10-05 12:03:54 +02:00
  • d5e43e1a83 Trajectory record (messages, raw tool results, metadata; reasoning apart) and trajectory runner Kral 2026-10-05 12:03:54 +02:00
  • a4eb567e4c Proxy: local abaplint messages on a bare 'save failed' (EPOD syntax check cannot see the rejected source) Kral 2026-10-05 12:03:54 +02:00
  • 229f862600 Generator training mode: train pool, eval overlap check, category mix, error-targeted tasks Kral 2026-10-05 11:53:07 +02:00
  • 3a9f9e75fc Step 400 check: Mac -13.4 % vs GPU -35.0 %; next GPU run bf16 mixed with stage 2; findings in STATE and roadmap Kral 2026-10-05 11:46:21 +02:00
  • 39dcdd9846 Step 200 check: GPU -26.0 % vs Mac -10.7 %, full run cancelled at step ~418 Kral 2026-10-04 20:48:49 +02:00
  • 0c0982f666 Overfit conversion test passed (Mac -98.5 %, GPU -99.97 %); full run started, alpha 32 Kral 2026-10-04 19:34:24 +02:00
  • 20ef7b2331 Stage 1 pipeline test: 30.5 s/step, converter checked, valid loss 0.848 with test adapter Kral 2026-10-04 19:24:25 +02:00
  • 5447874fd3 Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs Kral 2026-10-04 18:51:40 +02:00
  • a2ba9e7b44 serve.sh back to Qwen 3.8 27B 4-bit Kral 2026-10-04 17:02:59 +02:00
  • 9f85a71ca5 Roadmap: Qwen baseline result Kral 2026-10-04 17:02:02 +02:00
  • a368233d26 Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral Kral 2026-10-04 17:01:52 +02:00
  • dbe6077a48 Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further) Kral 2026-10-04 13:11:53 +02:00
  • 5532a5b226 baseline: setup failure is not recorded as a result; repair_stats tolerates a missing trajectory Kral 2026-10-04 12:07:33 +02:00
  • d15e190f72 baseline.py: --model, --base-url, --enable-thinking-false (Qwen run on MacBook) Kral 2026-10-04 11:39:33 +02:00
  • 0af2d6458e Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped Kral 2026-10-04 11:20:43 +02:00
  • 8d1db9c67c Devstral Small 2: serve.sh, baseline.py without thinking args, repair rate, activation error message fix, stop rule in chain (run base 21000) Kral 2026-10-04 08:57:12 +02:00
  • e84ea43a3d Base model decision: drop Qwen 3.8, candidate Devstral Small 2 (24B) Kral 2026-10-04 08:44:04 +02:00
  • ba0a7e6b3b STATE: correct start time Kral 2026-10-04 07:59:41 +02:00
  • 342eb6fc9a STATE: baseline restart time Kral 2026-10-04 07:59:38 +02:00
  • 5c68a1b1b7 Loop guard: also identical call with identical result 3 times in a row; baseline restarts (run base 20500) Kral 2026-10-04 07:59:28 +02:00
  • a94c2a3b5d Baseline restarted with loop guard, thinking off (run base 20400) Kral 2026-10-04 07:40:19 +02:00
  • 0401186a6c Loop guard (3 identical pushes), end_reason and activation error records per run; thinking off for the stage 1 baseline Kral 2026-10-04 07:07:44 +02:00
  • 040b9900fd STATE/handover: baseline restarted on the 11-task subset Kral 2026-10-03 22:36:30 +02:00
  • 689820af4d Baseline shortened: 11-task subset (1 per category + T01), max_tokens 16384, settings in README; baseline_chain.sh Kral 2026-10-03 22:36:00 +02:00
  • 7054d9fd2e STATE: correct timestamps Kral 2026-10-03 22:24:08 +02:00
  • 7b8ca01bde Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated Kral 2026-10-03 22:23:04 +02:00
  • 0677d035da Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations Kral 2026-10-03 22:09:15 +02:00
  • e083c9ca13 Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets) Kral 2026-10-03 21:58:04 +02:00
  • f0593933f9 Budget calendar: usage estimate, 20 USD paid for 60 USD usage Kral 2026-10-03 20:15:05 +02:00
  • eaaef125ff Budget guard 135 (ledger), about 50 USD real Kral 2026-10-03 20:12:38 +02:00
  • b0fd06259b Budget: real ratio ledger/2.7 (Ollama page 24.63 USD) Kral 2026-10-03 20:11:45 +02:00
  • aef4c6e461 Pending: A4H memory limit after the baseline Kral 2026-10-03 20:04:36 +02:00
  • 2f9037d756 Correct proxy finding: Z0FFK001 object belongs to the running T01 test Kral 2026-10-03 20:03:24 +02:00
  • 4409393302 Decision: no A4H move; memory limit after the baseline Kral 2026-10-03 19:46:15 +02:00
  • e3f5dcd668 G0181 failure analysis (int8 in SALV), W7B results Kral 2026-10-03 19:16:56 +02:00
  • d8cbf4de25 Decision: no separate SAP user for the harness (2026-10-03) Kral 2026-10-03 19:00:50 +02:00
  • 573707613e Eval review results, easy candidates, rerun queue; handover and STATE updated Kral 2026-10-03 18:57:23 +02:00
  • fb18dc8d9d Review: 9 random high-score checks; easy candidates list Kral 2026-10-03 18:56:12 +02:00
  • 22022ca6a5 Review: 10 remaining accepted tasks (checklist) Kral 2026-10-03 18:55:33 +02:00
  • 530590f8b3 Review: 10 remaining accepted tasks (checklist) Kral 2026-10-03 18:55:25 +02:00
  • 5e3baf43d3 Empirical filter results: DeepSeek reruns W6A/W7A (partial) Kral 2026-10-03 18:54:16 +02:00
  • d83a39a8d8 MLX server: prompt cache limit (memory); restart after the T01 test Kral 2026-10-03 18:47:03 +02:00
  • 81d8cf6f99 Handover: no new model runs during the baseline; rerun queue Kral 2026-10-03 18:43:25 +02:00
  • 9366ee9868 Handover notes, train/STATE.md, detached night chain for DeepSeek reruns and the stage 1 baseline Kral 2026-10-03 18:34:07 +02:00
  • d4ed32cf38 G0181: restore ALV output in K spec; K prompt keeps output form Kral 2026-10-03 18:26:24 +02:00
  • 92e30eb554 Runner: G2 finds the FUNCTION statement after local classes; budget floor applied to all tasks Kral 2026-10-03 18:25:30 +02:00
  • 7351849a76 Stage 1 step 0: train/.venv (py3.11, mlx-lm 0.32), MLX 4-bit Qwen3.8-27B, serve script, baseline runner, README Kral 2026-10-03 18:23:39 +02:00
  • cf8d6d5d07 Ignore train/.venv Kral 2026-10-03 17:34:02 +02:00
  • 65322f27ff Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset Kral 2026-10-03 17:33:55 +02:00
  • 01e99158a9 Agent: max_tokens 32k for cloud models and retry on empty turns (runaway reasoning); proxy: sap_activate fallback for PROG/FUNC Kral 2026-10-03 17:07:31 +02:00
  • 0e96c6708e Review fixes G0019, G0108: added hidden tests; revalidated Kral 2026-10-03 16:35:20 +02:00
  • 7510131a7d ADT fallback: Accept headers for FUNC; G0189 accepted; regen2 results Kral 2026-10-03 16:31:03 +02:00
  • 428b3565b2 Proxy: ADT REST write+activate fallback when EPOD does not activate a second PROG/FUNC write Kral 2026-10-03 14:29:37 +02:00
  • 39769afce8 Step F first pass: 81/92 accepted; budget floor 60 calls; rejected slots moved for regeneration Kral 2026-10-03 12:10:52 +02:00
  • d3250acfb8 Empirical filter runner (DeepSeek); real cost ratio in docs Kral 2026-10-03 09:29:20 +02:00
  • 03ab36149e Fix coverage with several contract classes; table annotation and file path checks; clean task dir on write Kral 2026-10-03 08:44:35 +02:00
  • fb576b0954 In-place E/I: G2 needs a changed seed contract object, G4 skips it; regenerate G0009/G0013/G0016/G0023 as G0188-G0191 Kral 2026-10-03 06:33:33 +02:00
  • 3550f0b651 Eval review v1: checklist, review helper, 24 tasks reviewed; G0002/G0022 tests fixed; I/E generator definitions Kral 2026-10-03 06:23:38 +02:00
  • a12fa4dff6 H and K tested (G0168, G0178 accepted); keyword fast path needs all groups Kral 2026-10-03 06:04:36 +02:00
  • f907351c0d Mutation check done on all accepted tasks; mutant fixes; dump filter in proxy Kral 2026-10-03 06:02:59 +02:00
  • 0ba25faeea Mutation in generator; category H (stop) scoring and judge; category K variants and tool schema variant; evalset plan Kral 2026-10-03 05:37:20 +02:00
  • 32f03fcad4 Hard batch G0022-G0026: 5/5 accepted; budget floor, CDS checks, mutation check Kral 2026-10-03 05:16:30 +02:00
  • 3570d7ef9b Harness: G6 ignores unknown standard superclass; dependency order for seed/reference; max_tokens 80k Kral 2026-10-03 05:06:04 +02:00