Commit Graph

62 Commits

Author SHA1 Message Date
Kral
39dcdd9846 Step 200 check: GPU -26.0 % vs Mac -10.7 %, full run cancelled at step ~418
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 20:48:49 +02:00
Kral
0c0982f666 Overfit conversion test passed (Mac -98.5 %, GPU -99.97 %); full run started, alpha 32
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 19:34:24 +02:00
Kral
20ef7b2331 Stage 1 pipeline test: 30.5 s/step, converter checked, valid loss 0.848 with test adapter
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 19:24:25 +02:00
Kral
5447874fd3 Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 18:51:40 +02:00
Kral
a2ba9e7b44 serve.sh back to Qwen 3.8 27B 4-bit
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:02:59 +02:00
Kral
9f85a71ca5 Roadmap: Qwen baseline result
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:02:02 +02:00
Kral
a368233d26 Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:01:52 +02:00
Kral
dbe6077a48 Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 13:11:53 +02:00
Kral
5532a5b226 baseline: setup failure is not recorded as a result; repair_stats tolerates a missing trajectory
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 12:07:33 +02:00
Kral
d15e190f72 baseline.py: --model, --base-url, --enable-thinking-false (Qwen run on MacBook)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 11:39:33 +02:00
Kral
0af2d6458e Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 11:20:43 +02:00
Kral
8d1db9c67c Devstral Small 2: serve.sh, baseline.py without thinking args, repair rate, activation error message fix, stop rule in chain (run base 21000)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 08:57:12 +02:00
Kral
e84ea43a3d Base model decision: drop Qwen 3.8, candidate Devstral Small 2 (24B)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 08:44:04 +02:00
Kral
ba0a7e6b3b STATE: correct start time 2026-10-04 07:59:41 +02:00
Kral
342eb6fc9a STATE: baseline restart time 2026-10-04 07:59:38 +02:00
Kral
5c68a1b1b7 Loop guard: also identical call with identical result 3 times in a row; baseline restarts (run base 20500)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 07:59:28 +02:00
Kral
a94c2a3b5d Baseline restarted with loop guard, thinking off (run base 20400) 2026-10-04 07:40:19 +02:00
Kral
0401186a6c Loop guard (3 identical pushes), end_reason and activation error records per run; thinking off for the stage 1 baseline
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 07:07:44 +02:00
Kral
040b9900fd STATE/handover: baseline restarted on the 11-task subset 2026-10-03 22:36:30 +02:00
Kral
689820af4d Baseline shortened: 11-task subset (1 per category + T01), max_tokens 16384, settings in README; baseline_chain.sh
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:36:00 +02:00
Kral
7054d9fd2e STATE: correct timestamps 2026-10-03 22:24:08 +02:00
Kral
7b8ca01bde Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:23:04 +02:00
Kral
0677d035da Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:09:15 +02:00
Kral
e083c9ca13 Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 21:58:04 +02:00
Kral
f0593933f9 Budget calendar: usage estimate, 20 USD paid for 60 USD usage 2026-10-03 20:15:05 +02:00
Kral
eaaef125ff Budget guard 135 (ledger), about 50 USD real 2026-10-03 20:12:38 +02:00
Kral
b0fd06259b Budget: real ratio ledger/2.7 (Ollama page 24.63 USD) 2026-10-03 20:11:45 +02:00
Kral
aef4c6e461 Pending: A4H memory limit after the baseline 2026-10-03 20:04:36 +02:00
Kral
2f9037d756 Correct proxy finding: Z0FFK001 object belongs to the running T01 test 2026-10-03 20:03:24 +02:00
Kral
4409393302 Decision: no A4H move; memory limit after the baseline 2026-10-03 19:46:15 +02:00
Kral
e3f5dcd668 G0181 failure analysis (int8 in SALV), W7B results 2026-10-03 19:16:56 +02:00
Kral
d8cbf4de25 Decision: no separate SAP user for the harness (2026-10-03) 2026-10-03 19:00:50 +02:00
Kral
573707613e Eval review results, easy candidates, rerun queue; handover and STATE updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:57:23 +02:00
Kral
fb18dc8d9d Review: 9 random high-score checks; easy candidates list 2026-10-03 18:56:12 +02:00
Kral
22022ca6a5 Review: 10 remaining accepted tasks (checklist) 2026-10-03 18:55:33 +02:00
Kral
530590f8b3 Review: 10 remaining accepted tasks (checklist) 2026-10-03 18:55:25 +02:00
Kral
5e3baf43d3 Empirical filter results: DeepSeek reruns W6A/W7A (partial) 2026-10-03 18:54:16 +02:00
Kral
d83a39a8d8 MLX server: prompt cache limit (memory); restart after the T01 test
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:47:03 +02:00
Kral
81d8cf6f99 Handover: no new model runs during the baseline; rerun queue
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:43:25 +02:00
Kral
9366ee9868 Handover notes, train/STATE.md, detached night chain for DeepSeek reruns and the stage 1 baseline
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:34:07 +02:00
Kral
d4ed32cf38 G0181: restore ALV output in K spec; K prompt keeps output form
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:26:24 +02:00
Kral
92e30eb554 Runner: G2 finds the FUNCTION statement after local classes; budget floor applied to all tasks
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:25:30 +02:00
Kral
7351849a76 Stage 1 step 0: train/.venv (py3.11, mlx-lm 0.32), MLX 4-bit Qwen3.8-27B, serve script, baseline runner, README
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:23:39 +02:00
Kral
cf8d6d5d07 Ignore train/.venv
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:34:02 +02:00
Kral
65322f27ff Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:33:55 +02:00
Kral
01e99158a9 Agent: max_tokens 32k for cloud models and retry on empty turns (runaway reasoning); proxy: sap_activate fallback for PROG/FUNC
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:07:31 +02:00
Kral
0e96c6708e Review fixes G0019, G0108: added hidden tests; revalidated
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 16:35:20 +02:00
Kral
7510131a7d ADT fallback: Accept headers for FUNC; G0189 accepted; regen2 results
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 16:31:03 +02:00
Kral
428b3565b2 Proxy: ADT REST write+activate fallback when EPOD does not activate a second PROG/FUNC write
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 14:29:37 +02:00
Kral
39769afce8 Step F first pass: 81/92 accepted; budget floor 60 calls; rejected slots moved for regeneration
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 12:10:52 +02:00