Kral
|
4432ac896c
|
Only kinds below their target share are generated and run; one-controller flock guard
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 19:39:37 +02:00 |
|
Kral
|
976274d48a
|
Balanced generation under the pipeline, per-kind brake; STATE: type mix fix
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 19:27:43 +02:00 |
|
Kral
|
55f4068330
|
STRU and exception tasks accepted; G6 cascade fix, exception class mutants, RTTI unit notes
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 19:25:09 +02:00 |
|
Kral
|
b119f1afac
|
Object type mix: INTF, TABL, STRU, MSAG, exception tasks (harness G2, mutants, generator notes), balanced generator, kind-deficit job order, dashboard mix card
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 18:52:29 +02:00 |
|
Kral
|
454db6626c
|
Review notes: token length for bf16 test (p95 48k), DDLS reject reasons, G1034 lock analysis (not reproduced)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 18:21:44 +02:00 |
|
Kral
|
224ba1a5f8
|
Summary: correct the guard line; pipeline summary reads the limit from .env
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 18:05:49 +02:00 |
|
Kral
|
5bcd1bf9cd
|
Stage 2 summary at 50 accepted trajectories
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 18:02:52 +02:00 |
|
Kral
|
332fb4601d
|
Live status page (SAP colors), rewritten every 5 minutes
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 17:56:30 +02:00 |
|
Kral
|
bc6827b744
|
Budget guard from panel 41.07: ratio 1.2, limit 121
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 17:47:39 +02:00 |
|
Kral
|
704cfa3ede
|
Budget guard from the panel: ratio 1.35 for trajectory runs, limit 123
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 15:54:10 +02:00 |
|
Kral
|
ae37617a0d
|
Budget guard: ratio 2.07 measured, limit 142; limit read from .env at each check
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 14:47:54 +02:00 |
|
Kral
|
c062df2ec5
|
Trajectory runs: tool-call budget 100 for tasks with a CDS contract object (eval keeps 60)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 14:08:33 +02:00 |
|
Kral
|
13360292dd
|
Pipeline controller: plan 2 (hard +, error +50 %), backlog throttle, 2 trajectory workers, stop rules, summaries every 50; budget reserve
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 13:35:54 +02:00 |
|
Kral
|
1c4667a0fc
|
MCP load test (read, heavy read, write modes)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 13:26:13 +02:00 |
|
Kral
|
9de3911578
|
K variants for training (free text, EPOD tool names), second attempt only for failed tasks, docs/epod-syntax-hint.md
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 12:50:43 +02:00 |
|
Kral
|
35f0eb0e0c
|
Overlap limit for whole spec 0.75, slot logs; STATE and roadmap: stage 2 data progress
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 12:04:11 +02:00 |
|
Kral
|
61b903f7b3
|
Acceptance filter: score, end reason, harness errors, loop trimming, repair marker
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 12:03:54 +02:00 |
|
Kral
|
c2d4997566
|
Converter to the Qwen 3.8 chat template (thinking off) with tokenizer round-trip check
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 12:03:54 +02:00 |
|
Kral
|
d5e43e1a83
|
Trajectory record (messages, raw tool results, metadata; reasoning apart) and trajectory runner
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 12:03:54 +02:00 |
|
Kral
|
a4eb567e4c
|
Proxy: local abaplint messages on a bare 'save failed' (EPOD syntax check cannot see the rejected source)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 12:03:54 +02:00 |
|
Kral
|
229f862600
|
Generator training mode: train pool, eval overlap check, category mix, error-targeted tasks
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 11:53:07 +02:00 |
|
Kral
|
3a9f9e75fc
|
Step 400 check: Mac -13.4 % vs GPU -35.0 %; next GPU run bf16 mixed with stage 2; findings in STATE and roadmap
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-05 11:46:21 +02:00 |
|
Kral
|
39dcdd9846
|
Step 200 check: GPU -26.0 % vs Mac -10.7 %, full run cancelled at step ~418
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-04 20:48:49 +02:00 |
|
Kral
|
0c0982f666
|
Overfit conversion test passed (Mac -98.5 %, GPU -99.97 %); full run started, alpha 32
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-04 19:34:24 +02:00 |
|
Kral
|
20ef7b2331
|
Stage 1 pipeline test: 30.5 s/step, converter checked, valid loss 0.848 with test adapter
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-04 19:24:25 +02:00 |
|
Kral
|
5447874fd3
|
Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
2026-10-04 18:51:40 +02:00 |
|
Kral
|
a2ba9e7b44
|
serve.sh back to Qwen 3.8 27B 4-bit
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 17:02:59 +02:00 |
|
Kral
|
9f85a71ca5
|
Roadmap: Qwen baseline result
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 17:02:02 +02:00 |
|
Kral
|
a368233d26
|
Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 17:01:52 +02:00 |
|
Kral
|
dbe6077a48
|
Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 13:11:53 +02:00 |
|
Kral
|
5532a5b226
|
baseline: setup failure is not recorded as a result; repair_stats tolerates a missing trajectory
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 12:07:33 +02:00 |
|
Kral
|
d15e190f72
|
baseline.py: --model, --base-url, --enable-thinking-false (Qwen run on MacBook)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 11:39:33 +02:00 |
|
Kral
|
0af2d6458e
|
Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 11:20:43 +02:00 |
|
Kral
|
8d1db9c67c
|
Devstral Small 2: serve.sh, baseline.py without thinking args, repair rate, activation error message fix, stop rule in chain (run base 21000)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 08:57:12 +02:00 |
|
Kral
|
e84ea43a3d
|
Base model decision: drop Qwen 3.8, candidate Devstral Small 2 (24B)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 08:44:04 +02:00 |
|
Kral
|
ba0a7e6b3b
|
STATE: correct start time
|
2026-10-04 07:59:41 +02:00 |
|
Kral
|
342eb6fc9a
|
STATE: baseline restart time
|
2026-10-04 07:59:38 +02:00 |
|
Kral
|
5c68a1b1b7
|
Loop guard: also identical call with identical result 3 times in a row; baseline restarts (run base 20500)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 07:59:28 +02:00 |
|
Kral
|
a94c2a3b5d
|
Baseline restarted with loop guard, thinking off (run base 20400)
|
2026-10-04 07:40:19 +02:00 |
|
Kral
|
0401186a6c
|
Loop guard (3 identical pushes), end_reason and activation error records per run; thinking off for the stage 1 baseline
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-04 07:07:44 +02:00 |
|
Kral
|
040b9900fd
|
STATE/handover: baseline restarted on the 11-task subset
|
2026-10-03 22:36:30 +02:00 |
|
Kral
|
689820af4d
|
Baseline shortened: 11-task subset (1 per category + T01), max_tokens 16384, settings in README; baseline_chain.sh
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-03 22:36:00 +02:00 |
|
Kral
|
7054d9fd2e
|
STATE: correct timestamps
|
2026-10-03 22:24:08 +02:00 |
|
Kral
|
7b8ca01bde
|
Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-03 22:23:04 +02:00 |
|
Kral
|
0677d035da
|
Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-03 22:09:15 +02:00 |
|
Kral
|
e083c9ca13
|
Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
|
2026-10-03 21:58:04 +02:00 |
|
Kral
|
f0593933f9
|
Budget calendar: usage estimate, 20 USD paid for 60 USD usage
|
2026-10-03 20:15:05 +02:00 |
|
Kral
|
eaaef125ff
|
Budget guard 135 (ledger), about 50 USD real
|
2026-10-03 20:12:38 +02:00 |
|
Kral
|
b0fd06259b
|
Budget: real ratio ledger/2.7 (Ollama page 24.63 USD)
|
2026-10-03 20:11:45 +02:00 |
|
Kral
|
aef4c6e461
|
Pending: A4H memory limit after the baseline
|
2026-10-03 20:04:36 +02:00 |
|