Kral
8ef85c2713
Decisions of 2026-10-06: 60 % rule in the training script, own-test weights (fractional), 64k in the sweep, foreign-read trajectories back to the pending pool, memory test waits, foreign object scan of baselines and eval runs, 11 October check list
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 09:04:09 +02:00
Kral
a465c33e3f
D: own-test mutation scores of the accepted trajectories (metadata only), stage 2 set rebuilt with the score
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 07:36:39 +02:00
Kral
c5f1a36df0
stage 2 builder: scrub of other runs' leftover objects, drop of foreign reads; rebuild (36 train + 4 valid), doc generator, HF dataset updated
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 06:16:02 +02:00
Kral
c2c3803b1b
G: EPOD acceptance tests; harness: hide and clean other runs' mid-name objects (proxy, teardown, sweep); lock leak cause (Eclipse restart during a write)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 06:15:08 +02:00
Kral
0c0e07fe96
F: eval slots for INTF/TABL/STRU/MSAG/exception (+K), Kral spot-check sheet, step 0 in the restart plan; D: own-test mutation scoring (running)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 06:06:41 +02:00
Kral
7f03849a86
C+E: stage 2 builder (mask, 48k, CLAS cap, family split, hook), private HF dataset, bf16 mixed training script with memory test, memory table, ratio proposal
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 06:01:49 +02:00
Kral
f5611cbf56
data analysis doc: corrected two counts
2026-10-06 05:56:31 +02:00
Kral
542fc8ec31
B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 05:56:23 +02:00
Kral
1ddb9a7c65
Series A result: local Qwen 3 of 20 accepted, failures are loops and search loops; dashboard public copy
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-06 05:45:24 +02:00
Kral
5c07026bbb
serve_remote.sh: pick ~/qwen-venv, refuse an old mlx-lm with a clear message; docs for mlx-lm 0.32.0 on the MacBook
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 22:16:56 +02:00
Kral
eb13453097
Series A: local Qwen on the MacBook (remote server), 12 h window with hard stop, clean pause on server loss, dashboard card, docs/remote-model.md
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 21:57:01 +02:00
Kral
59496cb08b
Lock leak: controlled reproduction attempts and method (not reproduced), enqueue reader; restart plan for 12 October
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 19:47:16 +02:00
Kral
454db6626c
Review notes: token length for bf16 test (p95 48k), DDLS reject reasons, G1034 lock analysis (not reproduced)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 18:21:44 +02:00
Kral
224ba1a5f8
Summary: correct the guard line; pipeline summary reads the limit from .env
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 18:05:49 +02:00
Kral
5bcd1bf9cd
Stage 2 summary at 50 accepted trajectories
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 18:02:52 +02:00
Kral
13360292dd
Pipeline controller: plan 2 (hard +, error +50 %), backlog throttle, 2 trajectory workers, stop rules, summaries every 50; budget reserve
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 13:35:54 +02:00
Kral
9de3911578
K variants for training (free text, EPOD tool names), second attempt only for failed tasks, docs/epod-syntax-hint.md
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 12:50:43 +02:00
Kral
35f0eb0e0c
Overlap limit for whole spec 0.75, slot logs; STATE and roadmap: stage 2 data progress
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 12:04:11 +02:00
Kral
3a9f9e75fc
Step 400 check: Mac -13.4 % vs GPU -35.0 %; next GPU run bf16 mixed with stage 2; findings in STATE and roadmap
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-05 11:46:21 +02:00
Kral
5447874fd3
Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
2026-10-04 18:51:40 +02:00
Kral
9f85a71ca5
Roadmap: Qwen baseline result
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:02:02 +02:00
Kral
a368233d26
Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:01:52 +02:00
Kral
dbe6077a48
Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 13:11:53 +02:00
Kral
0af2d6458e
Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 11:20:43 +02:00
Kral
040b9900fd
STATE/handover: baseline restarted on the 11-task subset
2026-10-03 22:36:30 +02:00
Kral
7b8ca01bde
Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:23:04 +02:00
Kral
0677d035da
Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:09:15 +02:00
Kral
e083c9ca13
Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 21:58:04 +02:00
Kral
f0593933f9
Budget calendar: usage estimate, 20 USD paid for 60 USD usage
2026-10-03 20:15:05 +02:00
Kral
eaaef125ff
Budget guard 135 (ledger), about 50 USD real
2026-10-03 20:12:38 +02:00
Kral
b0fd06259b
Budget: real ratio ledger/2.7 (Ollama page 24.63 USD)
2026-10-03 20:11:45 +02:00
Kral
aef4c6e461
Pending: A4H memory limit after the baseline
2026-10-03 20:04:36 +02:00
Kral
2f9037d756
Correct proxy finding: Z0FFK001 object belongs to the running T01 test
2026-10-03 20:03:24 +02:00
Kral
4409393302
Decision: no A4H move; memory limit after the baseline
2026-10-03 19:46:15 +02:00
Kral
e3f5dcd668
G0181 failure analysis (int8 in SALV), W7B results
2026-10-03 19:16:56 +02:00
Kral
d8cbf4de25
Decision: no separate SAP user for the harness (2026-10-03)
2026-10-03 19:00:50 +02:00
Kral
573707613e
Eval review results, easy candidates, rerun queue; handover and STATE updated
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:57:23 +02:00
Kral
d83a39a8d8
MLX server: prompt cache limit (memory); restart after the T01 test
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:47:03 +02:00
Kral
81d8cf6f99
Handover: no new model runs during the baseline; rerun queue
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:43:25 +02:00
Kral
9366ee9868
Handover notes, train/STATE.md, detached night chain for DeepSeek reruns and the stage 1 baseline
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:34:07 +02:00
Kral
65322f27ff
Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:33:55 +02:00
Kral
d3250acfb8
Empirical filter runner (DeepSeek); real cost ratio in docs
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 09:29:20 +02:00
Kral
3550f0b651
Eval review v1: checklist, review helper, 24 tasks reviewed; G0002/G0022 tests fixed; I/E generator definitions
...
- docs/eval-inceleme.md checklist; harness/review.py
- 19 accept, 1 fix pending (G0019), 4 flagged (I/E tasks solvable without legacy code)
- generator: categories I and E fix/refactor the seed object in place
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:23:38 +02:00
Kral
f907351c0d
Mutation check done on all accepted tasks; mutant fixes; dump filter in proxy
...
- skip sy-subrc lines and CDS type lengths; CDS literal mutants; mutate helper classes
- proxy: sap_short_dumps shows only dumps of the current run
- 26/26 tasks pass; results in docs/faz1-tasarim.md 11g
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:02:59 +02:00
Kral
0ba25faeea
Mutation in generator; category H (stop) scoring and judge; category K variants and tool schema variant; evalset plan
...
- mutation.py: negation mutants, skip WHILE/DO blocks (endless loop blocked RFC ~8 min)
- generator: accept only after mutation check; survivors go back as repair feedback
- runner/judge.py: stop tasks scored 100/30/0; gap by keywords, else judge model
- proxy: tool schema variant generic_v0 (draft); generator make_k_variant
- evalset.py: 88 slots for step F (A-I, H, K)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
2026-10-03 05:37:20 +02:00
Kral
32f03fcad4
Hard batch G0022-G0026: 5/5 accepted; budget floor, CDS checks, mutation check
...
- generator: budget floor (2x oracle activations, 3x calls), static checks for CDS
$parameters and UNION annotation
- harness/mutation.py: deterministic mutants of the reference; hidden tests must fail
- G0022 revalidated with harness fixes: oracle 100, null 0
- results in docs/faz1-tasarim.md 11f
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
2026-10-03 05:16:30 +02:00
Kral
3570d7ef9b
Harness: G6 ignores unknown standard superclass; dependency order for seed/reference; max_tokens 80k
...
- abaplint does not know CX_STATIC_CHECK: exception hierarchies failed G6 (G0022, G0005)
- generator orders seed and reference objects by name references (G0022 attempt 1)
- max_tokens from pilot distribution (accepted calls 4k-76k)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
2026-10-03 05:06:04 +02:00
Kral
5de5e2851c
Generator: pilot analysis and fixes; G0002-G0021 tasks; handover notes merged
...
- Static bundle checks (name length, seed type, reserved words, contract test classes,
testclasses_file), local abaplint parser check before SAP, max_tokens, robust JSON parse
- Runner: G2 for CDS views with parameters, contract_check detail in report
- G0020 and G0013 fixed by review and revalidated (oracle 100, null 0)
- Pilot results in docs/faz1-tasarim.md 11e; handover notes merged into CLAUDE.md and docs
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
2026-10-03 05:04:23 +02:00