Commit Graph

41 Commits

Author SHA1 Message Date
Kral
542fc8ec31 B: stage 2 data analysis (repair taxonomy, teacher vs Qwen behaviors, duplicates, empty_response); empty_response fix (cap 24000, retry temperature, stream guard)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 05:56:23 +02:00
Kral
1ddb9a7c65 Series A result: local Qwen 3 of 20 accepted, failures are loops and search loops; dashboard public copy
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-06 05:45:24 +02:00
Kral
5c07026bbb serve_remote.sh: pick ~/qwen-venv, refuse an old mlx-lm with a clear message; docs for mlx-lm 0.32.0 on the MacBook
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 22:16:56 +02:00
Kral
eb13453097 Series A: local Qwen on the MacBook (remote server), 12 h window with hard stop, clean pause on server loss, dashboard card, docs/remote-model.md
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 21:57:01 +02:00
Kral
59496cb08b Lock leak: controlled reproduction attempts and method (not reproduced), enqueue reader; restart plan for 12 October
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 19:47:16 +02:00
Kral
454db6626c Review notes: token length for bf16 test (p95 48k), DDLS reject reasons, G1034 lock analysis (not reproduced)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 18:21:44 +02:00
Kral
224ba1a5f8 Summary: correct the guard line; pipeline summary reads the limit from .env
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 18:05:49 +02:00
Kral
5bcd1bf9cd Stage 2 summary at 50 accepted trajectories
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 18:02:52 +02:00
Kral
13360292dd Pipeline controller: plan 2 (hard +, error +50 %), backlog throttle, 2 trajectory workers, stop rules, summaries every 50; budget reserve
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 13:35:54 +02:00
Kral
9de3911578 K variants for training (free text, EPOD tool names), second attempt only for failed tasks, docs/epod-syntax-hint.md
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:50:43 +02:00
Kral
35f0eb0e0c Overlap limit for whole spec 0.75, slot logs; STATE and roadmap: stage 2 data progress
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 12:04:11 +02:00
Kral
3a9f9e75fc Step 400 check: Mac -13.4 % vs GPU -35.0 %; next GPU run bf16 mixed with stage 2; findings in STATE and roadmap
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-05 11:46:21 +02:00
Kral
5447874fd3 Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 18:51:40 +02:00
Kral
9f85a71ca5 Roadmap: Qwen baseline result
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:02:02 +02:00
Kral
a368233d26 Stage 1 reference baseline: Qwen 3.8 27B mean 15.8 (3/11 above 0), comparison with Devstral
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 17:01:52 +02:00
Kral
dbe6077a48 Decision 2026-10-04: base model Qwen 3.8 27B (Devstral not used further)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 13:11:53 +02:00
Kral
0af2d6458e Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 11:20:43 +02:00
Kral
040b9900fd STATE/handover: baseline restarted on the 11-task subset 2026-10-03 22:36:30 +02:00
Kral
7b8ca01bde Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:23:04 +02:00
Kral
0677d035da Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:09:15 +02:00
Kral
e083c9ca13 Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 21:58:04 +02:00
Kral
f0593933f9 Budget calendar: usage estimate, 20 USD paid for 60 USD usage 2026-10-03 20:15:05 +02:00
Kral
eaaef125ff Budget guard 135 (ledger), about 50 USD real 2026-10-03 20:12:38 +02:00
Kral
b0fd06259b Budget: real ratio ledger/2.7 (Ollama page 24.63 USD) 2026-10-03 20:11:45 +02:00
Kral
aef4c6e461 Pending: A4H memory limit after the baseline 2026-10-03 20:04:36 +02:00
Kral
2f9037d756 Correct proxy finding: Z0FFK001 object belongs to the running T01 test 2026-10-03 20:03:24 +02:00
Kral
4409393302 Decision: no A4H move; memory limit after the baseline 2026-10-03 19:46:15 +02:00
Kral
e3f5dcd668 G0181 failure analysis (int8 in SALV), W7B results 2026-10-03 19:16:56 +02:00
Kral
d8cbf4de25 Decision: no separate SAP user for the harness (2026-10-03) 2026-10-03 19:00:50 +02:00
Kral
573707613e Eval review results, easy candidates, rerun queue; handover and STATE updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:57:23 +02:00
Kral
d83a39a8d8 MLX server: prompt cache limit (memory); restart after the T01 test
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:47:03 +02:00
Kral
81d8cf6f99 Handover: no new model runs during the baseline; rerun queue
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:43:25 +02:00
Kral
9366ee9868 Handover notes, train/STATE.md, detached night chain for DeepSeek reruns and the stage 1 baseline
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:34:07 +02:00
Kral
65322f27ff Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:33:55 +02:00
Kral
d3250acfb8 Empirical filter runner (DeepSeek); real cost ratio in docs
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 09:29:20 +02:00
Kral
3550f0b651 Eval review v1: checklist, review helper, 24 tasks reviewed; G0002/G0022 tests fixed; I/E generator definitions
- docs/eval-inceleme.md checklist; harness/review.py
- 19 accept, 1 fix pending (G0019), 4 flagged (I/E tasks solvable without legacy code)
- generator: categories I and E fix/refactor the seed object in place

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:23:38 +02:00
Kral
f907351c0d Mutation check done on all accepted tasks; mutant fixes; dump filter in proxy
- skip sy-subrc lines and CDS type lengths; CDS literal mutants; mutate helper classes
- proxy: sap_short_dumps shows only dumps of the current run
- 26/26 tasks pass; results in docs/faz1-tasarim.md 11g

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:02:59 +02:00
Kral
0ba25faeea Mutation in generator; category H (stop) scoring and judge; category K variants and tool schema variant; evalset plan
- mutation.py: negation mutants, skip WHILE/DO blocks (endless loop blocked RFC ~8 min)
- generator: accept only after mutation check; survivors go back as repair feedback
- runner/judge.py: stop tasks scored 100/30/0; gap by keywords, else judge model
- proxy: tool schema variant generic_v0 (draft); generator make_k_variant
- evalset.py: 88 slots for step F (A-I, H, K)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 05:37:20 +02:00
Kral
32f03fcad4 Hard batch G0022-G0026: 5/5 accepted; budget floor, CDS checks, mutation check
- generator: budget floor (2x oracle activations, 3x calls), static checks for CDS
  $parameters and UNION annotation
- harness/mutation.py: deterministic mutants of the reference; hidden tests must fail
- G0022 revalidated with harness fixes: oracle 100, null 0
- results in docs/faz1-tasarim.md 11f

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 05:16:30 +02:00
Kral
3570d7ef9b Harness: G6 ignores unknown standard superclass; dependency order for seed/reference; max_tokens 80k
- abaplint does not know CX_STATIC_CHECK: exception hierarchies failed G6 (G0022, G0005)
- generator orders seed and reference objects by name references (G0022 attempt 1)
- max_tokens from pilot distribution (accepted calls 4k-76k)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 05:06:04 +02:00
Kral
5de5e2851c Generator: pilot analysis and fixes; G0002-G0021 tasks; handover notes merged
- Static bundle checks (name length, seed type, reserved words, contract test classes,
  testclasses_file), local abaplint parser check before SAP, max_tokens, robust JSON parse
- Runner: G2 for CDS views with parameters, contract_check detail in report
- G0020 and G0013 fixed by review and revalidated (oracle 100, null 0)
- Pilot results in docs/faz1-tasarim.md 11e; handover notes merged into CLAUDE.md and docs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-03 05:04:23 +02:00