Kral
d15e190f72
baseline.py: --model, --base-url, --enable-thinking-false (Qwen run on MacBook)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 11:39:33 +02:00
Kral
0af2d6458e
Devstral Small 2 baseline (11 tasks): README, roadmap, results; Qwen 3.8 dropped
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 11:20:43 +02:00
Kral
8d1db9c67c
Devstral Small 2: serve.sh, baseline.py without thinking args, repair rate, activation error message fix, stop rule in chain (run base 21000)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 08:57:12 +02:00
Kral
e84ea43a3d
Base model decision: drop Qwen 3.8, candidate Devstral Small 2 (24B)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 08:44:04 +02:00
Kral
ba0a7e6b3b
STATE: correct start time
2026-10-04 07:59:41 +02:00
Kral
342eb6fc9a
STATE: baseline restart time
2026-10-04 07:59:38 +02:00
Kral
5c68a1b1b7
Loop guard: also identical call with identical result 3 times in a row; baseline restarts (run base 20500)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 07:59:28 +02:00
Kral
a94c2a3b5d
Baseline restarted with loop guard, thinking off (run base 20400)
2026-10-04 07:40:19 +02:00
Kral
0401186a6c
Loop guard (3 identical pushes), end_reason and activation error records per run; thinking off for the stage 1 baseline
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-04 07:07:44 +02:00
Kral
040b9900fd
STATE/handover: baseline restarted on the 11-task subset
2026-10-03 22:36:30 +02:00
Kral
689820af4d
Baseline shortened: 11-task subset (1 per category + T01), max_tokens 16384, settings in README; baseline_chain.sh
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:36:00 +02:00
Kral
7054d9fd2e
STATE: correct timestamps
2026-10-03 22:24:08 +02:00
Kral
7b8ca01bde
Stage 1 data: strict version dedup (older v* vs next newer v*); counts updated
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:23:04 +02:00
Kral
0677d035da
Stage 1 data: version dedup, splitting of long documents, family split; 748 iterations
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 22:09:15 +02:00
Kral
e083c9ca13
Stage 1 step 1: prepare.py, real corpus numbers (SAP-samples/abap-cheat-sheets)
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 21:58:04 +02:00
Kral
f0593933f9
Budget calendar: usage estimate, 20 USD paid for 60 USD usage
2026-10-03 20:15:05 +02:00
Kral
eaaef125ff
Budget guard 135 (ledger), about 50 USD real
2026-10-03 20:12:38 +02:00
Kral
b0fd06259b
Budget: real ratio ledger/2.7 (Ollama page 24.63 USD)
2026-10-03 20:11:45 +02:00
Kral
aef4c6e461
Pending: A4H memory limit after the baseline
2026-10-03 20:04:36 +02:00
Kral
2f9037d756
Correct proxy finding: Z0FFK001 object belongs to the running T01 test
2026-10-03 20:03:24 +02:00
Kral
4409393302
Decision: no A4H move; memory limit after the baseline
2026-10-03 19:46:15 +02:00
Kral
e3f5dcd668
G0181 failure analysis (int8 in SALV), W7B results
2026-10-03 19:16:56 +02:00
Kral
d8cbf4de25
Decision: no separate SAP user for the harness (2026-10-03)
2026-10-03 19:00:50 +02:00
Kral
573707613e
Eval review results, easy candidates, rerun queue; handover and STATE updated
...
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:57:23 +02:00
Kral
fb18dc8d9d
Review: 9 random high-score checks; easy candidates list
2026-10-03 18:56:12 +02:00
Kral
22022ca6a5
Review: 10 remaining accepted tasks (checklist)
2026-10-03 18:55:33 +02:00
Kral
530590f8b3
Review: 10 remaining accepted tasks (checklist)
2026-10-03 18:55:25 +02:00
Kral
5e3baf43d3
Empirical filter results: DeepSeek reruns W6A/W7A (partial)
2026-10-03 18:54:16 +02:00
Kral
d83a39a8d8
MLX server: prompt cache limit (memory); restart after the T01 test
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:47:03 +02:00
Kral
81d8cf6f99
Handover: no new model runs during the baseline; rerun queue
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:43:25 +02:00
Kral
9366ee9868
Handover notes, train/STATE.md, detached night chain for DeepSeek reruns and the stage 1 baseline
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:34:07 +02:00
Kral
d4ed32cf38
G0181: restore ALV output in K spec; K prompt keeps output form
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:26:24 +02:00
Kral
92e30eb554
Runner: G2 finds the FUNCTION statement after local classes; budget floor applied to all tasks
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:25:30 +02:00
Kral
7351849a76
Stage 1 step 0: train/.venv (py3.11, mlx-lm 0.32), MLX 4-bit Qwen3.8-27B, serve script, baseline runner, README
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 18:23:39 +02:00
Kral
cf8d6d5d07
Ignore train/.venv
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:34:02 +02:00
Kral
65322f27ff
Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:33:55 +02:00
Kral
01e99158a9
Agent: max_tokens 32k for cloud models and retry on empty turns (runaway reasoning); proxy: sap_activate fallback for PROG/FUNC
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 17:07:31 +02:00
Kral
0e96c6708e
Review fixes G0019, G0108: added hidden tests; revalidated
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 16:35:20 +02:00
Kral
7510131a7d
ADT fallback: Accept headers for FUNC; G0189 accepted; regen2 results
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 16:31:03 +02:00
Kral
428b3565b2
Proxy: ADT REST write+activate fallback when EPOD does not activate a second PROG/FUNC write
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 14:29:37 +02:00
Kral
39769afce8
Step F first pass: 81/92 accepted; budget floor 60 calls; rejected slots moved for regeneration
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 12:10:52 +02:00
Kral
d3250acfb8
Empirical filter runner (DeepSeek); real cost ratio in docs
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 09:29:20 +02:00
Kral
03ab36149e
Fix coverage with several contract classes; table annotation and file path checks; clean task dir on write
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 08:44:35 +02:00
Kral
fb576b0954
In-place E/I: G2 needs a changed seed contract object, G4 skips it; regenerate G0009/G0013/G0016/G0023 as G0188-G0191
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:33:33 +02:00
Kral
3550f0b651
Eval review v1: checklist, review helper, 24 tasks reviewed; G0002/G0022 tests fixed; I/E generator definitions
...
- docs/eval-inceleme.md checklist; harness/review.py
- 19 accept, 1 fix pending (G0019), 4 flagged (I/E tasks solvable without legacy code)
- generator: categories I and E fix/refactor the seed object in place
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:23:38 +02:00
Kral
a12fa4dff6
H and K tested (G0168, G0178 accepted); keyword fast path needs all groups
...
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:04:36 +02:00
Kral
f907351c0d
Mutation check done on all accepted tasks; mutant fixes; dump filter in proxy
...
- skip sy-subrc lines and CDS type lengths; CDS literal mutants; mutate helper classes
- proxy: sap_short_dumps shows only dumps of the current run
- 26/26 tasks pass; results in docs/faz1-tasarim.md 11g
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
2026-10-03 06:02:59 +02:00
Kral
0ba25faeea
Mutation in generator; category H (stop) scoring and judge; category K variants and tool schema variant; evalset plan
...
- mutation.py: negation mutants, skip WHILE/DO blocks (endless loop blocked RFC ~8 min)
- generator: accept only after mutation check; survivors go back as repair feedback
- runner/judge.py: stop tasks scored 100/30/0; gap by keywords, else judge model
- proxy: tool schema variant generic_v0 (draft); generator make_k_variant
- evalset.py: 88 slots for step F (A-I, H, K)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
2026-10-03 05:37:20 +02:00
Kral
32f03fcad4
Hard batch G0022-G0026: 5/5 accepted; budget floor, CDS checks, mutation check
...
- generator: budget floor (2x oracle activations, 3x calls), static checks for CDS
$parameters and UNION annotation
- harness/mutation.py: deterministic mutants of the reference; hidden tests must fail
- G0022 revalidated with harness fixes: oracle 100, null 0
- results in docs/faz1-tasarim.md 11f
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
2026-10-03 05:16:30 +02:00
Kral
3570d7ef9b
Harness: G6 ignores unknown standard superclass; dependency order for seed/reference; max_tokens 80k
...
- abaplint does not know CX_STATIC_CHECK: exception hierarchies failed G6 (G0022, G0005)
- generator orders seed and reference objects by name references (G0022 attempt 1)
- max_tokens from pilot distribution (accepted calls 4k-76k)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com >
2026-10-03 05:06:04 +02:00