Files
abap-llm/docs/devir-notlari.md

7.2 KiB
Raw Permalink Blame History

Handover notes (2026-10-03, 22:40; baseline shortened, see train/STATE.md)

Read this file, CLAUDE.md, train/STATE.md and train/README.md first. Merge the durable parts into CLAUDE.md and the docs when they are done, then empty this file.

0. Rule for the next session (Kral, 2026-10-03)

Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue reruns for after the baseline in runs/stage1/rerun_queue.txt (one task id per line, with the reason). The night chain still starts W7B and then the baseline as planned.

1. Running jobs

All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.

Job Command PID Log Expected end
MLX server (base model, no adapter) train/serve.sh 59352 runs/stage1/server.log runs until stopped (kill 59352)
T01 test run (budget 40, harness check only) python3 train/baseline.py --label baseline --only T01 59412 (sh), 59414 (python) runs/stage1/baseline_t01.log about 19:00
DeepSeek filter W6A (last task G0125) xargs python3 -m harness.empirical ... --run-base 15000 57246 / 57253 runs/emp_deepseek6a.log about 18:45
DeepSeek filter W7A (11 tasks, budget-60 reruns) xargs ... --run-base 17000 (list runs/stage1/w7a.txt) 59503 / 59510 runs/emp_deepseek7a.log about 19:45
Night chain train/night_chain.sh 60017 runs/stage1/night_chain.log see below
Server restart with prompt cache limit (after the T01 test; new server PID in the log) train/restart_server.sh 60447 runs/stage1/restart_server.log after the T01 test

Night chain (train/night_chain.sh):

  1. When W6A ends: starts W7B (13 DeepSeek reruns, list runs/stage1/w7b.txt, log runs/emp_deepseek7b.log, run base 18000). Expected end about 20:00.
  2. When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks T01 in runs/stage1/baseline.json. If there is no harness error, it moves the T01 entry to t01_test_budget40 and starts the baseline on all 25 subset tasks (budget 60): python3 train/baseline.py --label baseline --run-base 20100, log runs/stage1/baseline.log. If T01 failed: notification "T01 test failed: baseline not started", and it stops.
  3. Notification "Baseline (11 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October morning to noon.

After the jobs:

  • Read only the last lines of runs/stage1/night_chain.log and runs/stage1/baseline.log.
  • Then stage 1 prompt B2 (docs/stage1-step-prompts.md): write the summary from runs/stage1/baseline.json.
  • Stop the MLX server before training (step 3); A4H must be stopped by Kral.

2. Open decisions

Item Status
G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset.
Second model for the empirical filter (easy tasks) Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks.
EPOD server bug: a second write of a PROG/FUNC is not activated (outcome: notExecuted), also not by sap_activate. Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work.
Separate SAP user for the harness (dumps under user KESELI) Decided 2026-10-03 (Kral): not needed.
G0122 (C FUNC): not accepted after 4 generations Left out. C has 14 tasks (target 15).
A4H move to the MacBook (Podman) Decided 2026-10-03 (Kral): not needed (2-3 h, risk). Instead set a Docker memory limit for A4H after the baseline (container restart; ask Kral first).

3. Pending follow-ups

Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in docs/faz1-tasarim.md 11h and docs/eval-inceleme.md.

Done:

  • Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2. G0107 (old G2 failure) is in W7B.
  • Item 2: K specs checked, done.
  • Item 5, review part: the 10 unreviewed tasks are reviewed and accept. Low scores: G0108 model error, G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop.
  • Item 6: 10 random high-score tasks (seed 20261003) reviewed, all accept.
  • Item 7: tasks_gen/eval/easy_candidates.json (39 tasks; DeepSeek >= 95 and <= 20 tool calls). Candidates only; confirm with the Qwen baseline.
  • Item 8: written to docs/faz1-tasarim.md 11h and docs/eval-inceleme.md.

Waiting for jobs (do these in the next session):

  • After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
  • After "Baseline (11 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the Kral spot-check list (flagged tasks + one per category, ~10).
  • Rerun queue runs/stage1/rerun_queue.txt (G0119, G0162): run after the baseline. Cost: DeepSeek, about 0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting.
  • A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (11 tasks) ended"; no second question needed): check docker stats and the HANA allocation limit first; choose the limit with the data (about 32 GB is the SAP minimum for the trial image; too low gives an OOM kill); clean stop of a4h (give HANA time), apply the limit, start, test localhost:50000 and one MCP call. The Mac mini has 64 GB; swap was 9 of 10 GB during the Qwen run.
  • Proxy finding: sap_inactive_objects shows objects of other runs (no prefix filter). Z0FFK001_IBAN_VALIDATOR was the object of the running T01 test (run prefix Z0FFK001), not a leftover: teardown removes it. Fix the proxy after the baseline; no A4H cleanup needed (check cleanup-list after the T01 test).

4. Stage 1 status

Done (step 0): train/.venv (Python 3.11.17 with uv), mlx-lm 0.32.0, weights mlx-community/Qwen3.8-27B-4bit in ~/models/Qwen3.8-27B-4bit, train/serve.sh with the Ollama settings, tool-call test passed, fixed subset train/subset.json, runner train/baseline.py, train/README.md. Running: baseline (night chain). Not done: step 1 (data preparation, train/prepare.py); the valid loss of the base model (step 2) needs step 1. Details and next steps: train/STATE.md.

5. Budget

  • Kral pays 20 USD per month for 60 USD of DeepSeek usage (real money = usage / 3).
  • Real usage (Ollama page, Kral, 2026-10-03 20:15): 24.63 of 60 USD. Ledger (list prices): 66.97 USD. New ratio: real ≈ ledger / 2.7 (the old ratio 1.47 was wrong). About 35 USD real left until 12 October.
  • Guard: BUDGET_LIMIT_USD raised from 75 to 135 (ledger) on 2026-10-03 (Kral): real cap about 50 of 60 USD.
  • Planned cloud work is small: reruns (G0119, G0162, ≈ 0.15 USD real) and judge calls for H tasks in the baseline. The baseline itself is local and free.
  • The agent limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of the run cost.