Files
abap-llm/docs/devir-notlari.md

100 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Handover notes (2026-10-03, 19:10)
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
CLAUDE.md and the docs when they are done, then empty this file.
## 0. Rule for the next session (Kral, 2026-10-03)
Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue
reruns for after the baseline in `runs/stage1/rerun_queue.txt` (one task id per line, with the reason).
The night chain still starts W7B and then the baseline as planned.
## 1. Running jobs
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
| Job | Command | PID | Log | Expected end |
|---|---|---|---|---|
| MLX server (base model, no adapter) | `train/serve.sh` | 59352 | `runs/stage1/server.log` | runs until stopped (`kill 59352`) |
| T01 test run (budget 40, harness check only) | `python3 train/baseline.py --label baseline --only T01` | 59412 (sh), 59414 (python) | `runs/stage1/baseline_t01.log` | about 19:00 |
| DeepSeek filter W6A (last task G0125) | `xargs python3 -m harness.empirical ... --run-base 15000` | 57246 / 57253 | `runs/emp_deepseek6a.log` | about 18:45 |
| DeepSeek filter W7A (11 tasks, budget-60 reruns) | `xargs ... --run-base 17000 (list runs/stage1/w7a.txt)` | 59503 / 59510 | `runs/emp_deepseek7a.log` | about 19:45 |
| Night chain | `train/night_chain.sh` | 60017 | `runs/stage1/night_chain.log` | see below |
| Server restart with prompt cache limit (after the T01 test; new server PID in the log) | `train/restart_server.sh` | 60447 | `runs/stage1/restart_server.log` | after the T01 test |
Night chain (`train/night_chain.sh`):
1. When W6A ends: starts W7B (13 DeepSeek reruns, list `runs/stage1/w7b.txt`, log `runs/emp_deepseek7b.log`,
run base 18000). Expected end about 20:00.
2. When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks
T01 in `runs/stage1/baseline.json`. If there is no harness error, it moves the T01 entry to
`t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60):
`python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`.
If T01 failed: notification "T01 test failed: baseline not started", and it stops.
3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October
morning to noon.
After the jobs:
- Read only the last lines of `runs/stage1/night_chain.log` and `runs/stage1/baseline.log`.
- Then stage 1 prompt B2 (`docs/stage1-step-prompts.md`): write the summary from `runs/stage1/baseline.json`.
- Stop the MLX server before training (step 3); A4H must be stopped by Kral.
## 2. Open decisions
| Item | Status |
|---|---|
| G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. | Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset. |
| Second model for the empirical filter (easy tasks) | Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks. |
| EPOD server bug: a second write of a PROG/FUNC is not activated (`outcome: notExecuted`), also not by `sap_activate`. | Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work. |
| Separate SAP user for the harness (dumps under user KESELI) | Decided 2026-10-03 (Kral): not needed. |
| G0122 (C FUNC): not accepted after 4 generations | Left out. C has 14 tasks (target 15). |
| A4H move to the MacBook (Podman) | Decided 2026-10-03 (Kral): not needed (2-3 h, risk). Instead set a Docker memory limit for A4H after the baseline (container restart; ask Kral first). |
## 3. Pending follow-ups
Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in
`docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
Done:
- Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2.
G0107 (old G2 failure) is in W7B.
- Item 2: K specs checked, done.
- Item 5, review part: the 10 unreviewed tasks are reviewed and `accept`. Low scores: G0108 model error,
G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop.
- Item 6: 10 random high-score tasks (seed 20261003) reviewed, all `accept`.
- Item 7: `tasks_gen/eval/easy_candidates.json` (39 tasks; DeepSeek >= 95 and <= 20 tool calls).
Candidates only; confirm with the Qwen baseline.
- Item 8: written to `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
Waiting for jobs (do these in the next session):
- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and
the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the
Kral spot-check list (flagged tasks + one per category, ~10).
- Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about
0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting.
- A4H memory limit (Kral 2026-10-03: do it when the system is idle, so after "Baseline (25 tasks) ended"; no
second question needed): check `docker stats` and the HANA allocation limit first; choose the limit with the
data (about 32 GB is the SAP minimum for the trial image; too low gives an OOM kill); clean stop of `a4h`
(give HANA time), apply the limit, start, test `localhost:50000` and one MCP call. The Mac mini has 64 GB; swap
was 9 of 10 GB during the Qwen run.
- Proxy finding: `sap_inactive_objects` shows objects of other runs (no prefix filter). `Z0FFK001_IBAN_VALIDATOR`
was the object of the running T01 test (run prefix Z0FFK001), not a leftover: teardown removes it. Fix the proxy
after the baseline; no A4H cleanup needed (check `cleanup-list` after the T01 test).
## 4. Stage 1 status
Done (step 0): `train/.venv` (Python 3.11.17 with uv), `mlx-lm` 0.32.0, weights
`mlx-community/Qwen3.8-27B-4bit` in `~/models/Qwen3.8-27B-4bit`, `train/serve.sh` with the Ollama settings,
tool-call test passed, fixed subset `train/subset.json`, runner `train/baseline.py`, `train/README.md`.
Running: baseline (night chain). Not done: step 1 (data preparation, `train/prepare.py`); the valid loss of the
base model (step 2) needs step 1. Details and next steps: `train/STATE.md`.
## 5. Budget
- Real usage (Ollama page, Kral, 2026-10-03 20:15): 24.63 of 60 USD. Ledger (list prices): 66.97 USD.
New ratio: real ≈ ledger / 2.7 (the old ratio 1.47 was wrong). About 35 USD real left until 12 October.
- Guard: `BUDGET_LIMIT_USD` raised from 75 to 135 (ledger) on 2026-10-03 (Kral): real cap about 50 of 60 USD.
- Planned cloud work is small: reruns (G0119, G0162, ≈ 0.15 USD real) and judge calls for H tasks in the baseline.
The baseline itself is local and free.
- The agent limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of the run cost.