Files
abap-llm/docs/devir-notlari.md

93 lines
6.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Handover notes (2026-10-03, 18:35)
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
CLAUDE.md and the docs when they are done, then empty this file.
## 0. Rule for the next session (Kral, 2026-10-03)
Do not start new model runs while the baseline runs. Use only stored results for the follow-ups. Queue
reruns for after the baseline in `runs/stage1/rerun_queue.txt` (one task id per line, with the reason).
The night chain still starts W7B and then the baseline as planned.
## 1. Running jobs
All jobs are detached (parent PID 1). They do not stop when the Claude Code session ends.
| Job | Command | PID | Log | Expected end |
|---|---|---|---|---|
| MLX server (base model, no adapter) | `train/serve.sh` | 59352 | `runs/stage1/server.log` | runs until stopped (`kill 59352`) |
| T01 test run (budget 40, harness check only) | `python3 train/baseline.py --label baseline --only T01` | 59412 (sh), 59414 (python) | `runs/stage1/baseline_t01.log` | about 19:00 |
| DeepSeek filter W6A (last task G0125) | `xargs python3 -m harness.empirical ... --run-base 15000` | 57246 / 57253 | `runs/emp_deepseek6a.log` | about 18:45 |
| DeepSeek filter W7A (11 tasks, budget-60 reruns) | `xargs ... --run-base 17000 (list runs/stage1/w7a.txt)` | 59503 / 59510 | `runs/emp_deepseek7a.log` | about 19:45 |
| Night chain | `train/night_chain.sh` | 60017 | `runs/stage1/night_chain.log` | see below |
| Server restart with prompt cache limit (after the T01 test; new server PID in the log) | `train/restart_server.sh` | 60447 | `runs/stage1/restart_server.log` | after the T01 test |
Night chain (`train/night_chain.sh`):
1. When W6A ends: starts W7B (13 DeepSeek reruns, list `runs/stage1/w7b.txt`, log `runs/emp_deepseek7b.log`,
run base 18000). Expected end about 20:00.
2. When the T01 test and all DeepSeek workers end: macOS notification "DeepSeek reruns ended". Then it checks
T01 in `runs/stage1/baseline.json`. If there is no harness error, it moves the T01 entry to
`t01_test_budget40` and starts the baseline on all 25 subset tasks (budget 60):
`python3 train/baseline.py --label baseline --run-base 20100`, log `runs/stage1/baseline.log`.
If T01 failed: notification "T01 test failed: baseline not started", and it stops.
3. Notification "Baseline (25 tasks) ended". Expected: 25 tasks × 20–40 min ≈ 8–16 h, end on 4 October
morning to noon.
After the jobs:
- Read only the last lines of `runs/stage1/night_chain.log` and `runs/stage1/baseline.log`.
- Then stage 1 prompt B2 (`docs/stage1-step-prompts.md`): write the summary from `runs/stage1/baseline.json`.
- Stop the MLX server before training (step 3); A4H must be stopped by Kral.
## 2. Open decisions
| Item | Status |
|---|---|
| G0170 (H): DeepSeek did not stop; it assumed the rounding (gap: rounding target and mode of the loyalty discount). Gap maybe not critical. | Waiting for the DeepSeek rerun with budget 60 (in W7A). If it again implements: suggest a stronger gap, as for G0174. G0170 is in the stage 1 subset. |
| Second model for the empirical filter (easy tasks) | Qwen baseline gives 25 tasks only. Decide later if more models run on all 109 tasks. |
| EPOD server bug: a second write of a PROG/FUNC is not activated (`outcome: notExecuted`), also not by `sap_activate`. | Workaround in the proxy (ADT REST lock/write/activate). The server fix is the other project's work. |
| Separate SAP user for the harness (dumps under user KESELI) | Proposed, no decision. |
| G0122 (C FUNC): not accepted after 4 generations | Left out. C has 14 tasks (target 15). |
## 3. Pending follow-ups
0. Allowed while the baseline runs (stored results only): items 2, 5 (review part), 6, 7, 8, and the G2 check
of stored FUNC sources (item 1, first part). Everything that needs a model run or SAP goes to
`runs/stage1/rerun_queue.txt`.
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
G0107). With stored results: re-check G2 on the saved sources (`sources/` in the run directories). Where
G2 now passes, the hidden tests never ran: queue a rerun. G0107 is already queued in W7B.
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
response"). Queued in `runs/stage1/rerun_queue.txt`.
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks
(G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189).
6. **10 random high-score reviews** from G0100 and later (Kral decision 2): not done.
7. **Mark easy tasks** (Kral decision 3: mark, do not remove): not done. Rule proposal: DeepSeek 100 with few
tool calls; confirm with the Qwen baseline.
8. Write the empirical filter results and the review into `docs/faz1-tasarim.md` (new section) and
`docs/eval-inceleme.md`.
## 4. Stage 1 status
Done (step 0): `train/.venv` (Python 3.11.17 with uv), `mlx-lm` 0.32.0, weights
`mlx-community/Qwen3.8-27B-4bit` in `~/models/Qwen3.8-27B-4bit`, `train/serve.sh` with the Ollama settings,
tool-call test passed, fixed subset `train/subset.json`, runner `train/baseline.py`, `train/README.md`.
Running: baseline (night chain). Not done: step 1 (data preparation, `train/prepare.py`); the valid loss of the
base model (step 2) needs step 1. Details and next steps: `train/STATE.md`.
## 5. Budget
- Cycle budget 60 USD until 12 October 2026. Ledger 63.3 USD (list prices) ≈ 43 USD real (ratio 1.47,
measured 2026-10-03). About 17 USD real left.
- Guard: `BUDGET_LIMIT_USD=75` (ledger) ≈ 51 USD real.
- Tonight: about 25 DeepSeek reruns (≈ 1.5 USD real) and judge calls for H tasks in the baseline (small).
The baseline itself is local and free.
- The agent now limits a cloud turn to 32k output tokens. Before, runaway reasoning (393k tokens) took 46 % of
the run cost.