Restore faz1/yol-haritasi docs (overwritten by copy); stage1 docs; G0174 new gap; stage1 25-task subset
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
83
docs/stage1-step-prompts.md
Normal file
83
docs/stage1-step-prompts.md
Normal file
@@ -0,0 +1,83 @@
|
||||
# Stage 1: step prompts
|
||||
|
||||
Preparation (one time): copy `stage1-training-task.md` to
|
||||
`~/projects/abap-llm/harness/docs/stage1-training-task.md`.
|
||||
|
||||
Use Sonnet. Before each prompt: `/clear`. Each prompt is one session.
|
||||
|
||||
---
|
||||
|
||||
## A — Setup and data (A4H: no change needed)
|
||||
|
||||
```
|
||||
Read CLAUDE.md and docs/stage1-training-task.md. Do only Step 0 and Step 1.
|
||||
Create train/STATE.md: write what you did, the model id, the paths, the
|
||||
report of Step 1, and the next step. Commit. Then stop.
|
||||
```
|
||||
|
||||
## B1 — Baseline, start (A4H must run)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Do Step 2.
|
||||
Start the eval runs in the background with a log file and a macOS
|
||||
notification at the end. Measure the valid loss of the base model in the
|
||||
same background job, after the eval runs. Update train/STATE.md with the
|
||||
log path and the job command. Then stop. Do not poll.
|
||||
```
|
||||
|
||||
## B2 — Baseline, results (after the notification)
|
||||
|
||||
```
|
||||
Read CLAUDE.md and train/STATE.md. Read only the last 50 lines of the log.
|
||||
Write runs/stage1/baseline.json. Update train/STATE.md with the results.
|
||||
Commit. Show me a short summary. Then stop.
|
||||
```
|
||||
|
||||
## C — Training test (stop A4H first)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. A4H is
|
||||
stopped. Check that no other model is loaded. Do only the short test of
|
||||
Step 3 (20 iterations). Report peak memory, time per iteration, and the
|
||||
time estimate for the full run. If memory is not sufficient, give me
|
||||
options and do not change the settings. Update train/STATE.md. Then stop.
|
||||
```
|
||||
|
||||
## D1 — Training, start (after my approval)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Start the
|
||||
full training run of Step 3 in the background with a log file and a macOS
|
||||
notification at the end. Use the settings from train/config.yaml. Update
|
||||
train/STATE.md with the log path and the adapter path. Then stop. Do not
|
||||
poll.
|
||||
```
|
||||
|
||||
## D2 — Training, results (after the notification)
|
||||
|
||||
```
|
||||
Read CLAUDE.md and train/STATE.md. Read only the valid loss lines and the
|
||||
last 30 lines of the log. Select the adapter with the lowest valid loss.
|
||||
Update train/STATE.md. Commit. Show me the valid loss curve as a short
|
||||
table. Then stop.
|
||||
```
|
||||
|
||||
## E1 — Measure again, start (start A4H first)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. A4H runs.
|
||||
Do Step 4: serve the base model with the selected adapter, connect the
|
||||
harness llm agent to it (smallest change), and test it with 1 task. Then
|
||||
start the same eval runs as in Step 2 and the valid loss measurement in the
|
||||
background with a log file and a macOS notification at the end. Update
|
||||
train/STATE.md. Commit. Then stop. Do not poll.
|
||||
```
|
||||
|
||||
## E2 — Report (after the notification)
|
||||
|
||||
```
|
||||
Read CLAUDE.md, docs/stage1-training-task.md and train/STATE.md. Read only
|
||||
the result files and the last 50 lines of the log. Write
|
||||
runs/stage1/report.md as described in the task. Update train/STATE.md.
|
||||
Commit. Show me the report. Then stop.
|
||||
```
|
||||
Reference in New Issue
Block a user