Eval review results, easy candidates, rerun queue; handover and STATE updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -1,4 +1,4 @@
|
||||
# Handover notes (2026-10-03, 18:35)
|
||||
# Handover notes (2026-10-03, 19:10)
|
||||
|
||||
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
|
||||
CLAUDE.md and the docs when they are done, then empty this file.
|
||||
@@ -50,28 +50,29 @@ After the jobs:
|
||||
|
||||
## 3. Pending follow-ups
|
||||
|
||||
0. Allowed while the baseline runs (stored results only): items 2, 5 (review part), 6, 7, 8, and the G2 check
|
||||
of stored FUNC sources (item 1, first part). Everything that needs a model run or SAP goes to
|
||||
`runs/stage1/rerun_queue.txt`.
|
||||
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
|
||||
G0107). With stored results: re-check G2 on the saved sources (`sources/` in the run directories). Where
|
||||
G2 now passes, the hidden tests never ran: queue a rerun. G0107 is already queued in W7B.
|
||||
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
|
||||
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
|
||||
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
|
||||
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
|
||||
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
|
||||
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
|
||||
response"). Queued in `runs/stage1/rerun_queue.txt`.
|
||||
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
|
||||
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
|
||||
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks
|
||||
(G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189).
|
||||
6. **10 random high-score reviews** from G0100 and later (Kral decision 2): not done.
|
||||
7. **Mark easy tasks** (Kral decision 3: mark, do not remove): not done. Rule proposal: DeepSeek 100 with few
|
||||
tool calls; confirm with the Qwen baseline.
|
||||
8. Write the empirical filter results and the review into `docs/faz1-tasarim.md` (new section) and
|
||||
`docs/eval-inceleme.md`.
|
||||
Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in
|
||||
`docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
|
||||
|
||||
Done:
|
||||
- Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2.
|
||||
G0107 (old G2 failure) is in W7B.
|
||||
- Item 2: K specs checked, done.
|
||||
- Item 5, review part: the 10 unreviewed tasks are reviewed and `accept`. Low scores: G0108 model error,
|
||||
G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop.
|
||||
- Item 6: 10 random high-score tasks (seed 20261003) reviewed, all `accept`.
|
||||
- Item 7: `tasks_gen/eval/easy_candidates.json` (39 tasks; DeepSeek >= 95 and <= 20 tool calls).
|
||||
Candidates only; confirm with the Qwen baseline.
|
||||
- Item 8: written to `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
|
||||
|
||||
Waiting for jobs (do these in the next session):
|
||||
- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and
|
||||
the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
|
||||
- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the
|
||||
Kral spot-check list (flagged tasks + one per category, ~10).
|
||||
- Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about
|
||||
0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting.
|
||||
- Proxy finding: `sap_inactive_objects` shows objects of other runs (no prefix filter); an inactive object
|
||||
`Z0FFK001_IBAN_VALIDATOR` stays on A4H. Fix the proxy and clean up after the baseline (A4H change: ask Kral).
|
||||
|
||||
## 4. Stage 1 status
|
||||
|
||||
|
||||
Reference in New Issue
Block a user