Eval review results, easy candidates, rerun queue; handover and STATE updated

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 18:57:23 +02:00
parent fb18dc8d9d
commit 573707613e
4 changed files with 54 additions and 24 deletions

View File

@@ -1,4 +1,4 @@
# Handover notes (2026-10-03, 18:35)
# Handover notes (2026-10-03, 19:10)
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
CLAUDE.md and the docs when they are done, then empty this file.
@@ -50,28 +50,29 @@ After the jobs:
## 3. Pending follow-ups
0. Allowed while the baseline runs (stored results only): items 2, 5 (review part), 6, 7, 8, and the G2 check
of stored FUNC sources (item 1, first part). Everything that needs a model run or SAP goes to
`runs/stage1/rerun_queue.txt`.
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
G0107). With stored results: re-check G2 on the saved sources (`sources/` in the run directories). Where
G2 now passes, the hidden tests never ran: queue a rerun. G0107 is already queued in W7B.
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
response"). Queued in `runs/stage1/rerun_queue.txt`.
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks
(G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189).
6. **10 random high-score reviews** from G0100 and later (Kral decision 2): not done.
7. **Mark easy tasks** (Kral decision 3: mark, do not remove): not done. Rule proposal: DeepSeek 100 with few
tool calls; confirm with the Qwen baseline.
8. Write the empirical filter results and the review into `docs/faz1-tasarim.md` (new section) and
`docs/eval-inceleme.md`.
Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in
`docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
Done:
- Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2.
G0107 (old G2 failure) is in W7B.
- Item 2: K specs checked, done.
- Item 5, review part: the 10 unreviewed tasks are reviewed and `accept`. Low scores: G0108 model error,
G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop.
- Item 6: 10 random high-score tasks (seed 20261003) reviewed, all `accept`.
- Item 7: `tasks_gen/eval/easy_candidates.json` (39 tasks; DeepSeek >= 95 and <= 20 tool calls).
Candidates only; confirm with the Qwen baseline.
- Item 8: written to `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
Waiting for jobs (do these in the next session):
- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and
the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the
Kral spot-check list (flagged tasks + one per category, ~10).
- Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about
0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting.
- Proxy finding: `sap_inactive_objects` shows objects of other runs (no prefix filter); an inactive object
`Z0FFK001_IBAN_VALIDATOR` stays on A4H. Fix the proxy and clean up after the baseline (A4H change: ask Kral).
## 4. Stage 1 status

View File

@@ -28,3 +28,10 @@ Yardımcı: `python3 -m harness.review G0100` görevin spec'ini, sözleşmesini,
- **ret:** 2, 3 veya 8'de büyük sorun; düzeltme görevi yeniden yazmak demek.
Eval görevleri eğitim verisine girmez; Claude'un yaptığı düzeltmeler de sadece eval setinde kalır.
## Durum (2026-10-03)
- 113 kabul edilen görev; 4 ret (G0009 G0013 G0016 G0023, yerine G0188–G0191). Hepsi incelendi (`harness.review --list`: incelenmemiş yok).
- Rastgele yüksek puan kontrolü (karar 2): 10 görev, hepsi kabul. Tohum 20261003.
- Kolay aday işareti: `tasks_gen/eval/easy_candidates.json` (39 görev; Qwen baseline ile doğrulanacak).
- Kral'a gidecek: işaretlenen görev yok; kategori başına 1 örnek seçimi Qwen baseline sonrası.

View File

@@ -620,6 +620,25 @@ Ayarlar (deneyle):
Sonuç: T01, T13, T14, T15 ve kabul edilen 22 üretilmiş görevin hepsi geçti. Hayatta kalan mutantlar gerçek test boşluğu gösterir: `inner join`→`left outer join` (G0017, G0020: test verisinde eşleşmeyen satır yok), `sum(`→`max(` (G0019: grup başına tek satır). Süre: mutant başına ~10–50 s.
## 11h. Ampirik süzgeç ve inceleme sonuçları (2026-10-03)
DeepSeek (bütçe 60 çağrı) 109 eval görevinde koşuyor, 103'ünün güncel sonucu var; sonuç `tasks_gen/eval/<id>/empirical.json` (eski koşular `_superseded`). Kalan 6 görev (G0107, G0173, G0175, G0177, G0181, G0191) W7B'de koşuyor, bu satırlar o bitince güncellenir.
Kategori ortalaması (güncel koşu): A 96.3, B 88.9, C 88.7, D 97.7, E 93.6, F 91.5, G 84.0, H 71.4, I 91.0, K 93.4.
Dikkat gerektiren görevler:
- G0108 (53.2, 5/11): model hatası, görev geçerli (CDS float sabiti kesilmesi).
- G0119 (49.8): altyapı hatası (3 boş tur); `rerun_queue.txt`.
- G0162 (0, 3 koşuda G1 düştü): 7.02 PROG, SELECT/katı SQL modu aktivasyon hataları; son koşu iki boş 32k-token turuyla bitti. Rerun kuyruğunda; yine G1 düşerse 7.02 kısıtı ile sistemin katı SQL modu çelişiyor mu bakılır.
- G0170, G0176 (H, 0): DeepSeek durmadı, varsayımla uyguladı. G0170 için boşluk (yuvarlama hedefi ve modu) sınırda kritik; G0176 boşluğu (hangi oturumlar sayılır) gerçek. Qwen baseline sonucu karar verir.
- G0172, G0174 (H, 100): durdu.
İnceleme (kontrol listesi, `docs/eval-inceleme.md`): kabul edilen 109 görevin tamamı incelendi. Son 10 (G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189) hepsi `accept`; küçük notlar: G0121 `5->6` mutantı (üst öncelik sınırı veri satırında) hayatta; G0161 TOTAL_COST eşitliğinde sıra tanımsız (G0014/G0015 gibi), spec'te boşluk, test verisinde eşitlik yok. Kural 6 (rastgele 10 yüksek puan, tohum 20261003, DeepSeek ≥ 95 olan 48 görevden): G0100 G0101 G0111 G0114 G0130 G0133 G0136 G0146 G0149 G0164 - sızıntı, ipucu, kural–test eşlemesi tamam, hepsi `accept`. G0111'de yalnız 2 geçerli mutant var (zayıf mutasyon tabanı, kabul sınırında).
Kolay görev işareti (karar 3: işaretle, çıkarma): kural = DeepSeek ≥ 95 ve ≤ 20 araç çağrısı, uygulama görevi. 39 aday, liste `tasks_gen/eval/easy_candidates.json` (C 9, E 7, F 5, A 4, G 4, K 4, D 3, I 2, B 1). Aday kalır; Qwen baseline (25 görev) ile doğrulanır, DeepSeek ile ikisi de yüksekse "kolay" olur.
Bulgu (proxy): `sap_inactive_objects` başka koşunun nesnesini gösteriyor (`Z0FFK001_IBAN_VALIDATOR`, G0176 koşusunda): ön ek filtresi bu araçta yok ve A4H'de etkin olmayan bir artık nesne duruyor. Proxy düzeltmesi ve temizlik model koşusu/A4H değişikliği, baseline sonrası.
## 12. Açık sorular
Cevaplananlar (Adım 1.2):