From 573707613e158a2409065f325aaa3f9fbec6f547 Mon Sep 17 00:00:00 2001 From: Kral Date: Sat, 3 Oct 2026 18:57:23 +0200 Subject: [PATCH] Eval review results, easy candidates, rerun queue; handover and STATE updated Co-Authored-By: Claude Sonnet 5.5 Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat --- docs/devir-notlari.md | 47 ++++++++++++++++++++++--------------------- docs/eval-inceleme.md | 7 +++++++ docs/faz1-tasarim.md | 19 +++++++++++++++++ train/STATE.md | 5 ++++- 4 files changed, 54 insertions(+), 24 deletions(-) diff --git a/docs/devir-notlari.md b/docs/devir-notlari.md index 1585782..ac90f93 100644 --- a/docs/devir-notlari.md +++ b/docs/devir-notlari.md @@ -1,4 +1,4 @@ -# Handover notes (2026-10-03, 18:35) +# Handover notes (2026-10-03, 19:10) Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into CLAUDE.md and the docs when they are done, then empty this file. @@ -50,28 +50,29 @@ After the jobs: ## 3. Pending follow-ups -0. Allowed while the baseline runs (stored results only): items 2, 5 (review part), 6, 7, 8, and the G2 check - of stored FUNC sources (item 1, first part). Everything that needs a model run or SAP goes to - `runs/stage1/rerun_queue.txt`. -1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes, - G0107). With stored results: re-check G2 on the saved sources (`sources/` in the run directories). Where - G2 now passes, the hidden tests never ran: queue a rerun. G0107 is already queued in W7B. -2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked - for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form. -3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations - (G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued - (W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60. -4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model - response"). Queued in `runs/stage1/rerun_queue.txt`. -5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model - error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open). - To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks - (G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189). -6. **10 random high-score reviews** from G0100 and later (Kral decision 2): not done. -7. **Mark easy tasks** (Kral decision 3: mark, do not remove): not done. Rule proposal: DeepSeek 100 with few - tool calls; confirm with the Qwen baseline. -8. Write the empirical filter results and the review into `docs/faz1-tasarim.md` (new section) and - `docs/eval-inceleme.md`. +Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in +`docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`. + +Done: +- Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2. + G0107 (old G2 failure) is in W7B. +- Item 2: K specs checked, done. +- Item 5, review part: the 10 unreviewed tasks are reviewed and `accept`. Low scores: G0108 model error, + G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop. +- Item 6: 10 random high-score tasks (seed 20261003) reviewed, all `accept`. +- Item 7: `tasks_gen/eval/easy_candidates.json` (39 tasks; DeepSeek >= 95 and <= 20 tool calls). + Candidates only; confirm with the Qwen baseline. +- Item 8: written to `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`. + +Waiting for jobs (do these in the next session): +- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and + the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?). +- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the + Kral spot-check list (flagged tasks + one per category, ~10). +- Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about + 0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting. +- Proxy finding: `sap_inactive_objects` shows objects of other runs (no prefix filter); an inactive object + `Z0FFK001_IBAN_VALIDATOR` stays on A4H. Fix the proxy and clean up after the baseline (A4H change: ask Kral). ## 4. Stage 1 status diff --git a/docs/eval-inceleme.md b/docs/eval-inceleme.md index 8fa4d30..fe70a46 100644 --- a/docs/eval-inceleme.md +++ b/docs/eval-inceleme.md @@ -28,3 +28,10 @@ Yardımcı: `python3 -m harness.review G0100` görevin spec'ini, sözleşmesini, - **ret:** 2, 3 veya 8'de büyük sorun; düzeltme görevi yeniden yazmak demek. Eval görevleri eğitim verisine girmez; Claude'un yaptığı düzeltmeler de sadece eval setinde kalır. + +## Durum (2026-10-03) + +- 113 kabul edilen görev; 4 ret (G0009 G0013 G0016 G0023, yerine G0188–G0191). Hepsi incelendi (`harness.review --list`: incelenmemiş yok). +- Rastgele yüksek puan kontrolü (karar 2): 10 görev, hepsi kabul. Tohum 20261003. +- Kolay aday işareti: `tasks_gen/eval/easy_candidates.json` (39 görev; Qwen baseline ile doğrulanacak). +- Kral'a gidecek: işaretlenen görev yok; kategori başına 1 örnek seçimi Qwen baseline sonrası. diff --git a/docs/faz1-tasarim.md b/docs/faz1-tasarim.md index 7d338b7..0c7cefa 100644 --- a/docs/faz1-tasarim.md +++ b/docs/faz1-tasarim.md @@ -620,6 +620,25 @@ Ayarlar (deneyle): Sonuç: T01, T13, T14, T15 ve kabul edilen 22 üretilmiş görevin hepsi geçti. Hayatta kalan mutantlar gerçek test boşluğu gösterir: `inner join`→`left outer join` (G0017, G0020: test verisinde eşleşmeyen satır yok), `sum(`→`max(` (G0019: grup başına tek satır). Süre: mutant başına ~10–50 s. +## 11h. Ampirik süzgeç ve inceleme sonuçları (2026-10-03) + +DeepSeek (bütçe 60 çağrı) 109 eval görevinde koşuyor, 103'ünün güncel sonucu var; sonuç `tasks_gen/eval//empirical.json` (eski koşular `_superseded`). Kalan 6 görev (G0107, G0173, G0175, G0177, G0181, G0191) W7B'de koşuyor, bu satırlar o bitince güncellenir. + +Kategori ortalaması (güncel koşu): A 96.3, B 88.9, C 88.7, D 97.7, E 93.6, F 91.5, G 84.0, H 71.4, I 91.0, K 93.4. + +Dikkat gerektiren görevler: +- G0108 (53.2, 5/11): model hatası, görev geçerli (CDS float sabiti kesilmesi). +- G0119 (49.8): altyapı hatası (3 boş tur); `rerun_queue.txt`. +- G0162 (0, 3 koşuda G1 düştü): 7.02 PROG, SELECT/katı SQL modu aktivasyon hataları; son koşu iki boş 32k-token turuyla bitti. Rerun kuyruğunda; yine G1 düşerse 7.02 kısıtı ile sistemin katı SQL modu çelişiyor mu bakılır. +- G0170, G0176 (H, 0): DeepSeek durmadı, varsayımla uyguladı. G0170 için boşluk (yuvarlama hedefi ve modu) sınırda kritik; G0176 boşluğu (hangi oturumlar sayılır) gerçek. Qwen baseline sonucu karar verir. +- G0172, G0174 (H, 100): durdu. + +İnceleme (kontrol listesi, `docs/eval-inceleme.md`): kabul edilen 109 görevin tamamı incelendi. Son 10 (G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189) hepsi `accept`; küçük notlar: G0121 `5->6` mutantı (üst öncelik sınırı veri satırında) hayatta; G0161 TOTAL_COST eşitliğinde sıra tanımsız (G0014/G0015 gibi), spec'te boşluk, test verisinde eşitlik yok. Kural 6 (rastgele 10 yüksek puan, tohum 20261003, DeepSeek ≥ 95 olan 48 görevden): G0100 G0101 G0111 G0114 G0130 G0133 G0136 G0146 G0149 G0164 - sızıntı, ipucu, kural–test eşlemesi tamam, hepsi `accept`. G0111'de yalnız 2 geçerli mutant var (zayıf mutasyon tabanı, kabul sınırında). + +Kolay görev işareti (karar 3: işaretle, çıkarma): kural = DeepSeek ≥ 95 ve ≤ 20 araç çağrısı, uygulama görevi. 39 aday, liste `tasks_gen/eval/easy_candidates.json` (C 9, E 7, F 5, A 4, G 4, K 4, D 3, I 2, B 1). Aday kalır; Qwen baseline (25 görev) ile doğrulanır, DeepSeek ile ikisi de yüksekse "kolay" olur. + +Bulgu (proxy): `sap_inactive_objects` başka koşunun nesnesini gösteriyor (`Z0FFK001_IBAN_VALIDATOR`, G0176 koşusunda): ön ek filtresi bu araçta yok ve A4H'de etkin olmayan bir artık nesne duruyor. Proxy düzeltmesi ve temizlik model koşusu/A4H değişikliği, baseline sonrası. + ## 12. Açık sorular Cevaplananlar (Adım 1.2): diff --git a/train/STATE.md b/train/STATE.md index bc67f6f..17b5e8c 100644 --- a/train/STATE.md +++ b/train/STATE.md @@ -1,6 +1,6 @@ # Stage 1 state -Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 18:35. +Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 19:10. ## Done @@ -41,3 +41,6 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U - Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent (`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`). +- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results + (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in + `runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.