Eval review results, easy candidates, rerun queue; handover and STATE updated
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
@@ -1,4 +1,4 @@
|
|||||||
# Handover notes (2026-10-03, 18:35)
|
# Handover notes (2026-10-03, 19:10)
|
||||||
|
|
||||||
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
|
Read this file, `CLAUDE.md`, `train/STATE.md` and `train/README.md` first. Merge the durable parts into
|
||||||
CLAUDE.md and the docs when they are done, then empty this file.
|
CLAUDE.md and the docs when they are done, then empty this file.
|
||||||
@@ -50,28 +50,29 @@ After the jobs:
|
|||||||
|
|
||||||
## 3. Pending follow-ups
|
## 3. Pending follow-ups
|
||||||
|
|
||||||
0. Allowed while the baseline runs (stored results only): items 2, 5 (review part), 6, 7, 8, and the G2 check
|
Session of 2026-10-03 (evening): item 0 done with stored results only; nothing was run. Details in
|
||||||
of stored FUNC sources (item 1, first part). Everything that needs a model run or SAP goes to
|
`docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
|
||||||
`runs/stage1/rerun_queue.txt`.
|
|
||||||
1. **Re-score FUNC tasks with the new G2** (runner now finds the `FUNCTION` statement after local classes,
|
Done:
|
||||||
G0107). With stored results: re-check G2 on the saved sources (`sources/` in the run directories). Where
|
- Item 1, first part: all current FUNC runs (20 tasks) have hidden tests that ran; no FUNC task has a failed G2.
|
||||||
G2 now passes, the hidden tests never ran: queue a rerun. G0107 is already queued in W7B.
|
G0107 (old G2 failure) is in W7B.
|
||||||
2. **K specs: lost output format.** G0181 lost the ALV (fixed, queued in W7B). The other K tasks were checked
|
- Item 2: K specs checked, done.
|
||||||
for "CL_SALV_TABLE", "ALV", "global test class", "local test": no loss. The K prompt now keeps the output form.
|
- Item 5, review part: the 10 unreviewed tasks are reviewed and `accept`. Low scores: G0108 model error,
|
||||||
3. **Tasks with a call budget below 60:** none left. 26 were raised to 60 calls / 15 activations
|
G0119 infrastructure (queue), G0162 queued (see queue), G0170/G0176 H: DeepSeek did not stop.
|
||||||
(G0155–G0177 except G0160/G0161/G0168, G0190, G0191, T01, T13, T14, T15). Their DeepSeek runs are queued
|
- Item 6: 10 random high-score tasks (seed 20261003) reviewed, all `accept`.
|
||||||
(W7A, W7B). The T01 test run used the old budget 40; the baseline reruns T01 with 60.
|
- Item 7: `tasks_gen/eval/easy_candidates.json` (39 tasks; DeepSeek >= 95 and <= 20 tool calls).
|
||||||
4. **Rerun G0119 as an infrastructure failure:** DeepSeek stopped after 3 empty turns ("Stopped: empty model
|
Candidates only; confirm with the Qwen baseline.
|
||||||
response"). Queued in `runs/stage1/rerun_queue.txt`.
|
- Item 8: written to `docs/faz1-tasarim.md` 11h and `docs/eval-inceleme.md`.
|
||||||
5. **Deep review, rest:** done: G0107 (harness G2 bug), G0108 (valid; CDS float literal truncation is a model
|
|
||||||
error), G0181 (K lost ALV, fixed), G0176 (budget 20, rerun queued), G0119 (model failure), G0170 (open).
|
Waiting for jobs (do these in the next session):
|
||||||
To do: re-check low scores after the reruns; review the 10 accepted but not reviewed tasks
|
- After "DeepSeek reruns ended": update the 6 W7B tasks in 11h (G0107 G0173 G0175 G0177 G0181 G0191) and
|
||||||
(G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189).
|
the category means; check G0107 G2 and G0181 (K, ALV) results; G0170 decision (stronger gap?).
|
||||||
6. **10 random high-score reviews** from G0100 and later (Kral decision 2): not done.
|
- After "Baseline (25 tasks) ended": B2 summary; confirm the easy marks with the Qwen scores; decide the
|
||||||
7. **Mark easy tasks** (Kral decision 3: mark, do not remove): not done. Rule proposal: DeepSeek 100 with few
|
Kral spot-check list (flagged tasks + one per category, ~10).
|
||||||
tool calls; confirm with the Qwen baseline.
|
- Rerun queue `runs/stage1/rerun_queue.txt` (G0119, G0162): run after the baseline. Cost: DeepSeek, about
|
||||||
8. Write the empirical filter results and the review into `docs/faz1-tasarim.md` (new section) and
|
0.05-0.08 USD real per run, so about 0.15 USD for two runs. Give the estimate again before starting.
|
||||||
`docs/eval-inceleme.md`.
|
- Proxy finding: `sap_inactive_objects` shows objects of other runs (no prefix filter); an inactive object
|
||||||
|
`Z0FFK001_IBAN_VALIDATOR` stays on A4H. Fix the proxy and clean up after the baseline (A4H change: ask Kral).
|
||||||
|
|
||||||
## 4. Stage 1 status
|
## 4. Stage 1 status
|
||||||
|
|
||||||
|
|||||||
@@ -28,3 +28,10 @@ Yardımcı: `python3 -m harness.review G0100` görevin spec'ini, sözleşmesini,
|
|||||||
- **ret:** 2, 3 veya 8'de büyük sorun; düzeltme görevi yeniden yazmak demek.
|
- **ret:** 2, 3 veya 8'de büyük sorun; düzeltme görevi yeniden yazmak demek.
|
||||||
|
|
||||||
Eval görevleri eğitim verisine girmez; Claude'un yaptığı düzeltmeler de sadece eval setinde kalır.
|
Eval görevleri eğitim verisine girmez; Claude'un yaptığı düzeltmeler de sadece eval setinde kalır.
|
||||||
|
|
||||||
|
## Durum (2026-10-03)
|
||||||
|
|
||||||
|
- 113 kabul edilen görev; 4 ret (G0009 G0013 G0016 G0023, yerine G0188–G0191). Hepsi incelendi (`harness.review --list`: incelenmemiş yok).
|
||||||
|
- Rastgele yüksek puan kontrolü (karar 2): 10 görev, hepsi kabul. Tohum 20261003.
|
||||||
|
- Kolay aday işareti: `tasks_gen/eval/easy_candidates.json` (39 görev; Qwen baseline ile doğrulanacak).
|
||||||
|
- Kral'a gidecek: işaretlenen görev yok; kategori başına 1 örnek seçimi Qwen baseline sonrası.
|
||||||
|
|||||||
@@ -620,6 +620,25 @@ Ayarlar (deneyle):
|
|||||||
|
|
||||||
Sonuç: T01, T13, T14, T15 ve kabul edilen 22 üretilmiş görevin hepsi geçti. Hayatta kalan mutantlar gerçek test boşluğu gösterir: `inner join`→`left outer join` (G0017, G0020: test verisinde eşleşmeyen satır yok), `sum(`→`max(` (G0019: grup başına tek satır). Süre: mutant başına ~10–50 s.
|
Sonuç: T01, T13, T14, T15 ve kabul edilen 22 üretilmiş görevin hepsi geçti. Hayatta kalan mutantlar gerçek test boşluğu gösterir: `inner join`→`left outer join` (G0017, G0020: test verisinde eşleşmeyen satır yok), `sum(`→`max(` (G0019: grup başına tek satır). Süre: mutant başına ~10–50 s.
|
||||||
|
|
||||||
|
## 11h. Ampirik süzgeç ve inceleme sonuçları (2026-10-03)
|
||||||
|
|
||||||
|
DeepSeek (bütçe 60 çağrı) 109 eval görevinde koşuyor, 103'ünün güncel sonucu var; sonuç `tasks_gen/eval/<id>/empirical.json` (eski koşular `_superseded`). Kalan 6 görev (G0107, G0173, G0175, G0177, G0181, G0191) W7B'de koşuyor, bu satırlar o bitince güncellenir.
|
||||||
|
|
||||||
|
Kategori ortalaması (güncel koşu): A 96.3, B 88.9, C 88.7, D 97.7, E 93.6, F 91.5, G 84.0, H 71.4, I 91.0, K 93.4.
|
||||||
|
|
||||||
|
Dikkat gerektiren görevler:
|
||||||
|
- G0108 (53.2, 5/11): model hatası, görev geçerli (CDS float sabiti kesilmesi).
|
||||||
|
- G0119 (49.8): altyapı hatası (3 boş tur); `rerun_queue.txt`.
|
||||||
|
- G0162 (0, 3 koşuda G1 düştü): 7.02 PROG, SELECT/katı SQL modu aktivasyon hataları; son koşu iki boş 32k-token turuyla bitti. Rerun kuyruğunda; yine G1 düşerse 7.02 kısıtı ile sistemin katı SQL modu çelişiyor mu bakılır.
|
||||||
|
- G0170, G0176 (H, 0): DeepSeek durmadı, varsayımla uyguladı. G0170 için boşluk (yuvarlama hedefi ve modu) sınırda kritik; G0176 boşluğu (hangi oturumlar sayılır) gerçek. Qwen baseline sonucu karar verir.
|
||||||
|
- G0172, G0174 (H, 100): durdu.
|
||||||
|
|
||||||
|
İnceleme (kontrol listesi, `docs/eval-inceleme.md`): kabul edilen 109 görevin tamamı incelendi. Son 10 (G0121 G0124 G0129 G0130 G0141 G0143 G0144 G0160 G0161 G0189) hepsi `accept`; küçük notlar: G0121 `5->6` mutantı (üst öncelik sınırı veri satırında) hayatta; G0161 TOTAL_COST eşitliğinde sıra tanımsız (G0014/G0015 gibi), spec'te boşluk, test verisinde eşitlik yok. Kural 6 (rastgele 10 yüksek puan, tohum 20261003, DeepSeek ≥ 95 olan 48 görevden): G0100 G0101 G0111 G0114 G0130 G0133 G0136 G0146 G0149 G0164 - sızıntı, ipucu, kural–test eşlemesi tamam, hepsi `accept`. G0111'de yalnız 2 geçerli mutant var (zayıf mutasyon tabanı, kabul sınırında).
|
||||||
|
|
||||||
|
Kolay görev işareti (karar 3: işaretle, çıkarma): kural = DeepSeek ≥ 95 ve ≤ 20 araç çağrısı, uygulama görevi. 39 aday, liste `tasks_gen/eval/easy_candidates.json` (C 9, E 7, F 5, A 4, G 4, K 4, D 3, I 2, B 1). Aday kalır; Qwen baseline (25 görev) ile doğrulanır, DeepSeek ile ikisi de yüksekse "kolay" olur.
|
||||||
|
|
||||||
|
Bulgu (proxy): `sap_inactive_objects` başka koşunun nesnesini gösteriyor (`Z0FFK001_IBAN_VALIDATOR`, G0176 koşusunda): ön ek filtresi bu araçta yok ve A4H'de etkin olmayan bir artık nesne duruyor. Proxy düzeltmesi ve temizlik model koşusu/A4H değişikliği, baseline sonrası.
|
||||||
|
|
||||||
## 12. Açık sorular
|
## 12. Açık sorular
|
||||||
|
|
||||||
Cevaplananlar (Adım 1.2):
|
Cevaplananlar (Adım 1.2):
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Stage 1 state
|
# Stage 1 state
|
||||||
|
|
||||||
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 18:35.
|
Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. Updated 2026-10-03 19:10.
|
||||||
|
|
||||||
## Done
|
## Done
|
||||||
|
|
||||||
@@ -41,3 +41,6 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
|||||||
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
|
- Harness changes that matter for stage 1 runs: ADT activation fallback for PROG/FUNC (EPOD bug), G2 finds
|
||||||
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
|
the FUNCTION statement after local classes, call budget floor 60, empty-turn retry in the agent
|
||||||
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).
|
(`max_tokens` only for cloud models; the MLX server limit is `--max-tokens 32768`).
|
||||||
|
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
|
||||||
|
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
|
||||||
|
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
|
||||||
|
|||||||
Reference in New Issue
Block a user