Empirical filter runner (DeepSeek); real cost ratio in docs

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014aUaQeLnwbb1zTpN7kHeat
This commit is contained in:
Kral
2026-10-03 09:29:20 +02:00
parent 03ab36149e
commit d3250acfb8
3 changed files with 83 additions and 3 deletions

View File

@@ -45,7 +45,7 @@
- A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between - A4H is small: MCP limit is 8 sessions (by design). The server shares ONE RFC connection between
sessions; a parallel call gets "[LOCK] Concurrent call detected". `mcp_client.py` retries this. sessions; a parallel call gets "[LOCK] Concurrent call detected". `mcp_client.py` retries this.
Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs. Keep parallel runs low (max 3). DDIC activation during setup is sensitive to parallel runs.
- Budget: `runs/ledger.jsonl` (list prices, upper bound). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`. - Budget: `runs/ledger.jsonl` (list prices, upper bound; real cost ≈ ledger / 1.47, measured 2026-10-03). `.env`: `BUDGET_LIMIT_USD`, `BUDGET_CYCLE_START`.
Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work. Ollama usage resets on 12 October 2026, then +60 USD per month. At the limit, stop cloud work.
- Objects: package `$TMP` only. Prefix `Z` + run (4 chars base36) + task (3 chars base36) + `_` - Objects: package `$TMP` only. Prefix `Z` + run (4 chars base36) + task (3 chars base36) + `_`
(`harness/task.py`). Teardown after each run with the ADT deletion API (`adt_client.py`, credentials in `.env`). (`harness/task.py`). Teardown after each run with the ADT deletion API (`adt_client.py`, credentials in `.env`).

View File

@@ -93,7 +93,7 @@ Adım 3 görev ister; görevler el ile yazılırsa darboğaz olur. Bu yüzden il
## Bütçe takvimi (Ollama cloud, 2026-10-02) ## Bütçe takvimi (Ollama cloud, 2026-10-02)
Kalan: ~50 $ (12 Ekim'e kadar). 12 Ekim'de +60 $, sonra her ay +60 $. Birim maliyet (ölçülen): DeepSeek koşusu ≈ 0,05–0,08 $; görev üretimi ≈ 0,13 $ / görev, 0,16 $ / kabul edilen görev (pilot, liste fiyatı). Yerel modeller (Qwen vb.) ücretsiz; gece koşuları. Dönem bütçesi 60 $. 2026-10-03 09:20: gerçek harcama 5,20 $ (Ollama hesabı), kalan 54,80 $ (12 Ekim'e kadar). Ledger aynı anda 7,64 $ gösteriyordu: liste fiyatı gerçek harcamanın ~1,47 katı (cache indirimi). Gerçek harcama ≈ ledger / 1,47. 12 Ekim'de +60 $, sonra her ay +60 $. Birim maliyet (ölçülen): DeepSeek koşusu ≈ 0,05–0,08 $; görev üretimi ≈ 0,13 $ / görev, 0,16 $ / kabul edilen görev (pilot, liste fiyatı). Yerel modeller (Qwen vb.) ücretsiz; gece koşuları.
| Dönem | Bütçe | İş | Tahmini harcama | | Dönem | Bütçe | İş | Tahmini harcama |
|---|---|---|---| |---|---|---|---|
@@ -106,7 +106,7 @@ Kalan: ~50 $ (12 Ekim'e kadar). 12 Ekim'de +60 $, sonra her ay +60 $. Birim mali
| **12 Kasım sonrası** | 60 $/ay + ek kredi | Adım 3'ün tamamı, ardından Adım 4 (SFT) | | | **12 Kasım sonrası** | 60 $/ay + ek kredi | Adım 3'ün tamamı, ardından Adım 4 (SFT) | |
Kurallar: Kurallar:
- Harness her bulut isteğinin token kullanımını `runs/ledger.jsonl`'e yazar (liste fiyatı, cache indirimi olmadan = üst sınır). Dönem harcaması `BUDGET_LIMIT_USD`'ye (şu an 45 $) ulaşınca yeni üretim veya bulut koşusu başlamaz. ✓ - Harness her bulut isteğinin token kullanımını `runs/ledger.jsonl`'e yazar (liste fiyatı, cache indirimi olmadan = üst sınır). Dönem harcaması `BUDGET_LIMIT_USD`'ye (şu an 45 $, ledger cinsinden ≈ 30 $ gerçek) ulaşınca yeni üretim veya bulut koşusu başlamaz. ✓
- Eval görevleri ve eğitim görevleri ayrı havuzlar; eval görevleri eğitim verisine girmez. - Eval görevleri ve eğitim görevleri ayrı havuzlar; eval görevleri eğitim verisine girmez.
- Bulut ve yerel model aynı Ollama sunucusunda kuyrukta bekleyebilir; ölçümler sırasında ikisi aynı anda koşturulmaz. - Bulut ve yerel model aynı Ollama sunucusunda kuyrukta bekleyebilir; ölçümler sırasında ikisi aynı anda koşturulmaz.

80
harness/empirical.py Normal file
View File

@@ -0,0 +1,80 @@
"""Empirical filter (step F, second layer): run accepted eval tasks with real models.
A task where a strong model fails although the reference passes can have an unclear spec (flag it).
A task that every model passes with ease does not separate models (candidate for removal).
python3 -m harness.empirical --model deepseek-v4.1-flash:cloud --run-base 10000 [G0100 ...]
Results: runs/emp/results.jsonl and <task>/empirical.json (one entry per model). Resumable: a task that
already has a result for the model is skipped. Tasks with review decision "reject" are skipped.
"""
import argparse
import glob
import json
import os
from .adt_client import load_env
from .agents import LlmAgent
from .ledger import BudgetExceeded
from .runner import Runner
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
POOL = os.path.join(ROOT, "tasks_gen", "eval")
def _json(path, default):
try:
return json.load(open(path))
except (OSError, ValueError):
return default
def candidates():
out = []
for d in sorted(glob.glob(os.path.join(POOL, "G*"))):
if not _json(os.path.join(d, "generation.json"), {}).get("accepted"):
continue
if _json(os.path.join(d, "review.json"), {}).get("decision") == "reject":
continue
out.append(os.path.basename(d))
return out
def main():
load_env(os.path.join(ROOT, ".env"))
ap = argparse.ArgumentParser()
ap.add_argument("tasks", nargs="*")
ap.add_argument("--model", required=True)
ap.add_argument("--base-url")
ap.add_argument("--run-base", type=int, required=True)
a = ap.parse_args()
runs_root = os.path.join(ROOT, "runs", "emp")
os.makedirs(runs_root, exist_ok=True)
runner = Runner(POOL, runs_root)
tasks = a.tasks or candidates()
for i, tid in enumerate(tasks):
path = os.path.join(POOL, tid, "empirical.json")
res = _json(path, {})
if a.model in res:
continue
try:
rep, run_dir = runner.run(tid, LlmAgent(a.model, a.base_url), a.run_base + i)
except BudgetExceeded as e:
print("BUDGET", e, flush=True)
break
except Exception as e: # noqa: BLE001 one broken run must not stop the series
print(json.dumps({"task": tid, "error": str(e)[:300]}), flush=True)
continue
h = rep.get("hidden_tests") or {}
entry = {"score": (rep.get("score") or {}).get("total"), "parts": rep.get("score"),
"hidden": f"{h.get('passed')}/{h.get('total')}", "tool_calls": rep.get("tool_calls"),
"seconds": rep.get("seconds"), "run_dir": os.path.relpath(run_dir, ROOT)}
res[a.model] = entry
json.dump(res, open(path, "w"), indent=1)
line = dict(task=tid, model=a.model, **{k: entry[k] for k in ("score", "hidden", "tool_calls")})
open(os.path.join(runs_root, "results.jsonl"), "a").write(json.dumps(line) + "\n")
print(json.dumps(line), flush=True)
if __name__ == "__main__":
main()