Trajectory runs: tool-call budget 100 for tasks with a CDS contract object (eval keeps 60)

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-05 14:08:33 +02:00
parent 13360292dd
commit c062df2ec5
4 changed files with 12 additions and 3 deletions

View File

@@ -140,3 +140,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
round-trip check of every tool call; assistant spans for loss masking). B2d filter: `train/accept.py`.
- Test: scripted fake model on T01 (A4H, no cloud): score 100, 1 syntax hint, record, filter and converter OK.
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.