Trajectory runs: tool-call budget 100 for tasks with a CDS contract object (eval keeps 60)
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -140,3 +140,5 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
|
||||
round-trip check of every tool call; assistant spans for loss masking). B2d filter: `train/accept.py`.
|
||||
- Test: scripted fake model on T01 (A4H, no cloud): score 100, 1 syntax hint, record, filter and converter OK.
|
||||
- B3 trajectory runner: `harness/trajectories.py` (6 workers, DeepSeek V4.1 Flash, output `runs/traj/`). Starts after generation.
|
||||
|
||||
- 2026-10-05 Kral: tool-call budget 100 (eval stays 60) for trajectory runs on tasks with a CDS contract object (`Runner.run(cds_calls=100)`, `trajectories.CDS_CALLS`). Reason: first 3 CDS runs were correct (hidden tests all passed) but ended at 60 calls or by an empty response while writing the own CDS test class (score 80, not accepted). `metadata.tool_budget` is in every record. Failed CDS tasks get the second attempt with the new budget.
|
||||
|
||||
Reference in New Issue
Block a user