Files
abap-llm/docs/foreign-objects-report.md

3.0 KiB

Other runs' objects in tool results: which runs are affected (2026-10-06)

Cause: the model sometimes names its own helper or test class ZCL_<run prefix>_... (prefix inside the name). The teardown looked only for names that start with the prefix, so such classes stayed in A4H and showed up in sap_inactive_objects, searches and sometimes in a read of the next runs. Fixed on 2026-10-06 (proxy hides them, teardown finds them, harness/sweep.py). Scan: train/foreign_scan.py (read only; result runs/analysis/foreign_scan.json). A run is "affected" when a tool result contains a name of another run; a "foreign read" is a read tool call on such an object.

group runs with foreign names in a tool result with a foreign read tools
official baseline Qwen (MacBook, 11 tasks) 11 0 0 -
Devstral baseline (dropped) 11 0 0 -
older Qwen baselines (archive) 8 0 0 -
smoke 1 0 0 -
eval empirical filter (DeepSeek on eval candidates) 178 25 4 sap_inactive_objects 23, sap_search_object 4, sap_pull_source 5
eval generation validation (oracle, null, mutants) 1344 0 0 -
early pilot runs 16 1 0 sap_inactive_objects 1
training trajectories (DeepSeek) 117 67 15 sap_inactive_objects 66, sap_search_object 16, sap_pull_source 31, sap_run_unit_test 7, sap_element_info 1, sap_check_object 1
series A (local Qwen) 20 3 0 sap_inactive_objects 2, sap_search_object 1

Reading

  • Official baseline (Qwen, 11 tasks, 2026-10-04): not affected (0 of 11; also not Devstral 0 of 11, the older Qwen baselines 0 of 8, smoke 0 of 1). At that time A4H had few leftovers, and none of the 11 baseline runs called sap_inactive_objects (checked: 0 calls). The baseline numbers stay as they are. Nothing was changed.
  • Eval generation (1344 validation runs: oracle, null, mutants): not affected (they call no list or search tool).
  • Eval empirical filter (DeepSeek on eval candidates): 25 of 178 runs saw foreign names, 4 read one (G0183, G0180, G0143, G0185; scores 80, 85, 85, 85; all four are accepted in the review). The empirical filter decides which eval candidates look too easy or too hard; whether the read changed a result there was not examined (a read costs a few calls, the scores are in the normal range). Not changed; if Opus wants to be strict: rerun these four tasks after the reset (4 runs).
  • Early pilot runs: 1 of 16 saw a foreign name in an inactive list.
  • Training trajectories (DeepSeek, 117 runs incl. rejected): 67 saw foreign names, 15 read one. Of the accepted ones, 9 had a foreign read: they are moved back to the pending pool (train/requeue_foreign.py, rows in runs/traj/summary_excluded.jsonl) and run again after the reset in the normal order. List results of the other accepted ones are scrubbed in the builder.
  • Series A (local Qwen): 3 of 20 saw foreign names (inactive list, one search), 0 reads. The scores (3 of 20) are not affected by a read; the series ran before the fix.