Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-04 18:51:40 +02:00
parent a2ba9e7b44
commit 5447874fd3
8 changed files with 196 additions and 29 deletions

View File

@@ -109,10 +109,12 @@ python3 -c "from harness.ledger import spent; print(spent())"
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
must fix, and he does not like Devstral. Devstral is not used further.
- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0.
`runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`.
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference.
- Devstral Small 2 tested and dropped (same settings, 11 tasks): mean **6.8 vs Qwen 15.8**; 1/11 vs 3/11 tasks above 0;
end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: `runs/archive/devstral/`.
The Devstral weights were deleted. No further base model tests.
- **Official Qwen baseline** (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook,
run base 22000): `runs/stage1/baseline.json` (= `baseline_qwen.json`). Comparison: `docs/stage1-baseline.md`.
Older partial Qwen runs (Mac mini, aborted or thinking on): `runs/archive/qwen38/`, not comparable.
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
## 7. Next steps

View File

@@ -115,7 +115,7 @@ Kurallar:
## Adım 2 — Baz model seçimi
Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Eski Qwen sonuçları `runs/archive/qwen38/`; tam Qwen baseline'ı MacBook'ta koşuyor (`baseline_qwen.json`).
Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral denendi ve elendi (11 görev, `runs/archive/devstral/`): ortalama 6,8 (Qwen 15,8); 1/11 görev 0'dan yüksek (Qwen 3/11); loop 6 (Qwen 7), tool_budget 3, report 2; onarım oranı 0,84. Devstral ağırlıkları silindi; baz model testi yok. Eski kısmi Qwen sonuçları `runs/archive/qwen38/` (karşılaştırılamaz).
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.

View File

@@ -10,17 +10,16 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
## Base model
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
(Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again).
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
cannot load them without a re-quantization.
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
- Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0,
loops 6 vs 7. Results: `runs/archive/devstral/`. Weights deleted. No more base model tests.
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
`train/serve.sh` serves Qwen (commit a2ba9e7).
## Serving (`train/serve.sh`)
@@ -29,17 +28,14 @@ training the same script with `--adapter-path`.
| Setting | Value |
|---|---|
| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent |
| Thinking | server default on (`--chat-template-args`); `train/baseline.py` sends `enable_thinking=false` per request |
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
| presence / repeat penalty | not set (neutral) |
| Output limit | server `--max-tokens 32768`; each request sends 16384 |
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`);
`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors.
Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report
(G0128, G0174 ended with `report`).
Tool calls: `mlx_lm` returns OpenAI `tool_calls` with JSON arguments (tool-call test passed).
## Eval subset
@@ -52,8 +48,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
| Setting | Value |
|---|---|
| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) |
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
| Thinking | off (`enable_thinking=false` per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (`runs/archive/qwen38/baseline_thinking_on.json`) |
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
@@ -68,12 +64,11 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
## Baseline
- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered).
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`.
- Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
(the model changes the source, but the changes do not remove the cause).
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`).
- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.
- Official baseline: Qwen 3.8 27B, `python3 train/baseline.py --label baseline_qwen --run-base 22000` (MacBook, thinking off).
Results `runs/stage1/baseline.json` (= `baseline_qwen.json`), run directories `runs/stage1/baseline_qwen/`.
Mean **15.8**; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1.
Repair rate 39/54 = 0.72.
- Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84.
Results in `runs/archive/devstral/`.
- Per run record has `end_reason`, `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
- The stage 1 training test (step 3) uses Qwen (`config_test.yaml`).

View File

@@ -26,7 +26,12 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
`test.jsonl` = copy of valid.
## Running (detached)
## Base model
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
## Running (detached) — historical, all ended
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
@@ -40,8 +45,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit.
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
iterations first.
@@ -56,3 +60,15 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started.
- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost.

25
train/config.yaml Normal file
View File

@@ -0,0 +1,25 @@
# Stage 1 full run (start values of docs/stage1-training-task.md). Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config.yaml
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
train: true
data: train/data
fine_tune_type: lora
num_layers: -1
lora_parameters:
rank: 16
scale: 20.0
dropout: 0.0
batch_size: 1
grad_checkpoint: true
iters: 748
learning_rate: 5.0e-5
lr_schedule:
name: cosine_decay
warmup: 30
arguments: [5.0e-5, 748, 0.0]
max_seq_length: 16384
steps_per_report: 10
steps_per_eval: 200
val_batches: -1
save_every: 200
adapter_path: train/adapters
seed: 20261003

25
train/config_test.yaml Normal file
View File

@@ -0,0 +1,25 @@
# Stage 1 short test: 20 iterations, same settings as config.yaml. Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config_test.yaml
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
train: true
data: train/data
fine_tune_type: lora
num_layers: -1
lora_parameters:
rank: 16
scale: 20.0
dropout: 0.0
batch_size: 1
grad_checkpoint: true
iters: 20
learning_rate: 5.0e-5
lr_schedule:
name: cosine_decay
warmup: 2
arguments: [5.0e-5, 20, 0.0]
max_seq_length: 16384
steps_per_report: 1
steps_per_eval: 20
val_batches: 3
save_every: 20
adapter_path: train/adapters_test
seed: 20261003

60
train/hf_train.py Normal file
View File

@@ -0,0 +1,60 @@
# /// script
# requires-python = ">=3.10"
# dependencies = ["unsloth", "datasets", "trl", "huggingface_hub"]
# ///
"""Stage 1 LoRA training with Unsloth on Hugging Face Jobs (1x A100 80GB).
Run (pipeline test, 10 steps):
hf jobs uv run --flavor a100-large --timeout 45m --secrets HF_TOKEN train/hf_train.py -- --max-steps 10
Full run: no --max-steps (2 epochs). Settings follow train/config.yaml; alpha is a proposal, see STATE.md.
"""
import argparse, json, os, time
ap = argparse.ArgumentParser()
ap.add_argument("--model", default="unsloth/Qwen3.8-27B-unsloth-bnb-4bit")
ap.add_argument("--data", default="erhankeseli/abap-stage1-data")
ap.add_argument("--out", default="erhankeseli/abap-stage1-adapter-test")
ap.add_argument("--max-steps", type=int, default=-1)
ap.add_argument("--epochs", type=float, default=2.0)
ap.add_argument("--rank", type=int, default=16)
ap.add_argument("--alpha", type=int, default=16)
ap.add_argument("--lr", type=float, default=5e-5)
ap.add_argument("--max-seq-length", type=int, default=16384)
a = ap.parse_args()
from unsloth import FastLanguageModel # noqa: E402 (import first)
from datasets import load_dataset # noqa: E402
from trl import SFTConfig, SFTTrainer # noqa: E402
tok_hf = os.environ["HF_TOKEN"]
model, tok = FastLanguageModel.from_pretrained(
a.model, max_seq_length=a.max_seq_length, load_in_4bit=True, token=tok_hf)
model = FastLanguageModel.get_peft_model(
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj",
"in_proj_qkv", "in_proj_z", "out_proj"],
use_gradient_checkpointing="unsloth", random_state=20261003)
ds = load_dataset(a.data, data_files={"train": "train.jsonl", "valid": "valid.jsonl"}, token=tok_hf)
cfg = SFTConfig(
output_dir="out", per_device_train_batch_size=1, per_device_eval_batch_size=1,
gradient_accumulation_steps=1, num_train_epochs=a.epochs, max_steps=a.max_steps,
learning_rate=a.lr, lr_scheduler_type="cosine", warmup_steps=min(30, max(1, a.max_steps // 3)) if a.max_steps > 0 else 30,
optim="adamw_8bit", weight_decay=0.0, logging_steps=1, eval_strategy="no", save_strategy="no",
max_length=a.max_seq_length, dataset_text_field="text", packing=False, seed=20261003, report_to="none")
tr = SFTTrainer(model=model, processing_class=tok, train_dataset=ds["train"], eval_dataset=ds["valid"], args=cfg)
t0 = time.time()
tr.train()
train_s = time.time() - t0
steps = tr.state.global_step
ev = tr.evaluate() # valid loss of the adapter, for the conversion check
info = {"steps": steps, "train_seconds": round(train_s, 1), "sec_per_step": round(train_s / max(steps, 1), 2),
"eval_loss": ev.get("eval_loss"), "rank": a.rank, "alpha": a.alpha, "lr": a.lr,
"peak_gpu_gb": round(__import__("torch").cuda.max_memory_allocated() / 2**30, 1)}
print("RESULT", json.dumps(info))
model.save_pretrained("adapter")
json.dump(info, open("adapter/job_result.json", "w"))
from huggingface_hub import HfApi # noqa: E402
HfApi(token=tok_hf).upload_folder(folder_path="adapter", repo_id=a.out, repo_type="model")
print("PUSHED", a.out)

44
train/peft_to_mlx.py Normal file
View File

@@ -0,0 +1,44 @@
"""Convert a PEFT LoRA adapter (Unsloth / HF) to the mlx-lm adapter format.
train/.venv/bin/python train/peft_to_mlx.py <peft_adapter_dir> <mlx_adapter_dir>
PEFT key : base_model.model.model.language_model.layers.N.<mod>.lora_A.weight shape (r, in)
base_model.model.model.language_model.layers.N.<mod>.lora_B.weight shape (out, r)
mlx key : language_model.model.layers.N.<mod>.lora_a shape (in, r)
language_model.model.layers.N.<mod>.lora_b shape (r, out)
Scale : PEFT scaling = alpha / r; mlx `scale` is used directly (y + scale * x @ A @ B), so scale = alpha / r.
mlx loads only the modules listed in `lora_parameters.keys` (relative to the layer); they are taken from the file.
"""
import json, re, sys
from pathlib import Path
import mlx.core as mx
src, dst = Path(sys.argv[1]), Path(sys.argv[2])
cfg = json.load(open(src / "adapter_config.json"))
r, alpha = cfg["r"], cfg["lora_alpha"]
if cfg.get("use_rslora") or cfg.get("use_dora"):
sys.exit("rsLoRA / DoRA: scale is not alpha / r, not supported")
w = mx.load(str(src / "adapter_model.safetensors"))
pat = re.compile(r"^base_model\.model\.model\.language_model\.layers\.(\d+)\.(.+)\.lora_([AB])\.weight$")
out, keys, layers = {}, set(), set()
for k, v in w.items():
m = pat.match(k)
if not m:
sys.exit(f"unexpected key: {k}")
n, mod, ab = m.groups()
layers.add(int(n))
keys.add(mod)
out[f"language_model.model.layers.{n}.{mod}.lora_{ab.lower()}"] = mx.transpose(v).astype(mx.float32)
if ab == "A":
assert v.shape[0] == r, (k, v.shape)
dst.mkdir(parents=True, exist_ok=True)
mx.save_safetensors(str(dst / "adapters.safetensors"), out)
json.dump({
"fine_tune_type": "lora",
"num_layers": max(layers) + 1,
"lora_parameters": {"rank": r, "scale": alpha / r, "dropout": 0.0, "keys": sorted(keys)},
}, open(dst / "adapter_config.json", "w"), indent=1)
print(f"{len(out)} tensors, {len(layers)} layers, rank {r}, alpha {alpha} -> scale {alpha / r}, modules {sorted(keys)}")