Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -10,17 +10,16 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
|
||||
## Base model
|
||||
|
||||
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
|
||||
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
|
||||
(Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again).
|
||||
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
|
||||
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
|
||||
cannot load them without a re-quantization.
|
||||
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
|
||||
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
|
||||
- Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0,
|
||||
loops 6 vs 7. Results: `runs/archive/devstral/`. Weights deleted. No more base model tests.
|
||||
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
|
||||
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
|
||||
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
|
||||
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
|
||||
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
|
||||
`train/serve.sh` serves Qwen (commit a2ba9e7).
|
||||
|
||||
## Serving (`train/serve.sh`)
|
||||
|
||||
@@ -29,17 +28,14 @@ training the same script with `--adapter-path`.
|
||||
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent |
|
||||
| Thinking | server default on (`--chat-template-args`); `train/baseline.py` sends `enable_thinking=false` per request |
|
||||
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
|
||||
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
|
||||
| presence / repeat penalty | not set (neutral) |
|
||||
| Output limit | server `--max-tokens 32768`; each request sends 16384 |
|
||||
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
|
||||
|
||||
Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`);
|
||||
`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors.
|
||||
Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report
|
||||
(G0128, G0174 ended with `report`).
|
||||
Tool calls: `mlx_lm` returns OpenAI `tool_calls` with JSON arguments (tool-call test passed).
|
||||
|
||||
## Eval subset
|
||||
|
||||
@@ -52,8 +48,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
|
||||
|
||||
| Setting | Value |
|
||||
|---|---|
|
||||
| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||
| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) |
|
||||
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||
| Thinking | off (`enable_thinking=false` per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (`runs/archive/qwen38/baseline_thinking_on.json`) |
|
||||
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
|
||||
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
|
||||
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
|
||||
@@ -68,12 +64,11 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
|
||||
|
||||
## Baseline
|
||||
|
||||
- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by
|
||||
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered).
|
||||
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories
|
||||
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`.
|
||||
- Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
|
||||
(the model changes the source, but the changes do not remove the cause).
|
||||
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
|
||||
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`).
|
||||
- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.
|
||||
- Official baseline: Qwen 3.8 27B, `python3 train/baseline.py --label baseline_qwen --run-base 22000` (MacBook, thinking off).
|
||||
Results `runs/stage1/baseline.json` (= `baseline_qwen.json`), run directories `runs/stage1/baseline_qwen/`.
|
||||
Mean **15.8**; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1.
|
||||
Repair rate 39/54 = 0.72.
|
||||
- Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84.
|
||||
Results in `runs/archive/devstral/`.
|
||||
- Per run record has `end_reason`, `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
|
||||
- The stage 1 training test (step 3) uses Qwen (`config_test.yaml`).
|
||||
|
||||
@@ -26,7 +26,12 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
|
||||
`test.jsonl` = copy of valid.
|
||||
|
||||
## Running (detached)
|
||||
## Base model
|
||||
|
||||
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
|
||||
|
||||
## Running (detached) — historical, all ended
|
||||
|
||||
|
||||
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
|
||||
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
|
||||
@@ -40,8 +45,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
|
||||
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
||||
(`t01_test_budget40` holds the T01 test result). Commit.
|
||||
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
|
||||
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
||||
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
||||
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
||||
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
||||
iterations first.
|
||||
@@ -56,3 +60,15 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
||||
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
|
||||
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
|
||||
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
|
||||
|
||||
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
|
||||
|
||||
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
|
||||
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
|
||||
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
|
||||
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
|
||||
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
|
||||
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
|
||||
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
|
||||
- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started.
|
||||
- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost.
|
||||
|
||||
25
train/config.yaml
Normal file
25
train/config.yaml
Normal file
@@ -0,0 +1,25 @@
|
||||
# Stage 1 full run (start values of docs/stage1-training-task.md). Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config.yaml
|
||||
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
|
||||
train: true
|
||||
data: train/data
|
||||
fine_tune_type: lora
|
||||
num_layers: -1
|
||||
lora_parameters:
|
||||
rank: 16
|
||||
scale: 20.0
|
||||
dropout: 0.0
|
||||
batch_size: 1
|
||||
grad_checkpoint: true
|
||||
iters: 748
|
||||
learning_rate: 5.0e-5
|
||||
lr_schedule:
|
||||
name: cosine_decay
|
||||
warmup: 30
|
||||
arguments: [5.0e-5, 748, 0.0]
|
||||
max_seq_length: 16384
|
||||
steps_per_report: 10
|
||||
steps_per_eval: 200
|
||||
val_batches: -1
|
||||
save_every: 200
|
||||
adapter_path: train/adapters
|
||||
seed: 20261003
|
||||
25
train/config_test.yaml
Normal file
25
train/config_test.yaml
Normal file
@@ -0,0 +1,25 @@
|
||||
# Stage 1 short test: 20 iterations, same settings as config.yaml. Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config_test.yaml
|
||||
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
|
||||
train: true
|
||||
data: train/data
|
||||
fine_tune_type: lora
|
||||
num_layers: -1
|
||||
lora_parameters:
|
||||
rank: 16
|
||||
scale: 20.0
|
||||
dropout: 0.0
|
||||
batch_size: 1
|
||||
grad_checkpoint: true
|
||||
iters: 20
|
||||
learning_rate: 5.0e-5
|
||||
lr_schedule:
|
||||
name: cosine_decay
|
||||
warmup: 2
|
||||
arguments: [5.0e-5, 20, 0.0]
|
||||
max_seq_length: 16384
|
||||
steps_per_report: 1
|
||||
steps_per_eval: 20
|
||||
val_batches: 3
|
||||
save_every: 20
|
||||
adapter_path: train/adapters_test
|
||||
seed: 20261003
|
||||
60
train/hf_train.py
Normal file
60
train/hf_train.py
Normal file
@@ -0,0 +1,60 @@
|
||||
# /// script
|
||||
# requires-python = ">=3.10"
|
||||
# dependencies = ["unsloth", "datasets", "trl", "huggingface_hub"]
|
||||
# ///
|
||||
"""Stage 1 LoRA training with Unsloth on Hugging Face Jobs (1x A100 80GB).
|
||||
|
||||
Run (pipeline test, 10 steps):
|
||||
hf jobs uv run --flavor a100-large --timeout 45m --secrets HF_TOKEN train/hf_train.py -- --max-steps 10
|
||||
Full run: no --max-steps (2 epochs). Settings follow train/config.yaml; alpha is a proposal, see STATE.md.
|
||||
"""
|
||||
import argparse, json, os, time
|
||||
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--model", default="unsloth/Qwen3.8-27B-unsloth-bnb-4bit")
|
||||
ap.add_argument("--data", default="erhankeseli/abap-stage1-data")
|
||||
ap.add_argument("--out", default="erhankeseli/abap-stage1-adapter-test")
|
||||
ap.add_argument("--max-steps", type=int, default=-1)
|
||||
ap.add_argument("--epochs", type=float, default=2.0)
|
||||
ap.add_argument("--rank", type=int, default=16)
|
||||
ap.add_argument("--alpha", type=int, default=16)
|
||||
ap.add_argument("--lr", type=float, default=5e-5)
|
||||
ap.add_argument("--max-seq-length", type=int, default=16384)
|
||||
a = ap.parse_args()
|
||||
|
||||
from unsloth import FastLanguageModel # noqa: E402 (import first)
|
||||
from datasets import load_dataset # noqa: E402
|
||||
from trl import SFTConfig, SFTTrainer # noqa: E402
|
||||
|
||||
tok_hf = os.environ["HF_TOKEN"]
|
||||
model, tok = FastLanguageModel.from_pretrained(
|
||||
a.model, max_seq_length=a.max_seq_length, load_in_4bit=True, token=tok_hf)
|
||||
model = FastLanguageModel.get_peft_model(
|
||||
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
|
||||
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj",
|
||||
"in_proj_qkv", "in_proj_z", "out_proj"],
|
||||
use_gradient_checkpointing="unsloth", random_state=20261003)
|
||||
|
||||
ds = load_dataset(a.data, data_files={"train": "train.jsonl", "valid": "valid.jsonl"}, token=tok_hf)
|
||||
cfg = SFTConfig(
|
||||
output_dir="out", per_device_train_batch_size=1, per_device_eval_batch_size=1,
|
||||
gradient_accumulation_steps=1, num_train_epochs=a.epochs, max_steps=a.max_steps,
|
||||
learning_rate=a.lr, lr_scheduler_type="cosine", warmup_steps=min(30, max(1, a.max_steps // 3)) if a.max_steps > 0 else 30,
|
||||
optim="adamw_8bit", weight_decay=0.0, logging_steps=1, eval_strategy="no", save_strategy="no",
|
||||
max_length=a.max_seq_length, dataset_text_field="text", packing=False, seed=20261003, report_to="none")
|
||||
tr = SFTTrainer(model=model, processing_class=tok, train_dataset=ds["train"], eval_dataset=ds["valid"], args=cfg)
|
||||
|
||||
t0 = time.time()
|
||||
tr.train()
|
||||
train_s = time.time() - t0
|
||||
steps = tr.state.global_step
|
||||
ev = tr.evaluate() # valid loss of the adapter, for the conversion check
|
||||
info = {"steps": steps, "train_seconds": round(train_s, 1), "sec_per_step": round(train_s / max(steps, 1), 2),
|
||||
"eval_loss": ev.get("eval_loss"), "rank": a.rank, "alpha": a.alpha, "lr": a.lr,
|
||||
"peak_gpu_gb": round(__import__("torch").cuda.max_memory_allocated() / 2**30, 1)}
|
||||
print("RESULT", json.dumps(info))
|
||||
model.save_pretrained("adapter")
|
||||
json.dump(info, open("adapter/job_result.json", "w"))
|
||||
from huggingface_hub import HfApi # noqa: E402
|
||||
HfApi(token=tok_hf).upload_folder(folder_path="adapter", repo_id=a.out, repo_type="model")
|
||||
print("PUSHED", a.out)
|
||||
44
train/peft_to_mlx.py
Normal file
44
train/peft_to_mlx.py
Normal file
@@ -0,0 +1,44 @@
|
||||
"""Convert a PEFT LoRA adapter (Unsloth / HF) to the mlx-lm adapter format.
|
||||
|
||||
train/.venv/bin/python train/peft_to_mlx.py <peft_adapter_dir> <mlx_adapter_dir>
|
||||
|
||||
PEFT key : base_model.model.model.language_model.layers.N.<mod>.lora_A.weight shape (r, in)
|
||||
base_model.model.model.language_model.layers.N.<mod>.lora_B.weight shape (out, r)
|
||||
mlx key : language_model.model.layers.N.<mod>.lora_a shape (in, r)
|
||||
language_model.model.layers.N.<mod>.lora_b shape (r, out)
|
||||
Scale : PEFT scaling = alpha / r; mlx `scale` is used directly (y + scale * x @ A @ B), so scale = alpha / r.
|
||||
mlx loads only the modules listed in `lora_parameters.keys` (relative to the layer); they are taken from the file.
|
||||
"""
|
||||
import json, re, sys
|
||||
from pathlib import Path
|
||||
|
||||
import mlx.core as mx
|
||||
|
||||
src, dst = Path(sys.argv[1]), Path(sys.argv[2])
|
||||
cfg = json.load(open(src / "adapter_config.json"))
|
||||
r, alpha = cfg["r"], cfg["lora_alpha"]
|
||||
if cfg.get("use_rslora") or cfg.get("use_dora"):
|
||||
sys.exit("rsLoRA / DoRA: scale is not alpha / r, not supported")
|
||||
|
||||
w = mx.load(str(src / "adapter_model.safetensors"))
|
||||
pat = re.compile(r"^base_model\.model\.model\.language_model\.layers\.(\d+)\.(.+)\.lora_([AB])\.weight$")
|
||||
out, keys, layers = {}, set(), set()
|
||||
for k, v in w.items():
|
||||
m = pat.match(k)
|
||||
if not m:
|
||||
sys.exit(f"unexpected key: {k}")
|
||||
n, mod, ab = m.groups()
|
||||
layers.add(int(n))
|
||||
keys.add(mod)
|
||||
out[f"language_model.model.layers.{n}.{mod}.lora_{ab.lower()}"] = mx.transpose(v).astype(mx.float32)
|
||||
if ab == "A":
|
||||
assert v.shape[0] == r, (k, v.shape)
|
||||
|
||||
dst.mkdir(parents=True, exist_ok=True)
|
||||
mx.save_safetensors(str(dst / "adapters.safetensors"), out)
|
||||
json.dump({
|
||||
"fine_tune_type": "lora",
|
||||
"num_layers": max(layers) + 1,
|
||||
"lora_parameters": {"rank": r, "scale": alpha / r, "dropout": 0.0, "keys": sorted(keys)},
|
||||
}, open(dst / "adapter_config.json", "w"), indent=1)
|
||||
print(f"{len(out)} tensors, {len(layers)} layers, rank {r}, alpha {alpha} -> scale {alpha / r}, modules {sorted(keys)}")
|
||||
Reference in New Issue
Block a user