diff --git a/CLAUDE.md b/CLAUDE.md index 6acdfb4..87f00a3 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -109,10 +109,12 @@ python3 -c "from harness.ledger import spent; print(spent())" - Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training must fix, and he does not like Devstral. Devstral is not used further. -- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0. - `runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`. -- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the - MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference. +- Devstral Small 2 tested and dropped (same settings, 11 tasks): mean **6.8 vs Qwen 15.8**; 1/11 vs 3/11 tasks above 0; + end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: `runs/archive/devstral/`. + The Devstral weights were deleted. No further base model tests. +- **Official Qwen baseline** (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook, + run base 22000): `runs/stage1/baseline.json` (= `baseline_qwen.json`). Comparison: `docs/stage1-baseline.md`. + Older partial Qwen runs (Mac mini, aborted or thinking on): `runs/archive/qwen38/`, not comparable. - Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl). ## 7. Next steps diff --git a/docs/yol-haritasi.md b/docs/yol-haritasi.md index b610a02..83d26ae 100644 --- a/docs/yol-haritasi.md +++ b/docs/yol-haritasi.md @@ -115,7 +115,7 @@ Kurallar: ## Adım 2 — Baz model seçimi -Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Eski Qwen sonuçları `runs/archive/qwen38/`; tam Qwen baseline'ı MacBook'ta koşuyor (`baseline_qwen.json`). +Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral denendi ve elendi (11 görev, `runs/archive/devstral/`): ortalama 6,8 (Qwen 15,8); 1/11 görev 0'dan yüksek (Qwen 3/11); loop 6 (Qwen 7), tool_budget 3, report 2; onarım oranı 0,84. Devstral ağırlıkları silindi; baz model testi yok. Eski kısmi Qwen sonuçları `runs/archive/qwen38/` (karşılaştırılamaz). Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik. diff --git a/train/README.md b/train/README.md index 36658dd..e997233 100644 --- a/train/README.md +++ b/train/README.md @@ -10,17 +10,16 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST ## Base model - Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04 - (revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen). + (Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again). - Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path `~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm` cannot load them without a re-quantization. -- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`) - was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further. +- Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0, + loops 6 vs 7. Results: `runs/archive/devstral/`. Weights deleted. No more base model tests. - Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again), empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1. - Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`. - `train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or - restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64). + `train/serve.sh` serves Qwen (commit a2ba9e7). ## Serving (`train/serve.sh`) @@ -29,17 +28,14 @@ training the same script with `--adapter-path`. | Setting | Value | |---|---| -| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent | +| Thinking | server default on (`--chat-template-args`); `train/baseline.py` sends `enable_thinking=false` per request | | temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) | | top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) | | presence / repeat penalty | not set (neutral) | | Output limit | server `--max-tokens 32768`; each request sends 16384 | | Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` | -Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`); -`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors. -Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report -(G0128, G0174 ended with `report`). +Tool calls: `mlx_lm` returns OpenAI `tool_calls` with JSON arguments (tool-call test passed). ## Eval subset @@ -52,8 +48,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra | Setting | Value | |---|---| -| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` | -| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) | +| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` | +| Thinking | off (`enable_thinking=false` per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (`runs/archive/qwen38/baseline_thinking_on.json`) | | temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 | | max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 | | Tool-call budget per task | 60 calls, 15 activations (T01 too) | @@ -68,12 +64,11 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi ## Baseline -- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by - `train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered). - Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories - `runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`. -- Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84 - (the model changes the source, but the changes do not remove the cause). -- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`. -- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`). -- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test. +- Official baseline: Qwen 3.8 27B, `python3 train/baseline.py --label baseline_qwen --run-base 22000` (MacBook, thinking off). + Results `runs/stage1/baseline.json` (= `baseline_qwen.json`), run directories `runs/stage1/baseline_qwen/`. + Mean **15.8**; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1. + Repair rate 39/54 = 0.72. +- Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84. + Results in `runs/archive/devstral/`. +- Per run record has `end_reason`, `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`. +- The stage 1 training test (step 3) uses Qwen (`config_test.yaml`). diff --git a/train/STATE.md b/train/STATE.md index 3b155c1..64d5c43 100644 --- a/train/STATE.md +++ b/train/STATE.md @@ -26,7 +26,12 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U **748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence). `test.jsonl` = copy of valid. -## Running (detached) +## Base model + +Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests. + +## Running (detached) — historical, all ended + - MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`. - Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started @@ -40,8 +45,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U 1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json` (`t01_test_budget40` holds the T01 test result). Commit. -2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline. -3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the +2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the options with `--help` first). Add it to `runs/stage1/baseline.json`. 4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 iterations first. @@ -56,3 +60,15 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U - Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in `runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends. + +## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5) + +Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training. +- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`). +- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`; + `train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters. +- Key names: PEFT `base_model.model.model.language_model.layers.N..lora_A/B.weight` (A: r x in) -> + mlx `language_model.model.layers.N..lora_a/b` (transposed). mlx scale = alpha / r. +- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320. +- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started. +- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost. diff --git a/train/config.yaml b/train/config.yaml new file mode 100644 index 0000000..cbfb501 --- /dev/null +++ b/train/config.yaml @@ -0,0 +1,25 @@ +# Stage 1 full run (start values of docs/stage1-training-task.md). Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config.yaml +model: /Users/erhankeseli/models/Qwen3.8-27B-4bit +train: true +data: train/data +fine_tune_type: lora +num_layers: -1 +lora_parameters: + rank: 16 + scale: 20.0 + dropout: 0.0 +batch_size: 1 +grad_checkpoint: true +iters: 748 +learning_rate: 5.0e-5 +lr_schedule: + name: cosine_decay + warmup: 30 + arguments: [5.0e-5, 748, 0.0] +max_seq_length: 16384 +steps_per_report: 10 +steps_per_eval: 200 +val_batches: -1 +save_every: 200 +adapter_path: train/adapters +seed: 20261003 diff --git a/train/config_test.yaml b/train/config_test.yaml new file mode 100644 index 0000000..cb13602 --- /dev/null +++ b/train/config_test.yaml @@ -0,0 +1,25 @@ +# Stage 1 short test: 20 iterations, same settings as config.yaml. Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config_test.yaml +model: /Users/erhankeseli/models/Qwen3.8-27B-4bit +train: true +data: train/data +fine_tune_type: lora +num_layers: -1 +lora_parameters: + rank: 16 + scale: 20.0 + dropout: 0.0 +batch_size: 1 +grad_checkpoint: true +iters: 20 +learning_rate: 5.0e-5 +lr_schedule: + name: cosine_decay + warmup: 2 + arguments: [5.0e-5, 20, 0.0] +max_seq_length: 16384 +steps_per_report: 1 +steps_per_eval: 20 +val_batches: 3 +save_every: 20 +adapter_path: train/adapters_test +seed: 20261003 diff --git a/train/hf_train.py b/train/hf_train.py new file mode 100644 index 0000000..c4df384 --- /dev/null +++ b/train/hf_train.py @@ -0,0 +1,60 @@ +# /// script +# requires-python = ">=3.10" +# dependencies = ["unsloth", "datasets", "trl", "huggingface_hub"] +# /// +"""Stage 1 LoRA training with Unsloth on Hugging Face Jobs (1x A100 80GB). + +Run (pipeline test, 10 steps): + hf jobs uv run --flavor a100-large --timeout 45m --secrets HF_TOKEN train/hf_train.py -- --max-steps 10 +Full run: no --max-steps (2 epochs). Settings follow train/config.yaml; alpha is a proposal, see STATE.md. +""" +import argparse, json, os, time + +ap = argparse.ArgumentParser() +ap.add_argument("--model", default="unsloth/Qwen3.8-27B-unsloth-bnb-4bit") +ap.add_argument("--data", default="erhankeseli/abap-stage1-data") +ap.add_argument("--out", default="erhankeseli/abap-stage1-adapter-test") +ap.add_argument("--max-steps", type=int, default=-1) +ap.add_argument("--epochs", type=float, default=2.0) +ap.add_argument("--rank", type=int, default=16) +ap.add_argument("--alpha", type=int, default=16) +ap.add_argument("--lr", type=float, default=5e-5) +ap.add_argument("--max-seq-length", type=int, default=16384) +a = ap.parse_args() + +from unsloth import FastLanguageModel # noqa: E402 (import first) +from datasets import load_dataset # noqa: E402 +from trl import SFTConfig, SFTTrainer # noqa: E402 + +tok_hf = os.environ["HF_TOKEN"] +model, tok = FastLanguageModel.from_pretrained( + a.model, max_seq_length=a.max_seq_length, load_in_4bit=True, token=tok_hf) +model = FastLanguageModel.get_peft_model( + model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none", + target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj", + "in_proj_qkv", "in_proj_z", "out_proj"], + use_gradient_checkpointing="unsloth", random_state=20261003) + +ds = load_dataset(a.data, data_files={"train": "train.jsonl", "valid": "valid.jsonl"}, token=tok_hf) +cfg = SFTConfig( + output_dir="out", per_device_train_batch_size=1, per_device_eval_batch_size=1, + gradient_accumulation_steps=1, num_train_epochs=a.epochs, max_steps=a.max_steps, + learning_rate=a.lr, lr_scheduler_type="cosine", warmup_steps=min(30, max(1, a.max_steps // 3)) if a.max_steps > 0 else 30, + optim="adamw_8bit", weight_decay=0.0, logging_steps=1, eval_strategy="no", save_strategy="no", + max_length=a.max_seq_length, dataset_text_field="text", packing=False, seed=20261003, report_to="none") +tr = SFTTrainer(model=model, processing_class=tok, train_dataset=ds["train"], eval_dataset=ds["valid"], args=cfg) + +t0 = time.time() +tr.train() +train_s = time.time() - t0 +steps = tr.state.global_step +ev = tr.evaluate() # valid loss of the adapter, for the conversion check +info = {"steps": steps, "train_seconds": round(train_s, 1), "sec_per_step": round(train_s / max(steps, 1), 2), + "eval_loss": ev.get("eval_loss"), "rank": a.rank, "alpha": a.alpha, "lr": a.lr, + "peak_gpu_gb": round(__import__("torch").cuda.max_memory_allocated() / 2**30, 1)} +print("RESULT", json.dumps(info)) +model.save_pretrained("adapter") +json.dump(info, open("adapter/job_result.json", "w")) +from huggingface_hub import HfApi # noqa: E402 +HfApi(token=tok_hf).upload_folder(folder_path="adapter", repo_id=a.out, repo_type="model") +print("PUSHED", a.out) diff --git a/train/peft_to_mlx.py b/train/peft_to_mlx.py new file mode 100644 index 0000000..15e5ba8 --- /dev/null +++ b/train/peft_to_mlx.py @@ -0,0 +1,44 @@ +"""Convert a PEFT LoRA adapter (Unsloth / HF) to the mlx-lm adapter format. + + train/.venv/bin/python train/peft_to_mlx.py + +PEFT key : base_model.model.model.language_model.layers.N..lora_A.weight shape (r, in) + base_model.model.model.language_model.layers.N..lora_B.weight shape (out, r) +mlx key : language_model.model.layers.N..lora_a shape (in, r) + language_model.model.layers.N..lora_b shape (r, out) +Scale : PEFT scaling = alpha / r; mlx `scale` is used directly (y + scale * x @ A @ B), so scale = alpha / r. +mlx loads only the modules listed in `lora_parameters.keys` (relative to the layer); they are taken from the file. +""" +import json, re, sys +from pathlib import Path + +import mlx.core as mx + +src, dst = Path(sys.argv[1]), Path(sys.argv[2]) +cfg = json.load(open(src / "adapter_config.json")) +r, alpha = cfg["r"], cfg["lora_alpha"] +if cfg.get("use_rslora") or cfg.get("use_dora"): + sys.exit("rsLoRA / DoRA: scale is not alpha / r, not supported") + +w = mx.load(str(src / "adapter_model.safetensors")) +pat = re.compile(r"^base_model\.model\.model\.language_model\.layers\.(\d+)\.(.+)\.lora_([AB])\.weight$") +out, keys, layers = {}, set(), set() +for k, v in w.items(): + m = pat.match(k) + if not m: + sys.exit(f"unexpected key: {k}") + n, mod, ab = m.groups() + layers.add(int(n)) + keys.add(mod) + out[f"language_model.model.layers.{n}.{mod}.lora_{ab.lower()}"] = mx.transpose(v).astype(mx.float32) + if ab == "A": + assert v.shape[0] == r, (k, v.shape) + +dst.mkdir(parents=True, exist_ok=True) +mx.save_safetensors(str(dst / "adapters.safetensors"), out) +json.dump({ + "fine_tune_type": "lora", + "num_layers": max(layers) + 1, + "lora_parameters": {"rank": r, "scale": alpha / r, "dropout": 0.0, "keys": sorted(keys)}, +}, open(dst / "adapter_config.json", "w"), indent=1) +print(f"{len(out)} tensors, {len(layers)} layers, rank {r}, alpha {alpha} -> scale {alpha / r}, modules {sorted(keys)}")