Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
10
CLAUDE.md
10
CLAUDE.md
@@ -109,10 +109,12 @@ python3 -c "from harness.ledger import spent; print(spent())"
|
|||||||
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
|
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
|
||||||
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
|
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
|
||||||
must fix, and he does not like Devstral. Devstral is not used further.
|
must fix, and he does not like Devstral. Devstral is not used further.
|
||||||
- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0.
|
- Devstral Small 2 tested and dropped (same settings, 11 tasks): mean **6.8 vs Qwen 15.8**; 1/11 vs 3/11 tasks above 0;
|
||||||
`runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`.
|
end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: `runs/archive/devstral/`.
|
||||||
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the
|
The Devstral weights were deleted. No further base model tests.
|
||||||
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference.
|
- **Official Qwen baseline** (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook,
|
||||||
|
run base 22000): `runs/stage1/baseline.json` (= `baseline_qwen.json`). Comparison: `docs/stage1-baseline.md`.
|
||||||
|
Older partial Qwen runs (Mac mini, aborted or thinking on): `runs/archive/qwen38/`, not comparable.
|
||||||
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
|
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
|
||||||
|
|
||||||
## 7. Next steps
|
## 7. Next steps
|
||||||
|
|||||||
@@ -115,7 +115,7 @@ Kurallar:
|
|||||||
|
|
||||||
## Adım 2 — Baz model seçimi
|
## Adım 2 — Baz model seçimi
|
||||||
|
|
||||||
Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Eski Qwen sonuçları `runs/archive/qwen38/`; tam Qwen baseline'ı MacBook'ta koşuyor (`baseline_qwen.json`).
|
Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral denendi ve elendi (11 görev, `runs/archive/devstral/`): ortalama 6,8 (Qwen 15,8); 1/11 görev 0'dan yüksek (Qwen 3/11); loop 6 (Qwen 7), tool_budget 3, report 2; onarım oranı 0,84. Devstral ağırlıkları silindi; baz model testi yok. Eski kısmi Qwen sonuçları `runs/archive/qwen38/` (karşılaştırılamaz).
|
||||||
|
|
||||||
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.
|
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.
|
||||||
|
|
||||||
|
|||||||
@@ -10,17 +10,16 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
|
|||||||
## Base model
|
## Base model
|
||||||
|
|
||||||
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
|
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
|
||||||
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen).
|
(Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again).
|
||||||
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
|
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
|
||||||
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
|
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
|
||||||
cannot load them without a re-quantization.
|
cannot load them without a re-quantization.
|
||||||
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`)
|
- Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0,
|
||||||
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further.
|
loops 6 vs 7. Results: `runs/archive/devstral/`. Weights deleted. No more base model tests.
|
||||||
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
|
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
|
||||||
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
|
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
|
||||||
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
|
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
|
||||||
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or
|
`train/serve.sh` serves Qwen (commit a2ba9e7).
|
||||||
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
|
|
||||||
|
|
||||||
## Serving (`train/serve.sh`)
|
## Serving (`train/serve.sh`)
|
||||||
|
|
||||||
@@ -29,17 +28,14 @@ training the same script with `--adapter-path`.
|
|||||||
|
|
||||||
| Setting | Value |
|
| Setting | Value |
|
||||||
|---|---|
|
|---|---|
|
||||||
| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent |
|
| Thinking | server default on (`--chat-template-args`); `train/baseline.py` sends `enable_thinking=false` per request |
|
||||||
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
|
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
|
||||||
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
|
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
|
||||||
| presence / repeat penalty | not set (neutral) |
|
| presence / repeat penalty | not set (neutral) |
|
||||||
| Output limit | server `--max-tokens 32768`; each request sends 16384 |
|
| Output limit | server `--max-tokens 32768`; each request sends 16384 |
|
||||||
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
|
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
|
||||||
|
|
||||||
Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`);
|
Tool calls: `mlx_lm` returns OpenAI `tool_calls` with JSON arguments (tool-call test passed).
|
||||||
`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors.
|
|
||||||
Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report
|
|
||||||
(G0128, G0174 ended with `report`).
|
|
||||||
|
|
||||||
## Eval subset
|
## Eval subset
|
||||||
|
|
||||||
@@ -52,8 +48,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
|
|||||||
|
|
||||||
| Setting | Value |
|
| Setting | Value |
|
||||||
|---|---|
|
|---|---|
|
||||||
| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
| Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
|
||||||
| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) |
|
| Thinking | off (`enable_thinking=false` per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (`runs/archive/qwen38/baseline_thinking_on.json`) |
|
||||||
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
|
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
|
||||||
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
|
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
|
||||||
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
|
| Tool-call budget per task | 60 calls, 15 activations (T01 too) |
|
||||||
@@ -68,12 +64,11 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
|
|||||||
|
|
||||||
## Baseline
|
## Baseline
|
||||||
|
|
||||||
- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by
|
- Official baseline: Qwen 3.8 27B, `python3 train/baseline.py --label baseline_qwen --run-base 22000` (MacBook, thinking off).
|
||||||
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered).
|
Results `runs/stage1/baseline.json` (= `baseline_qwen.json`), run directories `runs/stage1/baseline_qwen/`.
|
||||||
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories
|
Mean **15.8**; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1.
|
||||||
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`.
|
Repair rate 39/54 = 0.72.
|
||||||
- Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84
|
- Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84.
|
||||||
(the model changes the source, but the changes do not remove the cause).
|
Results in `runs/archive/devstral/`.
|
||||||
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
|
- Per run record has `end_reason`, `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
|
||||||
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`).
|
- The stage 1 training test (step 3) uses Qwen (`config_test.yaml`).
|
||||||
- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.
|
|
||||||
|
|||||||
@@ -26,7 +26,12 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
|||||||
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
|
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
|
||||||
`test.jsonl` = copy of valid.
|
`test.jsonl` = copy of valid.
|
||||||
|
|
||||||
## Running (detached)
|
## Base model
|
||||||
|
|
||||||
|
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
|
||||||
|
|
||||||
|
## Running (detached) — historical, all ended
|
||||||
|
|
||||||
|
|
||||||
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
|
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
|
||||||
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
|
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
|
||||||
@@ -40,8 +45,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
|||||||
|
|
||||||
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
|
||||||
(`t01_test_budget40` holds the T01 test result). Commit.
|
(`t01_test_budget40` holds the T01 test result). Commit.
|
||||||
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline.
|
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
||||||
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
|
|
||||||
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
options with `--help` first). Add it to `runs/stage1/baseline.json`.
|
||||||
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
|
||||||
iterations first.
|
iterations first.
|
||||||
@@ -56,3 +60,15 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
|
|||||||
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
|
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
|
||||||
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
|
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
|
||||||
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
|
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
|
||||||
|
|
||||||
|
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
|
||||||
|
|
||||||
|
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
|
||||||
|
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
|
||||||
|
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
|
||||||
|
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
|
||||||
|
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
|
||||||
|
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
|
||||||
|
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
|
||||||
|
- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started.
|
||||||
|
- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost.
|
||||||
|
|||||||
25
train/config.yaml
Normal file
25
train/config.yaml
Normal file
@@ -0,0 +1,25 @@
|
|||||||
|
# Stage 1 full run (start values of docs/stage1-training-task.md). Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config.yaml
|
||||||
|
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
|
||||||
|
train: true
|
||||||
|
data: train/data
|
||||||
|
fine_tune_type: lora
|
||||||
|
num_layers: -1
|
||||||
|
lora_parameters:
|
||||||
|
rank: 16
|
||||||
|
scale: 20.0
|
||||||
|
dropout: 0.0
|
||||||
|
batch_size: 1
|
||||||
|
grad_checkpoint: true
|
||||||
|
iters: 748
|
||||||
|
learning_rate: 5.0e-5
|
||||||
|
lr_schedule:
|
||||||
|
name: cosine_decay
|
||||||
|
warmup: 30
|
||||||
|
arguments: [5.0e-5, 748, 0.0]
|
||||||
|
max_seq_length: 16384
|
||||||
|
steps_per_report: 10
|
||||||
|
steps_per_eval: 200
|
||||||
|
val_batches: -1
|
||||||
|
save_every: 200
|
||||||
|
adapter_path: train/adapters
|
||||||
|
seed: 20261003
|
||||||
25
train/config_test.yaml
Normal file
25
train/config_test.yaml
Normal file
@@ -0,0 +1,25 @@
|
|||||||
|
# Stage 1 short test: 20 iterations, same settings as config.yaml. Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config_test.yaml
|
||||||
|
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
|
||||||
|
train: true
|
||||||
|
data: train/data
|
||||||
|
fine_tune_type: lora
|
||||||
|
num_layers: -1
|
||||||
|
lora_parameters:
|
||||||
|
rank: 16
|
||||||
|
scale: 20.0
|
||||||
|
dropout: 0.0
|
||||||
|
batch_size: 1
|
||||||
|
grad_checkpoint: true
|
||||||
|
iters: 20
|
||||||
|
learning_rate: 5.0e-5
|
||||||
|
lr_schedule:
|
||||||
|
name: cosine_decay
|
||||||
|
warmup: 2
|
||||||
|
arguments: [5.0e-5, 20, 0.0]
|
||||||
|
max_seq_length: 16384
|
||||||
|
steps_per_report: 1
|
||||||
|
steps_per_eval: 20
|
||||||
|
val_batches: 3
|
||||||
|
save_every: 20
|
||||||
|
adapter_path: train/adapters_test
|
||||||
|
seed: 20261003
|
||||||
60
train/hf_train.py
Normal file
60
train/hf_train.py
Normal file
@@ -0,0 +1,60 @@
|
|||||||
|
# /// script
|
||||||
|
# requires-python = ">=3.10"
|
||||||
|
# dependencies = ["unsloth", "datasets", "trl", "huggingface_hub"]
|
||||||
|
# ///
|
||||||
|
"""Stage 1 LoRA training with Unsloth on Hugging Face Jobs (1x A100 80GB).
|
||||||
|
|
||||||
|
Run (pipeline test, 10 steps):
|
||||||
|
hf jobs uv run --flavor a100-large --timeout 45m --secrets HF_TOKEN train/hf_train.py -- --max-steps 10
|
||||||
|
Full run: no --max-steps (2 epochs). Settings follow train/config.yaml; alpha is a proposal, see STATE.md.
|
||||||
|
"""
|
||||||
|
import argparse, json, os, time
|
||||||
|
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument("--model", default="unsloth/Qwen3.8-27B-unsloth-bnb-4bit")
|
||||||
|
ap.add_argument("--data", default="erhankeseli/abap-stage1-data")
|
||||||
|
ap.add_argument("--out", default="erhankeseli/abap-stage1-adapter-test")
|
||||||
|
ap.add_argument("--max-steps", type=int, default=-1)
|
||||||
|
ap.add_argument("--epochs", type=float, default=2.0)
|
||||||
|
ap.add_argument("--rank", type=int, default=16)
|
||||||
|
ap.add_argument("--alpha", type=int, default=16)
|
||||||
|
ap.add_argument("--lr", type=float, default=5e-5)
|
||||||
|
ap.add_argument("--max-seq-length", type=int, default=16384)
|
||||||
|
a = ap.parse_args()
|
||||||
|
|
||||||
|
from unsloth import FastLanguageModel # noqa: E402 (import first)
|
||||||
|
from datasets import load_dataset # noqa: E402
|
||||||
|
from trl import SFTConfig, SFTTrainer # noqa: E402
|
||||||
|
|
||||||
|
tok_hf = os.environ["HF_TOKEN"]
|
||||||
|
model, tok = FastLanguageModel.from_pretrained(
|
||||||
|
a.model, max_seq_length=a.max_seq_length, load_in_4bit=True, token=tok_hf)
|
||||||
|
model = FastLanguageModel.get_peft_model(
|
||||||
|
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
|
||||||
|
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj",
|
||||||
|
"in_proj_qkv", "in_proj_z", "out_proj"],
|
||||||
|
use_gradient_checkpointing="unsloth", random_state=20261003)
|
||||||
|
|
||||||
|
ds = load_dataset(a.data, data_files={"train": "train.jsonl", "valid": "valid.jsonl"}, token=tok_hf)
|
||||||
|
cfg = SFTConfig(
|
||||||
|
output_dir="out", per_device_train_batch_size=1, per_device_eval_batch_size=1,
|
||||||
|
gradient_accumulation_steps=1, num_train_epochs=a.epochs, max_steps=a.max_steps,
|
||||||
|
learning_rate=a.lr, lr_scheduler_type="cosine", warmup_steps=min(30, max(1, a.max_steps // 3)) if a.max_steps > 0 else 30,
|
||||||
|
optim="adamw_8bit", weight_decay=0.0, logging_steps=1, eval_strategy="no", save_strategy="no",
|
||||||
|
max_length=a.max_seq_length, dataset_text_field="text", packing=False, seed=20261003, report_to="none")
|
||||||
|
tr = SFTTrainer(model=model, processing_class=tok, train_dataset=ds["train"], eval_dataset=ds["valid"], args=cfg)
|
||||||
|
|
||||||
|
t0 = time.time()
|
||||||
|
tr.train()
|
||||||
|
train_s = time.time() - t0
|
||||||
|
steps = tr.state.global_step
|
||||||
|
ev = tr.evaluate() # valid loss of the adapter, for the conversion check
|
||||||
|
info = {"steps": steps, "train_seconds": round(train_s, 1), "sec_per_step": round(train_s / max(steps, 1), 2),
|
||||||
|
"eval_loss": ev.get("eval_loss"), "rank": a.rank, "alpha": a.alpha, "lr": a.lr,
|
||||||
|
"peak_gpu_gb": round(__import__("torch").cuda.max_memory_allocated() / 2**30, 1)}
|
||||||
|
print("RESULT", json.dumps(info))
|
||||||
|
model.save_pretrained("adapter")
|
||||||
|
json.dump(info, open("adapter/job_result.json", "w"))
|
||||||
|
from huggingface_hub import HfApi # noqa: E402
|
||||||
|
HfApi(token=tok_hf).upload_folder(folder_path="adapter", repo_id=a.out, repo_type="model")
|
||||||
|
print("PUSHED", a.out)
|
||||||
44
train/peft_to_mlx.py
Normal file
44
train/peft_to_mlx.py
Normal file
@@ -0,0 +1,44 @@
|
|||||||
|
"""Convert a PEFT LoRA adapter (Unsloth / HF) to the mlx-lm adapter format.
|
||||||
|
|
||||||
|
train/.venv/bin/python train/peft_to_mlx.py <peft_adapter_dir> <mlx_adapter_dir>
|
||||||
|
|
||||||
|
PEFT key : base_model.model.model.language_model.layers.N.<mod>.lora_A.weight shape (r, in)
|
||||||
|
base_model.model.model.language_model.layers.N.<mod>.lora_B.weight shape (out, r)
|
||||||
|
mlx key : language_model.model.layers.N.<mod>.lora_a shape (in, r)
|
||||||
|
language_model.model.layers.N.<mod>.lora_b shape (r, out)
|
||||||
|
Scale : PEFT scaling = alpha / r; mlx `scale` is used directly (y + scale * x @ A @ B), so scale = alpha / r.
|
||||||
|
mlx loads only the modules listed in `lora_parameters.keys` (relative to the layer); they are taken from the file.
|
||||||
|
"""
|
||||||
|
import json, re, sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import mlx.core as mx
|
||||||
|
|
||||||
|
src, dst = Path(sys.argv[1]), Path(sys.argv[2])
|
||||||
|
cfg = json.load(open(src / "adapter_config.json"))
|
||||||
|
r, alpha = cfg["r"], cfg["lora_alpha"]
|
||||||
|
if cfg.get("use_rslora") or cfg.get("use_dora"):
|
||||||
|
sys.exit("rsLoRA / DoRA: scale is not alpha / r, not supported")
|
||||||
|
|
||||||
|
w = mx.load(str(src / "adapter_model.safetensors"))
|
||||||
|
pat = re.compile(r"^base_model\.model\.model\.language_model\.layers\.(\d+)\.(.+)\.lora_([AB])\.weight$")
|
||||||
|
out, keys, layers = {}, set(), set()
|
||||||
|
for k, v in w.items():
|
||||||
|
m = pat.match(k)
|
||||||
|
if not m:
|
||||||
|
sys.exit(f"unexpected key: {k}")
|
||||||
|
n, mod, ab = m.groups()
|
||||||
|
layers.add(int(n))
|
||||||
|
keys.add(mod)
|
||||||
|
out[f"language_model.model.layers.{n}.{mod}.lora_{ab.lower()}"] = mx.transpose(v).astype(mx.float32)
|
||||||
|
if ab == "A":
|
||||||
|
assert v.shape[0] == r, (k, v.shape)
|
||||||
|
|
||||||
|
dst.mkdir(parents=True, exist_ok=True)
|
||||||
|
mx.save_safetensors(str(dst / "adapters.safetensors"), out)
|
||||||
|
json.dump({
|
||||||
|
"fine_tune_type": "lora",
|
||||||
|
"num_layers": max(layers) + 1,
|
||||||
|
"lora_parameters": {"rank": r, "scale": alpha / r, "dropout": 0.0, "keys": sorted(keys)},
|
||||||
|
}, open(dst / "adapter_config.json", "w"), indent=1)
|
||||||
|
print(f"{len(out)} tensors, {len(layers)} layers, rank {r}, alpha {alpha} -> scale {alpha / r}, modules {sorted(keys)}")
|
||||||
Reference in New Issue
Block a user