Stage 1 on HF Jobs: Unsloth job script, PEFT to MLX converter, base valid loss 0.849, Qwen base model docs

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-04 18:51:40 +02:00
parent a2ba9e7b44
commit 5447874fd3
8 changed files with 196 additions and 29 deletions

View File

@@ -109,10 +109,12 @@ python3 -c "from harness.ledger import spent; print(spent())"
- Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking - Earlier the same day Qwen was dropped (no repair after activation errors, loops, empty responses at the thinking
limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training limit) and Devstral Small 2 (24B) was the candidate. Kral reversed this: the Qwen weaknesses are what training
must fix, and he does not like Devstral. Devstral is not used further. must fix, and he does not like Devstral. Devstral is not used further.
- Devstral baseline (11 tasks, thinking off n/a, guard 3, base 21000): mean 6.8, 1/11 above 0. - Devstral Small 2 tested and dropped (same settings, 11 tasks): mean **6.8 vs Qwen 15.8**; 1/11 vs 3/11 tasks above 0;
`runs/stage1/baseline_devstral.md`. Qwen results (older, partial): `runs/archive/qwen38/`. end reason loop 6 vs 7 (no lower loop rate), tool_budget 3 vs 3. Results: `runs/archive/devstral/`.
- A complete Qwen baseline with the same settings (thinking off, max_tokens 16384, budget 60, guard 3) runs on the The Devstral weights were deleted. No further base model tests.
MacBook (run base 22000, `baseline_qwen.json`); it becomes the stage 1 reference. - **Official Qwen baseline** (complete, 11 tasks, thinking off, max_tokens 16384, budget 60, guard 3, MacBook,
run base 22000): `runs/stage1/baseline.json` (= `baseline_qwen.json`). Comparison: `docs/stage1-baseline.md`.
Older partial Qwen runs (Mac mini, aborted or thinking on): `runs/archive/qwen38/`, not comparable.
- Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl). - Training tool is not chosen. Mac: `mlx_lm.lora` (small test). Rented GPU: open (Unsloth, TRL + PEFT, Axolotl).
## 7. Next steps ## 7. Next steps

View File

@@ -115,7 +115,7 @@ Kurallar:
## Adım 2 — Baz model seçimi ## Adım 2 — Baz model seçimi
Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral baseline (11 görev, `runs/stage1/baseline_devstral.md`): ortalama 6,8; 11 görevden 1'i 0'dan yüksek; bitiş nedeni loop 6, tool_budget 3, report 2; onarım oranı 0,84. Eski Qwen sonuçları `runs/archive/qwen38/`; tam Qwen baseline'ı MacBook'ta koşuyor (`baseline_qwen.json`). Durum (2026-10-04, güncel): baz model Qwen 3.8 27B (Kral kararı). Referans baseline Qwen (MacBook, aynı ayarlar, 11 görev): ortalama 15,8; 3 görev 0'dan yüksek (T01 48,3, G0157 51,0, G0185 75,0); ayrıntı `docs/stage1-baseline.md`. Devstral yalnız karşılaştırma: Devstral denendi ve elendi (11 görev, `runs/archive/devstral/`): ortalama 6,8 (Qwen 15,8); 1/11 görev 0'dan yüksek (Qwen 3/11); loop 6 (Qwen 7), tool_budget 3, report 2; onarım oranı 0,84. Devstral ağırlıkları silindi; baz model testi yok. Eski kısmi Qwen sonuçları `runs/archive/qwen38/` (karşılaştırılamaz).
Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik. Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisanslı modeller. Kriterler: puan, boyut, Mac mini'de çalışabilirlik.

View File

@@ -10,17 +10,16 @@ Task: `docs/stage1-training-task.md`. State of the work: this file and `train/ST
## Base model ## Base model
- Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04 - Base model: **`Qwen/Qwen3.8-27B`** (architecture `qwen3_5`, 27.8B, dense; Apache 2.0). Kral decision 2026-10-04
(revised the same day: Qwen was dropped for a few hours, Devstral Small 2 was a candidate; Kral chose Qwen). (Qwen was dropped for a few hours, Devstral Small 2 was tested; Kral chose Qwen again).
- Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path - Weights: **`mlx-community/Qwen3.8-27B-4bit`** (MLX affine, 4 bit, group size 64, 16.1 GB), local path
`~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm` `~/models/Qwen3.8-27B-4bit`. Why not the Ollama weights: they are NVFP4 with a global scale per layer; `mlx_lm`
cannot load them without a re-quantization. cannot load them without a re-quantization.
- Devstral Small 2 (`mlx-community/Devstral-Small-2-24B-Instruct-2512-4bit`, `~/models/Devstral-Small-2-24B-4bit`) - Devstral Small 2 was tested for the baseline only and dropped: mean 6.8 vs Qwen 15.8, 1 vs 3 of 11 tasks above 0,
was tested only for the baseline: `runs/stage1/baseline_devstral.md` (mean 6.8, 1 of 11 tasks above 0). Not used further. loops 6 vs 7. Results: `runs/archive/devstral/`. Weights deleted. No more base model tests.
- Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again), - Known Qwen weaknesses (the training target): no repair after activation errors, loops (same source pushed again),
empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1. empty responses at the thinking limit when thinking is on. Thinking stays off in stage 1.
- Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`. - Stage 1 reference baseline: **Qwen, mean 15.8, 3 of 11 tasks above 0** (T01 48.3, G0157 51.0, G0185 75.0), same settings as Devstral, run on the MacBook (`runs/stage1/baseline_qwen.json`, run base 22000; copy `baseline.json`). Comparison: `docs/stage1-baseline.md`.
`train/serve.sh` serves Devstral at the moment; for Qwen use `train/serve_qwen.sh` (in the MacBook package) or `train/serve.sh` serves Qwen (commit a2ba9e7).
restore the Qwen line (`--model ~/models/Qwen3.8-27B-4bit`, `--chat-template-args` as in git history before 0af2d64).
## Serving (`train/serve.sh`) ## Serving (`train/serve.sh`)
@@ -29,17 +28,14 @@ training the same script with `--adapter-path`.
| Setting | Value | | Setting | Value |
|---|---| |---|---|
| Thinking | none (Devstral has no thinking mode); no `chat_template_kwargs` are sent | | Thinking | server default on (`--chat-template-args`); `train/baseline.py` sends `enable_thinking=false` per request |
| temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) | | temperature | 0.2 (sent by the harness llm agent; the model card suggests 0.15) |
| top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) | | top_p / top_k / min_p | 0.95 / 20 / 0 (server flags) |
| presence / repeat penalty | not set (neutral) | | presence / repeat penalty | not set (neutral) |
| Output limit | server `--max-tokens 32768`; each request sends 16384 | | Output limit | server `--max-tokens 32768`; each request sends 16384 |
| Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` | | Prompt cache | `--prompt-cache-size 4 --prompt-cache-bytes 6000000000` |
Tool calls (2026-10-04): the chat template uses the Mistral format (`[AVAILABLE_TOOLS]`, `[TOOL_CALLS]name[ARGS]{json}`); Tool calls: `mlx_lm` returns OpenAI `tool_calls` with JSON arguments (tool-call test passed).
`mlx_lm` returns OpenAI `tool_calls` with JSON arguments. The smoke test T01 and the baseline had no parse errors.
Known behaviour: some turns have prose and no tool call; the harness takes such a turn as the final report
(G0128, G0174 ended with `report`).
## Eval subset ## Eval subset
@@ -52,8 +48,8 @@ G0139, G0151, G0157, G0174, G0167, G0185. Use the same list before and after tra
| Setting | Value | | Setting | Value |
|---|---| |---|---|
| Model | `~/models/Devstral-Small-2-24B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` | | Model | `~/models/Qwen3.8-27B-4bit` (MLX affine 4 bit); after training the same with `--adapter-path` |
| Thinking | none: Devstral has no thinking mode; no `enable_thinking` is sent. (Qwen: off, fixed; with thinking on, all Qwen runs ended with empty responses at the thinking limit, `runs/archive/qwen38/baseline_thinking_on.json`) | | Thinking | off (`enable_thinking=false` per request); with thinking on, all Qwen runs ended with empty responses at the thinking limit (`runs/archive/qwen38/baseline_thinking_on.json`) |
| temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 | | temperature / top_p / top_k / min_p | 0.2 / 0.95 / 20 / 0 |
| max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 | | max_tokens per turn | **16384**, sent in each request by `train/baseline.py` (`MAX_TOKENS`); the server limit stays 32768 |
| Tool-call budget per task | 60 calls, 15 activations (T01 too) | | Tool-call budget per task | 60 calls, 15 activations (T01 too) |
@@ -68,12 +64,11 @@ T01 test with 32768 tokens and budget 40 (`t01_test_budget40`: 40.0) and the thi
## Baseline ## Baseline
- Devstral baseline (2026-10-04): `python3 train/baseline.py --label baseline_devstral --run-base 21000`, started by - Official baseline: Qwen 3.8 27B, `python3 train/baseline.py --label baseline_qwen --run-base 22000` (MacBook, thinking off).
`train/baseline_chain.sh` (stop rule after 4 tasks: all loop and repair rate below 20 % → stop; not triggered). Results `runs/stage1/baseline.json` (= `baseline_qwen.json`), run directories `runs/stage1/baseline_qwen/`.
Results `runs/stage1/baseline_devstral.json` (copy: `runs/stage1/baseline.json`), run directories Mean **15.8**; 3 of 11 tasks above 0 (T01 48.3, G0157 51.0, G0185 75.0). End reasons: loop 7, tool_budget 3, report 1.
`runs/stage1/baseline_devstral/`, report `runs/stage1/baseline_devstral.md`. Repair rate 39/54 = 0.72.
- Devstral result (not used further): mean 6.8; 1 of 11 tasks above 0 (G0167: 75). Qwen: mean 15.8, 3 of 11. End reasons: loop 6, tool_budget 3, report 2. Repair rate 64/76 = 0.84 - Devstral Small 2 (dropped): mean 6.8; 1 of 11 above 0 (G0167: 75); loop 6, tool_budget 3, report 2; repair rate 0.84.
(the model changes the source, but the changes do not remove the cause). Results in `runs/archive/devstral/`.
- Per run record now also has `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`. - Per run record has `end_reason`, `pushes_after_error`, `pushes_changed_after_error`, `repair_rate`.
- The Qwen baseline was never completed (archive: `runs/archive/qwen38/`). - The stage 1 training test (step 3) uses Qwen (`config_test.yaml`).
- The stage 1 training test (step 3) now uses Devstral. `mlx_lm.lora` must be checked for `mistral3` before the test.

View File

@@ -26,7 +26,12 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
**748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence). **748 iterations for 2 epochs** (the count is the same as with the first dedup rule by coincidence).
`test.jsonl` = copy of valid. `test.jsonl` = copy of valid.
## Running (detached) ## Base model
Qwen 3.8 27B (Kral decision 2026-10-04). Devstral Small 2 tested and dropped (mean 6.8 vs 15.8; loops 6 vs 7; `runs/archive/devstral/`). Official baseline: `runs/stage1/baseline.json` (11 tasks, Qwen mean 15.8, 3/11 above 0). No more base model tests.
## Running (detached) — historical, all ended
- MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`. - MLX server (restarted 2026-10-04 06:38 after it had exited; PID in `pgrep -f mlx_lm`), log `runs/stage1/server.log`.
- Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started - Baseline on the 11-task subset (`train/subset.json`), thinking off, loop guard 3, max_tokens 16384, started
@@ -40,8 +45,7 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json` 1. B2: read the last lines of `runs/stage1/baseline.log`; summary from `runs/stage1/baseline.json`
(`t01_test_budget40` holds the T01 test result). Commit. (`t01_test_budget40` holds the T01 test result). Commit.
2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test`, no adapter) after the baseline. 2. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
3. Step 2, second part: valid loss of the base model (`mlx_lm.lora --test` without adapter; check the
options with `--help` first). Add it to `runs/stage1/baseline.json`. options with `--help` first). Add it to `runs/stage1/baseline.json`.
4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20 4. Step 3 (training): Kral stops A4H; stop the MLX server; no other model loaded. Short test of 20
iterations first. iterations first.
@@ -56,3 +60,15 @@ Task: `docs/stage1-training-task.md`. Settings and weights: `train/README.md`. U
- Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results - Session 2026-10-03 (evening): follow-ups of `docs/devir-notlari.md` section 3 done from stored results
(review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in (review of 10 + 10 tasks, easy candidates, docs). No model run was started. Reruns wait in
`runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends. `runs/stage1/rerun_queue.txt` (G0119, G0162) until the baseline ends.
## Stage 1 on Hugging Face Jobs (plan change 2026-10-04, Kral + Opus 5.5)
Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
- Base valid loss (Mac, MLX 4-bit, no adapter): **0.849**, ppl 2.337 (`runs/stage1/baseline.json`, key `valid_loss`).
- Done: private dataset `erhankeseli/abap-stage1-data`; private model repo `erhankeseli/abap-stage1-adapter-test`;
`train/hf_train.py` (Unsloth job), `train/peft_to_mlx.py` (converter, not yet tested). mlx-lm 0.32 does not load PEFT adapters.
- Key names: PEFT `base_model.model.model.language_model.layers.N.<mod>.lora_A/B.weight` (A: r x in) ->
mlx `language_model.model.layers.N.<mod>.lora_a/b` (transposed). mlx scale = alpha / r.
- Alpha proposal (open, Kral decides): 32 (scale 2); 16 (scale 1) is the safer option. mlx `scale: 20` would be alpha 320.
- Blocked: HF Jobs returns 402 (no prepaid credit). Test job (a100-large, 2.50 USD/h, 10 steps, timeout 45m, max 1.88 USD) not started.
- Next: credit -> test job -> convert -> valid loss on Mac vs the loss Unsloth reports -> report time per step and full-run cost.

25
train/config.yaml Normal file
View File

@@ -0,0 +1,25 @@
# Stage 1 full run (start values of docs/stage1-training-task.md). Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config.yaml
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
train: true
data: train/data
fine_tune_type: lora
num_layers: -1
lora_parameters:
rank: 16
scale: 20.0
dropout: 0.0
batch_size: 1
grad_checkpoint: true
iters: 748
learning_rate: 5.0e-5
lr_schedule:
name: cosine_decay
warmup: 30
arguments: [5.0e-5, 748, 0.0]
max_seq_length: 16384
steps_per_report: 10
steps_per_eval: 200
val_batches: -1
save_every: 200
adapter_path: train/adapters
seed: 20261003

25
train/config_test.yaml Normal file
View File

@@ -0,0 +1,25 @@
# Stage 1 short test: 20 iterations, same settings as config.yaml. Run: MLX_DISABLE_COMPILE=1 train/.venv/bin/mlx_lm.lora -c train/config_test.yaml
model: /Users/erhankeseli/models/Qwen3.8-27B-4bit
train: true
data: train/data
fine_tune_type: lora
num_layers: -1
lora_parameters:
rank: 16
scale: 20.0
dropout: 0.0
batch_size: 1
grad_checkpoint: true
iters: 20
learning_rate: 5.0e-5
lr_schedule:
name: cosine_decay
warmup: 2
arguments: [5.0e-5, 20, 0.0]
max_seq_length: 16384
steps_per_report: 1
steps_per_eval: 20
val_batches: 3
save_every: 20
adapter_path: train/adapters_test
seed: 20261003

60
train/hf_train.py Normal file
View File

@@ -0,0 +1,60 @@
# /// script
# requires-python = ">=3.10"
# dependencies = ["unsloth", "datasets", "trl", "huggingface_hub"]
# ///
"""Stage 1 LoRA training with Unsloth on Hugging Face Jobs (1x A100 80GB).
Run (pipeline test, 10 steps):
hf jobs uv run --flavor a100-large --timeout 45m --secrets HF_TOKEN train/hf_train.py -- --max-steps 10
Full run: no --max-steps (2 epochs). Settings follow train/config.yaml; alpha is a proposal, see STATE.md.
"""
import argparse, json, os, time
ap = argparse.ArgumentParser()
ap.add_argument("--model", default="unsloth/Qwen3.8-27B-unsloth-bnb-4bit")
ap.add_argument("--data", default="erhankeseli/abap-stage1-data")
ap.add_argument("--out", default="erhankeseli/abap-stage1-adapter-test")
ap.add_argument("--max-steps", type=int, default=-1)
ap.add_argument("--epochs", type=float, default=2.0)
ap.add_argument("--rank", type=int, default=16)
ap.add_argument("--alpha", type=int, default=16)
ap.add_argument("--lr", type=float, default=5e-5)
ap.add_argument("--max-seq-length", type=int, default=16384)
a = ap.parse_args()
from unsloth import FastLanguageModel # noqa: E402 (import first)
from datasets import load_dataset # noqa: E402
from trl import SFTConfig, SFTTrainer # noqa: E402
tok_hf = os.environ["HF_TOKEN"]
model, tok = FastLanguageModel.from_pretrained(
a.model, max_seq_length=a.max_seq_length, load_in_4bit=True, token=tok_hf)
model = FastLanguageModel.get_peft_model(
model, r=a.rank, lora_alpha=a.alpha, lora_dropout=0.0, bias="none",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj",
"in_proj_qkv", "in_proj_z", "out_proj"],
use_gradient_checkpointing="unsloth", random_state=20261003)
ds = load_dataset(a.data, data_files={"train": "train.jsonl", "valid": "valid.jsonl"}, token=tok_hf)
cfg = SFTConfig(
output_dir="out", per_device_train_batch_size=1, per_device_eval_batch_size=1,
gradient_accumulation_steps=1, num_train_epochs=a.epochs, max_steps=a.max_steps,
learning_rate=a.lr, lr_scheduler_type="cosine", warmup_steps=min(30, max(1, a.max_steps // 3)) if a.max_steps > 0 else 30,
optim="adamw_8bit", weight_decay=0.0, logging_steps=1, eval_strategy="no", save_strategy="no",
max_length=a.max_seq_length, dataset_text_field="text", packing=False, seed=20261003, report_to="none")
tr = SFTTrainer(model=model, processing_class=tok, train_dataset=ds["train"], eval_dataset=ds["valid"], args=cfg)
t0 = time.time()
tr.train()
train_s = time.time() - t0
steps = tr.state.global_step
ev = tr.evaluate() # valid loss of the adapter, for the conversion check
info = {"steps": steps, "train_seconds": round(train_s, 1), "sec_per_step": round(train_s / max(steps, 1), 2),
"eval_loss": ev.get("eval_loss"), "rank": a.rank, "alpha": a.alpha, "lr": a.lr,
"peak_gpu_gb": round(__import__("torch").cuda.max_memory_allocated() / 2**30, 1)}
print("RESULT", json.dumps(info))
model.save_pretrained("adapter")
json.dump(info, open("adapter/job_result.json", "w"))
from huggingface_hub import HfApi # noqa: E402
HfApi(token=tok_hf).upload_folder(folder_path="adapter", repo_id=a.out, repo_type="model")
print("PUSHED", a.out)

44
train/peft_to_mlx.py Normal file
View File

@@ -0,0 +1,44 @@
"""Convert a PEFT LoRA adapter (Unsloth / HF) to the mlx-lm adapter format.
train/.venv/bin/python train/peft_to_mlx.py <peft_adapter_dir> <mlx_adapter_dir>
PEFT key : base_model.model.model.language_model.layers.N.<mod>.lora_A.weight shape (r, in)
base_model.model.model.language_model.layers.N.<mod>.lora_B.weight shape (out, r)
mlx key : language_model.model.layers.N.<mod>.lora_a shape (in, r)
language_model.model.layers.N.<mod>.lora_b shape (r, out)
Scale : PEFT scaling = alpha / r; mlx `scale` is used directly (y + scale * x @ A @ B), so scale = alpha / r.
mlx loads only the modules listed in `lora_parameters.keys` (relative to the layer); they are taken from the file.
"""
import json, re, sys
from pathlib import Path
import mlx.core as mx
src, dst = Path(sys.argv[1]), Path(sys.argv[2])
cfg = json.load(open(src / "adapter_config.json"))
r, alpha = cfg["r"], cfg["lora_alpha"]
if cfg.get("use_rslora") or cfg.get("use_dora"):
sys.exit("rsLoRA / DoRA: scale is not alpha / r, not supported")
w = mx.load(str(src / "adapter_model.safetensors"))
pat = re.compile(r"^base_model\.model\.model\.language_model\.layers\.(\d+)\.(.+)\.lora_([AB])\.weight$")
out, keys, layers = {}, set(), set()
for k, v in w.items():
m = pat.match(k)
if not m:
sys.exit(f"unexpected key: {k}")
n, mod, ab = m.groups()
layers.add(int(n))
keys.add(mod)
out[f"language_model.model.layers.{n}.{mod}.lora_{ab.lower()}"] = mx.transpose(v).astype(mx.float32)
if ab == "A":
assert v.shape[0] == r, (k, v.shape)
dst.mkdir(parents=True, exist_ok=True)
mx.save_safetensors(str(dst / "adapters.safetensors"), out)
json.dump({
"fine_tune_type": "lora",
"num_layers": max(layers) + 1,
"lora_parameters": {"rank": r, "scale": alpha / r, "dropout": 0.0, "keys": sorted(keys)},
}, open(dst / "adapter_config.json", "w"), indent=1)
print(f"{len(out)} tensors, {len(layers)} layers, rank {r}, alpha {alpha} -> scale {alpha / r}, modules {sorted(keys)}")