Step 400 check: Mac -13.4 % vs GPU -35.0 %; next GPU run bf16 mixed with stage 2; findings in STATE and roadmap

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This commit is contained in:
Kral
2026-10-05 11:46:21 +02:00
parent 39dcdd9846
commit 3a9f9e75fc
2 changed files with 23 additions and 1 deletions

View File

@@ -1,6 +1,6 @@
# ABAP Danışman Modeli — Yol Haritası # ABAP Danışman Modeli — Yol Haritası
Durum: v20 · 2026-10-04 Durum: v21 · 2026-10-05
Şu anki konum: **Adım 1.3** Şu anki konum: **Adım 1.3**
## Adım 0 — Çerçeve ve kararlar ## Adım 0 — Çerçeve ve kararlar
@@ -121,6 +121,13 @@ Adım 1.6 tablosundan seçilir. Aday havuzu: sadece Apache 2.0 veya MIT lisansl
Bitti sayılır: tek baz model seçildi, gerekçesi yazıldı. Bitti sayılır: tek baz model seçildi, gerekçesi yazıldı.
## Aşama 1 GPU sonucu ve karar (2026-10-05)
- QLoRA (nf4) aşama 1 koşusu ~418. adımda durduruldu. Valid loss, Mac (MLX 4-bit) / GPU (nf4): adım 200 0,758 (-%10,7) / -%26,0; adım 400 0,735 (-%13,4) / -%35,0; baz Mac 0,849.
- PEFT→MLX dönüştürücü overfit testiyle kanıtlandı (Mac -%98,5). nf4 adaptörü MLX 4-bit baza iyi taşınmıyor.
- **Karar:** şimdilik GPU koşusu yok. Aşama 1 tek başına eğitilmez; aşama 1 korpusu, aşama 2 yörüngeleri hazır olunca tek bir bf16 LoRA koşusunda karıştırılır. Sonraki GPU koşusu bf16 baz kullanır; öncesinde bf16 bellek testi gerekir. Araç isimleri EPOD ABAP MCP'deki gibi kalır (generic_v0 kullanılmaz). ADT "save failed" yanıtı sunucuda düzeltilemez: proxy sözdizimi kontrolü ekler.
- HF test ve checkpoint repoları silindi; aşama 1 veri seti kaldı.
## Adım 3 — Eğitim verisi üretimi ## Adım 3 — Eğitim verisi üretimi
| Alt adım | İş | | Alt adım | İş |

View File

@@ -105,3 +105,18 @@ Training runs on HF Jobs with Unsloth, not on the Mac. No `mlx_lm` training.
on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right. on Mac) is not equal. The overfit test (one document) passed, so key names, transpose and scale are right.
- Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA, - Open: decide what to do. Options: (a) accept the base mismatch and train further; (b) train on a higher-precision base (bf16 LoRA,
larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base. larger GPU) so the Mac conversion matches; (c) evaluate on the GPU with the MLX-equivalent base.
## Findings and next GPU run (2026-10-05)
| Step | GPU valid loss (nf4 base) | GPU drop | Mac valid loss (MLX 4-bit, converted adapter) | Mac drop |
|---|---|---|---|---|
| 0 / base | 0.9228 | | 0.849 | |
| 200 | 0.6830 | -26.0 % | 0.758 | -10.7 % |
| 400 | 0.6003 | -35.0 % | 0.735 | -13.4 % |
- Step 400 Mac log: `runs/stage1/full_step400_valid_loss.log`. The Mac drop grows with the steps (-10.7 % to -13.4 %), but stays far below the GPU drop.
- Converter `train/peft_to_mlx.py` is proven (overfit test: Mac -98.5 %, GPU -99.97 %). The gap is the base mismatch: an nf4 adapter does not transfer well to the MLX 4-bit base.
- Decisions (Kral, 2026-10-05): no more GPU runs now. Stage 1 is not trained alone. The stage 1 corpus is mixed with the stage 2 trajectories in one bf16 LoRA run later. The next GPU run uses a bf16 base (no nf4). A bf16 memory test is needed before it (27B bf16 weights are about 54 GB; check GPU size, sequence length 16384, gradient checkpointing).
- Tool names stay as in the EPOD ABAP MCP server (generic_v0 not used). The ADT "save failed" response cannot be fixed on the server side; the proxy adds a syntax check instead.
- HF cleanup done: test, overfit and full adapter repos deleted (step 200 and 400 checkpoints too). Kept: dataset `erhankeseli/abap-stage1-data`. Local adapter folders and `data_overfit` deleted.
- `BUDGET_LIMIT_USD` = 161 (ledger 66.97 + 35 usage x 2.7), until 12 October.