Skip to content

fix(train): move the model to the device before the base evaluation - #1052

Merged
NandhaKishorM merged 1 commit into
NandhaKishorM:mainfrom
FerryQ:fix/finetune-device-before-eval
Oct 8, 2026
Merged

NandhaKishorM merged 1 commit into
NandhaKishorM:mainfrom
FerryQ:fix/finetune-device-before-eval

Conversation

@FerryQ

@FerryQ FerryQ commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

What

Unable to finetune on CUDA or MPS without workarounds.finetune now moves the model to the training device right after load_checkpoint, instead of only inside train_model. One line in laya/train.py, plus a regression test.

Why

Since 0.3.29 (#967), finetune scores the base checkpoint on eval_items before training. The model is still on CPU at that point, but calibration_records sends its inputs to the device, so every GPU fine-tune with calibration items (the default) or --eval crashes before training starts:

  • CUDA: RuntimeError: Expected all tensors to be on the same device, but got index is on cuda:0, different from other tensors on cpu
  • MPS: RuntimeError: Placeholder storage has not been allocated on MPS device!

CPU is unaffected, which is why the test suite did not catch it. notebooks/laya_finetune_typed_decisions_mps.py hits it too, since it calls finetune(device="mps") with calibration items.

I found it while fine-tuning multilingual on Czech text categorization, on a CUDA VM and on an Apple Silicon Mac. Minimal repro: any finetune(data, "multilingual", out, TrainConfig(epochs=1)) with 10 or more rows on CUDA or MPS.

How it was verified

New test EndToEndTests.test_model_is_on_the_device_before_the_base_evaluation. It records the order of model.to and calibration_records calls, so it catches the bug on CPU: it fails on main and passes with the fix.

python tests/test_train.py         # Ran 54 tests, OK
python tests/test_training.py      # 61 passed, 0 failed (same as main)
ruff check laya/ --select=E9,F63,F7,F82,F401,F811 --line-length=120   # All checks passed
python -m compileall -q laya/ tests/

finetune scores the base checkpoint on eval_items before train_model runs
(NandhaKishorM#967), but the model was only moved to the device inside train_model.
calibration_records sends its inputs to the device, so on CUDA and MPS the
pre-training pass mixed devices and finetune crashed:

  CUDA: Expected all tensors to be on the same device, but got index is
        on cuda:0, different from other tensors on cpu
  MPS:  Placeholder storage has not been allocated on MPS device!

CPU was unaffected, which is why the test suite did not catch it. The
regression test records the order of events, so it fails on CPU too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@NandhaKishorM
NandhaKishorM merged commit d7f63b4 into NandhaKishorM:main Oct 8, 2026
@NandhaKishorM

Copy link
Copy Markdown
Owner

Merged.

Moving the model to the device before the base evaluation is a real fix with a narrow window, and that window is new: #967 added the before-and-after comparison in the last release, which is the only thing that evaluates the base checkpoint before training moves the model. So the bug arrived with the feature, and nothing else would have hit it.

The symptom is a device mismatch on the base pass, which means the before column either crashes or never gets measured, and the whole point of that table is the comparison.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants