Repository navigation
fix(train): move the model to the device before the base evaluation - #1052
Merged
NandhaKishorM merged 1 commit intoOct 8, 2026
Merged
Conversation
finetune scores the base checkpoint on eval_items before train_model runs (NandhaKishorM#967), but the model was only moved to the device inside train_model. calibration_records sends its inputs to the device, so on CUDA and MPS the pre-training pass mixed devices and finetune crashed: CUDA: Expected all tensors to be on the same device, but got index is on cuda:0, different from other tensors on cpu MPS: Placeholder storage has not been allocated on MPS device! CPU was unaffected, which is why the test suite did not catch it. The regression test records the order of events, so it fails on CPU too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Owner
|
Merged. Moving the model to the device before the base evaluation is a real fix with a narrow window, and that window is new: #967 added the before-and-after comparison in the last release, which is the only thing that evaluates the base checkpoint before training moves the model. So the bug arrived with the feature, and nothing else would have hit it. The symptom is a device mismatch on the base pass, which means the before column either crashes or never gets measured, and the whole point of that table is the comparison. |
4 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Unable to
finetuneon CUDA or MPS without workarounds.finetunenow moves the model to the training device right afterload_checkpoint, instead of only insidetrain_model. One line inlaya/train.py, plus a regression test.Why
Since 0.3.29 (#967),
finetunescores the base checkpoint oneval_itemsbefore training. The model is still on CPU at that point, butcalibration_recordssends its inputs to the device, so every GPU fine-tune with calibration items (the default) or--evalcrashes before training starts:RuntimeError: Expected all tensors to be on the same device, but got index is on cuda:0, different from other tensors on cpuRuntimeError: Placeholder storage has not been allocated on MPS device!CPU is unaffected, which is why the test suite did not catch it.
notebooks/laya_finetune_typed_decisions_mps.pyhits it too, since it callsfinetune(device="mps")with calibration items.I found it while fine-tuning
multilingualon Czech text categorization, on a CUDA VM and on an Apple Silicon Mac. Minimal repro: anyfinetune(data, "multilingual", out, TrainConfig(epochs=1))with 10 or more rows on CUDA or MPS.How it was verified
New test
EndToEndTests.test_model_is_on_the_device_before_the_base_evaluation. It records the order ofmodel.toandcalibration_recordscalls, so it catches the bug on CPU: it fails onmainand passes with the fix.