Repository navigation
Replies: 1 comment
|
Hi @polarlily, It looks like your training config for the second run was set to use To fix this from the GUI:
That should let training pick up your GPU correctly. Let us know if that works for you! Thanks, |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi, I had migrated my project from Sleap v.1.4.2 to the current updated version. Training a new model using the frames I had labeled worked well, but when I tried to train the model again off another set of labels, the prompt "An error occurred while training centroid. Your command line terminal may have more information about the error." I'm currently running Sleap through the GUI and I installed it through Anaconda Terminal. This is what showed up with the error:
val_labels_path: null
validation_fraction: 0.1
use_same_data_for_val: false
test_file_path: null
provider: LabelsReader
user_instances_only: true
data_pipeline_fw: torch_dataset_cache_img_memory
cache_img_path: null
use_existing_imgs: false
delete_cache_imgs_after_training: true
parallel_caching: true
cache_workers: 0
preprocessing:
ensure_rgb: false
ensure_grayscale: true
max_height: 1200
max_width: 3200
scale: 0.25
crop_size: null
min_crop_size: 100
crop_padding: null
use_augmentations_train: true
augmentation_config:
intensity:
uniform_noise_min: 0.0
uniform_noise_max: 0.04
uniform_noise_p: 0.0
gaussian_noise_mean: 0.0
gaussian_noise_std: 0.02
gaussian_noise_p: 0.0
contrast_min: 0.9
contrast_max: 1.1
contrast_p: 0.0
brightness_min: 0.9
brightness_max: 1.1
brightness_p: 0.0
geometric:
rotation_min: -180.0
rotation_max: 180.0
rotation_p: 1.0
scale_min: 0.9
scale_max: 1.1
scale_p: 1.0
translate_width: 0.0
translate_height: 0.0
translate_p: null
affine_p: 0.0
erase_scale_min: 0.0001
erase_scale_max: 0.01
erase_ratio_min: 1.0
erase_ratio_max: 1.0
erase_p: 0.0
mixup_lambda_min: 0.01
mixup_lambda_max: 0.05
mixup_p: 0.0
use_negative_frames: false
negative_loss_weight: 1.0
skeletons: []
model_config:
init_weights: default
pretrained_backbone_weights: D:/Tobi Sleap Vids/2026_data/models/NewAttemptOne.centroid.n=601/best.ckpt
pretrained_head_weights: D:/Tobi Sleap Vids/2026_data/models/NewAttemptOne.centroid.n=601/best.ckpt
backbone_config:
unet:
in_channels: 1
kernel_size: 3
filters: 16
filters_rate: 2.0
max_stride: 16
stem_stride: null
middle_block: true
up_interpolate: true
stacks: 1
convs_per_block: 2
output_stride: 2
convnext: null
swint: null
head_configs:
single_instance: null
centroid:
confmaps:
anchor_part: mid_body
sigma: 3.0
output_stride: 2
centered_instance: null
bottomup: null
multi_class_bottomup: null
multi_class_topdown: null
total_params: 1953105
trainer_config:
train_data_loader:
batch_size: 5
num_workers: 0
shuffle: true
val_data_loader:
batch_size: 5
num_workers: 0
shuffle: false
model_ckpt:
save_top_k: 1
save_last: false
trainer_devices: 1
trainer_device_indices: null
trainer_accelerator: mps
profiler: null
trainer_strategy: auto
enable_progress_bar: true
min_train_steps_per_epoch: 200
train_steps_per_epoch: 200
visualize_preds_during_training: true
keep_viz: false
max_epochs: 200
seed: null
use_wandb: false
save_ckpt: true
ckpt_dir: D:/Tobi Sleap Vids/2026_data/models
run_name: 260717_161045.centroid.n=634
resume_ckpt_path: null
wandb:
entity: null
project: null
name: null
save_viz_imgs_wandb: false
api_key: null
wandb_mode: null
prv_runid: null
group: null
current_run_id: null
viz_enabled: true
viz_boxes: false
viz_masks: false
viz_box_size: 5.0
viz_confmap_threshold: 0.1
log_viz_table: false
delete_local_logs: null
optimizer_name: Adam
optimizer:
lr: 0.0001
amsgrad: false
lr_scheduler:
step_lr: null
reduce_lr_on_plateau:
threshold: 1.0e-06
threshold_mode: abs
cooldown: 3
patience: 5
factor: 0.5
min_lr: 1.0e-08
cosine_annealing_warmup: null
linear_warmup_linear_decay: null
early_stopping:
min_delta: 1.0e-08
patience: 10
stop_training_on_plateau: true
online_hard_keypoint_mining:
online_mining: false
hard_to_easy_ratio: 2.0
min_hard_keypoints: 2
max_hard_keypoints: null
loss_scale: 5.0
zmq:
controller_port: 9000
controller_polling_timeout: 10
publish_port: 9001
eval:
enabled: true
frequency: 1
oks_stddev: 0.025
oks_scale: null
match_threshold: 50.0
name: ''
description: ''
sleap_nn_version: 0.2.0
filename: ''
2026-07-17 16:10:55 | Started training at: 2026-07-17 16:10:55.518100
2026-07-17 16:10:55 | sleap-nn 0.2.0 | Python 3.13.14 | PyTorch 2.13.0+cu132 | CUDA 13.2 | 1 GPU(s)
2026-07-17 16:10:55 | Creating train-val split...
2026-07-17 16:11:00 | # Train Labeled frames: 570
2026-07-17 16:11:00 | # Val Labeled frames: 63
2026-07-17 16:11:00 | Setting up config...
2026-07-17 16:11:00 | Setting up for training...
2026-07-17 16:11:00 | Setting up model ckpt dir:
D:/Tobi Sleap Vids/2026_data/models/260717_161045.centroid.n=634...2026-07-17 16:11:02 | Setting up visualization train and val datasets...
2026-07-17 16:11:02 | Setting up Trainer...
2026-07-17 16:11:02 | Setting up callbacks and loggers...
2026-07-17 16:11:02 | Training controller subscribed to: tcp://127.0.0.1:9000 (topic: )
2026-07-17 16:11:02 | ProgressReporterZMQ publishing to tcp://127.0.0.1:9001 for ''
2026-07-17 16:11:02 | Trainer devices: 1
Traceback (most recent call last):
File "", line 203, in _run_module_as_main
File "", line 88, in _run_code
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\sleap\cli.py", line 1183, in
cli()
~~~^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\rich_click\rich_command.py", line 402, in call
return super().call(*args, **kwargs)
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\click\core.py", line 1569, in call
return self.main(*args, **kwargs)
~~~~~~~~~^^^^^^^^^^^^^^^^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\rich_click\rich_command.py", line 216, in main
rv = self.invoke(ctx)
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\click\core.py", line 1970, in invoke
return _process_result(sub_ctx.command.invoke(sub_ctx))
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\click\core.py", line 1353, in invoke
return ctx.invoke(self.callback, **ctx.params)
~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\click\core.py", line 907, in invoke
return callback(*args, **kwargs)
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\sleap_nn\cli.py", line 534, in train
run_training(config=cfg, train_labels=train_labels, val_labels=val_labels)
~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\sleap_nn\train.py", line 44, in run_training
trainer.train()
~~~~~~~~~~~~~^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\sleap_nn\training\model_trainer.py", line 1191, in train
self.trainer = L.Trainer(
~~~~~~~~~^
callbacks=callbacks,
^^^^^^^^^^^^^^^^^^^^
...<8 lines>...
log_every_n_steps=1,
^^^^^^^^^^^^^^^^^^^^
)
^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\lightning\pytorch\utilities\argparse.py", line 70, in insert_env_defaults
return fn(self, **kwargs)
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\lightning\pytorch\trainer\trainer.py", line 417, in init
self._accelerator_connector = _AcceleratorConnector(
~~~~~~~~~~~~~~~~~~~~~^
devices=devices,
^^^^^^^^^^^^^^^^
...<8 lines>...
plugins=plugins,
^^^^^^^^^^^^^^^^
)
^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\lightning\pytorch\trainer\connectors\accelerator_connector.py", line 145, in init
self._set_parallel_devices_and_init_accelerator()
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~^^
File "C:\Users\tobi\AppData\Roaming\uv\tools\sleap\Lib\site-packages\lightning\pytorch\trainer\connectors\accelerator_connector.py", line 356, in _set_parallel_devices_and_init_accelerator
raise MisconfigurationException(
...<4 lines>...
)
lightning.fabric.utilities.exceptions.MisconfigurationException:
MPSAcceleratorcan not run on your system since the accelerator is not available. The following accelerator(s) is available and can be passed intoacceleratorargument ofTrainer: ['cpu', 'cuda'].Run Path: D:/Tobi Sleap Vids/2026_data/models/260717_161045.centroid.n=634
I wasn't sure but it seems like I would need to change the 'accelerator' argument to either cpu or cuda, but I'm not sure how to do that since I work mainly off the gui. I would love some guidance on this. I attached photos as well.
All reactions