Distributed¶
.. py:class:: TeacherPlacement(teacher_only_devices: list[int] =
- module:
silverspoon_kd.distributed
- canonical:
silverspoon_kd.distributed.teacher_placement.TeacherPlacement
Bases: :py:class:
objectConfiguration for dedicated-GPU teacher placement.
Specifies which physical GPUs are reserved for the teacher model and which distribution strategy to use across those GPUs.
- param teacher_only_devices:
Absolute physical GPU IDs (matching
nvidia-smioutput) dedicated to the teacher. These GPUs will be hidden from the student’s distributed setup.- param strategy:
Distribution strategy for the teacher across its dedicated GPUs. One of
"pp"(pipeline parallel),"tp"(tensor parallel), or"sharded"(FSDP full_shard).- param device_type:
Accelerator type. Auto-detected as
"cuda"if not set.- param wrap_cls:
Module class name(s) for FSDP wrapping granularity (
strategy="sharded"only). IfNone, uses the model’s_no_split_modulesor falls back to size-based policy.
.. py:attribute:: TeacherPlacement.teacher_only_devices
- module:
silverspoon_kd.distributed
- type:
list[int]
.. py:attribute:: TeacherPlacement.strategy
- module:
silverspoon_kd.distributed
- type:
str
- value:
‘pp’
.. py:attribute:: TeacherPlacement.device_type
- module:
silverspoon_kd.distributed
- type:
str | None
- value:
None
.. py:attribute:: TeacherPlacement.wrap_cls
- module:
silverspoon_kd.distributed
- type:
str | list[str] | None
- value:
None
.. py:method:: TeacherPlacement.init(teacher_only_devices: list[int] =
, strategy: str = ‘pp’, device_type: str | None = None, wrap_cls: str | list[str] | None = None) -> None - module:
silverspoon_kd.distributed
.. py:function:: setup_split_gpu(placement: ~silverspoon_kd.distributed.teacher_placement.TeacherPlacement) -> tuple[list[int], list[int]]
- module:
silverspoon_kd.distributed
Reorder
CUDA_VISIBLE_DEVICESfor split-GPU teacher placement.Must be called before
import torch(or at least before any CUDA context is created).The environment variable is set so that student GPUs occupy the lowest CUDA indices and teacher GPUs follow:
.. code-block:: text
teacher_only_devices=[0, 1] on a 4-GPU system → student physical GPUs: [2, 3] → CUDA_VISIBLE_DEVICES=2,3,0,1 cuda:0 → physical 2 (student) cuda:1 → physical 3 (student) cuda:2 → physical 0 (teacher) cuda:3 → physical 1 (teacher)- param placement:
TeacherPlacementwithteacher_only_devicesset.- returns:
Tuple of
(student_physical_gpus, remapped_teacher_cuda_indices).