Distributed

.. py:class:: TeacherPlacement(teacher_only_devices: list[int] = , strategy: str = ‘pp’, device_type: str | None = None, wrap_cls: str | list[str] | None = None)

module:

silverspoon_kd.distributed

canonical:

silverspoon_kd.distributed.teacher_placement.TeacherPlacement

Bases: :py:class:object

Configuration for dedicated-GPU teacher placement.

Specifies which physical GPUs are reserved for the teacher model and which distribution strategy to use across those GPUs.

param teacher_only_devices:

Absolute physical GPU IDs (matching nvidia-smi output) dedicated to the teacher. These GPUs will be hidden from the student’s distributed setup.

param strategy:

Distribution strategy for the teacher across its dedicated GPUs. One of "pp" (pipeline parallel), "tp" (tensor parallel), or "sharded" (FSDP full_shard).

param device_type:

Accelerator type. Auto-detected as "cuda" if not set.

param wrap_cls:

Module class name(s) for FSDP wrapping granularity (strategy="sharded" only). If None, uses the model’s _no_split_modules or falls back to size-based policy.

.. py:attribute:: TeacherPlacement.teacher_only_devices

module:

silverspoon_kd.distributed

type:

list[int]

.. py:attribute:: TeacherPlacement.strategy

module:

silverspoon_kd.distributed

type:

str

value:

‘pp’

.. py:attribute:: TeacherPlacement.device_type

module:

silverspoon_kd.distributed

type:

str | None

value:

None

.. py:attribute:: TeacherPlacement.wrap_cls

module:

silverspoon_kd.distributed

type:

str | list[str] | None

value:

None

.. py:method:: TeacherPlacement.init(teacher_only_devices: list[int] = , strategy: str = ‘pp’, device_type: str | None = None, wrap_cls: str | list[str] | None = None) -> None

module:

silverspoon_kd.distributed

.. py:function:: setup_split_gpu(placement: ~silverspoon_kd.distributed.teacher_placement.TeacherPlacement) -> tuple[list[int], list[int]]

module:

silverspoon_kd.distributed

Reorder CUDA_VISIBLE_DEVICES for split-GPU teacher placement.

Must be called before import torch (or at least before any CUDA context is created).

The environment variable is set so that student GPUs occupy the lowest CUDA indices and teacher GPUs follow:

.. code-block:: text

teacher_only_devices=[0, 1]  on a 4-GPU system
→ student physical GPUs: [2, 3]
→ CUDA_VISIBLE_DEVICES=2,3,0,1
    cuda:0 → physical 2  (student)
    cuda:1 → physical 3  (student)
    cuda:2 → physical 0  (teacher)
    cuda:3 → physical 1  (teacher)
param placement:

TeacherPlacement with teacher_only_devices set.

returns:

Tuple of (student_physical_gpus, remapped_teacher_cuda_indices).