This chapter explains how to manage LXD cluster hosts within a CloudCIX region.
To take an LXD cluster node out of service for maintenance (without removing it from the cluster):
1. Disable scheduling
Prevent new VMs and containers from being placed on the node:
lxc cluster set <node> scheduler.instance=manual
2. Evacuate existing workloads
Migrate all running instances off the node to other cluster members:
lxc cluster evacuate <node>
3. Power off the host
The node can now be safely powered down for maintenance.
Once maintenance is complete and the node is back online:
1. Restore evacuated instances
Migrate the node’s instances back from their temporary hosts:
lxc cluster restore <node>
2. Re-enable scheduling
Allow new instances to be placed on the node again:
lxc cluster set <node> scheduler.instance=all
New and restored instances will then be eligible for placement on the node.
These checks and commands apply to the H100 NVSwitch host. For initial driver, Fabric Manager, and partition setup, see LXD Cluster Installation with Remote Ceph.
Check Fabric Manager:
systemctl status nvidia-fabricmanager --no-pager
systemctl is-enabled nvidia-fabricmanager
grep '^FABRIC_MODE=' /usr/share/nvidia/nvswitch/fabricmanager.cfg
The expected configuration is FABRIC_MODE=1.
Use the local partition helper at /opt/Fabric-Manager-Client/fmpm-lite to
list or change partition state:
cd /opt/Fabric-Manager-Client
./fmpm-lite list
./fmpm-lite activate <partition_id>
./fmpm-lite deactivate <partition_id>
Always run ./fmpm-lite list after activation or deactivation to confirm the
result. On this host, partitions 7–14 map to the eight single-GPU partitions:
Partition |
Physical GPU ID |
PCI address |
|---|---|---|
7 |
1 |
|
8 |
2 |
|
9 |
3 |
|
10 |
4 |
|
11 |
5 |
|
12 |
6 |
|
13 |
7 |
|
14 |
8 |
|
The expected state on this host is for partitions 7–14 to be active. Activate all eight manually if required:
cd /opt/Fabric-Manager-Client
for id in 7 8 9 10 11 12 13 14; do
./fmpm-lite activate "$id"
done
./fmpm-lite list
The activation command may return a non-zero value even when activation
succeeds; trust the state shown by list.
The cloudcix-h100-fabric.service systemd unit activates these partitions
at startup. Check its state and logs:
systemctl status cloudcix-h100-fabric.service --no-pager
systemctl is-enabled cloudcix-h100-fabric.service
journalctl -u cloudcix-h100-fabric.service --no-pager
It should be enabled and show active (exited) after running. Start it
manually if needed:
sudo systemctl start cloudcix-h100-fabric.service
Quick verification:
systemctl is-active nvidia-fabricmanager
systemctl is-enabled nvidia-fabricmanager
systemctl is-active cloudcix-h100-fabric.service
systemctl is-enabled cloudcix-h100-fabric.service
grep '^FABRIC_MODE=' /usr/share/nvidia/nvswitch/fabricmanager.cfg
/opt/Fabric-Manager-Client/fmpm-lite list
See also