Managing LXD Hosts

This chapter explains how to manage LXD cluster hosts within a CloudCIX region.

Taking a Cluster Member Out of Service

To take an LXD cluster node out of service for maintenance (without removing it from the cluster):

1. Disable scheduling

Prevent new VMs and containers from being placed on the node:

lxc cluster set <node> scheduler.instance=manual

2. Evacuate existing workloads

Migrate all running instances off the node to other cluster members:

lxc cluster evacuate <node>

3. Power off the host

The node can now be safely powered down for maintenance.

Returning a Node to Service

Once maintenance is complete and the node is back online:

1. Restore evacuated instances

Migrate the node’s instances back from their temporary hosts:

lxc cluster restore <node>

2. Re-enable scheduling

Allow new instances to be placed on the node again:

lxc cluster set <node> scheduler.instance=all

New and restored instances will then be eligible for placement on the node.

Managing H100 NVSwitch Fabric Manager

These checks and commands apply to the H100 NVSwitch host. For initial driver, Fabric Manager, and partition setup, see LXD Cluster Installation with Remote Ceph.

Check Fabric Manager:

systemctl status nvidia-fabricmanager --no-pager
systemctl is-enabled nvidia-fabricmanager
grep '^FABRIC_MODE=' /usr/share/nvidia/nvswitch/fabricmanager.cfg

The expected configuration is FABRIC_MODE=1.

Use the local partition helper at /opt/Fabric-Manager-Client/fmpm-lite to list or change partition state:

cd /opt/Fabric-Manager-Client
./fmpm-lite list
./fmpm-lite activate <partition_id>
./fmpm-lite deactivate <partition_id>

Always run ./fmpm-lite list after activation or deactivation to confirm the result. On this host, partitions 7–14 map to the eight single-GPU partitions:

H100 host single-GPU partition mapping

Partition

Physical GPU ID

PCI address

7

1

0a:00.0

8

2

18:00.0

9

3

3c:00.0

10

4

45:00.0

11

5

87:00.0

12

6

90:00.0

13

7

b9:00.0

14

8

c2:00.0

The expected state on this host is for partitions 7–14 to be active. Activate all eight manually if required:

cd /opt/Fabric-Manager-Client
for id in 7 8 9 10 11 12 13 14; do
    ./fmpm-lite activate "$id"
done
./fmpm-lite list

The activation command may return a non-zero value even when activation succeeds; trust the state shown by list.

The cloudcix-h100-fabric.service systemd unit activates these partitions at startup. Check its state and logs:

systemctl status cloudcix-h100-fabric.service --no-pager
systemctl is-enabled cloudcix-h100-fabric.service
journalctl -u cloudcix-h100-fabric.service --no-pager

It should be enabled and show active (exited) after running. Start it manually if needed:

sudo systemctl start cloudcix-h100-fabric.service

Quick verification:

systemctl is-active nvidia-fabricmanager
systemctl is-enabled nvidia-fabricmanager
systemctl is-active cloudcix-h100-fabric.service
systemctl is-enabled cloudcix-h100-fabric.service
grep '^FABRIC_MODE=' /usr/share/nvidia/nvswitch/fabricmanager.cfg
/opt/Fabric-Manager-Client/fmpm-lite list

See also