This page lists the most common High Availability problems and how to resolve them. Before you work through these checks, confirm that your cluster meets the
requirements and limitations.
HA does not trigger an evacuation event after a compute resource goes offline#
Verify that you enabled HA globally in Settings > High Availability.
Verify that the affected compute resource belongs to a failover domain with HA enabled.
Verify that the configured grace period has elapsed. HA does not trigger an event until the compute resource has been unreachable for the full grace period.
Check whether HA has reached the Maximum concurrent failures threshold. If HA has already reached the maximum number of concurrent evacuation events, it suppresses new events.
Enabling High Availability for a virtual server fails (NFS storage)#
On NFS failover domains, each HA-protected virtual server needs a storage lease. The system grants the lease only when the virtual server starts after you enable HA on its compute resource. If the virtual server was already running, HA may fail to protect it.
To resolve this: stop the virtual server, enable High Availability, then start it again.
A compute resource shows “Not protected” HA status#
A compute resource shows the Not protected status in one of the following cases:
The task that adds the compute resource to the failover domain failed or was cancelled. Go to Tasks, find the failed join failover domain task, and check the details.
On a Shared LVM failover domain, SolusVM has not read back the storage lock host ID of the compute resource yet. On Shared LVM, you configure storage locking on each compute resource in advance, and joining a failover domain only records the host ID. Verify the storage lock configuration on the compute resource.
The compute resource is the only member of its failover domain. A failover domain needs at least two compute resources to provide failover. Add at least one more compute resource to the domain.
HA marks virtual servers as “Not evacuated” after an event#
Verify that at least one destination compute resource in the failover domain is online with sufficient free CPU, memory, and storage.
Expand the failed event row on High Availability > Activities, then review the per-server error details.