High Availability
On this page
- Overview
- How High Availability Works
- Known Limitations
- Configure High Availability
- Failover Domains
- Compute Resources
- Failover Behavior
- High Availability Activities
- After a Failover Event
- Webhook Notifications for High Availability Events
- Offer High Availability to Users
- Troubleshooting
- HA does not trigger an evacuation event after a compute resource goes offline
- Enabling High Availability for a virtual server fails (NFS storage)
- A compute resource shows “Not protected” HA status
- HA marks virtual servers as “Not evacuated” after an event
- A “Move back” action fails
- HA configuration screens display a warning that HA is off
Overview
High Availability (HA) automatically recovers protected virtual servers when a compute resource fails. HA uses the shared-storage lock to confirm a failure and prevent false failover events. After HA confirms the failure and the grace period expires, it places the affected virtual servers on a healthy compute resources within the same failover domain.
A failover domain is a named group of compute resources that share storage and provide capacity for HA recovery.
Warning:
SolusVM supports High Availability only for KVM virtual servers that use Shared LVM (iSCSI) or NFS (Network File System) storage. Containers are not supported.
This documentation explains how to enable the High Availability protection feature, create and configure a failover domain, and prepare compute resources for automatic failover during a failure.
Before you offer High Availability protection to users, read the following documentation and review the requirements and limitations.
How High Availability Works
The following steps describe how High Availability protects a virtual server when a compute resource fails:
- Compute resources in the failover domain continuously report their shared-storage lock status.
- SolusVM confirms a compute resource failure when the other compute resources in the failover domain confirm that its storage lock has expired.
- After the grace period expires, SolusVM places the affected virtual servers on the healthy compute resources in the same failover domain. It places each virtual server individually. They may all land on one compute resource or spread across several, depending on available resources.
- SolusVM places each protected virtual server on the selected compute resource. The virtual server disks stay on the shared storage and do not move.
- If a virtual server was running when the failure occurred, HA starts it on the new compute resource.
- The virtual server keeps the same IP address, so services remain available through the same network address after a failover.
- When the original compute resource comes back online, the virtual server continues to run on the new compute resource. You must manually move the virtual server back to the original compute resource.
Note:
- High Availability does not rely on the SolusVM agent or the connector between the management node and a compute resource to confirm a failure. If the management node loses network access to a compute resource but that compute resource still holds its storage lock, HA considers it active and does not start an evacuation. This prevents false failover events when a network connection fails, but the compute resource and its virtual servers continue to run.
- High Availability and Disaster Recovery can coexist. If HA cannot migrate a virtual server and that virtual server has a valid backup, you can perform Disaster Recovery manually.
Known Limitations
The following limitations apply to the High Availability feature:
- A failover domain requires at least two compute resources. A domain with one compute resource cannot provide failover, and the system reports it as Not protected.
- All compute resources in the failover domain must run KVM virtual servers and use Shared LVM (iSCSI) or NFS (Network File System) storage.
- SolusVM does not support Virtuozzo compute resources for High Availability, even when they host KVM virtual servers.
- Each compute resource must have enough CPU and memory to run virtual servers after a failover.
- A failover domain can contain up to 2,000 compute resources.
- Use dedicated shared storage for each failover domain. Use one storage for each domain. Do not use the same storage for multiple failover domains. SolusVM does not block this setup, but it does not support it. Shared storage creates one lock space, which can cause the domains to interfere with each other’s failover decisions.
- You cannot change the location or the storage of a failover domain after you create it. To use a different storage, create a new failover domain and move the virtual servers to it.
- Manage failover domain membership and the IP blocks of compute resources from the failover domain pages. You cannot manage these settings from the compute resource creation or edit forms. You also cannot attach IP blocks to a compute resource that already belongs to a failover domain.
- You cannot enable HA for a plan that is linked to an offer with a storage tag, because High Availability places virtual server disks on the failover domain storage. Remove the storage tag from the linked offers first.
- SolusVM places additional disks that you add to an HA virtual server after creation on the failover domain storage, and rejects storage tags for those disks.
Configure High Availability
To move virtual servers to a healthy compute resource after a failure, High Availability requires the following:
- Create and configure a failover domain to use the High Availability feature.
- At least one other healthy compute resource exists in the same failover domain.
- Destination compute resources have enough free CPU and memory to host the virtual servers during evacuation.
- All compute resources in the failover domain use the KVM virtualization type.
- All compute resources in the failover domain use Shared LVM (iSCSI) or NFS (Network File System) storage.
The following documentation explains how to configure High Availability protection and manage failover domains.
Enable High Availability
You must enable High Availability protection globally before any failover domain configuration takes effect.
Note:
The account must have the MANAGE_HIGH_AVAILABILITY permission to access the Settings > High Availability interface. For more information, see User Access Control.
Go to Settings > High Availability.
Select the Enable checkbox.
Configure the following settings:
- Grace period: The number of seconds a compute resource must remain unreachable before HA starts an evacuation.
Note:
- The default value is 300 seconds, and the minimum value is 140 seconds. HA needs at least 140 seconds to confirm that a compute resource is down.
- Set a longer grace period if you want to give the compute resource more time to recover. For example, if you set the grace period to 160 seconds, HA starts an evacuation only if the compute resource remains unreachable for 160 seconds.
- Maximum concurrent failures: The maximum number of compute resources that can fail at the same time in a single failover domain before the system stops automatic failover.
Note:
- The default value is 3, and the minimum value is 1.
- If more compute resources than this fail at the same time, SolusVM treats the failover domain as the likely cause of the failures. In this case, the system stops failover for all compute resources in the domain until an administrator investigates the issue. Evacuation would not help in this case, because the virtual server disks use the same shared storage.
Click Save.
Disabling High Availability globally preserves all failover domains and plan configurations. The configurations take effect again when you re-enable HA. When you disable HA, the system shows a warning on the Failover Domains interface. The HA status of each compute resource also shows a warning. The warning includes a link to Settings > High Availability, where you can re-enable HA.
Failover Domains
A failover domain groups compute resources that can host and recover virtual servers with High Availability protection. Each compute resource can belong to only one failover domain at a time. HA evacuates virtual servers only to compute resources within the same failover domain.
To view and manage your failover domains, go to High Availability > Failover Domains. The Failover Domains interface lists all configured failover domains, including the number of compute resources in each domain, their locations, storage types, assigned IP blocks, and active status.
Create a Failover Domain
You must create and configure a failover domain to use High Availability.
To create a failover domain, perform the following steps:
Go to High Availability > Failover Domains.
Click Add Domain. The New Failover Domain interface appears.
Enter a name for the failover domain. The name must be unique. Optionally, enter a description. The description can span multiple lines and contain up to 10,000 characters.
Select a location.
Select the network storage at the location you chose. SolusVM displays a list of all compute resources that match your location and storage selection.
Select at least two compute resources to include in the failover domain.
Click Next.
Select IP blocks.
Note:
All compute resources in a failover domain must use matching IP address blocks. Mismatched or incompatible IP ranges can cause virtual servers to lose network connectivity after a failover. Verify that all IP blocks you select can route traffic to the internet. When you click Create, SolusVM applies these IP blocks to all compute resources in the failover domain.
Click Create.
After you create a failover domain, you can remove compute resources from it until it has one or no members. To delete the failover domain, remove all compute resources first.
Manage a Failover Domain
To manage a failover domain, go to High Availability > Failover Domains and click the domain. The failover domain page displays a health and resource summary for that domain.

The summary includes the following information:
- Compute Resources — Shows the total number of compute resources in the failover domain and their status. For more information, see High Availability Protection Status.
- RAM — Shows the total RAM across all compute resources in the domain and their status:
- Online RAM
- Offline RAM
- RAM in use by HA-protected virtual servers
- RAM in use by non-HA virtual servers
- Free RAM
- Storage — Shows the failover domain’s shared storage and its free space. Click the storage name to manage storage for this domain.
- IP Blocks — Lists the IP blocks assigned to the failover domain. Click IP Blocks to manage them.
- Recent activities — Shows the number of evacuation events that occurred in the last seven days. Click the arrow to go to the Activities interface.
Below the summary, the compute resource table lists all compute resources in the failover domain with the following columns:
| Column | Description |
|---|---|
| ID | Unique identifier of the compute resource. |
| Compute Resource | Hostname and IP address of the compute resource. |
| Status | Current connection status of the compute resource, such as Active, Running, or Unavailable. |
| HA Protection | HA protection status of the compute resource. See High Availability Protection Status. |
| Version | SolusVM Agent version running on the compute resource. |
| Servers (HA) | Total virtual servers on the compute resource. |
| RAM Usage | RAM in use out of the total RAM available on the compute resource. |
| Locked | Whether the compute resource is locked. |
| Maintenance | Whether the compute resource is in maintenance mode. |
Disable a Failover Domain
You can disable HA for a specific failover domain in two ways, without removing its configuration.
From the Failover Domains list:
- Go to High Availability > Failover Domains.
- Locate the failover domain and set its Active toggle to inactive.
From the failover domain detail page:
- Go to High Availability > Failover Domains.
- Click the failover domain you want to disable.
- In the upper right corner, set the Active toggle to inactive.
Note:
When you disable High Availability for a domain, the system does not trigger evacuation events for compute resources in that domain, even when the global HA setting is enabled.
Remove a Compute Resource from a Failover Domain
When you remove a compute resource from a failover domain, the compute resource stays in the cluster, but High Availability no longer protects it or its virtual servers.
Before you remove a compute resource, verify that it does not host any virtual servers with High Availability enabled. Disable HA on those virtual servers or migrate them to another compute resource first. If HA-protected virtual servers remain on the compute resource, SolusVM cancels the operation and displays the following message:
Compute resource "<hostname>" still has high-availability VMs. Disable high availability on them or migrate them to another compute resource before removing it from the failover domain.To remove a compute resource from a failover domain, perform the following steps:
- Go to High Availability > Failover Domains.
- Click the failover domain that contains the compute resource.
- In the compute resource table, find the compute resource and click the Exclude from domain
icon. - Confirm the action.
Note:
You cannot move a compute resource directly from one failover domain to another. Remove the compute resource from its current failover domain first, and then add it to the new failover domain.
Delete a Failover Domain
- Go to High Availability > Failover Domains.
- Remove all compute resources from the failover domain.
- Locate the failover domain and click the Delete failover domain
icon.
Warning:
Deleting a failover domain removes High Availability protection from all virtual servers in that domain. If a host fails, SolusVM does not automatically restart the affected virtual servers on another host. You cannot undo this action.
Compute Resources
Go to Compute Resources to view the High Availability protection status of all compute resources in your cluster. This interface shows which compute resources HA protects, which ones are at risk, and which ones need attention.
High Availability Protection Status
The HA Protection column shows one of the following statuses for each compute resource:
| Status | Description |
|---|---|
| HA covers this compute resource. The system can evacuate its virtual servers in the event of a failure. | |
| The compute resource does not have enough available resources, such as RAM, to guarantee the successful evacuation of all virtual servers. The system recalculates this status on a schedule and caches the result for 10 minutes, so the status can lag behind a recent capacity change by up to 10 minutes. | |
| HA does not currently protect this compute resource. For the possible causes and how to resolve them, see A compute resource shows “Not protected” HA status. | |
| This compute resource meets the High Availability requirements, and you can add it to a failover domain. | |
| This compute resource does not meet one or more High Availability requirements, and you cannot add it to a failover domain. For example, compute resources that host VZ containers are not eligible because HA supports only KVM virtual servers. | |
| You have disabled High Availability for this compute resource’s failover domain. See how to disable a failover domain. | |
| The system cannot reach the compute resource, usually because the host is offline. This status can also appear temporarily while a compute resource joins or leaves a failover domain. |
Maintenance Mode
Use maintenance mode when you need to perform planned work on a compute resource, such as hardware upgrades or software updates, without triggering or interfering with HA events.
When you put a compute resource in maintenance mode, SolusVM excludes it from HA evacuation in both directions:
- The system does not select the compute resource as a destination for virtual server migration.
- If the compute resource goes down while in maintenance mode, the system does not trigger an HA evacuation event for the virtual servers it hosts.
Note:
Before you put a compute resource in maintenance mode, verify that the remaining compute resources in the failover domain have enough capacity to handle a failover. If capacity is insufficient, the At risk status appears for the affected virtual servers.
To enable or disable maintenance mode for a compute resource, use the Maintenance toggle on the Compute Resources interface. For more information, see Compute Resource Maintenance Mode documentation.
Failover Behavior
High Availability starts an evacuation event automatically when all of the following conditions apply:
- The other compute resources in the failover domain confirm that the storage lock of the affected compute resource has expired, which means its watchdog has fired.
- The confirmed failure persists for longer than the grace period.
- The compute resource completed its failover domain join successfully, and it is not in maintenance mode.
- The compute resource belongs to an active failover domain.
- The HA global setting is enabled.
HA does not evacuate a compute resource that the management node cannot reach if that compute resource still holds its storage lock.
When a failure occurs, HA migrates each protected virtual server to a healthy compute resource within the failover domain. IP addresses do not change after a failover. Users can access all services through the same address regardless of which compute resource hosts the virtual server.
High Availability Activities
Go to High Availability > Activities to view all past and ongoing evacuation events. These include moves to a healthy compute resource after a failure and moves back to the original compute resource after recovery.
Each row represents one evacuation event — one compute resource whose failure triggered a recovery. Expand a row to view the migration status for individual virtual servers in that event.

The event row shows the following information:
| Column | Description |
|---|---|
| ID | Unique identifier for the evacuation event. |
| Compute Resource / Server | The source compute resource whose failure triggered the event. |
| Status | The event status. When an event finishes with failed or cancelled virtual servers, a second line shows how many. See Event Statuses. |
| Start time | Timestamp when the system started the event. After the event finishes, a second line shows when it ended. |
| Duration | Elapsed time for the event. |
| Processes | The number of virtual servers included in the event. |
| Failover Domain | The failover domain associated with the event. Click the name to open the activities of that failover domain. |
Expand an event row to see the following information about each virtual server:
| Column | Description |
|---|---|
| ID | Unique identifier for the process. |
| Server name | The virtual server that the system is recovering. |
| Status | The recovery status of this virtual server. See Per-Server Statuses. |
| Start time | Timestamp when the recovery of this virtual server started. |
| Duration | Elapsed time for this virtual server. |
| Destination CR | The compute resource on which the system recreated this virtual server. |
| Failover Domain | The failover domain associated with the event. |
Note:
When you open the Recent activities interface for a single failover domain, SolusVM hides the Failover Domain column because all events belong to that domain.
Event Statuses
An event status describes the evacuation of an entire compute resource. You can filter the list by any of these statuses.
| Status | Description |
|---|---|
| The evacuation event is in queue and waiting to begin. | |
| Evacuation is in progress. | |
| Evacuation is in progress, and one or more virtual servers have already failed. | |
| The system evacuated all virtual servers successfully. | |
| The system successfully evacuated some virtual servers, but not others. | |
| The system did not evacuate any virtual servers. | |
| The system evacuated none of the virtual servers in the event. A second line shows the counts. | |
| The system cancelled the event. For example, the source compute resource came back online before the evacuation started. |
Per-Server Statuses
| Status | Description |
|---|---|
| The evacuation event is in queue and waiting to begin. | |
| The system is recreating the virtual server on the destination compute resource. A second line shows the progress as a percentage. | |
| The system recreated the virtual server on the destination compute resource successfully. If the virtual server was running before the failure, the system also started it. | |
| The system could not recover the virtual server. Point to the status to see the reason for the failure. | |
| The system cancelled the recovery of this virtual server before it ran. For example, the source compute resource came back online before the evacuation started. | |
| The moving back event is queued and waiting to begin. | |
| A moving back migration is in progress. | |
| The system successfully moved back all virtual servers to their source compute resource. | |
| The move back failed. The virtual server continues to run on the destination compute resource. |
Available Actions
Use the following actions to retry failed evacuations or move virtual servers back to their source compute resource after recovery:
- Retry: Starts a new evacuation attempt for virtual servers that failed during the evacuation event. For more information, see Retry a Failed Evacuation.
- Move back: Migrates evacuated virtual servers back to the source compute resource. Available only after the source compute resource is back online and reachable. For more information, see Move Virtual Servers Back to the Source Compute Resource.
After a Failover Event
After an evacuation event completes, the next steps depend on whether you can recover the source compute resource:
- If you cannot recover the source compute resource, delete it from Compute Resources to avoid IP and network conflicts.
- If the source compute resource comes back online, you can manually migrate the evacuated virtual servers back to it.
Warning:
High Availability does not automatically return virtual servers to the source compute resource when it comes back online. Moving virtual servers back is always a manual action.
Move Virtual Servers Back to the Source Compute Resource
Once the source compute resource is back online, you can migrate the evacuated virtual servers back to it.
- Go to High Availability > Activities.
- Locate the evacuation event.
- Select the compute resource or specific virtual servers that you want to move back.
- Click Move back
.
Note:
If you cannot recover the source compute resource, migrate the evacuated virtual servers to a permanent destination using the standard Migration of Servers Between Compute Resources procedure.
Retry a Failed Evacuation
If a compute resource (or a virtual server) failed to migrate during an evacuation event, you can retry the migration.
Before retrying, verify that at least one destination compute resource in the failover domain is online and has sufficient CPU and memory.
- Go to High Availability > Activities.
- Locate the event with virtual servers that failed to migrate.
- Select the checkboxes for the virtual servers you want to retry, or select the event row to retry all failed migrations.
- Click Retry
.
Webhook Notifications for High Availability Events
You can configure one or more webhook URLs to receive notifications when a High Availability event starts or completes.
HA supports the following webhook events:
- HA incident recovery started: SolusVM detects a compute resource failure and starts the recovery process.
- HA incident recovery completed: SolusVM completes the recovery process for all virtual servers in the failover domain.
- HA VM recovery started: SolusVM starts recovery for an individual virtual server and moves it to a healthy compute resource.
- HA VM recovery completed: SolusVM completes recovery for an individual virtual server and starts it on the destination compute resource.
Read the Using Event Handlers documentation to configure webhooks.
Offer High Availability to Users
No additional subscription is required for administrators to use High Availability directly. You can also include HA in plans offered to end users, either for free or as a paid add-on.
Note:
A virtual server can use High Availability only if its plan offers High Availability. If the plan does not offer HA, the HA toggle is unavailable for that server. To protect such a server, first move it to a plan that offers HA.
An HA-capable plan makes HA available, but does not enable it automatically. HA is enabled per virtual server, and you can turn it on or off at any time in the server’s settings.
How HA is set at creation depends on how you create the virtual server:
- In the SolusVM interface: select the Include in failover domain option when you create the virtual server. If you do not select it, the virtual server starts without HA protection, and you can enable it later.
- In WHMCS: configure HA in the product settings.
HA does not evacuate virtual servers that do not have HA enabled. Those virtual servers remain on the offline compute resource until an administrator intervenes manually.
Add High Availability to a Plan
To include the High Availability feature in a plan:
- Go to Compute Resources > Plans.
- Click Add Plan to create a new plan, or click the corresponding
button to edit an existing plan. - Under Additional Offers, select the Offer high availability per CR checkbox.
- (Optional) To make High Availability a paid feature, enter a value in the High availability price in % field. This is the percentage of the virtual server price that customers pay additionally for enabling High Availability per compute resource. By default, the value is zero.
- Click Save.
Warning:
Enabling High Availability in a plan does not guarantee that evacuation will succeed. You must configure a failover domain, ensure the destination compute resources have sufficient resources, and enable HA globally in Settings > High Availability.
Enable High Availability for a Virtual Server
Before you can enable High Availability for a virtual server, the following conditions must be met:
- The virtual server’s plan offers High Availability.
- The virtual server’s compute resource belongs to a failover domain.
- That failover domain is active.
- The virtual server uses the same storage as the failover domain.
- The virtual server does not use a primary disk offer, and any additional disk offer it uses does not have a storage tag.
To enable or disable HA protection for an individual virtual server, perform the following steps:
- Go to Virtual Servers and select the virtual server. The Settings tab opens.
- In the High availability card, set the toggle to on or off.
- Confirm the action in the interface that appears.
Note:
- If the plan charges an extra fee for HA, the confirmation interface displays the additional cost as a percentage of the virtual server price.
- Disabling HA always requires confirmation. After you disable HA, SolusVM no longer recovers the virtual server automatically if its compute resource fails.
You can also change the setting from the Virtual Servers list. Open the Actions menu for the virtual server, click Edit Server, and then set the Enable HA toggle.
The Virtual Servers interface displays the HA status of each virtual server with an icon:
- Virtual servers with active HA protection display the
icon. - Virtual servers eligible for HA that do not yet have protection enabled display the
icon. - Virtual servers not eligible for HA protection display the
icon. - Virtual servers that lack sufficient resources to guarantee successful evacuation display the
icon. - Virtual servers with HA protection that an administrator has disabled display the
icon. - Virtual servers that SolusVM cannot currently reach display the
icon.
High Availability Priority
Within a single evacuation event, SolusVM recovers virtual servers in descending order of their High Availability priority. The system recovers servers with a higher priority first, and recovers servers with equal priority in order of their ID.
You set the priority with the ha_priority parameter, which accepts a value from 0 to 65535. To set it, send a PATCH request to /servers/{id}. You can set the priority only for virtual servers whose plan offers High Availability. For more information, see the
RESTful API documentation.
Note:
High Availability priority is available through the API only. You cannot set it in the SolusVM interface.
Troubleshooting
HA does not trigger an evacuation event after a compute resource goes offline
- Verify that you enabled HA globally in Settings > High Availability.
- Verify that the affected compute resource belongs to a failover domain with HA enabled.
- Verify that the configured grace period has elapsed. HA does not trigger an event until the compute resource has been unreachable for the full grace period.
- Check whether HA has reached the Maximum concurrent failures threshold. If HA has already reached the maximum number of concurrent evacuation events, it suppresses new events.
Enabling High Availability for a virtual server fails (NFS storage)
On NFS failover domains, each HA-protected virtual server needs a storage lease. The system grants the lease only when the virtual server starts after you enable HA on its compute resource. If the virtual server was already running, HA may fail to protect it.
To resolve this: stop the virtual server, enable High Availability, then start it again.
A compute resource shows “Not protected” HA status
A compute resource shows the Not protected status in one of the following cases:
- The task that adds the compute resource to the failover domain failed or was cancelled. Go to Tasks, find the failed join failover domain task, and check the details.
- On a Shared LVM failover domain, SolusVM has not read back the storage lock host ID of the compute resource yet. On Shared LVM, you configure storage locking on each compute resource in advance, and joining a failover domain only records the host ID. Verify the storage lock configuration on the compute resource.
- The compute resource is the only member of its failover domain. A failover domain needs at least two compute resources to provide failover. Add at least one more compute resource to the domain.
HA marks virtual servers as “Not evacuated” after an event
- Verify that at least one destination compute resource in the failover domain is online with sufficient free CPU, memory, and storage.
- Expand the failed event row on High Availability > Activities, then review the per-server error details.
A “Move back” action fails
- Confirm that the source compute resource is back online.
- Verify that the source compute resource has sufficient free resources to accommodate the virtual servers you want to move back.
- Check High Availability > Activities for error details about the move-back operation.
- Retry the Move back.
HA configuration screens display a warning that HA is off
- Verify that you enabled HA globally in Settings > High Availability.
- Go to High Availability > Failover Domains, locate the failover domain, and verify that the Active toggle is on.
- SolusVM preserves the failover domain and plan configuration while HA is off. You do not need to reconfigure anything after you re-enable HA.