Skip to content
Need help?

Get support from our team or browse troubleshooting resources.

High availability and recovery ​

The HA & Recovery page answers one question first: can Quantix recover your VMs by itself if a host fails, and if not, what is it still waiting on? Below that you choose which VMs should be recovered and in what order, set up fencing for each host, handle a recovery in progress, and review past recoveries.

Open it from Admin Panel → System → Resilience & recovery.

Nothing here changes how your VMs run

While automatic recovery is not ready, your VMs keep running exactly as they do today. The page only prepares recovery and shows what is missing.

Prerequisites ​

  • An administrator account on QvDC.
  • Shared storage (for example an NFS pool) for any VM you want recovered. A VM whose disks live on one host cannot be started anywhere else.
  • For fencing: each host's management controller address and credentials (Redfish), or a person who can physically switch a failed host off.

Read the readiness summary ​

The card at the top of the page shows one of three states:

  • Automatic recovery is ready: every safety requirement is complete.
  • Automatic recovery is not ready: Quantix is holding back automatic recovery until the cluster meets every safety requirement. The card shows how many requirements are complete and lists the ones still open.
  • Status unavailable: Quantix could not read the answer just now. Select Retry. Nothing is shown as ready until the answer can be read.

Each open requirement names the problem once, with the hosts or VMs it affects, for example Hardware watchdogs need verification · 6 hosts. The colour tells you who acts:

ColourMeaning
RedSomething is unsafe right now. Fix it first.
AmberYou can fix this, usually on this page or on the affected host or VM.
GreyNot something to fix today: a capability still to come in a later release, or a host you put in maintenance yourself.
GreenComplete.

To see the details, select Review requirements. Each requirement explains what it means and, where this page can help, offers a button such as Configure fencing. Technical details under each requirement shows the exact codes Quantix reported, which support may ask for.

The line at the top right, for example Live · Updated 8 s ago, shows how fresh the page is. Select Refresh to read everything again.

Choose which VMs are recovered, and in what order ​

The Protected virtual machines table lists every VM with its host, whether it asks for recovery (HA), its Priority and its Status.

  • Turn recovery on or off on the VM's own page, with the HA Auto-Restart switch.
  • To change the order, pick a new value in the Priority column. Critical VMs are recovered first, then High, Medium and Low. The same priority decides how long a host spares the VM when it runs out of memory. Default: Medium recovery priority recovers in the Medium group but gets no such memory protection.
  • Use the search box and the filters to find VMs by name, host, priority, or whether they ask for recovery. Select View all to see more than the first rows.

The Status column tells you what will happen to each VM:

  • Waiting for readiness: the VM asks for recovery and will be protected once automatic recovery is ready.
  • Can't be recovered: the VM asks for recovery, but its storage is not shared, so no other host could start it. Move it to shared storage, or turn recovery off for it.
  • Not requested: the VM does not ask for recovery.

The Protected VMs count at the top only counts VMs that are actually protected, so it reads 0 while automatic recovery is not ready.

Set up fencing for each host ​

Fencing makes certain a failed host is really off before its VMs start somewhere else. The Fencing section shows your coverage, for example Fence coverage · 2 / 3 hosts, with one row per host.

  1. In the host's row, select Add provider.
  2. Choose the Type:
    • Redfish BMC: enter the controller Endpoint, System ID, Username and Password. Add a Private CA certificate if the controller uses one.
    • Operator attestation: a named person physically isolates the host when it fails. This is a human confirmation, not certified fencing.
  3. Select Create.
  4. Select Preflight in the host's row to check the provider. The Verification column changes to Verified once Quantix accepts it.

Verification expired means the last successful check is too old: run Preflight again. To change a provider, select Edit provider; leave the password empty to keep the current one. To remove it, open the ⋯ menu, select Delete provider… and confirm.

Handle a recovery in progress ​

When Quantix detects a failed host, the recovery appears in the section after Protected virtual machines. Each step waits for you:

  1. Select Prepare recovery to collect current evidence about the host and its VMs.
  2. Review the evidence and choose which VMs to recover.
  3. Select Arm fence to fence the host. While the requirement Manual recovery certification incomplete is open, Arm fence stays disabled.
  4. If the host uses operator attestation, isolate it physically, type its host name, choose what you did, and select Complete attestation.

When nothing is in progress, that section says so and explains that recoveries appear there when Quantix detects a host failure.

Return a fenced host to service ​

After a recovery is finished and the host is safe, open Return a fenced host to service at the bottom of the Fencing section, type the host name, and select Unfence host.

Review past recoveries ​

Recovery history lists finished and cancelled recoveries, newest first, with when each started, the failed host, its VMs and the result. Use the search box and the filters to narrow the list. Select a row to open its full record: the outcome for each VM, timestamps, identifiers and the fencing receipt. These records cannot be changed.

Under Fence receipts by host, choose a host to see every fencing receipt it has, including ones no recovery refers to.

Need help? Open a ticket in the customer portal.