---
title: Troubleshooting
product: trilio-site-recovery-for-kubernetes-openshift-virtualization
doc_type: runbook
version: feature-tsr-24
source: git2docs (code-derived, validation-filtered)
canonical: https://git2docs.com/murali-balcha/docs/trilio-site-recovery-for-kubernetes-openshift-virtualization/site-recovery/troubleshooting
---

# Troubleshooting

_Common controller errors, debugging steps, and log analysis for site recovery failures_

## Objective

Use this runbook to diagnose and resolve common controller errors, replication failures, and degraded-node conditions in a Trilio Site Recovery for OpenShift Virtualization deployment.

## Scope

This runbook covers: collecting a support bundle with `tsr-gather`, interpreting bundle warnings and manifest status fields, diagnosing controller errors on the quorum and workload clusters, debugging DRBD replication problems, investigating degraded or unreachable nodes, and escalating unresolved issues. It does not cover initial deployment, planned or unplanned failover execution, or LINSTOR-model deployments.

## Prerequisites

Before starting, confirm the following:

- **Kubernetes ≥ 1.28** on all three clusters (quorum, primary, DR)
- **OpenShift ≥ 4.14** if running on OpenShift
- **KubeVirt ≥ 1.0** installed on primary and DR clusters
- **`kubectl`** or **`oc`** CLI configured with access to the quorum cluster
- **`oc adm must-gather`** available if collecting a full support bundle
- The `site-recovery-quorum-control-plane` Deployment is running in its zone namespace (for example, `dr-<name>`)
- The `site-recovery-workload-control-plane` Deployment and `site-recovery-workload-control-plane-agent` DaemonSet are running in `trilio-site-recovery-system` on the primary and DR clusters
- Kubeconfig credentials for the primary and DR clusters stored as Secrets on the quorum cluster
- TCP ports 7000–7999 (DRBD replication) and 6443 (API server) are open between clusters

## Steps

## Step 1 — Collect a support bundle with `tsr-gather`

The `tsr-gather` tool is packaged as an `oc adm must-gather`-compatible image. It collects logs, CRD manifests, and status from the quorum cluster and all connected workload clusters into a single read-only archive. Run it from a host with access to the quorum cluster:

```bash
oc adm must-gather --image=<tsr-gather-image> -- /usr/bin/gather
```

If you need to restrict collection to a single DR deployment (zone), use the `--zone` flag:

```bash
oc adm must-gather --image=<tsr-gather-image> -- /usr/bin/gather --zone <zone-name>
```

To collect only the local cluster without reaching through to workload clusters:

```bash
oc adm must-gather --image=<tsr-gather-image> -- /usr/bin/gather --local-only
```

The tool writes its output to `/must-gather` by default. The top-level `manifest.yaml` file is the entry point; read it first. The `SUMMARY.md` file provides a human-readable overview of replication health, protection group state, and what could not be collected.

**Success looks like:** The bundle directory contains `manifest.yaml`, `SUMMARY.md`, a `quorum/` subdirectory, and one `workload-<zone>-<cluster>/` subdirectory per connected workload cluster. The `manifest.yaml` lists each cluster with `status: ok` or `status: partial`.

**Warning:** A cluster recorded as `status: failed` in `manifest.yaml` was not reached. Check the `error` field for the exact dial or authentication error.

---

## Step 2 — Read the manifest and identify degraded clusters

Open `manifest.yaml` in the bundle root. Each entry in `clusters[]` has:

- `status`: `ok`, `partial`, or `failed`
- `warnings[]`: a list of warnings; lines prefixed with `degraded:` indicate that part of the bundle is missing
- `inventory`: a map of kind to object count; a value of `-1` means the LIST for that kind failed
- `error`: populated only when `status: failed`

A `partial` status means the cluster was reached but at least one subsystem could not be collected. A `failed` status means the cluster's API server was never reached.

Also open `SUMMARY.md` and look for the `⚠` marker in the **Replication health** and **Protection groups** sections. These rows call out unhealthy connection states, RPO violations, and failed or incomplete operations.

**Success looks like:** You can identify which clusters have problems and which sections of the bundle to investigate next.

---

## Step 3 — Diagnose quorum controller errors

The `site-recovery-quorum-control-plane` operator runs the failover controller, protection controller, pg-sync controller, and replication monitor. Controller errors appear in its pod logs.

Locate the quorum control-plane pod in the zone's namespace:

```bash
kubectl get pods -n dr-<zone-name> -l app.kubernetes.io/name=site-recovery-quorum-control-plane
```

Stream or inspect the `manager` container logs:

```bash
kubectl logs -n dr-<zone-name> <quorum-pod-name> -c manager --since=2h
```

In the support bundle, quorum logs are at:
```
quorum/namespaces/dr-<zone-name>/logs/<pod-name>/manager.log
```

Common error patterns to look for:

- **`forbidden listing ProtectionZones`** — the quorum ServiceAccount lacks RBAC. Check `ClusterRoleBinding` objects in the bundle under `quorum/cluster-scoped/`.
- **`listing ProtectionZones to infer the quorum role`** — the quorum cannot reach its own API server. Check network connectivity and API server health.
- **`no ProtectionZone CRs found`** combined with zone namespaces still existing — ProtectionZone CRs were deleted but their namespaces linger. This is itself a finding: re-create the ProtectionZone CR or clean up the orphaned namespace.
- **CRD schema drift** — if a controller logs unexpected field errors, compare the installed CRD definitions (in the bundle under `quorum/cluster-scoped/crds/`) against the expected versions from the Helm chart.

**Check admission webhooks:** A missing or misconfigured webhook causes invalid CRs to reach the controller silently. In the bundle, look for the expected webhook configurations under `quorum/cluster-scoped/webhooks/`. The bundle warns if an expected configuration is absent:

```
site-recovery-quorum-control-plane-<zone-ns>-validating-webhook-configuration
site-recovery-quorum-control-plane-<zone-ns>-mutating-webhook-configuration
```

If these are absent, reinstall the Helm chart or check the OLM CSV for the `site-recovery-protectionzone-controller`.

---

## Step 4 — Diagnose workload controller and agent errors

The `site-recovery-workload-control-plane` operator runs as a Deployment plus a DaemonSet on each workload cluster.

Check the manager Deployment:

```bash
kubectl get deployment site-recovery-workload-control-plane -n trilio-site-recovery-system
kubectl logs -n trilio-site-recovery-system deployment/site-recovery-workload-control-plane -c manager --since=2h
```

Check the DaemonSet and node-agent pods:

```bash
kubectl get daemonset site-recovery-workload-control-plane-agent -n trilio-site-recovery-system
kubectl get pods -n trilio-site-recovery-system -l control-plane=controller-manager,app.kubernetes.io/component=drbd-node-agent
```

In the support bundle, workload logs are at:
```
workload-<zone>-<cluster>/namespaces/trilio-site-recovery-system/logs/<pod-name>/manager.log
workload-<zone>-<cluster>/namespaces/trilio-site-recovery-system/logs/<pod-name>/agent.log
```

For node-agent pods, also check the `drbd-module-loader` init container log — a failed DRBD kernel module load explains an agent that never starts:

```bash
kubectl logs -n trilio-site-recovery-system <agent-pod-name> -c drbd-module-loader
```

In the bundle this log is at:
```
workload-<zone>-<cluster>/namespaces/trilio-site-recovery-system/logs/<pod-name>/drbd-module-loader.log
```

> **Note:** The `drbd-module-loader` log is collected without a `--since` time cap, because the init container runs only once at pod start (during node boot or a DaemonSet rollout), which is almost always outside a 48-hour window.

Common error patterns:

- **Agent pod not `Running`** — the bundle will contain a warning: `node agent pod on node <node-name> not ready; DRBD state for that node not collected`. The affected node's DRBD kernel state will be missing. See Step 5.
- **`forbidden getting Deployment site-recovery-workload-control-plane`** — RBAC is missing on the workload cluster. Check `ClusterRoleBinding` objects in the bundle under `workload-<zone>-<cluster>/cluster-scoped/`.
- **ResourceQuota exhaustion** — check `workload-<zone>-<cluster>/namespaces/trilio-site-recovery-system/resourcequotas.yaml` and `limitranges.yaml` in the bundle. A quota with no headroom explains pending anchor pods or virt-launchers.
- **NetworkPolicy blocking the webhook Service** — check `workload-<zone>-<cluster>/namespaces/trilio-site-recovery-system/networkpolicies/` in the bundle.

---

## Step 5 — Diagnose a degraded or unreachable node

When a node-agent pod is not `Running`, DRBD state for that node is unknown. The bundle records this as a `degraded:` warning and sets the cluster to `status: partial`.

First, check the node's status and taints:

```bash
kubectl get node <node-name> -o yaml
```

Look for the DRBD not-ready taint, which prevents VMs from being scheduled on the node until the agent is healthy:

```bash
kubectl describe node <node-name> | grep -i drbd
```

In the bundle, node information is summarized at:
```
workload-<zone>-<cluster>/nodes.yaml
```

This projection includes kernel version, kubelet version, and any DRBD-related taints — the fields most relevant to DR triage.

Check the agent pod on the affected node:

```bash
kubectl get pods -n trilio-site-recovery-system -l control-plane=controller-manager,app.kubernetes.io/component=drbd-node-agent --field-selector spec.nodeName=<node-name>
kubectl describe pod -n trilio-site-recovery-system <agent-pod-name>
kubectl logs -n trilio-site-recovery-system <agent-pod-name> -c drbd-module-loader
kubectl logs -n trilio-site-recovery-system <agent-pod-name> -c agent
```

If the DRBD module failed to load, verify that the DRBD kernel module (≥ 9.0) is installed on the node:

```bash
oc debug node/<node-name> -- chroot /host drbdadm --version
```

If the module is loaded but the agent is still failing, check `dmesg` on the node for DRBD-specific messages:

```bash
oc debug node/<node-name> -- chroot /host dmesg | grep -i drbd | tail -50
```

In the bundle, per-node DRBD kernel state is collected at:
```
workload-<zone>-<cluster>/drbd/nodes/<node-name>/drbdsetup-status.json
workload-<zone>-<cluster>/drbd/nodes/<node-name>/drbdadm-dump.txt
workload-<zone>-<cluster>/drbd/nodes/<node-name>/drbd-version.txt
workload-<zone>-<cluster>/drbd/nodes/<node-name>/kernel-drbd.txt
workload-<zone>-<cluster>/drbd/nodes/<node-name>/res-files.txt
workload-<zone>-<cluster>/drbd/nodes/<node-name>/drbdsetup-events2.txt
```

**Interpreting `drbdsetup-status.json`:** This is JSON output from `drbdsetup status --json --verbose`. Each resource entry shows `role` (Primary/Secondary), `disk-state`, and `replication-state`. A resource in `Inconsistent`, `Outdated`, or `StandAlone` state needs attention.

**Interpreting `kernel-drbd.txt`:** This contains `dmesg` output filtered for DRBD. Split-brain declarations, IO errors, and connection drops appear here first.

---

## Step 6 — Diagnose replication issues

### Check the ProtectionRequest status (DRBD Operator model)

A `ProtectionRequest` tracks the full protection lifecycle for a single VM. Inspect its status:

```bash
kubectl get protectionrequest -n dr-<zone-name> <pr-name> -o yaml
```

Key status fields:
- `status.phase` — progresses from `Pending` through to `Protected`; `Failed` or `Syncing` for extended periods indicates a problem
- `status.slowestVolumeSyncProgress.percentage` — sync progress across all of the VM's volumes; a value that is not advancing indicates a stalled sync
- `status.conditions[]` where `type: Consistent` — `status: True` means all volumes are synchronized; `False` or `Unknown` means they are not
- `status.consistencySummary` — human-readable description of per-volume state

In the bundle, `ProtectionRequest` CRs are at:
```
quorum/namespaces/dr-<zone-name>/crs/protectionrequest/
```

The `SUMMARY.md` **VM protection requests** section surfaces the `Consistent` condition, sync progress, and any non-healthy phases with a `⚠` marker.

### Check the DRBDResource status (DRBD Operator model)

A `DRBDResource` represents all disks of one VM as a single DRBD replication unit. One exists per VM per workload cluster side:

```bash
kubectl get drbdresource -n <vm-namespace> -o yaml
```

Key status fields:
- Per-volume sync progress
- Connection state between the two cluster sides

In the bundle, `DRBDResource` CRs are at:
```
workload-<zone>-<cluster>/namespaces/<vm-namespace>/crs/drbdresource/
```

### Check the ReplicationGroupStatus

The `ReplicationGroupStatus` provides an aggregated health view per Protection Group. Inspect it:

```bash
kubectl get replicationgroupstatus -n dr-<zone-name> -o yaml
```

Key fields:
- `status.aggregated.connectionState` — `Connected` is healthy; `Degraded` or `Disconnected` requires investigation
- `status.rpoStatus` — `Met` is healthy; `Violated` means the replication lag has exceeded the configured RPO threshold
- `status.currentRPOSeconds` — current observed replication lag in seconds
- `status.lastFullySyncedAt` — timestamp of the last full synchronization; a stale timestamp alongside a `Connected` state may indicate a silent replication stall

In the bundle:
```
quorum/namespaces/dr-<zone-name>/crs/replicationgroupstatus/
```

The `SUMMARY.md` **Replication health** section shows these fields with `⚠` markers on unhealthy values.

### Check RPOEvents for violation history

`RPOEvent` resources record replication lag violations. Query recent events:

```bash
kubectl get rpoevent -n dr-<zone-name> --sort-by=.metadata.creationTimestamp
```

In the bundle, RPO events are collected as a single newest-first list (capped at 2000 most recent events):
```
quorum/namespaces/dr-<zone-name>/crs/rpoevent.yaml
```

The `tsr-gather/totalInNamespace` field in the file header tells you the total number of RPOEvents in the namespace. If the file is truncated, the bundle will contain a warning naming how far back the history goes.

### Diagnose a Protocol C (synchronous) replication stall

Protocol C requires network round-trip latency below approximately 50 ms between primary and DR worker nodes. If latency exceeds this threshold, writes on the primary will stall waiting for DR acknowledgment.

Verify current latency:
```bash
# From a primary worker node to a DR worker node
ping -c 20 <dr-worker-node-ip>
```

Check for DRBD split-brain or connection loss on the affected node:
```bash
# In the bundle, or live:
dmesg | grep -i drbd | grep -i 'split\|disconnect\|connect'
```

If split-brain is detected, the DRBD resource will be in `StandAlone` role on one side. Do not attempt manual resolution if VMs are running; escalate to Trilio Support with the full bundle.

### Diagnose a Protocol A (asynchronous) replication lag

Protocol A acknowledges writes after local disk commit and replicates in the background. The observed RPO window equals the in-flight buffer. Investigate:

1. Check `status.currentRPOSeconds` on the `ReplicationGroupStatus` for the affected Protection Group.
2. Check network bandwidth utilization between primary and DR sites. A congested WAN link causes the replication buffer to grow.
3. Check disk I/O on the primary worker nodes. A write-heavy workload that exceeds replication throughput causes growing lag.

---

## Step 7 — Investigate a FailoverRequest that is stuck or failed

A `FailoverRequest` triggers planned or unplanned failover for a Protection Group. If a failover is stuck:

```bash
kubectl get failoverrequest -n dr-<zone-name> -o yaml
```

Key status fields:
- `status.state` and `status.phase` — the current phase of orchestration
- `status.conditions[]` where `type: Ready` — the `reason` and `message` fields explain what the controller is waiting for. A common blocking state is `WaitingForLocks`, which means the controller is waiting for every member `ProtectionRequest` to confirm the role flip (its `status.activeCluster` reaching the target cluster)
- `status.startTime` and `status.completionTime` — if `completionTime` is absent and `startTime` is old, the failover is stalled

Check the quorum control-plane logs for failover controller messages (see Step 3).

In the bundle, `FailoverRequest` CRs are at:
```
quorum/namespaces/dr-<zone-name>/crs/failoverrequest/
```

The `SUMMARY.md` **Recent operations** section shows `FailoverRequest` state and the `Ready` condition reason.

---

## Step 8 — Investigate a TestFailover that is stuck or failed

> **Note:** Test failover is only supported for DRBD Operator deployment models.

A `TestFailover` creates snapshot-based VM copies on the DR cluster in an isolated namespace and runs verification checks. If it is stuck or failed:

```bash
kubectl get testfailover -n dr-<zone-name> -o yaml
```

Key status fields:
- `status.phase` — `Completed` is healthy; `Failed` or a phase that has not advanced indicates a problem
- `status.testNamespace` — the isolated namespace on the DR cluster where test VMs were created
- `status.targetCluster` — which cluster the test ran on

If the test namespace is stuck (not cleaned up after a failed test), check the DR workload cluster:

```bash
# On the DR cluster
kubectl get namespace <test-namespace>
kubectl get pods -n <test-namespace>
kubectl get pvc -n <test-namespace>
```

In the bundle, test failover namespaces are discovered and collected automatically. They are identified by the label `site-recovery.trilio.io/test-failover=true`. Look for them under:
```
workload-<zone>-<dr-cluster>/namespaces/<test-namespace>/
```

The bundle also collects the network isolation policies (named `tsr-test-isolation-*`) applied to the test namespace. If the test VM had no connectivity, these policies are the place to start.

---

## Step 9 — Check for ProtectionZone controller issues

The `site-recovery-protectionzone-controller` manages the ProtectionZone lifecycle and must be deployed before the quorum control plane. Verify it is running:

```bash
kubectl get deployment protection-zone-controller -n trilio-site-recovery-system
kubectl logs -n trilio-site-recovery-system deployment/protection-zone-controller -c manager --since=2h
```

If the ProtectionZone controller is missing or not `Available`, ProtectionZone CRs cannot be created or updated, which prevents any DR deployment from being configured.

Check OLM state if using the operator marketplace:

```bash
kubectl get csv -n trilio-site-recovery-system
kubectl get subscription -n trilio-site-recovery-system
kubectl get installplan -n trilio-site-recovery-system
```

In the bundle, OLM state is at:
```
quorum/namespaces/trilio-site-recovery-system/olm/
```

---

## Step 10 — Verify CRD installation and schema

Schema drift between what a controller expects and what is installed is a common failure mode after upgrades. Check the installed Site Recovery CRDs:

```bash
kubectl get crd -l app.kubernetes.io/part-of=site-recovery
```

Verify that all expected CRDs are present:
- `protectiongroups.siterecovery.trilio.io`
- `protectionrequests.siterecovery.trilio.io`
- `failoverrequests.siterecovery.trilio.io`
- `testfailovers.siterecovery.trilio.io`
- `drbdreplicationpolicies.siterecovery.trilio.io`
- `drbdresources.siterecovery.trilio.io`
- `rpoevents.siterecovery.trilio.io`
- `replicationgroupstatuses.siterecovery.trilio.io`
- `protectionzones.siterecovery.trilio.io`

In the bundle, CRD definitions are at:
```
quorum/cluster-scoped/crds/
workload-<zone>-<cluster>/cluster-scoped/crds/
```

If a controller logs field validation errors that did not exist before an upgrade, compare the `spec.versions[].schema` in the installed CRD against the chart-rendered version.

## Verification

After completing the diagnostic steps, confirm the following to consider the issue resolved or escalated:

1. **Bundle is complete:** `manifest.yaml` shows `status: ok` or `status: partial` (not `failed`) for the quorum cluster and all reachable workload clusters. Investigate any `status: failed` clusters separately.

2. **Replication is healthy:** Run the following on the quorum cluster and confirm all `ReplicationGroupStatus` resources show `status.aggregated.connectionState: Connected` and `status.rpoStatus: Met`:
   ```bash
   kubectl get replicationgroupstatus -A -o custom-columns=NAMESPACE:.metadata.namespace,NAME:.metadata.name,CONNECTION:.status.aggregated.connectionState,RPO:.status.rpoStatus
   ```

3. **All ProtectionRequests are Protected:** Confirm that all `ProtectionRequest` resources have reached phase `Protected` and the `Consistent` condition is `True`:
   ```bash
   kubectl get protectionrequest -A -o custom-columns=NAMESPACE:.metadata.namespace,NAME:.metadata.name,PHASE:.status.phase,CONSISTENT:.status.conditions
   ```

4. **All node-agent pods are Running:** On each workload cluster, confirm all agent pods are in the `Running` phase:
   ```bash
   kubectl get pods -n trilio-site-recovery-system -l control-plane=controller-manager,app.kubernetes.io/component=drbd-node-agent
   ```

5. **No active RPO violations:** Confirm no new `RPOEvent` resources have been created recently:
   ```bash
   kubectl get rpoevent -A --sort-by=.metadata.creationTimestamp | tail -10
   ```

6. **SUMMARY.md shows no `⚠` markers** in the Replication health or Protection groups sections of the most recently collected bundle.

## Rollback

This runbook is read-only and diagnostic in nature — the steps above do not modify cluster state. There is nothing to undo.

If, in the course of diagnosing an issue, you performed any of the following actions and need to reverse them:

- **Deleted a `ProtectionRequest`:** Re-create the resource from your source manifest. The controller will re-validate and re-provision DRBD volumes. Be aware that data may need to re-synchronize fully, which can take time depending on VM disk size.
- **Deleted a `FailoverRequest`:** A `FailoverRequest` that completed (either successfully or unsuccessfully) can be safely deleted as it is a record of a completed operation. A `FailoverRequest` that is in progress should not be deleted; contact Trilio Support before intervening.
- **Manually modified a DRBD resource configuration:** Revert the change and allow the `drbd-node-agent` on the affected node to reconcile. The agent continuously reconciles local DRBD state on behalf of the quorum control plane. Check the agent log after reverting to confirm reconciliation succeeds.

If a node is left with the DRBD not-ready taint after resolving an agent issue, the taint will be removed automatically once the agent reports the node as healthy. You can check taint status with:
```bash
kubectl describe node <node-name> | grep -i taint
```

## Escalation

If the steps in this runbook do not resolve the issue, escalate to Trilio Support with the following information:

**1. Full support bundle**

Collect a complete bundle from the quorum cluster using `tsr-gather`:
```bash
oc adm must-gather --image=<tsr-gather-image> -- /usr/bin/gather
```
Attach the entire output directory. The bundle includes logs, CRD manifests, DRBD kernel state, event history, and the `SUMMARY.md` and `manifest.yaml` files that Trilio Support will read first.

**2. Summary of findings**

Provide:
- Which cluster or clusters are affected (quorum, primary, or DR)
- The name of the DR deployment (zone) and the affected Protection Group or VM names
- The `status.phase` and `status.conditions` from any relevant `FailoverRequest`, `ProtectionRequest`, or `TestFailover` CRs
- The `status.aggregated.connectionState` and `status.rpoStatus` from the `ReplicationGroupStatus` for the affected Protection Group
- Any `degraded:` prefixed warnings from `manifest.yaml`
- The DRBD kernel version from `drbd-version.txt` in the bundle (under `drbd/nodes/<node-name>/`), if a node-level issue is suspected

**3. Timeline**

Describe when the issue was first observed relative to any recent changes: upgrades, node reboots, network changes, or failover events.

**4. What has already been tried**

List any steps from this runbook you have already completed and their outcomes.

> **Note for reviewer:** This runbook does not have access to Trilio's support portal URL, SLA tiers, or on-call contact information. The escalation section should be updated with the actual support contact details before publishing.
