[Post-Mortem] Stakefish Lido Validator Connectivity Loss - August 7-8, 2026

Status: Resolved

Incident Date: August 7-8, 2026

Duration: Approximately 21 hours 36 minutes (August 8, 00:04 UTC - 21:40 UTC)

Service recovery began: August 8, 2026, 05:40 UTC

Published: September 20, 2026


Executive Summary

On August 7-8, 2026, our Lido validator infrastructure hosted in the latitude data center (LON2) experienced a major outage. The incident was triggered by planned upstream facility maintenance at the latitude data center.

During the maintenance window, the A power distribution line of the main power distribution system was isolated for electrical engineering work. The PDU supplying power to high-density validator cabinets experienced a circuit breaker trip under single-line operation, causing multiple validator servers to lose power.

This incident exposed several critical issues:

  1. Lack of direct feedback from the data center - When making decisions at 00:30 UTC, the team did not receive real-time updates on incident handling progress from latitude, leading to decisions based on past experience assumptions rather than actual information
  2. Passive waiting strategy - Based on incorrect estimates of recovery time, the team decided to passively wait for the data center to recover on its own rather than actively initiate failover
  3. Validator infrastructure lacks high-availability design - Any single point of failure renders validators inoperable
  4. All physical servers concentrated in a single cabinet - A cabinet failure affected the majority of Lido validators

In total, over 6000 Lido validators in LON2 were affected. After phased recovery, by 10:04 UTC over 5000 validators had recovered, but 1,286 remained offline. Full recovery was not achieved until 21:40 UTC.

Stakefish has compensated users with 6.9613 ETH to cover losses incurred during this incident.


Infrastructure Architecture

Stakefish’s validator infrastructure follows institutional-grade security practices to protect signing keys:

  • Validator Nodes: Deployed on dedicated bare-metal servers in the latitude data center (LON2), running consensus and execution layer clients
  • Remote Signer: A hardened environment containing validator private keys, accessible exclusively via secure network connections
  • Service Management: All critical validator processes managed by dedicated operations tools
  • Monitoring & Alerting: Prometheus and Grafana for real-time monitoring, Alertmanager for incident alerting

Timeline of Events

All timestamps are in UTC on August 8, 2026.

Time Event
00:04 TargetDown alerts cascade across multiple Lido validator instances. Network connectivity loss confirmed.
00:06 Stakefish operations team intervenes; latitude data center failure confirmed.
00:12 ValidatorsMissedAttestationsCritical alert escalates on Lido validators; multiple validators show infinite attestation miss rates.
00:13 Scope of affected Lido validators determined: all Lido validators in LON2 affected, totaling over 6000 validators.
00:30 Critical decision point - Team decision: Without direct feedback from latitude on incident handling progress, unable to accurately estimate fault recovery time. Based on past experience (“faults recover quickly as before”), team considers enabling new servers, but due to slashing risk concerns, decides to passively wait for latitude service to recover.
05:40 LON2 maintenance window completed. Power distribution lines restored to dual-line operation. Partial Lido validator instance connectivity restored; execution client begins syncing from 6,301 blocks behind.
07:46 Partial Lido server recovery.
08:00 Downtime far exceeds estimates. Operations team reassesses situation, considers setting up new servers, but still believes slashing protection (slashing protection) is more important than attestation and proposal penalties, decides not to take the risk.
08:49 Confirms over 2786 validators still offline.
10:04 Partial recovery progress confirmed: over 5000 validators recovered, confirms 1,286 validators still offline.
21:40 Final connectivity restored; all validators back online and syncing normally. Full recovery confirmed.

Root Cause Analysis (RCA)

Triggering Event

Planned upstream facility maintenance at Telehouse South (LON2 data center) / latitude data center required isolation of the A distribution line of the main power distribution system. This isolation forced the PDU into single-line operation mode, shifting 100% of the load to distribution line B. During single-line operation, the electrical load distribution between validator cabinets was not balanced to accommodate the full load on a single power phase. The PDU supplying power to the high-density cabinet experienced a circuit breaker trip, triggering on-site technicians to perform controlled load shedding of non-critical compute nodes. Unfortunately, our servers were among those shed.

Root Cause of Validator Failure

The current architecture lacks high-availability design.

Root Cause of Decision Delay

Passive waiting strategy resulting from lack of direct feedback from the data center:

At 00:30 UTC when the critical decision was made, the Stakefish operations team did not receive from latitude:

  • Real-time updates on incident handling progress
  • Expected recovery time estimates (ETA)
  • Critical event notifications

This information vacuum caused the team to make assumptions based on past experience: “Like before, the fault will recover quickly.” Under this mistaken assumption, and considering that enabling new validator servers would introduce slashing risk, the team chose to wait passively.

However, actual recovery took much longer than expected. For 7.5 hours from 00:30 to 08:00, there were no progress updates from latitude, until 08:00 UTC when operations personnel realized “downtime far exceeds estimates” and reassessed the situation.

This 5-hour passive wait delayed the initiation of recovery, ultimately causing the entire incident to extend until 21:40 UTC for complete recovery.


Impact

Affected Validators and Services:

  • All Lido validators in LON2 (totaling over 6000)

Downtime and Service Degradation:

  • Total network downtime: Approximately 21 hours 36 minutes from initial alert to complete recovery
  • Direct downtime: Approximately 5 hours 36 minutes (00:04 UTC – 05:40 UTC)
  • Secondary recovery delay: From 05:40 UTC through 21:40 UTC, with partial validators remaining offline
  • Peak offline validator count: Over 6000 Lido validators displayed offline in monitoring systems
  • Phased recovery: By 10:04 UTC, over 5000 validators had recovered, but 1,286 remained offline
  • Missed attestations: Validators showed infinite miss rates (~+Inf) or near 100% attestation loss during peak degradation
  • Missed proposals: Multiple proposal misses recorded
  • Execution client lag: Execution client instance lagged 6,301 blocks; completed sync during recovery

Financial Impact:

  • Stakefish Compensation: 6.9613 ETH (to compensate Lido users for losses incurred during this incident)

Actions Taken & Improvement Plan

Immediate Actions Taken

  1. Establish Failure Communication Protocol with latitude

    • Require regular progress updates during critical failures
    • Establish ETA (expected recovery time) provision mechanism
    • Set automatic alert threshold for >2 hours recovery time
  2. Pre-maintenance Coordination Process

    • Establish formal maintenance coordination agreement with latitude
    • Require >48 hours maintenance notice with mandatory pre-migration window
    • Pre-migrate critical workloads before all circuit-isolation maintenance that could impact validator connectivity
  3. Deploy High-Availability Architecture (Remediation Completed)

    • Validators must connect to multiple beacon and execution layer nodes
    • Nodes must be distributed across different geographic regions
    • Servers in the same region cannot be placed in the same cabinet

Lessons Learned

Key Lessons

1. Single Point of Failure Risk in Geographically Colocated Infrastructure

This incident demonstrates the criticality of server infrastructure redundancy across multiple data centers.

2. Real-time Failure Information and Expected Recovery Time are Critical to Recovery Decisions

The most critical lesson from this incident: under high-pressure fault scenarios, lack of real-time vendor feedback and ETA leads to costly passive decisions.

At 00:30 UTC, the Stakefish team made a decision to passively wait based on the assumption that “faults recover quickly as before,” but actual recovery took much longer than expected. This information vacuum delayed action for 5 hours until 08:00 UTC when operations personnel realized the severity of the situation.

Establishing a continuous communication protocol with the IDC is key to the future:

  • Regular progress updates (at least every 1-2 hours)
  • ETA provision
  • Critical event threshold alerts (e.g., >2 hours without progress)

Such a protocol enables operations teams to make proactive decisions at earlier timepoints rather than passive waiting based on incorrect assumptions.

3. Monitoring Can Quickly Detect Problems, But Cannot Replace Redundancy

Although monitoring immediately flagged the problem at 00:04 UTC, validators remained offline for over 5 hours due to high dependence on LON2 for critical services. Monitoring is necessary but not sufficient.

4. Slashing Protection Takes Priority Over Short-term Recovery Speed, But Know When to Stop Waiting

At 00:30, Stakefish’s decision to refuse enabling new servers to avoid slashing risk was correct. However, this decision also led to adoption of a passive waiting strategy. Under high-pressure conditions, balancing two conflicting objectives is critical:

  • Slashing protection > short-term metrics
  • But adjust strategy promptly when actual conditions exceed estimates

Recommendations & Next Steps

  1. Immediately Implement IDC Communication Protocol - Establish incident handling SLA with latitude requiring regular updates and ETA
  2. Establish Failure Decision Thresholds - Define clear time thresholds (e.g., >2 hours without progress) to trigger automatic failover procedures

This incident report is published by Stakefish, committed to providing complete transparency to the Lido community. We pledge to continue improving infrastructure resilience and preventing similar incidents from recurring.