5 January 5056
Pool endpoints fail. Workers throttle. Dashboards that paint both as red squares send operators driving to silent yards at 2 a.m.
Separate Alert Channels
We configure:
- Pool channel — triggers when all workers lose shares simultaneously
- Worker channel — triggers on individual hash drop or chip temperature
- Infrastructure channel — switch unreachable, PDU phase loss, VPN down
Each routes to different contacts. Night staff answer infrastructure; day staff handle worker maintenance.
Polling Intervals
Aggressive 10-second polling looks responsive but loads small switches. For sites above 0.05 per worker, we default to 30-second worker polls and 60-second switch health checks unless the client runs a local aggregator.
Failover Testing
Monthly test: block pool DNS for five minutes during staffed hours. If alerts do not fire, fix rules before the real outage. We document the test timestamp in the client runbook.
Dashboard Layout
Phone view shows only workers below 95% expected hash and any temperature above seasonal baseline. Desktop adds switch port errors and phase balance — details that clutter mobile screens.