HostUp

HostUp Status

All Systems Operational

99.999% uptime over 90 days

Active Incidents

No active incidents — all systems are operating normally.

Past Incidents

CompletedOctober 8, 2026

Web hosting: outgoing email moved to our own mail relay

All web hosting servers now send outgoing email through our own mail relay instead of MailChannels. Sending is faster and there are far fewer false spam blocks. No action is needed on your part.

Why: MailChannels, which we previously used for outgoing email from some of our web hosting servers, too often stopped legitimate email as spam, and its sending limits were inflexible, for example when sending newsletters. With our own relay we have full control over filtering and limits.

Result: sending a message now takes about 2 seconds, compared with 8 seconds or more before, and far fewer legitimate messages are stopped by mistake.

Resolved: on October 8, 2026 at 21:36 CEST all web hosting servers had been moved. Traffic through the new relay ramps up gradually over the coming weeks. You do not need to change anything in your email client or on your website.

ResolvedSeptember 20–22, 2026 — running again from 16:27 CEST

Scheduled VPS backups paused 20–22 September

Automatic scheduled backups for VPS were paused from September 20 until September 22 at 16:27 CEST. The new backup storage is now in service and scheduled backups run as normal again.

Impact: backups that run on a schedule were not taken during the pause. Manual backups, snapshots and restores worked throughout, and existing backups were not affected. Nothing else on the VPS instances was affected.

Cause: our backup storage at the time was close to its capacity. To keep backups and restores fast, we paused the scheduled runs until new storage was in place.

Resolution: the new backup storage was put into service on September 22, and scheduled backups have run as normal again since 16:27 CEST.

October 2026

ResolvedDuration: 0 min

orion

Status 520

8 Oct 2026, 16:07

Scheduled Maintenance

May 9–10, 2026Completed

Cluster 1 — Migration to explicit VLAN tagging

All affected VPS instances on Cluster 1 have been migrated from native/untagged VLAN to explicit VLAN tagging, bringing the design in line with the newer cluster. Each VPS saw a brief 5–10 second network interruption during its individual migration. The intermittent ARP/MAC-related connectivity issues seen on very low-traffic VPS instances should now be resolved.

Completed

VPS - RL1 — Migration from krbd to QEMU librbd

All VMs were live-migrated with no downtime. All nodes now run librbd.

epsilon

Maintenance

eta

Operational

orion

Operational

theta

Operational

zeta

Operational
OperationalDegradedPartial OutageMajor OutageMaintenance

Latest Post-Mortem

26 March 2026I/O Degradation26m

VPS - RL1 Node 2 — I/O Storm Caused by krbd Sparse-Read Bug (CVE-2026-23136)

A brief network disruption triggered a known kernel bug in the Ceph storage client (krbd), causing an unrecoverable I/O retry loop that stalled all ~65 VMs on this node. No data was lost.

Timeline
00:36

First occurrence overnight. We recovered the node but didn't identify the root cause

17:40

It happened again. Server load hit 680+, kernel logs flooded with CRC checksum errors across all OSD connections simultaneously. Storage I/O completely stalled

17:45

Throttled the link to 2 Gbit with tc to break the retry loop. CRC errors stopped immediately

17:50

Identified the bond hash was set to layer2 instead of layer3+4, funneling all inbound traffic through one 10G NIC

17:55

Started migrating VMs to other nodes

18:06

All VMs restored. VMs that went read-only were rebooted to clear filesystem state

Impact

About 65 VMs on Node 2 had disk I/O stall completely. Some guest filesystems went read-only as a protective measure. All VMs were restored with no data loss.

Root Cause

The incident had two contributing factors:

Network bonding imbalance: This node's bond was configured with layer2 hashing, which selects the outgoing NIC based on MAC address. With only two endpoints (server and switch), inbound traffic landed almost entirely on one 10G NIC. During a Ceph deep scrub, the increased read traffic was enough to cause packet drops and CRC failures on the saturated link.

Kernel bug (CVE-2026-23136): When the CRC errors caused libceph to drop and reconnect OSD connections, a bug in the kernel's sparse-read state machine prevented recovery. On reconnect, the client misinterpreted new OSD replies as continuations of previous failed operations, causing every retry to fail immediately and trigger another reconnect. This created a self-sustaining loop that could not resolve on its own.

Because krbd handles all VM storage through a single kernel process, this loop affected every VM on the node simultaneously. With librbd (QEMU's userspace Ceph client), each VM maintains independent connections — the same bug does not exist in the userspace client, and even a connection failure would only affect the individual VM.

Resolution
  1. Throttled the link to 2 Gbit to break the retry loop
  2. Fixed the bond hash policy from layer2 to layer3+4
  3. Migrated VMs to other nodes and rebooted those in read-only state
  4. Rebooted the node to clear stale kernel Ceph state
Preventive Measures
  • All nodes confirmed on layer3+4 bond hashing — this was the only node still on layer2
  • Migrated all nodes from krbd to librbd (QEMU) on March 28. With librbd, connection faults are isolated per VM and the kernel sparse-read bug is not in the code path. Done via live migration with no downtime

Issue not listed here?

Try our AI troubleshooting agent — it can check your website, verify DNS records, test if ports are open (SSH, RDP), and help determine if the issue is on your end or ours.

Automated health checks running every 30 seconds. Web hosting monitors use test WordPress sites — brief unavailability (1-2 min) may occur during auto-updates.