Six of six services answer in every region. One incident is being watched after a fix.
Storage nodes in fra1 restart one at a time from 02:00 UTC. No downtime is expected.
See the planA replica fell behind after a schema change and was taken out of rotation. We are watching the fix.
A queue worker ran out of memory. Deliveries were retried and none were lost.
Nodes in iad1 were restarted one at a time. Finished ten minutes early.
Everything that went wrong or was planned since January, with how long it took to put right.
| Id | Incident | Service | State | Started | Minutes | Requests failed |
|---|---|---|---|---|---|---|
| INC-214 | Elevated query latency in fra1 | Query engine | Monitoring | 18 Mar, 09:12 | 21 | 0 |
| INC-213 | Webhook deliveries delayed | Webhooks | Resolved | 27 Feb, 22:40 | 48 | 0 |
| INC-212 | Write errors on large batches | Query engine | Resolved | 12 Feb, 14:03 | 40 | 1,284 |
| MNT-031 | Scheduled storage maintenance | Storage | Maintenance | 9 Feb, 02:00 | 35 | 0 |
| INC-211 | Dashboard sign-in failing | Dashboard | Resolved | 30 Jan, 11:26 | 9 | 312 |
| INC-210 | Gateway returning 502 in gru1 | API gateway | Outage | 17 Jan, 05:51 | 14 | 8,906 |
An incident opens with the first alert and closes when the error rate has been normal for thirty minutes.
Started 18 Mar 2025 at 09:12 UTC. Queries were slow for 21 minutes; none failed.
The replica is out of rotation and p95 is back under 150 ms. We will watch for thirty minutes before closing.
One replica in fra1 is 90 seconds behind after this morning's schema change. Reads sent to it wait for it to catch up.
Latency alerts fired for the query engine in fra1. Other regions are normal.
A schema change rewrote a large table, and one replica applied it more slowly than the others.
The router sends reads to any replica less than max_lag behind. That setting was 120 seconds, so a replica 90 seconds behind still took traffic.
max_lag drops to 10 secondsNothing. The new default applies to every project from 20 March.
412 subscribers get a last e-mail and the banner on the status page goes back to green.
A rule watches one measurement and tells a channel when it crosses a line for long enough.
Alerts are quiet during maintenance. Rules pause for the services named in a maintenance window.
The webhook for the on-call channel answered 500. Three alerts could not be delivered there today.
{
"rule": "replica-lag",
"state": "firing",
"value": 92,
"region": "fra1",
"since": "2025-03-18T09:12:04Z"
}