chronoskin
northwindproduction
Status / checked 40 seconds ago

All systems operational

Six of six services answer in every region. One incident is being watched after a fix.

  • uptime, 90 days99.982%
  • open incidents1 monitoring
  • last outage19 days ago
  • next maintenanceSun 23 Mar, 02:00 UTC

Services ?One bar is one day. A bar turns amber after five minutes of errors.

last 60 days, one bar a day
  • API gateway99.99%
  • Query engine99.91%
  • Storage99.99%
  • Webhooks99.94%
  • Dashboard100.00%
  • Scheduler99.97%
  • Edge cache New99.98%
60 days ago up degraded outage today

Latency, 24 hours

fra1
  • p50 38 ms
  • p95 142 ms
  • p99 310 ms

Maintenance on Sunday

Storage nodes in fra1 restart one at a time from 02:00 UTC. No downtime is expected.

See the plan

Regions

p95 from the nearest probe
fra1

Frankfurt

142 ms
lhr1

London

96 ms
iad1

Virginia

88 ms
gru1

Sao Paulo

231 ms
sin1

Singapore

118 ms
syd1

Sydney

104 ms

Recent incidents

All incidents
  1. 18 Mar 202509:12 UTC, 21 min
    Elevated query latency in fra1

    A replica fell behind after a schema change and was taken out of rotation. We are watching the fix.

    Monitoring
  2. 27 Feb 202522:40 UTC, 48 min
    Webhook deliveries delayed

    A queue worker ran out of memory. Deliveries were retried and none were lost.

    Resolved
  3. 9 Feb 202502:00 UTC, 35 min
    Scheduled storage maintenance

    Nodes in iad1 were restarted one at a time. Finished ten minutes early.

    Maintenance
  1. Status
  2. Incidents

Incidents

Everything that went wrong or was planned since January, with how long it took to put right.

Times in UTC. Durations count from first alert to resolution.
IdIncidentServiceStateStartedMinutesRequests failed
INC-214Elevated query latency in fra1Query engineMonitoring18 Mar, 09:12210
INC-213Webhook deliveries delayedWebhooksResolved27 Feb, 22:40480
INC-212Write errors on large batchesQuery engineResolved12 Feb, 14:03401,284
MNT-031Scheduled storage maintenanceStorageMaintenance9 Feb, 02:00350
INC-211Dashboard sign-in failingDashboardResolved30 Jan, 11:269312
INC-210Gateway returning 502 in gru1API gatewayOutage17 Jan, 05:51148,906
6 of 23 incidents

No maintenance after 23 March

Planned work is announced here at least seven days ahead.

Notify me

States

what each label means
Investigating Identified Monitoring Resolved Outage Maintenance

An incident opens with the first alert and closes when the error rate has been normal for thirty minutes.

Export as CSV Postmortems Uptime report Status feed

  1. Status
  2. Incidents
  3. INC-214

Elevated query latency in fra1 Monitoring

Started 18 Mar 2025 at 09:12 UTC. Queries were slow for 21 minutes; none failed.

Updates

newest first
  1. 09:33 UTCMonitoring

    The replica is out of rotation and p95 is back under 150 ms. We will watch for thirty minutes before closing.

  2. 09:24 UTCIdentified

    One replica in fra1 is 90 seconds behind after this morning's schema change. Reads sent to it wait for it to catch up.

  3. 09:12 UTCInvestigating

    Latency alerts fired for the query engine in fra1. Other regions are normal.

Postmortem

Draft

A replica that could not keep up

A schema change rewrote a large table, and one replica applied it more slowly than the others.

The router sends reads to any replica less than max_lag behind. That setting was 120 seconds, so a replica 90 seconds behind still took traffic.

What we are changing

  • The default for max_lag drops to 10 seconds
  • Schema changes on tables over 50 GB are throttled per replica
  • A new alert fires when any replica is behind for more than a minute

What you need to do

Nothing. The new default applies to every project from 20 March.

Facts

INC-214
  • Severityminor
  • Detected byalert, p95 latency
  • Time to identify12 min
  • On callImre Solvang
Error budget, March18% used
18%

Actions

  1. Status
  2. Alert rules

Alert rules

A rule watches one measurement and tells a channel when it crosses a line for long enough.

Alerts are quiet during maintenance. Rules pause for the services named in a maintenance window.

The webhook for the on-call channel answered 500. Three alerts could not be delivered there today.

New rule

7 of 20 used
Send a test alert

Rules

3 on, 1 off
  • Latency p95 over 400 ms any region, for 5mon-call
  • Error rate over 1% API gateway, for 2mon-call
  • Webhook queue over 10,000 for 10me-mail
  • Disk over 85% storage, for 30mpaused
What the webhook receivesjson
{
  "rule":    "replica-lag",
  "state":   "firing",
  "value":   92,
  "region":  "fra1",
  "since":   "2025-03-18T09:12:04Z"
}