chronoskin

refresh 5 s | 14:02:07

fleet 12 hosts, 3 groups

11 of 12 hosts answer; db-03 is out of disk.

One alert is firing and one is acknowledged. Everything else has been quiet for six days.

fleet average last 5 min

cpu
||||||||||||||..........................................
24.1%
mem
|||||||||||||||||||||||||||||||.........................
55.8%
swp
||......................................................
3.0%
dsk
||||||||||||||||||||||||||||||||||||||||||..............
74.6%
net
||||||||||..............................................
182 Mb/s
err
|.......................................................
0.4%
  • load, 1 min0.420.38 an hour ago
  • requests1,284/speak today 2,910/s
  • slowest 5%The slowest twentieth of requests in the last five minutes.212 mslimit 400 ms
  • uptime, 30 d99.94%26 min down

busiest hosts sorted by cpu

hoststatecpu%cpumem%loadup
web-02ok71.3||||||||||||||......48.02.1441 d
web-01ok64.9|||||||||||||.......51.21.8741 d
db-01ok38.2||||||||............82.51.02112 d
queue-01slow31.0||||||..............67.40.969 d
db-02ok22.7|||||...............79.10.61112 d
cache-01ok12.4||..................90.30.3063 d
mail-01ok4.8|...................33.60.08201 d
db-03down0.0....................0.00.000 d
8 of 12 hosts all hosts

log

  • 14:01 db-03 does not answer
  • 13:58 db-03 disk at 100%
  • 13:40 queue-01 slow, 9 s behind
  • 11:15 web-02 restarted by deploy
  • 09:00 daily report sent

db-03 disk##########100%

  1. fleet
  2. /
  3. hosts
  4. /
  5. all groups

hosts

Twelve hosts in three groups, asked every five seconds.

ok

web

web-01

cpu ||||||||||...... 64.9%

mem ||||||||........ 51.2%

ok

web

web-02

cpu |||||||||||..... 71.3%

mem |||||||......... 48.0%

ok

data

db-01

cpu ||||||.......... 38.2%

mem |||||||||||||... 82.5%

ok

data

db-02

cpu ||||............ 22.7%

mem ||||||||||||.... 79.1%

down

data

db-03

cpu ................ no answer

dsk |||||||||||||||| 100%

ok

data

cache-01

cpu ||.............. 12.4%

mem ||||||||||||||.. 90.3%

slow

mail and queue

queue-01

cpu |||||........... 31.0%

lag |||||||......... 9 s

ok

mail and queue

mail-01

cpu |............... 4.8%

mem ||||||.......... 33.6%

hosts 1 to 8 of 12

paused 0 hosts

No host is paused. A paused host is still asked, but raises no alert.

pause a host

legend

ok slow down data 2
ok
answered the last three questions in time.
slow
answered, but over a limit you set.
down
no answer for 15 seconds.
#group
the group a host belongs to; a count is its open alerts.

alerts

One firing, one acknowledged, none silenced.

Firing: db-03 has not answered since 14:01. Its last report showed the disk at 100%.

Acknowledged: queue-01 is 9 seconds behind. ops took it at 13:44.

today 4 alerts

runbook disk nearly full

Disk nearly full

A data host that runs out of disk stops writing and then stops answering. Free space first, find the cause after.

Free space

  • Remove rotated logs older than a week in /var/log.
  • Drop the oldest local dump; a copy is on the backup host.
If the disk fills again

Lower the days of logs kept, then ask whether the dump can move.

Who to tell

Whoever acknowledged the alert; the name is on its line.

Check it worked

Run df -h /var. The alert closes by itself under 90%.

settings: db-03

How often the host is asked, and when it raises an alert.

host

error: 300 is not a part of an address.

Down means no answer to three questions in a row.

alerts:

confirm

what is kept

One line per answer for a day, one per minute for a month, one per hour after that.