PMM - disk latency graph

In PMM there is a “Disk Latency” graph. Its description says:

Shows average latency for read and write IO operations. Higher than typical latency for highly loaded storage indicates saturation (overload) and is a frequent cause of performance problems. Higher than normal latency can also indicate internal storage problems.

Calculation formula:

sum(rate(node_disk_read_time_seconds_total{instance=~“instance”,device= “instance”,device= “device”}[instance", device=~“device”}[interval])) > 0 or sum(irate(node_disk_read_time_seconds_total{instance=~“instance”,device= “instance”,device= “device”}[5m]) / irate(node_disk_reads_completed_total{instance=~“instance”,device= “instance”,device= “device”}[5m])) > 0 or vector(0)

Using this calculation formula, I wrote a script:

#!/bin/bash
INTERVAL=300

DEV_REGEX=‘^(sd[a-z]+|nvme[0-9]+n[0-9]+|vd[a-z]+|xvd[a-z]+|mmcblk[0-9]+)$’

snapshot() {
awk -v re=“$DEV_REGEX” ‘$3 ~ re {print $3, $4, $7, $8, $11}’ /proc/diskstats
}

now_human() {
date ‘+%Y-%m-%d %H:%M:%S’
}

echo “================================================================”
echo " Disk Latency Report"
echo " Started: $(now_human)"
echo " Interval: ${INTERVAL} sec ($(echo “scale=1; $INTERVAL/60” | bc) min)"
echo “================================================================”
echo

snapshot > /tmp/disk_latency_a.$$

echo “Snapshot 1 taken at $(now_human). Waiting ${INTERVAL}s…”
echo

sleep “$INTERVAL”

snapshot > /tmp/disk_latency_b.$$

echo “Snapshot 2 taken at $(now_human).”
echo

printf “%-15s %12s %14s %12s %14s\n”
“DEVICE” “READS” “READ LAT(ms)” “WRITES” “WRITE LAT(ms)”
printf “%-15s %12s %14s %12s %14s\n”
“---------------” “------------” “--------------”
“------------” “--------------”

join /tmp/disk_latency_a.$$ /tmp/disk_latency_b.$$ |
awk ’
{

d_reads  = $6 - $2
d_rtime  = $7 - $3
d_writes = $8 - $4
d_wtime  = $9 - $5

if (d_reads < 0 || d_rtime < 0) { d_reads = 0; d_rtime = 0 }
if (d_writes < 0 || d_wtime < 0) { d_writes = 0; d_wtime = 0 }

r_lat = (d_reads  > 0) ? sprintf("%.3f", d_rtime  / d_reads)  : "-"
w_lat = (d_writes > 0) ? sprintf("%.3f", d_wtime  / d_writes) : "-"

printf "%-15s %12d %14s %12d %14s\n",
       $1, d_reads, r_lat, d_writes, w_lat

}’

echo
echo “================================================================”
echo " Finished: $(now_human)"
echo “================================================================”

rm -f /tmp/disk_latency_a.$$ /tmp/disk_latency_b.$$

What does Percona show

and what does my script show, over the same time interval

The disk latency data does not match any other monitoring system. Maybe I misread the description and calculation formula you provided?

Hi @Andrey_Virus,

Is it possible you have more than one item selected (or “ALL”) in the Device drop-down list at the top of the PMM dashboard?

If you have ALL or more than one selected, it is showing the sum of the latencies for each of them. Please select one and see if it matches what you expect to see, and let us know.