Monitoring Troubleshoot

This document describes steps that needed to be done to troubleshoot monitoring problems when using Grafana/Prometheus monitoring tool.

Problem

No Data Points

No data points on all data charts.

Solution

  • Prometheus may be pointing to the wrong target. Check your prometheus/scylla_servers.yml and prometheus/node_exporter_servers.yml. Make sure in both cases Prometheus is pulling data from the Scylla server.

Or

  • Your dashboard and Scylla version may not be aligned. If you are running Scylla 2.0.x, you need to start the monitoring server with ./start-all.sh.

For example:

./start-all.sh -v 2.0.1

More on start-all.sh options.

Grafana Chart Shows Error (!) Sign

Run this procedure on the Monitoring server.

All of Grafana chart shows error (!) sign. There is a problem with the connection between Grafana and Prometheus. On the monitoring server:

Solution

1. Check Prometheus is running using sudo docker ps. If it is not running check the prometheus.yml for errors.

For example:

CONTAINER ID  IMAGE    COMMAND                  CREATED         STATUS         PORTS                                                    NAMES
41bd3db26240  monitor  "/docker-entrypoin..."   25 seconds ago  Up 23 seconds  7000-7001/tcp, 9042/tcp, 9160/tcp, 9180/tcp, 10000/tcp   monitor
  1. If it is running, go to “Data Source” in the Grafana GUI, choose Prometheus and click Test Connection.

Grafana Shows Server Level Metrics, but not Scylla Metrics

Grafana shows server level metrics like disk usage, but not Scylla metrics. Prometheus fails to fetch metrics from Scylla servers.

Solution

  • use curl <scylla_node>:9180/metrics to fetch binary metric data from Scylla. If curl does not return data, the problem is the connectivity between the monitoring and Scylla server. Please check your IPs and firewalls.

For example

curl 172.17.0.2:9180/metrics

Grafana Shows Scylla Metrics, but not Server Level Metrics

Grafana dashboard shows Scylla metrics, such as load, but not server metrics like disk usage. Prometheus fail to fetch metrics from node_exporter.

Solution

1. Make sure node_exporter is running on each Scylla server. node_exporter is installed by scylla_setup. If it does not, make sure to install and run it.

  1. If is running, use curl <scylla_node>:9100/metrics (where 172.17.0.2 is a Scylla server IP) to fetch binary metric data from Scylla. If curl does not return data, the problem is the connectivity between the monitoring and Scylla server. Please check your IPs and firewalls.

Working with wire-shark

No metrics shown in Scylla monitor.

  1. Install wire-shark

2. Capture the traffic between Scylla monitor and Scylla node using the tshark command. tshark -i <network interface name> -f "dst port 9180"

For example:

tshark -i eth0 -f "dst port 9180"

Capture from Scylla node towards Scylla monitor server.

Scylla is running.

Monitor ip        Scylla node ip
199.203.229.89 -> 172.16.12.142 TCP 66 59212 > 9180 [ACK] Seq=317 Ack=78193 Win=158080 Len=0 TSval=79869679 TSecr=3347447210

Scylla is not running

Monitor ip        Scylla node ip
199.203.229.89 -> 172.16.12.142 TCP 74 60440 > 9180 [SYN] Seq=0 Win=29200 Len=0 MSS=1460 SACK_PERM=1 TSval=79988291 TSecr=0 WS=128

Monitoring Scylla

Troubleshoot