> For the complete documentation index, see [llms.txt](https://docs.starrocks.io/llms.txt). This page is also available as Markdown at its `.md` URL.

# Monitor and Alert with Prometheus and Grafana

StarRocks provides a monitor and alert solution by using Prometheus and Grafana. This allows you to visualize the running of your cluster, facilitating monitoring and troubleshooting.

## Overview[​](#overview "Direct link to Overview")

StarRocks provides a Prometheus-compatible information collection interface. Prometheus can retrieve metric information of StarRocks by connecting to the HTTP ports of BE and FE nodes and storing the information in its own time-series database. Grafana can then use Prometheus as a data source to visualize the metric information. By using the dashboard templates provided by StarRocks, you can easily monitor your StarRocks cluster and set alerts for it with Grafana.

![MA-1](/assets/images/monitor1-c933907ca154e75553acfa00488e5904.png)

Follow these steps to integrate your StarRocks cluster with Prometheus and Grafana:

1. Install necessary components - Prometheus and Grafana.
2. Understand the core monitoring metrics of StarRocks.
3. Set alert channel and alert rule.

## Step 1: Install Monitoring Components[​](#step-1-install-monitoring-components "Direct link to Step 1: Install Monitoring Components")

The default ports of Prometheus and Grafana do not conflict with those of StarRocks. However, it is recommended to deploy them on a different server from that of your StarRocks clusters for production. This reduces the risk of resource conflicts and avoids potential alert failure due to the server's abnormal shutdown.

Additionally, please note that Prometheus and Grafana cannot monitor their own service's availability. Therefore, in a production environment, it is recommended to use Supervisor to set up a heartbeat service for them.

The following tutorial deploys monitoring components on the monitoring node (IP: 192.168.110.23) using the root OS user. They monitor the following StarRocks cluster (which uses default ports). When setting up a monitoring service for your own StarRocks cluster based on this tutorial, you only need to replace the IP addresses.

| **Host** | **IP**          | **OS user** | **Services** |
| -------- | --------------- | ----------- | ------------ |
| node01   | 192.168.110.101 | root        | 1 FE + 1 BE  |
| node02   | 192.168.110.102 | root        | 1 FE + 1 BE  |
| node03   | 192.168.110.103 | root        | 1 FE + 1 BE  |

> **NOTE**
>
> Prometheus and Grafana can only monitor FE, BE, and CN nodes, not Broker nodes.

### 1.1 Deploy Prometheus[​](#11-deploy-prometheus "Direct link to 1.1 Deploy Prometheus")

#### 1.1.1 Download Prometheus[​](#111-download-prometheus "Direct link to 1.1.1 Download Prometheus")

For StarRocks, you only need to download the installation package of the Prometheus server. Download the package to the monitoring node.

[Click here to download Prometheus](https://prometheus.io/download/).

Take the LTS version v2.45.0 as an example, click the package to download it.

![MA-2](/assets/images/monitor2-c1ff652ccf35625c91e46ffcb24cd63e.png)

Alternatively, you can download it using the `wget` command:

```bash
# The following example downloads the LTS version v2.45.0.
# You can download other versions by replacing the version number in the command.
wget https://github.com/prometheus/prometheus/releases/download/v2.45.0/prometheus-2.45.0.linux-amd64.tar.gz

```

After the download is complete, upload or copy the installation package to the directory **/opt** on the monitoring node.

#### 1.1.2 Install Prometheus[​](#112-install-prometheus "Direct link to 1.1.2 Install Prometheus")

1. Navigate to **/opt** and decompress the Prometheus installation package.

   ```bash
   cd /opt
   tar xvf prometheus-2.45.0.linux-amd64.tar.gz

   ```

2. For ease of management, rename the decompressed directory to **prometheus**.

   ```bash
   mv prometheus-2.45.0.linux-amd64 prometheus

   ```

3. Create a data storage path for Prometheus.

   ```bash
   mkdir prometheus/data

   ```

4. For ease of management, you can create a system service startup file for Prometheus.

   ```bash
   vim /etc/systemd/system/prometheus.service

   ```

   Add the following content to the file:

   ```properties
   [Unit]
   Description=Prometheus service
   After=network.target

   [Service]
   User=root
   Type=simple
   ExecReload=/bin/sh -c "/bin/kill -1 `/usr/bin/pgrep prometheus`"
   ExecStop=/bin/sh -c "/bin/kill -9 `/usr/bin/pgrep prometheus`"
   ExecStart=/opt/prometheus/prometheus --config.file=/opt/prometheus/prometheus.yml --storage.tsdb.path=/opt/prometheus/data --storage.tsdb.retention.time=30d --storage.tsdb.retention.size=30GB

   [Install]
   WantedBy=multi-user.target

   ```

   Then, save and exit the editor.

   > **NOTE**
   >
   > If you deploy Prometheus under a different path, please make sure to synchronize the path in the ExecStart command in the file above. Additionally, the file configures the expiration conditions for Prometheus data storage to be "30 days or more" or "greater than 30 GB". You can modify this according to your needs.

5. Modify the Prometheus configuration file **prometheus/prometheus.yml**. This file has strict requirements for the format of the content. Please pay special attention to spaces and indentation when making modifications.

   ```bash
   vim prometheus/prometheus.yml

   ```

   Add the following content to the file:

   ```yaml
   global:
     scrape_interval: 15s # Set the global scrape interval to 15s. The default is 1 min.
     evaluation_interval: 15s # Set the global rule evaluation interval to 15s. The default is 1 min.
   scrape_configs:
     - job_name: 'StarRocks_Cluster01' # A cluster being monitored corresponds to a job. You can customize the StarRocks cluster name here.
       metrics_path: '/metrics'    # Specify the Restful API for retrieving monitoring metrics.
       static_configs:
       # The following configuration specifies an FE group, which includes 3 FE nodes.
       # Here, you need to fill in the IP and HTTP ports corresponding to each FE.
       # If you modified the HTTP ports during cluster deployment, make sure to adjust them accordingly.
         - targets: ['192.168.110.101:8030','192.168.110.102:8030','192.168.110.103:8030']
           labels:
             group: fe
       # The following configuration specifies a BE group, which includes 3 BE nodes.
       # Here, you need to fill in the IP and HTTP ports corresponding to each BE.
       # If you modified the HTTP ports during cluster deployment, make sure to adjust them accordingly.
         - targets: ['192.168.110.101:8040','192.168.110.102:8040','192.168.110.103:8040']
           labels:
             group: be

   ```

   note

   Please note that Prometheus is unable to detect the service changes (`targets`) after the cluster has been scaled in or out. For example, for clusters deployed on AWS, you can grant the EC2 instance that hosts the Prometheus service the `ec2:DescribeInstances` and `ec2:DescribeTags` permissions, and add the `ec2_sd_configs` and `relabel_configs` properties to **prometheus/prometheus.yml**. For detailed instructions, see [Appendix - Enable Service Detection for Prometheus](#enable-service-detection-for-prometheus).

   After you have modified the configuration file, you can use `promtool` to verify whether the modification is valid.

   ```bash
   ./prometheus/promtool check config prometheus/prometheus.yml

   ```

   The following prompt indicates that the check has passed. You can then proceed.

   ```bash
   SUCCESS: prometheus/prometheus.yml is valid prometheus config file syntax

   ```

6. Start Prometheus.

   ```bash
   systemctl daemon-reload
   systemctl start prometheus.service

   ```

7. Check the status of Prometheus.

   ```bash
   systemctl status prometheus.service

   ```

   If `Active: active (running)` is returned, it indicates that Prometheus has started successfully.

   You can also use `netstat` to check the status of the default Prometheus port (9090).

   ```bash
   netstat -nltp | grep 9090

   ```

8. Set Prometheus to start on boot.

   ```bash
   systemctl enable prometheus.service

   ```

**Other commands**:

* Stop Prometheus.

  ```bash
  systemctl stop prometheus.service

  ```

* Restart Prometheus.

  ```bash
  systemctl restart prometheus.service

  ```

* Reload configurations on runtime.

  ```bash
  systemctl reload prometheus.service

  ```

* Disable start on boot.

  ```bash
  systemctl disable prometheus.service

  ```

#### 1.1.3 Access Prometheus[​](#113-access-prometheus "Direct link to 1.1.3 Access Prometheus")

You can access the Prometheus Web UI through a browser, and the default port is 9090. For the monitoring node in this tutorial, you need to visit `192.168.110.23:9090`.

On the Prometheus homepage, navigate to **Status** --> **Targets** in the top menu. Here, you can see all the monitored nodes for each group job configured in the **prometheus.yml** file. Usually, the status of all nodes should be UP, indicating that the service communication is normal.

![MA-3](/assets/images/monitor3-0edea7a4319bf836da3e55c4a24559f9.jpeg)

At this point, Prometheus is configured and set up. For more detailed information, you can refer to the [Prometheus Documentation](https://prometheus.io/docs/).

### 1.2 Deploy Grafana[​](#12-deploy-grafana "Direct link to 1.2 Deploy Grafana")

#### 1.2.1 Download Grafana[​](#121-download-grafana "Direct link to 1.2.1 Download Grafana")

[Click here to download Grafana](https://grafana.com/grafana/download).

Alternatively, you can use the `wget` command to download the Grafana RPM installation package.

```bash
# The following example downloads the LTS version v10.0.3.
# You can download other versions by replacing the version number in the command.
wget https://dl.grafana.com/enterprise/release/grafana-enterprise-10.0.3-1.x86_64.rpm

```

#### 1.2.2 Install Grafana[​](#122-install-grafana "Direct link to 1.2.2 Install Grafana")

1. Use the `yum` command to install Grafana. This command will automatically install the dependencies required for Grafana.

   ```bash
   yum -y install grafana-enterprise-10.0.3-1.x86_64.rpm

   ```

2. Start Grafana.

   ```bash
   systemctl start grafana-server.service

   ```

3. Check the status of Grafana.

   ```bash
   systemctl status grafana-server.service

   ```

   If `Active: active (running)` is returned, it indicates that Grafana has started successfully.

   You can also use `netstat` to check the status of the default Grafana port (3000).

   ```bash
   netstat -nltp | grep 3000

   ```

4. Set Grafana to start on boot.

   ```bash
   systemctl enable grafana-server.service

   ```

**Other commands**:

* Stop Grafana.

  ```bash
  systemctl stop grafana-server.service

  ```

* Restart Grafana.

  ```bash
  systemctl restart grafana-server.service

  ```

* Disable start on boot.

  ```bash
  systemctl disable grafana-server.service

  ```

For more information, refer to the [Grafana Documentation](https://grafana.com/docs/grafana/latest/).

#### 1.2.3 Access Grafana[​](#123-access-grafana "Direct link to 1.2.3 Access Grafana")

You can access the Grafana Web UI through a browser, and the default port is 3000. For the monitoring node in this tutorial, you need to visit `192.168.110.23:3000`. The default username and password required for login are both set to `admin`. Upon the initial login, Grafana will prompt you to change the default login password. If you want to skip this for now, you can click `Skip`. Then, you will be re-directed to the Grafana Web UI homepage.

![MA-4](/assets/images/monitor4-769f8f814df66a7014c2baba78a300f0.png)

#### 1.2.4 Configure data sources[​](#124-configure-data-sources "Direct link to 1.2.4 Configure data sources")

Click on the menu button in the upper-left corner, expand **Administration**, and then click **Data sources**.

![MA-5](/assets/images/monitor5-ce53b5a566261841ef338f2dbd65e148.png)

On the page that appears, click **Add data source**, and then choose **Prometheus**.

![MA-6](/assets/images/monitor6-12deb6c827e59ca679b17ea3823564e5.png)

![MA-7](/assets/images/monitor7-c79999ed4ef04037c3c97153a52b7f6c.png)

To integrate Grafana with your Prometheus service, you need to modify the following configuration:

* **Name**: The name of the data source. You can customize the name for the data source.

  ![MA-8](data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAmwAAABoCAYAAABMm6waAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAABquSURBVHhe7d3/U1XnnQfw/glLtt1EK2C8UVtDivFLwKBGNmwhrKQkGiImBA1QFDQQQoAgguGqCGQSghZmDaQWthaiQeMAzTBki6bFYcMP5of1p+zMznZmZ2faaTc72yRtmnz283nOc+59zuUA96K5HK7vH15z7jnPc8659/LIefs85zl8K+6uuwkAAAAAvAuBDQAAAMDjENgAAAAAPA6BDQAAAMDjENgAAAAAPM4zge17319PL1S84tj29J79tHPnU45tR46cVHXNbRB9j+zIpEOHax1KSipo5b1rFXkdWp6S+ojrsQAAAGBunglsO9KzaHr6E8e2t9++oJjbpI7UNbd5QnYrDX0wRT2HU9zLY8xhDmDHj7+hlrb29m5avz5FkddmmV3X7ViLYWvpG9ReXeBaJnZXd5O/NMO1LGryGqm9uYq2upUBAMAd5c4IbGmV1HFxiianPlH7T09O0dCJMvK51XVTP6r2G6gPbqsblGONUp1dp7CfJrjO8Kns4H4xzA5i5rbQwGaWudWfXQYdaO5Wx7C1HGum4uxNLnUXBoENAACWEk8EtvgEH/X2XlChyBwWvT2BrYA6xiRc3aChMx3U2NBBXRdvqOOMvTn7BdshnMB2h4lGYLMD04r7d1BWQSP5209Tdd7tCW0IbAAAsJR4IrBVVNTTyMiv1VKC0abNW9X20MAm20ND3bzyrJ6vyd5KY3sRdX3AgevDQSpR6z7KaRilsUkJYWxinLpKH1V1rWBmGBymAXOdqSCnzzNxxgoB1n58nNfHaUL37I2dO0ppgffwKJWcnQocY+zsIA0b+8cl5FHj+WD55NgoNeb49L6LL5qBzbZp/0lqb6mnncv1tuU7aM9L7dTSavfC1VPuA8H6y3aUUe1x3UvX+gbVFuXQCl2mAltNGeVWtAfL9z9Gy3S5CmyHXqDKY/b+7VSZa4ZFH2UUNFOTPnf7sWZ6dkeiUb6JskqbyW+XH2+mfUa5dfwqqpb3Z4eyB/KD5+P6Bw7x5zUC21yfBwAAYpuHAtuHKohJOHELbNILJ2XvvPM+Xb16g4qKX3AcY1YbWmlIQs/kOPl3rXOtE3/4Mk1ynYnzXaoHrmfkJp/rGrVl302+1GxKPzGuzj10gl+nptHmjGzyX7ICmZ9fb17Dx3ENbBy0RgbJHzjmTeo7bJ0zs/2aS3lw/8Kz0gt4gwaaiiinuIuGJExODFKhfs+LreTHFVZwCGEGtlCyj9uxZnIPbHFrS+hI+2kqz7LWd1ac5qBURRmrOQgtT6U91bxet59Wq/pPUWVLNzUdzKHVyxNp9Q+rqIn3rcy1QpMKbLxeW/AYrU5MotRnmqml/STt22gdWwJVe3s7ledspRWJ6ynroAS7Ztqz1ipPekbC40kq/mESLVueRBmq/CQ9qwPjg4VS3kh7NkvI5nAn5cdrKEO9t+DxK3c/Rg8m8zHu2kr7Gjh01pXQjtU+Wr3lKSqX8BYIbHN/HgAAiG2eGRKVYCaBRcKbvd0MbPY9blJXQpvMFrXrzSetflQFMtl/euIaDZzx0y4JWao8hfxXeHugt43p8DX2Zp61Hs6Q6Cw9bI0Jurx4UL0H65jZ1DbC5ZOXqcouT7CCpWN/ozyt0M9hsoZ2Jen6i0xmgtrhzLSMw4RwK5N93I410yyB7a4CqubgV50nrzm0JKdSkoQ1u1yGENsbaXegLoe7bLs8iXY8uZ9ydySpdRXY6kt0uLPr87F3W+uqB+zgY7pMPEHlJ7vpyP5Ufm29v6Yi6z8Wlq1U3Bh8zyvuT6UH7zd6RLdXkZ8DWvHD1rocv+XQEzPLU/Q6W/Z0sxHY5v48AAAQ25bUpAMJalJHetginniwpoDqzozS8IQEKTY1RR35ckGtoT57KDREYHhywYFttnK/Nax6pYOS7XK9zd4/uVaHzKkbNHxxkNoOF4Q/SSIKJHwlcwgLZQc2t7LbG9hYYoZjSNRiB7Yk2nlIetE4WNXV04H8J2hTon0ct3vYnMdWgc1x/kTaU2Nv0+FJ9/TZdr54OnjM5am00xwSVd6gA9utujOOL2FThnvtdXtbILDN/XkAACC2eTawSU+aDJMKeW1vlzqhz2aLlK+wn8YkDI10Ubodnsb6qTAjm9INacn6vIsQ2IQvs5IazZDJ7zczUH9xLf6QaBKHHmtINMvuyXL0sFlWbH6Cdu+voVoZXmw5Sft0YPomAtvul/gc+pjq+DIkukWGO7lc9aDdSmCzzPZ5AAAgtnk2sHV2/lRNRBDy2t4udSLpXUvX94oN1BvDU3zBVZMOJi9TuT08KcOXG4L7paVnU7xd/7YHtvmGRMuo7fwoDZwqs8ru8lFJr9zjNkUdubr+Ilv8SQdWnUBvG1NDiHZg42Cz5+knaJMuCx2yDCuw3cKQ6IxAtmOewDbfkOg8nwcAAGKbZwObDHvK/WxCXtvbIw1scWl+GpAhzyn7sR6tHIasx3pMnverHq7A8OMHo9TG5f6z13hdJgjov6jA5VJ/op/3L85V2+rOSyC7QX0nWqkkk+tEFNg4EDbPNekghfe3AlpfbRGlP+unPnk0iRnwFlk0A5s81mNn6UlqaTcf65Fq3aRfX0Y77l9P3/v7EqrlQBUIbCllqjeutiBDzaRcsZnLWzhwFUrgCi+wzT/pQGaG6kkHJfL+gpMO1ISI5hrauYXfm0wg4HA155DojEkHBVRpziCd5/MAAEBs82xgk0kFsi7MCQayHlFgY/HpNc4H58p9YedaKScQfkIe6zE1RUOvVwYfwZFQRB3v67LzR61j5vfSmD7eQDXXiTCwzXisx0UJiUb5mgLHYz3kUSMdhe6zXBdDNAKbHMPm+uDc0MdgFDVaockOVdlVjsdgHCnLpyS9b1g9bCGP9Sh3nN9HGUXGPWryWI8Moxd3+WO0r47PofetLqhyzHCdGdjYPI/1mOvzAABAbFtSkw6E1Ik0sHnTo5SeblzgHbNIzXreJOHLf/wNx98KlSBhBjazTOqGH9gAAADA5JnAJn/Q3Xykh5DJBaHhLFb++Pv3G8ZpevIa9ZxopcaGLhpQEwusZ7+51fca+UPudq+ZTSYVyExQIa9DyyXIuR0LAAAA5uaZwHbHkb9kcO5a4K8gmH9dAQAAAMCEwAYAAADgcQhsAAAAAB6HwAYAAADgcQhsAAAAAB6HwAYAAADgcQhsAAAAAB6HwAYAAADgcbcc2MofWUW/b3yQqGUj20R00unrE7xkXx8Xm+lrPy/9stSatVctXx0TDwU1aY3aUctf1TKF/trArxt4eUSrd/ryFV6+kkpf1mm1IWq0l8UW+ks1L6tlqb2kVWkviocD/lypVWgvmNLoz4d5eTiNvjjEr9kX5aat9EUZL8t4edDy+QFeMlkqpdqPxTb6vMTps2KtSGynz57n5fOy1PZb/rRvO/12bxqVJq9x/TkCAACAd91yYPtd43oV1r62ndROCAlpvFRhTYLaJvpKBTZeNmuvCg5pslRhTWuy/LXRoIKaXqqgxo5YvpRlPS8VCWq85LD2ZZ2hNoX+wtSyRqTSX17mJYc1hcNawEuWP1eJLdaSw1pApVZh+UKWHNK+eIFfy5KDmnLIUP4wfc4krH2ughovDxoOGErFVvrsx7zksPaZKDEUa0VCBzYOa3+y7df2GQq303/uedj15wgAAADedcuBTfWunZLeNbuHjZfSs2b0sFnBTQvpYftK97BJaAv2sGmqd00vdQ+b1bNmLVXP2lw9bBLYbCq0pbr3rqngxkvdwzajd83oYZOwJr1samn3rs3WwyY9a0YPm2L2sOnQZgc3u4dtRu+ao4eNl9KzJj1sdu9aoIeNBXrYeCk9a0YPm+ply09z/TkCAACAd936kOj2e+l3HNqkd031tM3Vw8bMHjYrqNk9bJpLD9tXgV42HdRcetgsHMoCPWxmUDMYPWx2UHNw9LBxMKvipeph02bpYbM8rHrXVE+bWw+bCmy8NMKao4ct0MsmIY2XgaC2VfWwqaWjh01CGi/tsMakd031tDl62CSsbVNhDUOiAAAAS8+37lv7AAEAAACAd6nAds93EwEAAADAoxDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ0AAADA4xDYAAAAADwOgQ3uOM9997v0aUIcUeLfEMkyQZYAGreL/+N2cTL+Htf2AwCwGBDY4I4iYe1rt4s0wAxxdDr+btd2BAAQbVEJbCsS76N771tHq1bfDxA10uak7Zlt8U/oUYMIfM3tZcPmNIAlZ/3GLfx7cJ3j91+sXovdfteLlavW0rqkjZT0wOaYEJXAhrAGi0XantkWzYvxG/d8m9b+7d0UdxdAUM53vuNoJ251ALxu2fJEFdrM33+xfC0O/V0vJKzdsyzB9ftZiqIS2Ny+XIBoMduidd+aZSXCGswieG9jnGs5wFIgPW3m7z+334+xxPysQnql3L6XpQqBDWKe2RbtsCbc/kEACLQTiAUIbAhsEXP7YgGixWyLuBBDONBOIBYgsCGwRcztiwWIFrMt4kIM4UA7gViAwIbAFjG3LxYgWsy2iAsxhAPtBGIBAhsCW8TcvliAaDHbIi7EEA60E4gFCGwIbBFz+2IBosVsi7gQQzjQTiAWILDdWmC7NzmL0gr/iXJe/ZjyOv9AT5/53/Cc/pTr/5GeevP3tOv1/6Yn239LT7T+B+W2fEJZr/yaUp/ppJU/+KHrOeeCwAYxz2yLuBBDONBOIBYgsC08sElQcw1j81Fh7Q8qqP2o5d/pH5tuUGbdrymj+l8UeS3bpEyCm9u5Z4PABp61fUcW5T9TTC/XvqrIa9nmVncuZluM3Quxnwamb1BXnlvZEpGQR+UNNZST4FIWZbHbTuBOgsC2sMCWUTniHsbmY4e11/6Ldjb/mwpoP8g4QL77Uyhx5SpFXss2KZM66WUXXd+DG28EtucH6er0J3T9F8fogcD2g/TWrz6h6Xdfc9aFmPf9pA207/lyerPzbTpw8CXa9dRziryWbVLmtt9szLZ42y7E9aM0PT1KdW5liyIGAltuL41N36S+Mmu9bpD//Q/6Z9aLgtvWTgAWUcSBTV+Lp23XP6ILJ4qN6/Ickiup48rH1n5XOinVrU4Ymt7l/X81SM+p9dfoAh/vQvPMem7MzyoWEtgW3LPGZBhUetYkiG0t7rOCWuJKV1ImdaRuuD1tngps09Mf07mKh/R2BLY71f6iQ9TQ2EopW9JnlMk2KZM6oWWzMdvibbsQI7B94xDYAG7NQgPbe6/l0+NPFlNF5wRd5/XRzufc65tqh7nuTQ5XxfR45lb3OmFYzMAm96y5BbGwnP5U3bMmQ53SezZXWLNJHakr+4RzT5vHAhsb79c/qJDAlvIy/cRO7+zqlbepOFnqWT/Qq++O0fiULvt5C1V3/kY1tOmpj/mHvStwrqyqQXrvmlVvenyY6rOt7eANMuQpvWhmWPuHzB8p9rqUSZ1wh0fNthjJhTg+v4uGJnRbmbpBQ6+XkY+3qyAh27SBeqt+Wmm/s357EcXrY+06c4OmL/ZTx/tSbgU9Oc5Ebz8NyD4f9NMu3haffpR6RriuPsbw2aOUpo+h3lNOKw2MGec4U6nLnYEtPr9f9VYNNGRb7yGhiNou6eOyiUu9VLgmeFxXef00MT1OXe3jNKH/bU1c6qK9/B77Au9hivpefDS4z5oy8l+cCpxn7GIH7Q0McVrvsaepN/g9jV2mqjRdrs4nn0Hq6XLFDsY+Sq8epGF734lr1GOc2+07tssiFUk7AfCqhQa2YEB6iCp+xtfd63ytVNfbWa6hIT1zV88epFXJ+VT/84/ouv7dYV5vnzsr1/IxatLnNUNa4LXbMXX92ZifVUQa2G6pd+30/6jeNbk/TYY83QKaG6kr+4TTy+apwHbhbD+N89JK887AlnV0kC68M0hNT22lDWVDqv5oZz6XWYFtenyIqgvyqaLnI/XDvf5uJ+Xz/xDaLt1Uja1azpPdTaNcNt5zhP/3cITeGuf9rnTSNvO9wKLa+2yJGvoM3SbMbVIndNtszLYY/oW4hvomOXB0l1Fygo+Sn7cCUN+LPqs8tIcto4uGJSA15apQ5yuU+hxOCq1yFSa4fKi9knIythnBjwNPbRGlb0+h+IRK6uEwMtHvp0wOU75MDmf8HiZ6K63QtYGDTOA9cfmuDhri9aGmFD6HEdjSdL0zwcBYfo7/HYz1U2Gyj+KTy6iLA9dk/9FAuSsVoPg4vTWUk7yOkvN7+TPye+aQ1nOYP+eabVTSzeFs8jKVq30eJf8VPu4Ih0HjPNNXOoxQyetjg1SeuY58qfx5P5D3UWOcLxg6Q3vY4osHVXlfrXzH6yizaZQm5Tsutn4mbt+xvW+kwm8nAN5164GNNY3xv6uP6a3n+fVc19Bmo56sl3bT+XeGqe1QJm3I7aT3rlu3PUlZWIFNlUW3h01mg7qFsXDIcKjMBpVJBeH0rtmkruwjs0fd3pPJW4Gt2U7zE3Qie64hUavMSty6h81O36ENTjUiq2E83vkbfs3H1v9T2NYm68FGA4vv5dpmdb+auW39xi2KuU3qSF1z22zMthj+hVjCxU3qKdUB7a4U2lvdSlX5Eo54PTSwcXhJd4SEAuriMDL2Zp5aV2GCg0tyoNwKJJMcxuz1eDkmh58q46b7+Bcvcyjh8/C25OZrNP3hIJXoMpFZP0gDp8r4tQ5s+UVWGOOgY/bM2eeyA1py/lFqrC5yvJ8ZdA9bY+D9+KjxIv+bvNgaDHpmyFKvp6gj267PsuW+NDuE6e+0OFiuvpexXsoJPRavOwObz/oMjpDpo6p+DqJcR7a5fccLFX47AfCu2xLY7G1H57mGhga2ECqI6eu5VwNbRI/uCCH7yqM7ZIhTTTBwCWeuuK7sI4/8cHtPJo8FNn6dfIzOqyT+9pxDosEu0vADm9VIgvtbENi8xC2wufWwffOBLYXKe6XHhkPXpcvU1VSper0C5TPuYfNRepkxJKpNnClQ5SpMhNyPJQHELg/UCQ0cqufOCjGq3AxLDlZgG77CdabGyZ/uLE8+PGgNa46NU9+ZVirJXOcod6UClHNo0Rmi7Do6ZMl3EhIo7Z5Ka9hYh0odyIT6THo4eO7AZgXg4WYdmLX0dg6xen+373ihwm8nAN51u3vY5ryGhga20CFRgcA205IObLyedcK60dH8AR/8Kf+ApectVyYlLKyHLb9bhkv5tf7fAXiPBDNvDIlafJmVVNc+SEMytDd5jdrssBEa2FTYkCHRPDVcaQeMWw5sauZkuIFNhkHLVK9TaA+bsiaXSmq7qOeS3GN2k4ZPBc/t6rYENut9fVOBLedN/iwIbACubvc9bHNeQ0MC27aWCbV+rsyagLAUetgwJMrcvliHGY3kOer4Jf/QeJv9A67+hdyLNkFtBfmUXzuk7nW72lPJZeEHtlVPv22Nv//sGOU/mU/Fr03Q1Uud9Lg6J3iB26QD2WZOMIjKpAMOao0NlZQZ2JarbmYPBKzQwCbrdvBQiiIObAsZEk0r9FNjcS6/NsKQuteNw2O9fUM+B7UG6VUL7pf5ejDo2NtmiDSwqdfzDYkuNLCFOSSKwAYQsNDAFjpLdLxbX1/nuoaGBDZr+FSe/MDHKu1W97BNc115RMgDuu65WqNsjsA22llMWY/YT5CYnflZBSYdLIDbF+swI7AFtwWGRLNb6Lzc4CjbxifoPRkuVT/8CAIbc8xwuTZBb1UFZ5CCN6jHehw95QhtNtkmZd/4Yz2yrUkEQycKrEkEmdYN/sOnsq3yfB1OclMoOdlHcWUSrG5Qz+E82pyaR+VnrZmSkQS2OD3pQG7ytyYd+KlPJiFEOumA66Y1jNPk5Dg1qhmYudQ2wmHnUgftkmHdNbnkv8hBZ6SL0u1zu4k0sNmTDvg8e/Wkg44RPo9j0kH4gW1vN5e930s5yfwd8+e1Jh3oCQ8y6aD2sqrvmHSAwAYQsNDApq6PwuU5bLNeQ0MCm+O5bFzvvV/y7wIOYsV22TCv67IL8to1sGVS0ztWvfHu4sB7mI35WUWkgQ2P9WBuXyzAXCSQSS9a6cGX6MndBYq8lm2RhDVhtsVILsShj+lwPmKjgNquWL9IrIkFPtprPP5i7HwHdUl40ZMKwgpsTB7r0Temf5Hpc6abPW6Ox3pMcaC0HjUyMwxZ4SkQltIqqct8rMfIYPBxGrOJOLCxNRzSjPPIYz2Cjw+JLLDF5XFo/lCOI8FYtsljPS7TGIdUdXx5rEe1fmyJfSwENoCAiAPbEmd+VhFpYBN4cK7LFwswHxnylPvUZHKBkNfhDoOazLaIC7GTCmA6XDncpuCzVKGdQCxAYIs8sAn8aSqARWK2RVyIIRxoJxALENgWFtgE/vg7wCIw2yIuxBAOtBOIBQhsCw9sQu5pk+Ams0cjeuSHCm1/VPe0SXCT2aPyyA95dIfMBpWgFs49a6EQ2CDmmW0RF2IIB9oJxAIEtlsLbF6DwAYxz2yLuBBDONBOIBYgsCGwRcztiwWIFrMt4kIM4UA7gViAwIbAFjG3LxYgWsy2iAsxhAPtBGIBAhsCW8TcvliAaDHbIi7EEA60E4gFCGwIbBFz+2IBosVsi7gQQzjQTiAWILAhsEXs3vvWuX65AN80aXtmW1QX4URciGFuaCew1C1bvpLWb9zi+P0Xy9fi0N/1Yl3SRrpnWYLr97MURSWwrUi8D6ENok7anLQ9sy2qC7G2c8ND6n+gAKbUTQ872olbHQCvk7C2arUzxMTqtdjtd71YuWqtCm3S0xYLohLYALziK+NCTAlxxmsAm7NduLUjAIBoQ2CDO8o/x/+d42IMMJfr8d9xbUcAANGGwAZ3nEsrzNCGXjZwE0f/Gv9t1/YDALAYENgAAAAAPA6BDQAAAMDjENgAAAAAPA6BDQAAAMDjENgAAAAAPA6BDQAAAMDjENgAAAAAPA6BDQAAAMDjENgAAAAAPA6BDQAAAMDjENgAAAAAPA6BDQAAAMDjENgAAAAAPC2R/h/myl9s2S1DBgAAAABJRU5ErkJggg==)

* **Prometheus Server URL**: The URL of the Prometheus server, which, in this tutorial, is `http://192.168.110.23:9090`.

  ![MA-9](/assets/images/monitor9-0b3e34d7be2872123104e38932fcc362.png)

After the configuration is complete, click **Save & Test** to save and test the configuration. If **Successfully queried the Prometheus API** is displayed, it means the data source is accessible.

![MA-10](/assets/images/monitor10-2d17d7d8ca6a8d25231e381d60df8f39.png)

#### 1.2.5 Configure Dashboard[​](#125-configure-dashboard "Direct link to 1.2.5 Configure Dashboard")

1. Download the corresponding Dashboard template based on your StarRocks version.

   * [Dashboard template for All Architecture](https://releases.starrocks.io/resources/Dashboard-All-Arch-20260113.json)
   * [Dashboard template for Shared-data Cluster - General](https://releases.starrocks.io/resources/Dashboard-Shared-data-General-3.5.json)
   * [Dashboard template for Shared-data Cluster - Starlet](https://releases.starrocks.io/resources/Dashboard-Shared-data-Starlet-3.5.json)

   > **NOTE**
   >
   > The template file needs to be uploaded through the Grafana Web UI. Therefore, you need to download the template file to the machine you use to access Grafana, not the monitoring node itself.

2. Configure the Dashboard template.

   Click on the menu button in the upper-left corner and click **Dashboards**.

   ![MA-11](/assets/images/monitor11-e18189f47fceec908b80c7d127a51446.png)

   On the page that appears, expand the **New** button and click **Import**.

   ![MA-12](/assets/images/monitor12-40d0bdaad5a6097aaa9edbe3e19c2093.png)

   On the new page, click on **Upload Dashboard JSON file** and upload the template file you downloaded earlier.

   ![MA-13](/assets/images/monitor13-200a76c2389b70d7ad8799f3db961757.png)

   After uploading the file, you can rename the Dashboard. By default, it is named `StarRocks Overview`. Then, select the data source, which is the one you created earlier (`starrocks_monitor`). Then, click **Import**.

   ![MA-14](/assets/images/monitor14-d5df752b2e5dd0b98b25014c1dbf2f21.png)

   After the import is complete, you should see the StarRocks Dashboard displayed.

   ![MA-15](/assets/images/monitor15-210b53b9e827c45dffd4e1c1c37e415c.png)

#### 1.2.6 Monitor StarRocks via Grafana[​](#126-monitor-starrocks-via-grafana "Direct link to 1.2.6 Monitor StarRocks via Grafana")

Log in to the Grafana Web UI, click on the menu button in the upper-left corner, and click **Dashboards**.

![MA-16](/assets/images/monitor16-5a986d52c1c7f54049d8d545a6a2f4a6.png)

On the page that appears, select **StarRocks Overview** from the **General** directory.

![MA-17](/assets/images/monitor17-bda14636790d5f7f78a7f5565e6f771e.png)

After you enter the StarRocks monitoring Dashboard, you can manually refresh the page in the upper-right corner or set the automatic refresh interval for monitoring the StarRocks cluster status.

![MA-18](/assets/images/monitor18-9a6f4f7963351ba69d2147ee5a54d88d.png)

## Step 2: Understand the core monitoring metrics[​](#step-2-understand-the-core-monitoring-metrics "Direct link to Step 2: Understand the core monitoring metrics")

To accommodate the needs of development, operations, DBA, and more, StarRocks provides a wide range of monitoring metrics. This section only introduces some important metrics commonly used in business and their alert rules. For other metric details, please refer to [Monitoring Metrics](https://docs.starrocks.io/docs/administration/management/monitoring/metrics.md).

### 2.1 Metrics for FE and BE status[​](#21-metrics-for-fe-and-be-status "Direct link to 2.1 Metrics for FE and BE status")

| **Metric**       | **Description**                                                                                                             | **Alert rule**                                                                                             | **Note**                                                                                                                           |
| ---------------- | --------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Frontends Status | FE Node Status. The status of a live node is represented by `1`, while a node that is down (DEAD) will be displayed as `0`. | The status of all FE nodes should be alive, and any FE node with a status of DEAD should trigger an alert. | The failure of any FE or BE nodes is considered critical, and it requires prompt troubleshooting to identify the cause of failure. |
| Backends Status  | BE Node Status. The status of a live node is represented by `1`, while a node that is down (DEAD) will be displayed as `0`. | The status of all BE nodes should be alive, and any BE node with a status of DEAD should trigger an alert. |                                                                                                                                    |

### 2.2 Metrics for query failure[​](#22-metrics-for-query-failure "Direct link to 2.2 Metrics for query failure")

| **Metric**  | **Description**                                                                                                                                            | **Alert rule**                                                                                                                                               | **Note**                                                                                                                                                                                                                                                              |
| ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Query Error | The query failure (including timeout) rate within one minute. Its value is calculated as the number of failed queries in one minute divided by 60 seconds. | You can configure this based on the actual QPS of your business. 0.05, for example, can be used as a preliminary setting. You can adjust it later as needed. | Usually, the query failure rate should be kept low. Setting this threshold to 0.05 means allowing a maximum of 3 failed queries per minute. If you receive the alert from this item, you can check resource utilization or configure the query timeout appropriately. |

### 2.3 Metrics for external operation failure[​](#23-metrics-for-external-operation-failure "Direct link to 2.3 Metrics for external operation failure")

| **Metric**    | **Description**                           | **Alert rule**                                                                                               | **Note**                                                                                                                                                                                             |
| ------------- | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Schema Change | The Schema Change operation failure rate. | Schema Change is a low-frequency operation. You can set this item to send an alert immediately upon failure. | Usually, Schema Change operations should not fail. If an alert is triggered for this item, you can consider increasing the memory limit of Schema Change operations, which is set to 2GB by default. |

### 2.4 Metrics for internal operation failure[​](#24-metrics-for-internal-operation-failure "Direct link to 2.4 Metrics for internal operation failure")

| **Metric**          | **Description**                                                                              | **Alert rule**                                                                                                                                                                                                                                     | **Note**                                                                                                                                                                                   |
| ------------------- | -------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| BE Compaction Score | The highest Compaction Score among all BE nodes, indicating the current compaction pressure. | In typical offline scenarios, this value is usually lower than 100. However, when there are a large number of loading tasks, the Compaction Score may increase significantly. In most cases, intervention is required when this value exceeds 800. | Usually, if the Compaction Score is greater than 1000, StarRocks will return an error "Too many versions". In such cases, you may consider reducing the loading concurrency and frequency. |
| Clone               | The tablet clone operation failure rate.                                                     | You can set this item to send an alert immediately upon failure.                                                                                                                                                                                   | If an alert is triggered for this item, you can check the status of BE nodes, disk status, and network status.                                                                             |

### 2.5 Metrics for service availability[​](#25-metrics-for-service-availability "Direct link to 2.5 Metrics for service availability")

| **Metric**     | **Description**                                        | **Alert rule**                                                                                | **Note**                                                                                                                                                                                                                                                                                               |
| -------------- | ------------------------------------------------------ | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Meta Log Count | The number of BDB metadata log entries on the FE node. | It is recommended to configure this item to trigger an immediate alert if it exceeds 100,000. | By default, the leader FE node triggers a checkpoint to flush the log to disk when the number of logs exceeds 50,000. If this value exceeds 50,000 by a large margin, it usually indicates a checkpoint failure. You can check whether the Xmx heap memory configuration is reasonable in **fe.conf**. |

### 2.6 Metrics for system load[​](#26-metrics-for-system-load "Direct link to 2.6 Metrics for system load")

| **Metric**           | **Description**                                                             | **Alert rule**                                                                                                              | **Note**                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| -------------------- | --------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| BE CPU Idle          | CPU idle rate of the BE node.                                               | It is recommended to configure this item to trigger an alert if the idle rate is lower than 10% for 30 consecutive seconds. | This item is used to monitor CPU resource bottlenecks. CPU usage can fluctuate significantly, and setting a small polling interval may result in false alerts. Therefore, you need to adjust this item based on the actual business conditions. If you have multiple batch processing tasks or a large number of queries, you may consider setting a lower threshold.                                                                                 |
| BE Mem               | Memory usage for the BE node.                                               | It is recommended to configure this item to 90% of the available memory size for each BE.                                   | This value is equivalent to the value of Process Mem, and BE's default memory limit is 90% of the server's memory size (controlled by configuration `mem_limit` in **be.conf**). If you have deployed other services on the same server, be sure to adjust this value to avoid OOM. The alert threshold for this item should be set to 90% of BE's actual memory limit so that you can confirm whether BE memory resources have reached a bottleneck. |
| Disks Avail Capacity | Available disk space ratio (percentage) of the local disks on each BE node. | It is recommended to configure this item to trigger an alert if the value is less than 20%.                                 | It is recommended to reserve sufficient available space for StarRocks based on your business requirements.                                                                                                                                                                                                                                                                                                                                            |
| FE JVM Heap Stat     | JVM heap memory usage percentage for each FE node in the cluster.           | It is recommended to configure this item to trigger an alert if the value is greater than or equal to 80%.                  | If an alert is triggered for this item, it is recommended to increase the Xmx heap memory configuration in **fe.conf**; otherwise, it may affect query efficiency or lead to FE OOM issues.                                                                                                                                                                                                                                                           |

## Step 3: Configure alert via Email[​](#step-3-configure-alert-via-email "Direct link to Step 3: Configure alert via Email")

### 3.1 Configure SMTP service[​](#31-configure-smtp-service "Direct link to 3.1 Configure SMTP service")

Grafana supports various alerting solutions, such as email and webhooks. This tutorial uses email as an example.

To enable email alerting, you first need to configure SMTP information in Grafana, allowing Grafana to send emails to your mailbox. Most commonly used email providers support SMTP services, and you need to enable SMTP service for your email account and obtain an authorization code.

After completing these steps, modify the Grafana configuration file on the node where Grafana is deployed.

```bash
vim /usr/share/grafana/conf/defaults.ini

```

Example:

```properties
###################### SMTP / Emailing #####################
[smtp]
enabled = true
host = <smtp_server_address_and_port>
user = johndoe@gmail.com
# If the password contains # or ; you have to wrap it with triple quotes.Ex """#password;"""
password = ABCDEFGHIJKLMNOP  # The authorization password obtained after enabling SMTP.
cert_file =
key_file =
skip_verify = true  ## Verify SSL for SMTP server
from_address = johndoe@gmail.com  ## Address used when sending out emails.
from_name = Grafana
ehlo_identity =
startTLS_policy =

[emails]
welcome_email_on_sign_up = false
templates_pattern = emails/*.html, emails/*.txt
content_types = text/html

```

You need to modify the following configuration items:

* `enabled`: Whether to allow Grafana to send email alerts. Set this item to `true`.
* `host`: The SMTP server address and port for your email, separated by a colon (`:`). Example: `smtp.gmail.com:465`.
* `user`: SMTP username.
* `password`: The authorization password obtained after enabling SMTP.
* `skip_verify`: Whether to skip SSL verification for the SMTP server. Set this item to `true`.
* `from_address`: The email address used to send alert emails.

After the configuration is complete, restart Grafana.

```bash
systemctl daemon-reload
systemctl restart grafana-server.service

```

### 3.2 Create alert channel[​](#32-create-alert-channel "Direct link to 3.2 Create alert channel")

You need to create an alert channel (Contact Point) in Grafana to specify how to notify contacts when an alert is triggered.

1. Log in to the Grafana Web UI, click on the menu button in the upper-left corner, expand **Alerting**, and select **Contact Points**. On the **Contact points** page, click **Add contact point** to create a new alert channel.

   ![MA-19](/assets/images/monitor19-6109fd0b12144f79f8f68466b5b47507.png)

2. In the **Name** field, customize the name of the contact point. Then, in the **Integration** dropdown list, select **Email**.

   ![MA-20](/assets/images/monitor20-8e6a36e70958a93278b465e375aea321.png)

3. In the **Addresses** field, enter the email addresses of the contacts to receive the alert. If there are multiple email addresses, separate the addresses using semicolons (`;`), commas (`,`), or line breaks.

   The configurations on the page can be left with their default values except for the following two items:

   * **Single email**: When enabled, if there are multiple contacts, the alert will be sent to them through a single email. It's recommended to enable this item.
   * **Disable resolved message**: By default, when the issue causing the alert is resolved, Grafana sends another notification notifying the service recovery. If you don't need this recovery notification, you can disable this item. It's not recommended to disable this option.

4. After the configuration is complete, click the **Test** button in the upper-right corner of the page. In the prompt that appears, click **Sent test notification**. If your SMTP service and address configuration are correct, the target email account should receive a test email with the subject "TestAlert Grafana". Once you confirm that you can receive the test alert email successfully, click the **Save contact point** button at the bottom of the page to complete the configuration.

   ![MA-21](/assets/images/monitor21-d97b1c38652ebb72514b2b5df4579fc6.png)

   ![MA-22](/assets/images/monitor22-4dca103ae0c8a75cfd54e43c39337abb.png)

You can configure multiple notification methods for each contact point through "Add contact point integration", which will not be detailed here. For more details about Contact Points, you can refer to the [Grafana Documentation](https://grafana.com/docs/grafana-cloud/alerting-and-irm/alerting/fundamentals/notifications/contact-points/)

For subsequent demonstration, let's assume that in this step, you have created two contact points, "StarRocksDev" and "StarRocksOp", using different email addresses.

### 3.3 Set notification policies[​](#33-set-notification-policies "Direct link to 3.3 Set notification policies")

Grafana uses notification policies to associate contact points with alert rules. Notification policies use matching labels to provide a flexible way to route different alerts to different contacts, allowing for alert grouping during O\&M.

1. Log in to the Grafana Web UI, click on the menu button in the upper-left corner, expand **Alerting**, and select **Notification policies**.

   ![MA-23](/assets/images/monitor23-0a648b6966faea93ea94a1870594e82a.png)

2. On the **Notification policies** page, click the more (**...**) icon to the right of **Default policy** and click **Edit** to modify the Default policy.

   ![MA-24](/assets/images/monitor24-bf167112de0d0d807fa4e548e1d813fb.png)

   ![MA-25](/assets/images/monitor25-7f769039514cfe682b6360d88d18ca55.png)

   Notification policies use a tree-like structure, and the Default policy represents the default root policy for notification. When no other policies are set, all alert rules will default to matching this policy. It will then use the default contact point configured within it for notifications.

   1. In the **Default contact point** field, select the contact point you created previously, for example, "StarRocksOp".

   2. **Group by** is a key concept in Grafana Alerting, grouping alert instances with similar characteristics into a single funnel. This tutorial does not involve grouping, and you can use the default setting.

      ![MA-26](/assets/images/monitor26-f0f18c8a09e860caf8e1292f0a3df0ab.png)

   3. Expand the **Timing options** field and configure **Group wait**, **Group interval**, and **Repeat interval**.

      * **Group wait**: The time waiting for the initial notification to send after the new alert creates a new group. Default 30 seconds.
      * **Group interval**: The interval at which alerts are sent for an existing group. Defaults to 5 minutes, which means that notifications will not be sent to this group any sooner than 5 minutes since the previous alert was sent. This means that notifications will not be sent any sooner than 5 minutes (default) since the last batch of updates were delivered, regardless of whether the alert rule interval for those alert instances was lower. Default 5 minutes.
      * **Repeat interval**: The waiting time to resend an alert after they have successfully been sent. The interval at which alerts are sent for an existing group. Defaults to 5 minutes, which means that notifications will not be sent to this group any sooner than 5 minutes since the previous alert was sent.

      You can configure the parameters as shown below so that Grafana will send the alert by these rules: 0 seconds (Group wait) after the **alert conditions are met**, Grafana will send the first alert email. After that, Grafana will re-send the alert every 1 minute (Group interval + Repeat interval).

      ![MA-27](/assets/images/monitor27-1555bd354163a5ba006ff518c875ff62.png)

      > **NOTE**
      >
      > The previous paragraph uses "meeting the alert conditions" rather than "reaching the alert threshold" to avoid false alerts. It's recommended to set the alert to be triggered a certain duration of time after the threshold has been reached.

3. After the configuration is complete, click **Update default policy**.

4. If you need to create a nested policy, click on **New nested policy** on the **Notification policies** page.

   Nested policies use labels to define matching rules. The labels defined in a nested policy can be used as conditions to match when configuring alert rules later. The following example configures a label as `Group=Development_team`.

   ![MA-28](/assets/images/monitor28-4ee3adad9577249db57b9f77193d8ea9.png)

   In the **Contact point** field, select "StarRocksDev". This way, when configuring alert rules with the label `Group=Development_team`, "StarRocksDev" is set to receive the alerts.

   You can have the nested policy inherit the timing options from the parent policy. After the configuration is complete, click **Save policy** to save the policy.

   ![MA-29](/assets/images/monitor29-757db2ff569cc44b57305940cc844a67.png)

If you are interested in the details of notification policies or if your business has more complex alerting scenarios, you can refer to the [Grafana Documentation](https://grafana.com/docs/grafana-cloud/alerting-and-irm/alerting/fundamentals/notifications/contact-points/) for more information.

### 3.4 Define alert rules[​](#34-define-alert-rules "Direct link to 3.4 Define alert rules")

After setting up notification policies, you also need to define alert rules for StarRocks.

Log in to the Grafana Web UI, and search for and navigate to the previously configured StarRocks Overview Dashboard.

![MA-30](/assets/images/monitor30-a9c570096158e3748f7eaafb731ef8fe.png)

![MA-31](/assets/images/monitor31-39435dfb2795b8b967dbd7022ea11aee.jpeg)

#### 3.4.1 FE and BE status alert rule[​](#341-fe-and-be-status-alert-rule "Direct link to 3.4.1 FE and BE status alert rule")

For a StarRocks cluster, the status of all FE and BE nodes must be alive. Any node with a status of DEAD should trigger an alert.

The following example uses the Frontends Status and Backends Status metrics under StarRocks Overview to monitor FE and BE status. As you can configure multiple StarRocks clusters in Prometheus, note that the Frontends Status and Backends Status metrics are for all clusters that you have registered.

##### Configure the alert rule for FE[​](#configure-the-alert-rule-for-fe "Direct link to Configure the alert rule for FE")

Follow these procedures to configure alerts for **Frontends Status**:

1. Click on the More (...) icon to the right of the **Frontends Status** monitoring item, and click **Edit**.

   ![MA-32](/assets/images/monitor32-ffeb20bf4d65aec43d33d6e55c77bf28.jpeg)

2. On the new page, choose **Alert**, then click **Create alert rule** from this panel to enter the rule creation page.

   ![MA-33](/assets/images/monitor33-2dd84ee6a971dcad5cf34fbc3289d1be.jpeg)

3. Set the rule name in the **Rule name** field. The default value is the title of the monitoring metric. If you have multiple clusters, you can add the cluster name as a prefix for differentiation, for example, "\[PROD]Frontends Status".

   ![MA-34](/assets/images/monitor34-462961937451b355d3dec8634360782e.jpeg)

4. Configure the alert rule as follows.

   1. Choose **Grafana managed alert**.
   2. For section **B**, modify the rule as `(up{group="fe"})`.
   3. Click on the delete icon on the right of section **A** to remove section **A**.
   4. For section **C**, modify the **Input** field to **B**.
   5. For section **D**, modify the condition to `IS BELOW 1`.

   After completing these settings, the page will appear as shown below:

   ![MA-35](/assets/images/monitor35-9b9881f59ac229c3d6b340b22087cf62.jpeg)

   Click to view detailed instructions

   Configuring alert rules in Grafana typically involves three steps:

   1. Retrieve the metric values from Prometheus through PromQL queries. PromQL is a data query DSL language developed by Prometheus, and it is also used in the JSON templates of Dashboards. The `expr` property of each monitoring item corresponds to the respective PromQL. You can click **Run queries** on the rule settings page to view the query results.
   2. Apply functions and modes to process the result data from the above queries. Usually, you need to use the Last function to retrieve the latest value and use Strict mode to ensure that if the returned value is non-numeric data, it can be displayed as `NaN`.
   3. Set rules for the processed query results. Taking FE as an example, if the FE node status is alive, the output result is `1`. If the FE node is down, the result is `0`. Therefore, you can set the rule to `IS BELOW 1`, meaning an alert will be triggered when this condition occurs.

5. Set up alert evaluation rules.

   According to the Grafana documentation, you need to configure the frequency for evaluating alert rules and the frequency at which their status changes. In simple terms, this involves configuring "how often to check with the alert rules" and "how long the abnormal state must persist after detection before triggering the alert (to avoid false alerts caused by transient spikes)". Each Evaluation group contains an independent evaluation interval to determine the frequency of checking the alert rules. You can create a new folder named **PROD** specifically for the StarRocks production cluster and create a new Evaluation group `01` within it. Then, configure this group to check every `10` seconds, and trigger the alert if the anomaly persists for `30` seconds.

   ![MA-36](/assets/images/monitor36-45dbde604bf936329236090fe7560a61.png)

   > **NOTE**
   >
   > The previously mentioned "Disable resolved message" option in the alert channel configuration section, which controls the timing of sending emails for cluster service recovery, is also influenced by the "Evaluate every" parameter above. In other words, when Grafana performs a new check and detects that the service has recovered, it sends an email to notify the contacts.

6. Add alert annotations.

   In the **Add details for your alert rule** section, click **Add annotation** to configure the content of the alert email. Please note not to modify the **Dashboard UID** and **Panel ID** fields.

   ![MA-37](/assets/images/monitor37-af96a30ba8b1a14ec11923250de8aa6d.jpeg)

   In the **Choose** drop-down list, select **Description**, and add the descriptive content for the alert email, for example, "FE node in your StarRocks production cluster failed, please check!"

7. Match notification policies.

   Specify the notification policy for the alert rule. By default, all alert rules match the Default policy. When the alert condition is met, Grafana will use the "StarRocksOp" contact point in the Default policy to send alert messages to the configured email group.

   ![MA-38](/assets/images/monitor38-249ddb362026eabbad3bdd3dba5f8f06.jpeg)

   If you want to use a nested policy, set the **Label** field to the corresponding nested policy, for example, `Group=Development_team`.

   Example:

   ![MA-39](/assets/images/monitor39-bcb6c9f6d736823e0dc4550a73901a1a.jpeg)

   When the alert condition is met, emails will be sent to "StarRocksDev" instead of "StarRocksOp" in the Default policy.

8. Once all configurations are complete, click **Save rule and exit**.

   ![MA-40](data:image/jpeg;base64,/9j/4AAQSkZJRgABAQEAYABgAAD/2wBDAAMCAgMCAgMDAwMEAwMEBQgFBQQEBQoHBwYIDAoMDAsKCwsNDhIQDQ4RDgsLEBYQERMUFRUVDA8XGBYUGBIUFRT/2wBDAQMEBAUEBQkFBQkUDQsNFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBQUFBT/wAARCACkAZcDAREAAhEBAxEB/8QAHQABAAIDAQEBAQAAAAAAAAAAAAMEBQYIBwECCf/EAEQQAAEDAgIHBAUIBwkBAAAAAAABAgMEBQYRBxIVVJGT0hQXIdMTGDFBYQgWIjJRVXOyMzQ1ZHJ0gTY3RXGDobGzwSP/xAAcAQEAAQUBAQAAAAAAAAAAAAAAAgMEBQYHAQj/xABKEQACAQMBAggJBgoLAQEAAAAAAQIDBBEhBTEGEhdBUWGR0xMVUlRVgZPR0hYiMnGSlDQ1NkVzdKGxssIHFCNCcnWCs8Hw8WLh/9oADAMBAAIRAxEAPwD+dmzKbeZeSnUXBTyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMjZlNvMvJTqAyNmU28y8lOoDI2ZTbzLyU6gMkwPAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAb3hXQ3fMXWWG6Uk9FBTTK5GJUSPRy6qq1Vya1fei8DarDg5ebQoK4pSiovOMt8zxzJmIuNp0Lao6Uk210f+mX9XbEm+2nmyeWZH5HX/AJcO2Xwlt47t/Jl2L3j1dsSb7aebJ5Y+R1/5cO2Xwjx3b+TLsXvHq7Yk32082Tyx8jr/AMuHbL4R47t/Jl2L3j1dsSb7aebJ5Y+R1/5cO2Xwjx3b+TLsXvHq7Yk32082Tyx8jr/y4dsvhHju38mXYvePV2xJvtp5snlj5HX/AJcO2Xwjx3b+TLsXvHq7Yk32082Tyx8jr/y4dsvhHju38mXYvePV2xJvtp5snlj5HX/lw7ZfCPHdv5Muxe8ertiTfbTzZPLHyOv/AC4dsvhHju38mXYvePV2xJvtp5snlj5HX/lw7ZfCPHdv5Muxe8/EvyecSxxPe2qtkqtRVRjJpM3fBM2In+5CXBDaEU2pQfrfwnq21bN4w+xe880pqeSrqIoIm68srkYxueWaquSJ4mh1KkaUJVJvCSy/qRtNpa1r64p2lvHjVKklGK0WXJ4Sy8JZb53g3WLRHcZImvddLZEqoirG902afBco1T/cto7T2TJLN5Bf6avdHQuTLhn6OftKHen77oK/73tXGfyiXjLY/nsPs1u6HJnwz9Gy9pQ70d0Ff972rjP5Q8ZbH89h9mt3Q5M+Gfo2XtKHejugr/ve1cZ/KHjLY/nsPs1u6HJnwz9Gy9pQ70d0Ff8Ae9q4z+UPGWx/PYfZrd0OTPhn6Nl7Sh3o7oK/73tXGfyh4y2P57D7NbuhyZ8M/RsvaUO9HdBX/e9q4z+UPGWx/PYfZrd0OTPhn6Nl7Sh3o7oK/wC97Vxn8oeMtj+ew+zW7ocmfDP0bL2lDvR3QV/3vauM/lDxlsfz2H2a3dDkz4Z+jZe0od6O6Cv+97Vxn8oeMtj+ew+zW7ocmfDP0bL2lDvR3QV/3vauM/lDxlsfz2H2a3dDkz4Z+jZe0od6O6Cv+97Vxn8oeMtj+ew+zW7ocmfDP0bL2lDvR3QV/wB72rjP5Q8ZbH89h9mt3Q5M+Gfo2XtKHejugr/ve1cZ/KHjLY/nsPs1u6HJnwz9Gy9pQ70d0Ff972rjP5Q8ZbH89h9mt3Q5M+Gfo2XtKHejugr/AL3tXGfyh4y2P57D7NbuhyZ8M/RsvaUO9HdBX/e9q4z+UPGWx/PYfZrd0OTPhn6Nl7Sh3o7oK/73tXGfyh4y2P57D7NbuhyZ8M/RsvaUO9Hc9c3ZJFcrZNIq5MiY6bNyr7ETONETiVqF5s26rQt7e7hKc2opcWqstvCWXTS39JZXv9H3CvZ1rVvbuwcaVKMpyfHovEYptvCqtvCWcJN9CNDKpoBlvm3Ong6eBrve1Vd4cGgYHzcm3in4v6Qe4Hzcm3in4v6QMD5uTbxT8X9IGB83Jt4p+L+kDA+bk28U/F/SBgfNybeKfi/pAwPm5NvFPxf0gYHzcm3in4v6QMENVZJqSnfMssUjWZayMVc08cveifaDzBjwAAAAAAAAAAAdR6Fv7s7N/rf9zzunBr8VUf8AV/Ezn+1fwyfq/cjdzZzElmgttXdaj0FFSz1k+Su9FTxq92Se1ckTMpzqQprjTaS69CUYuTxFZZBJG6J7mParHtVUc1yZKi/YpNNNZR4008M/J6eAAAAAmpqSetl9HTwyTyZK7UiYrlyTxVckISlGCzJ4R6k5NJbyEmeAAAHGuHf7QWz+ai/Oh8q7Q/A63+GX7mfQPBL8otnfpqX8cTr+0fsmi/AZ+VD6H4N/iSx/RU/4Ec04cflVtX9Yrf7ki2bGaSW4bVW1FDNWRUdRLRwrqyVDInLGxfsc5EyT+pSlUhGShKSTe5Z1JqEpJtLcVCqQAAAAAAJkpJ3UzqlIZFp2uRjpkauojl9iKvsz+BDjRUuLnUkk2m0txCTIgAAAAmqqSeilWKohkp5URF1JWK12S+KLkpCMozWYvJJpreiEmRABZrrbV2uVkdZSzUkj2NlayeNWK5i+LXIip4ovuUpwqQqZcGnjTTp6CTi0k2t+q610lYqETE4g/wAN/nYv/TnPDP8ANn61R/mO2f0Yfn7/AC+5/kOOzlRiTdKj9Yl/iX/kEiMAs0tsrK6GeampJ6iKButLJFG5zY0+1yongn+Yeiy9wWrwisAAAAACRtNM+B87YnuhYqNdIjV1WqvsRV9wem8b9CMAr3L9l1f8LfztAZq4IgAAAAAAAAAA6j0Lf3Z2b/W/7nndODX4qo/6v4mc/wBq/hk/V+5G7mzmJPdNC2N6XDOje9W+rr7vgta+vY6LF1uoXTx6zWfq0ipk77XZNXPxXPw9uobXs53F3SqQjGrxU805PGcv6S5urXQzFhWjShVUsx42Fx0s46n9fabXddH81/0jXe6Y2pLZiOlhs9LVRXaGtS1UE7JFVsU9S9fpo9yNVMmoqqqfYYqlext7ONKycqcnOScWuPJNauMVuwusvXQlXrwnWfhI8TKkvm5XM5Pm6+fcQ4q0F4Wlr8RWWw0q7ansNLf7QkVW+eNPpKk8LHLl6RqoiK1zkz8Sdtti7Uada4l8yNR055STw0uK2tcNPfgjUtLfOEvp03KOHlKUW8pPn4yWmfUWaPQngWDEWLUmippaPClHRUc8dwur6Snqa6RF9LJJN4qxqL9FGtyTNMinLa99KhRlFvNaU2sR4zUI7klplvfrzEoWNCNV0qm+nCLlzJyl18yXUeRacMKYbwtimjbhitpqigq6KOpkp6WsSrZSyrmjo2y5Irm+CKir45KbPse5ubmlL+txacZNJtcVtczxzPmeDHX1GjS4kqTXzlqk84f19D957DNow0dv0nU+C4MLztkbZ1uc1e65TLrv7Kr0jazPwTW1XZ5555plkaq9o7RVjUvnWWkuKlxV5SWc/Vpu695koWtsq1pQlDPhEm3l+TLTHW1n9h+vk64MtdBhvDWJYqZzLvXsvFPNOsjlR8bIF1URqrkmS5+KIR4QXdadSvat/MioNLrckQ2TRi5Ua7XzvC49Xg5P958boN0eWixW213qutlLcayztr5LxUX30VVHM+NXt1KVU1XRJllmq5qmf2E57Y2jUq1KlvGTjCfFUVDKaTxrLemyVtZ2zo0XWa/tIptuWGs7sLnxz56zllUyVUzzy96HRzW2sPB8B4ca4d/tBbP5qL86HyrtD8Drf4ZfuZ9A8Evyi2d+mpfxxOv7R+yaL8Bn5UPofg3+JLH9FT/gRzThx+VW1f1it/uSLZsZpJ1tg7GVZeKfBeHrJfq3A+Iqa2spY8MXq2PW23ZVYq+lVzcvCTx+k5F+Hx5peWkITubmtTVam5NucZLjw6tfJ6F6+rZbW4SpUKUZunJYxp82eXo/Xu6OjpMJos0L4dvFLbocV2Cmprhd7jU0/p6m99ldkx6syo6dmayarkVF1/D/ADQuto7VuKbk7Oq3GEIvSHG3rPz5PCWV0HlG2puUpXEFlzcd/FW/DUcatp827ciphXQXhrFUVhqIY5I6S0Xiut+JZFmd9OKFHSsk9v0NZjNVdXLxUncbaubbjuX9+nGVNdDeE116vOvMQhs+nWl4CnnjRquLfTHV56sJNZ6TK4Z0QaOH4csNzu0dBDBiSWeeN1ffnUktDT+kVsbYI8l9M5qZZ66+3w95b3G09pKtUo0m80lHOIcZSljLcn/dT5scxVo29pOHh8LiylJJOTWIxeNOl8+vUeU6JsCWPEenSkwxXSJd7EtVUw+lhkVqTxsZIrHo5qp7dVF8FNl2jeV6GypXcVxanFTx0NtZWv1mMp0Kbv426fGg5Y+teoy+JMN4PvWibEGJbDh2Sy1FBeYLbCj66SoVY1Y5XOdnkmbly8MvDLwLO3r31LaFG2uKvHU4yk/mpdGEvq1LydC2lSuZQjh0+KlrvfGw361zcx7VWaHLNNZ7lge3JJbbTV4ktmv/APRXvajqP0kmSuVfFfHLP2Kpqkdq1lUp31T50owqdW6WFuMpG0hG3qUqenHjSz9bnhv/APDyvTHo8wHacGV1wsU1qtt2oK9tOyjob4twdVQqqoqyNciLHI1UzVG5plmbDsu/2hVuYQrqUoTjltw4qT36Nb092vUWF3a2tOnVUcKUGsfOy3rhprmfPp0Mi0G6NcPX7DUdyxJZYKiOtuTaGCruV67BCrck1mwMZm+WVFX2Kmr7Ez9pW2xf3NGt4G2qNNRcmow4z6m86KP7TH2NKnOE6tSOUmlq+Kt2frb/AGYMtctHGBNHVixTcLvYanEi23FbrNTRrXvp84VjRyayt9qp4+xEVVy8ci0o39/tCrb0qNRQ49PjN8VPXONMmTuLO3tFcTceMoOOFnH0lnDfr379C1jjRRgTRLTYju9wstZiWl21FbKGh7c6DszHU7J1c57Uzc5NfVRF+xM/eU7LaW0Npyo0KdRQk4OUpcVPOJOKwnpzZZCva29uqtdxzFcXEctfSjnV78cxpfysWsbplrGxo5GJQ0aNR/1kT0DMs/iZTgxl2Gu/jS/eUNrY41Hi7uJH/k9GuugrR5YLRsa5V1so7psdKxbxPfNSq7QseuiJSKmr6L3Z555cTCLbO0a1SVajGTip44qhlYTx9PfxvVgu6FjaxhSjXazOOW+NhrOdy3NLr6zWbzoxwhR6Ge8WOzVbW1lBFRwWp0kupT1qyOY+qV+eaxLq5tRVyVVyL6O0bx7S8WeEWks8bT6OM8XHldPPjUt6NrQnZ/1yUfopprXDlnClv3dK9R6Rdo7DZMOY0S42ea/OiwnaamSStuMrnvY5ckia5c1YiOTWzb9uRgIuvVrUvBTUP7aa0it+N76ejUv7LiKEHNcb+wzq+bXKXRnq3cxxqviq5Jkn2HWDUXv0MRiD/Df52L/05zwz/Nn61R/mO1/0Yfn7/L7n+Q47OVGJN0qP1iX+Jf8AkEiMA6W0i6SMWaIu7qx4ErJbZZJLHR10UdJEituVRLmsrpPD/wCiq76OqvsLupKX9fqUksxi1GMeZxwv3vOvOUYJK0hUejkm5PnTy9M//On/ACTYZ0dW3EV7xHe8cYIobLUVV6ZSLT1t62ZR0z3MR8kcLG60kkv0kdq/VTPL4FK3o0+LTpp/Sbx04UsYiv8A5emvUidarNudSX92KfQs4zmT5srXTrbIL7ovwDovs+Orjd7DV4nSy4oZaaOFbg+m1oXwq9Ee5ieOX2oiKqonsTNC3o1IulSlNZcpTi/9L346err6i5lTbq1VF6RhCX2ub9v7CxjvRFgLRHDijEVfZK3EtsZdaW32+0rXugSnbLStqHOklYms7LW1W/5eOZXklR4tOercpLPVF8y6dezUpqLqpVY6LiRljrk2tX0aftwVcEaIsDXGhxfimpoZIrRS18FDb7Riu57K9F6SNJHLLKxHKqongxE+snip6oqnSjx9W5SW/Gi6vK115ljJBPwlR8XRKKfTq3jf0aaPn3FzH2F8PYS0L6SKLC1xiuVmfd7VPEsNQlQkLnxvV0XpE+tqrmiL70yzLW4b4lGD1Uas0n0rwa1/4fWme0ceFqyW904Nrfh8d/8AvrOYCoCvcv2XV/wt/O0BmrgiAAAAAAAAAADqPQt/dnZv9b/ued04Nfiqj/q/iZz/AGr+GT9X7kbubOYk3PAml3EujqmqaS0VcLrfUvSSWhrKaOohc9EyR+q9Fyd8Uy9iGJvdl2t/JTrL5y3NNp9qLqhc1bZvwb0e9b0/UZCl0941p7/c7vNdI7hUXKJkFXDXUsU0EjGfUb6JW6qI1VzTJE/3UoPYlk6MaMYYUXlNNppve85zl85Wd/cOr4Vy1xjGFjHRjoKs+mzGtRjGhxW+9u2/Qwdlp6ttNC3UiycmrqIzUVPpO9qL7fghVjsiyjbztVT+ZJ5ay9+muc55luZCd5XqVIVZS1hu0Wm/q62U8N6VsTYVvtzu1DcEdVXRXLXsqYWTRVWs5XL6SNyK1fFVX2eGfhkVK+zbW4owoTj82H0cNprGmj3kY3daNZ11L5zznrzv03f90KGMsb3nH142le6tKqqSNsMerG2NkUbc9VjGNREa1M18EQr2lnRsafg6Cws5erbb6W3zka9xUuGnUe7RcyS9RkU0s4sTGdLixLu5t/po2QxVbYY25MazURqsRuqqavh4p4+8oeLLT+rytOJ8yTbay97eenO89lc1pSpzctYJKL00Sz73vM1W/KJ0gV1TBK+9sZ6D03oY46GnayJJWaj0aiR5ZK3w/wB/b4lnHYOzoxcfB78J/Olrh5XP0+7cXL2jcuUZcbc+MtFo8NZ3dDfbneQW/T3ja14cjs1PdWtp4oFpIah1NE6phgVMljZMrddrcvj4e7InV2LY1qzrzhq3lrLw30tZwyFG+uKEFCEt27RNr6m93/cHnpnTHgA41w7/AGgtn81F+dD5V2h+B1v8Mv3M+geCX5RbO/TUv44nX9o/ZNF+Az8qH0Pwb/Elj+ip/wACOacOPyq2r+sVv9yRbNjNJPS7N8orHdhsdPbKS7R6lLF6ClqpaSGSpp48stRkrmq5Ey/qnuyNfrbCsLiq6s4PMtWstJvpaTMjRv69CChB6Ldonj6slfDenzGuFbVTUFvukKNpJHS0tRUUcM89Ornaz0ZI9qqiOVVzT35lS42LZXNR1KkN6w0m0mlospNbuYhTva9JNRlz51SeG97Wed85hrZpRxRZ6LEVJR3Z8FPiHW2nG2KNUqM1dn7W/Qz1nfVy9uRdVNm2tXwTnDPgvo6vTGOvXct+SEbutCc6kZYc853a5z73uMnhHThjDBNljtdruELaSBzn0vaaSKd1I931nROe1VYq/AtrrZFneVPC1YvL0eG1ldDw9SVC7rW8eLB6ZzhpPD6Vnca7YMZ3rDOJ48RW2udDeWPfKlW9jZXaz0VHKqPRUVVRy+1PeZCvaULig7apH5jWMbtFu3FFV6ka3h0/nZznrMrgrS1inR7HcY7Hcm00VeqOnilp45mOemeq9GvaqI5M/ahbXezLW+4nh454u7Vp/VlNaFSldVqM5VIPWW/Ra+rcWrnpxxzeJZZKq/yPllq4K58jIIo3rPC1GRv1msRUVGoiZJ4L78yjT2NYUsKFPcmt73S3rV8//hUnfXFSLjOecpJ7v7ryux653nzGmmnFmPrS22XauhWi9Kk8sVLSRU/aJUTL0kmo1Nd3j7/D4HtpsizsqnhaMddyy28LoWXoe1b6vWg4Tej36LLxuy+c+YP0zYrwLZnWu1VsDaRJu0wJU0kU7qaZUyWSJXtXUcqeHgSutlWt5UVWrF5xjRtZXQ8PVFKhc1LeMoQ3PXVJ67s685Qvmk/E2JKGvo7jcu0U1dcNq1DPQRN16nV1fSZtaip4eGqmSfArUNn21tOE6UMOEeKtXos5xq+np1JVbyvXU1Ulnj4zotcLC/Z0GeoPlCY7objda3bEdRNc1Y+pbU0UEkbpGMRjJEYrNVrmtaiZoiexM8yxnsKwnCFPiYUc4w2nq8tZznDJxv7iM3PjZyknlLDS3abtDUsXYxvGPb5JeL9WLX3KRjGPnWNjFcjWo1vg1ET2InjkZO1tKFlT8Dbx4sdXjXn+st69xVuJKVV5aWPUbTDp8xvBhpLI27M7O2mWjbUrSxLVNp1TJYkm1ddG5fHP4mPnsWxnW8O4a5y1l4b6Ws4yXFK+r0aapwlotFospdT5v+4KFRpkxhVw1MMt31qaotzLVJTdlhSFaZv1WJHqarcvFUciI7NfaVlsqzTUuJqpcfOXnjdOc59W7qIRvK8FFRlok4pYWMPfpuf1vXrLNDpzxpQXaS4suzJJ5aGO2yNlpIXRyU7PqMcxWaq5fbln8SnU2NZVKbpuGnGct7zxnvaec/8AB5Su61GUJwl9GPFX+Ho6/XqaGq5qqmaLRvLyYjEH+G/zsX/pznhn+bP1qj/Mdr/ow/P3+X3P8hx2cqMSbpUfrEv8S/8AIJEYB6LhH5QmPMD2KGz2q9NZQUyqtKyppIah1KqrmqxOkY5Wf0XLxKkqk5JJvdpnnx0Z/wC9RGMVHOOfXHNn6iHD2nfGuGaathpbrHUJV1a3Bz6+khqnx1S+CzsdK1yseqe9CMZOEYwjpxd3SulZ6+clL58pTlrxt/Xjdp1cxjMRaVcU4rorrSXW6dqp7pXtudYzs8TPS1LWaiPza1FT6PhkmSfApqKjGMVui219ct/b/wCE+PLMpZ1kkn9S3Gbt/wAofHlvu9xuO2IqqW4NhbVQ1dDBLBL6JqNid6JzNRHNREyciIvh45lbwktdd7cvW+ddDKfFWIroXF9XQ+n1la06dca2m+Xi6pd21tReHNfcIq+miqIalW/VV0T2q1NX3ZImWWSeBGEnCHEW7OenXp15z2Xzpcd78Y9XR9RRxDpexdiujvFJdbw6sp7tPDUVkboYkR74m6seWTU1EangjW5J8Cm4qSSfM3L1tYb69NNT1PEnJc6UfUnlLq11NPJHhXuX7Lq/4W/naAzVwRAAAAAAAAAAB7Jo600WjCmEqO011HWvmp1k+nTsY5rkc9zve5MvrZf0Oj7H4S2thZQtq0JNxzuw97b52uk1i92XVuK8qsJLDxvz0Y6DZPWJw3uV25UfmGa+WNh5E+yPxFh4kuPKj2v3D1icN7lduVH5g+WNh5E+yPxDxJceVHtfuHrE4b3K7cqPzB8sbDyJ9kfiHiS48qPa/cPWJw3uV25UfmD5Y2HkT7I/EPElx5Ue1+4esThvcrtyo/MHyxsPIn2R+IeJLjyo9r9w9YnDe5XblR+YPljYeRPsj8Q8SXHlR7X7h6xOG9yu3Kj8wfLGw8ifZH4h4kuPKj2v3D1icN7lduVH5g+WNh5E+yPxDxJceVHtfuHrE4b3K7cqPzB8sbDyJ9kfiHiS48qPa/cPWJw3uV25UfmD5Y2HkT7I/EPElx5Ue1+4/MvyisPpE9YqC5ukyXVa+ONEVfdmuuuXAjLhjY4fFpzz9S+I9WxLjOsl+33Hg9hkZDfLdJI5rI2VMbnOcuSIiOTNVU4nfRlO0qxistxl+5nWuDFanb7esK1aSjCNam228JJTi223oklq29x0pQaUsPU1DTQvrotaNjWLqzR5eCZKqLrJ/wAG9bI4eWlhs62tK1nXcqcIxeILGYxSePnrTToRnuEf9Gs9r7bvdpW+2bJQrVak0nX1SnNySeINZw9cNrPOT97OG9+Zzo+sy3KPYeaXHs4/Ga3yS3Hpqw+8Pux3s4b35nOj6xyj2Hmlx7OPxjkluPTVh94fdjvZw3vzOdH1jlHsPNLj2cfjHJLcemrD7w+7Hezhvfmc6PrHKPYeaXHs4/GOSW49NWH3h92O9nDe/M50fWOUew80uPZx+Mcktx6asPvD7sd7OG9+Zzo+sco9h5pcezj8Y5Jbj01YfeH3Y72cN78znR9Y5R7DzS49nH4xyS3Hpqw+8Pux3s4b35nOj6xyj2Hmlx7OPxjkluPTVh94fdjvZw3vzOdH1jlHsPNLj2cfjHJLcemrD7w+7Hezhvfmc6PrHKPYeaXHs4/GOSW49NWH3h92O9nDe/M50fWOUew80uPZx+Mcktx6asPvD7sd7OG9+Zzo+sco9h5pcezj8Y5Jbj01YfeH3Y72cN78znR9Y5R7DzS49nH4xyS3Hpqw+8Pux3s4b35nOj6xyj2Hmlx7OPxjkluPTVh94fdjvZw3vzOdH1jlHsPNLj2cfjHJLcemrD7w+7Hezhvfmc6PrHKPYeaXHs4/GOSW49NWH3h92O9nDe/M50fWOUew80uPZx+Mcktx6asPvD7sxeItKNilo2TU9XHLJSypUei9NHnJqoqo1PpL4r7PYatt7hdQ2zKyhb2taLpV6dRuUElxY5zuk9defC6ze+DPA2PBO32vc3e1bSr4azr0oxp1uNJzkotaOMc/Ra0y8tJI5qMKclNjfiCjkcr1SdquXNURjVy/rrA9yfNu0X7xy29QGRt2i/eOW3qAyNu0X7xy29QGRt2i/eOW3qAyNu0X7xy29QGRt2i/eOW3qAyNu0X7xy29QGRt2i/eOW3qAyV6+8009FLDEkqvkREze1ERMlRftX7AMmFB4ACfs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gDs0e9Q8H9IA7NHvUPB/SAOzR71Dwf0gHx1MiRveyaOXVTNUbrIqJnl70T7UAIQAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAT036Gq/DT87QCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAE9N+hqvw0/O0AgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABPTfoar8NPztAIAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAT036Gq/DT87QCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAE9N+hqvw0/O0AgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABPTfoar8NPztAIAAAf/2Q==)

##### Test alert trigger[​](#test-alert-trigger "Direct link to Test alert trigger")

You can manually stop an FE node to test the alert. At this point, the heart-shaped symbol to the right of Frontends Status will change from green to yellow and then to red.

**Green**: Indicates that during the last periodic check, the status of each instance of the metric item was normal, and no alert was triggered. The green status does not guarantee that the current node is in normal status. There may be a delay in status change after a node service anomaly, but typically, the delay is not in the order of minutes.

![MA-41](/assets/images/monitor41-d6555dd45cee2a87e315e538a7fff036.png)

**Yellow**: Indicates that during the last periodic check, an instance of the metric item was found abnormal, but the abnormal state duration has not yet reached the "Duration" configured above. At this point, Grafana will not send an alert and will continue periodic checks until the abnormal state duration reaches the configured "Duration". During this period, if the status is restored, the symbol will change back to green.

![MA-42](/assets/images/monitor42-44ec0dd64f7e71b1f398710742990b94.jpeg)

**Red**: When the abnormal state duration reaches the configured "Duration", the symbol changes to red, and Grafana will send an email alert. The symbol will remain red until the abnormal state is resolved, at which point it will change back to green.

![MA-43](/assets/images/monitor43-859a3b060d8b224da48e1cf53f593e18.jpeg)

##### Manually pause Alerts[​](#manually-pause-alerts "Direct link to Manually pause Alerts")

Suppose the anomaly requires an extended period for resolution or the alerts are continuously triggered for some reasons other than anomaly. You can temporarily pause the evaluation of the alert rule to prevent Grafana from persistently sending alert emails.

Navigate to the Alert tab corresponding to the metric item on the Dashboard and click the edit icon:

![MA-44](/assets/images/monitor44-f370c012a1066414412bf5024cfba667.png)

In the **Alert Evaluation Behavior** section, toggle the **Pause Evaluation** switch to the ON position.

![MA-45](/assets/images/monitor45-2ef24b9bf18cc0429210e63aff8a9554.png)

> **NOTE**
>
> After pausing the evaluation, you will receive an email notifying you that the service is restored.

##### Configure the alert rule for BE[​](#configure-the-alert-rule-for-be "Direct link to Configure the alert rule for BE")

You can follow the above process to configure alert rules for BE.

Editing the configuration for the metric item Backends Status:

1. In the **Set an alert rule name** section, configure the name as "\[PROD]Backends Status".
2. In the **Set a query and alert condition** section, set PromSQL to `(up{group="be"})`, and use the same settings as those in the FE alert rule for other items.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "BE node in your StarRocks production cluster failed, please check! Stack information for BE failure will be printed in the BE log file **be.out**. You can identify the cause based on the logs".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

#### 3.4.2 Query alert rule[​](#342-query-alert-rule "Direct link to 3.4.2 Query alert rule")

The metric item for query failures is **Query Error** under **Query Statistic**.

Configure the alert rule for the metric item "Query Error" as follows:

1. In the **Set an alert rule name** section, configure the name as "\[PROD] Query Error".
2. In the **Set a query and alert condition** section, remove section **B**. Set **Input** in section **A** to **C**. In section **C**, use the default value for PromQL, which is `rate(starrocks_fe_query_err{job="StarRocks_Cluster01"}[1m])`, representing the number of failed queries per minute divided by 60s. This includes both failed queries and queries that exceeded the timeout limit. Then, in section **D**, configure the rule as `A IS ABOVE 0.05`.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "High query failure rate, please check the resource usage or configure query timeout reasonably. If queries are failing due to timeouts, you can adjust the query timeout by setting the system variable `query_timeout`".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

#### 3.4.3 User operation failure alert rule[​](#343-user-operation-failure-alert-rule "Direct link to 3.4.3 User operation failure alert rule")

This item monitors the rate of Schema Change operation failures, corresponding to the metric item **Schema Change** under **BE tasks**. It should be configured to alert when greater than 0.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] Schema Change".
2. In the **Set a query and alert condition** section, remove section **A**. Set **Input** in section **C** to **B**. In section **B**, use the default value for PromQL, which is `irate(starrocks_be_engine_requests_total{job="StarRocks_Cluster01", type="create_rollup", status="failed"}[1m])`, representing the number of failed Schema Change tasks per minute divided by 60s. Then, in section **D**, configure the rule as `C IS ABOVE 0`.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "Failed Schema Change tasks detected, please check promptly. You can increase the memory limit available for Schema Change by adjusting the BE configuration parameter `memory_limitation_per_thread_for_schema_change`, which is set to 2GB by default".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

#### 3.4.4 StarRocks operation failure alert rule[​](#344-starrocks-operation-failure-alert-rule "Direct link to 3.4.4 StarRocks operation failure alert rule")

##### BE Compaction Score[​](#be-compaction-score "Direct link to BE Compaction Score")

This item corresponds to **BE Compaction Score** under **Cluster Overview**, and is used to monitor the compaction pressure on the cluster.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] BE Compaction Score".
2. In the **Set a query and alert condition** section, configure the rule in section C as `B IS ABOVE 0`. You can use default values for other items.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "High compaction pressure. Please check whether there are high-frequency or high concurrency loading tasks and reduce the loading frequency. If the cluster has sufficient CPU, memory, and I/O resources, consider adjusting the cluster compaction strategy".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

##### Clone[​](#clone "Direct link to Clone")

This item corresponds to **Clone** in **BE tasks** and is mainly used to monitor replica balancing or replica repair operations within StarRocks, which usually should not fail.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] Clone".
2. In the **Set a query and alert condition** section, remove section A. Set **Input** in section **C** to **B**. In section **B**, use the default value for PromQL, which is `irate(starrocks_be_engine_requests_total{job="StarRocks_Cluster01", type="clone", status="failed"}[1m])`, representing the number of failed Clone tasks per minute divided by 60s. Then, in section **D**, configure the rule as `C IS ABOVE 0`.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "Detected a failure in the clone task. Please check the cluster BE status, disk status, and network status".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

#### 3.4.5 Service availability alert rule[​](#345-service-availability-alert-rule "Direct link to 3.4.5 Service availability alert rule")

This item monitors the metadata log count in BDB, corresponding to the **Meta Log Count** monitoring item under **Cluster Overview**.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] Meta Log Count".
2. In the **Set a query and alert condition** section, configure the rule in section **C** as `B IS ABOVE 100000`. You can use default values for other items.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "Detected that the metadata count in FE BDB is significantly higher than the expected value, which can indicate a failed Checkpoint operation. Please check whether the Xmx heap memory configuration in the FE configuration file **fe.conf** is reasonable".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

#### 3.4.6 System overload alert rule[​](#346-system-overload-alert-rule "Direct link to 3.4.6 System overload alert rule")

##### BE CPU Idle[​](#be-cpu-idle "Direct link to BE CPU Idle")

This item monitors the CPU idle rate on BE nodes.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] BE CPU Idle".
2. In the **Set a query and alert condition** section, configure the rule in section C as `B IS BELOW 10`. You can use default values for other items.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "Detected that BE CPU load is consistently high. It will impact other tasks in the cluster. Please check whether the cluster is abnormal or if there is a CPU resource bottleneck".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

##### BE Memory[​](#be-memory "Direct link to BE Memory")

This item corresponds to **BE Mem** under **BE**, monitoring the memory usage on BE nodes.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] BE Mem".
2. In the **Set a query and alert condition** section, configure PromSQL as `starrocks_be_process_mem_bytes{job="StarRocks_Cluster01"}/(<be_mem_limit>*1024*1024*1024)`, where `<be_mem_limit>` needs to be replaced with the current BE node's available memory limit, that is, the server's memory size multiplied by the value of the BE configuration item `mem_limit`. Example: `starrocks_be_process_mem_bytes{job="StarRocks_Cluster01"}/(49*1024*1024*1024)`. Then, in section **C**, configure the rule as `B IS ABOVE 0.9`.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "Detected that BE memory usage is consistently high. To prevent query failure, please consider expanding memory size or adding BE nodes".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

##### Disks Avail Capacity[​](#disks-avail-capacity "Direct link to Disks Avail Capacity")

This item corresponds to **Disk Usage** under **BE**, monitoring the remaining space ratio in the directory where the BE storage path is located.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] Disks Avail Capacity".
2. In the **Set a query and alert condition** section, configure the rule in section **C** as `B`` IS BELOW 0.2`. You can use default values for other items.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "Detected that BE disk available space is below 20%, please release disk space or expand the disk".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

##### FE JVM Heap Stat[​](#fe-jvm-heap-stat "Direct link to FE JVM Heap Stat")

This item corresponds to **Cluster FE JVM Heap Stat** under **Overview**, monitoring the proportion of FE's JVM memory usage to FE heap memory limit.

1. In the **Set an alert rule name** section, configure the name as "\[PROD] FE JVM Heap Stat".
2. In the **Set a query and alert condition** section, configure the rule in section **C** as `B IS ABOVE 80`. You can use default values for other items.
3. In the **Alert evaluation behavior** section, choose the **PROD** directory and Evaluation group **01** created earlier, and set the duration to 30 seconds.
4. In the **Add details for your alert rule** section, click **Add annotation**, select **Description**, and input the alert content, for example, "Detected that FE heap memory usage is high, please adjust the heap memory limit in the FE configuration file **fe.conf**".
5. In the **Notifications** section, configure **Labels** the same as the FE alert rule. If Labels are not configured, Grafana will use the Default policy and send alert emails to the "StarRocksOp" alert channel.

## Appendix[​](#appendix "Direct link to Appendix")

### Enable Service Detection for Prometheus[​](#enable-service-detection-for-prometheus "Direct link to Enable Service Detection for Prometheus")

You can enable Service Detection for Prometheus so that it can automatically detect the services (nodes) after the cluster is scaled in or out.

note

The following section uses AWS as an example.

1. Grant the EC2 instance that hosts your Prometheus service the following permissions using IAM Policy:

   ```json
   {
         "Version": "2012-10-17",
         "Statement": [
                  {
                           "Effect": "Allow",
                           "Action": [
                                 "ec2:DescribeInstances",
                                 "ec2:DescribeTags"
                           ],
                           "Resource": "*"
                  }
         ]
   }

   ```

   For detailed instructions of authentication to AWS resources, see [Authenticate to AWS resources](https://docs.starrocks.io/docs/integrations/csp_auth/authenticate_to_aws_resources.md).

   With these permissions, Prometheus is able to list the instances and their tags in the region.

2. Add the `ec2_sd_configs` and `relabel_configs` sections to **prometheus/prometheus.yml**.

   Example:

   ```yaml
   global:
   scrape_interval: 15s # Set the global scrape interval to 15s. The default is 1 min.
   evaluation_interval: 15s # Set the global rule evaluation interval to 15s. The default is 1 min.
   scrape_configs:
   - job_name: 'StarRocks_Cluster01'
      metrics_path: '/metrics'
      ec2_sd_configs:
         - region: us-west-2
         port: 8030
         filters:
            - name: tag:ClusterName
               values: ['test-stage-20251021']
            - name: tag:ProcessType
               values: ['FE']
         - region: us-west-2
         port: 8040
         filters:
            - name: tag:ClusterName
               values: ['test-stage-20251021']
            - name: tag:ProcessType
               values: ['BE']
      relabel_configs:
         - source_labels: [__meta_ec2_tag_ClusterName]
         regex: test-stage-20251021
         target_label: cluster
         replacement: test-stage-20251021
         - source_labels: [__meta_ec2_tag_ProcessType]
         regex: FE
         target_label: group
         replacement: fe
         - source_labels: [__meta_ec2_tag_ProcessType]
         regex: BE
         target_label: group
         replacement: be

   ```

## Q\&A[​](#qa "Direct link to Q\&A")

### Q: Why can't the Dashboard detect anomalies?[​](#q-why-cant-the-dashboard-detect-anomalies "Direct link to Q: Why can't the Dashboard detect anomalies?")

A: Grafana Dashboard relies on the system time of the server it is hosted on to fetch values for monitoring items. If the Grafana Dashboard page remains unchanged after a cluster anomaly, you can check if the servers' system clock is synchronized and then perform cluster time calibration.

### Q: How can I implement alert grading?[​](#q-how-can-i-implement-alert-grading "Direct link to Q: How can I implement alert grading?")

A: Taking the Query Error item as an example, you can create two alert rules for it with different alert thresholds. For example:

* **Risk Level**: Set the failure rate greater than 0.05, indicating a risk. Send the alert to the development team.
* **Severity Level**: Set the failure rate greater than 0.20, indicating a severity. At this point, the alert notification will be sent to both the development and operations teams simultaneously.

### Q: How can I retrieve more detailed metrics, including table-level metrics, materialized view metrics, and connection statistics with user labels?[​](#q-how-can-i-retrieve-more-detailed-metrics-including-table-level-metrics-materialized-view-metrics-and-connection-statistics-with-user-labels "Direct link to Q: How can I retrieve more detailed metrics, including table-level metrics, materialized view metrics, and connection statistics with user labels?")

A: By default, the `/metrics` endpoint collects metrics in a minified mode to minimize performance impact. To retrieve detailed metrics, you need to add specific parameters to the request and provide Basic Authentication credentials for a user with ADMIN privileges.

**Supported Parameters:**

* `with_table_metrics=all`: Collects all table-level metrics.
* `with_materialized_view_metrics=all`: Collects all materialized view metrics.
* `with_user_connections=all`: Collects connection statistics categorized by user labels.

**Authentication Requirement:**

These parameters take effect only when the request includes valid Basic Authentication credentials for an ADMIN user. If the request is anonymous or the user lacks ADMIN privileges, these parameters are ignored, and only default metrics are returned.

**Example Curl Command:**

```bash
curl -u <admin_username>:<admin_password> \
"http://<fe_host>:<fe_http_port>/metrics?with_table_metrics=all&with_materialized_view_metrics=all&with_user_connections=all"

```

**Prometheus Configuration Example:**

To enable detailed metric collection in Prometheus, configure `params` and `basic_auth` in your `prometheus.yml`:

```yaml
scrape_configs:
  - job_name: 'StarRocks_Detailed_Metrics'
    metrics_path: '/metrics'
    params:
      with_table_metrics: ['all']
      with_materialized_view_metrics: ['all']
      with_user_connections: ['all']
    basic_auth:
      username: '<admin_username>'
      password: '<admin_password>'
    static_configs:
      - targets: ['<fe_host>:<fe_http_port>']

```

note

Collecting all table and materialized view metrics may increase the load on the FE node. Use these parameters with caution in large-scale environments.
