---
title: Monitoring the Network Observability Operator
---

# Monitoring the Network Observability Operator {#network-observability-operator-monitoring}

Use the OpenShift Container Platform web console to monitor alerts related to the Network Observability Operator’s health. This helps you maintain system stability and quickly detect operational issues.

## Health dashboards {#network-observability-health-dashboard-overview_network_observability}

View the Network Observability Operator health dashboards in the OpenShift Container Platform web console to monitor the health status, resource usage, and internal statistics of the operator and its components.

Metrics are located in the **Observe** → **Dashboards** page in the OpenShift Container Platform web console. You can view metrics about the health of the Network Observability Operator in the following categories:

- **Flows per second**
- **Sampling**
- **Errors last minute**
- **Dropped flows per second**
- **Flowlogs-pipeline statistics**
- **Flowlogs-pipleine statistics views**
- **eBPF agent statistics views**
- **Operator statistics**
- **Resource usage**

## Health alerts {#network-observability-health-alert-overview_network_observability}

Understand the health alerts generated by the Network Observability Operator, which trigger banners when conditions like Loki ingestion errors, zero flow ingestion, or dropped eBPF flows occur.

A health alert banner that directs you to the dashboard can appear on the **Network Traffic** and **Home** pages if an alert is triggered. Alerts are generated in the following cases:

- The `NetObservLokiError` alert occurs if the `flowlogs-pipeline` workload is dropping flows because of Loki errors, such as if the Loki ingestion rate limit has been reached.
- The `NetObservNoFlows` alert occurs if no flows are ingested for a certain amount of time.
- The `NetObservFlowsDropped` alert occurs if the Network Observability eBPF agent hashmap table is full, and the eBPF agent processes flows with degraded performance, or when the capacity limiter is triggered.

## Viewing health information {#network-observability-dashboard-view_network_observability}

View the **Netobserv/Health** dashboard within the OpenShift Container Platform web console to monitor the health status and resource usage of the Network Observability Operator and its components.

**Prerequisites**

- You have the Network Observability Operator installed.
- You have access to the cluster as a user with the `cluster-admin` role or with view permissions for all projects.

**Procedure**

1. From the **Administrator** perspective in the web console, navigate to **Observe** → **Dashboards**.
2. From the **Dashboards** dropdown, select **Netobserv/Health**.
3. View the metrics about the health of the Operator that are displayed on the page.

### Disabling health alerts {#network-observability-disable-alerts_network_observability}

Disable specific health alerts, such as `NetObservLokiError` or `NetObservNoFlows`, by editing the `FlowCollector` resource and using the `spec.processor.metrics.disableAlerts` specification.

**Procedure**

1. In the web console, navigate to **Ecosystem** → **Installed Operators**.
2. Under the **Provided APIs** heading for the **NetObserv Operator**, select **Flow Collector**.
3. Select **cluster** then select the **YAML** tab.
4. Add `spec.processor.metrics.disableAlerts` to disable health alerts, as in the following YAML sample:

   ```yaml
   apiVersion: flows.netobserv.io/v1beta2
   kind: FlowCollector
   metadata:
     name: cluster
   spec:
     processor:
       metrics:
         disableAlerts: [NetObservLokiError, NetObservNoFlows]
   ```

   where:

   `spec.processor.metrics.disableAlerts`
   :   Specifies one or more types of alerts to disable.

## Creating Loki rate limit alerts for the NetObserv dashboard {#network-observability-netobserv-dashboard-rate-limit-alerts_network_observability}

Create a custom `AlertingRule` resource based on Loki metrics to monitor for and trigger alerts when the Loki ingestion rate limits are reached, indicated by HTTP 429 errors.

You can create custom alerting rules for the **Netobserv** dashboard metrics to trigger alerts when Loki rate limits have been reached.

**Prerequisites**

- You have access to the cluster as a user with the cluster-admin role or with view permissions for all projects.
- You have the Network Observability Operator installed.

**Procedure**

1. Create a YAML file by clicking the import icon, **+**.
2. Add an alerting rule configuration to the YAML file. In the YAML sample that follows, an alert is created for when Loki rate limits have been reached:

   ```yaml
   apiVersion: monitoring.openshift.io/v1
   kind: AlertingRule
   metadata:
     name: loki-alerts
     namespace: openshift-monitoring
   spec:
     groups:
     - name: LokiRateLimitAlerts
       rules:
       - alert: LokiTenantRateLimit
         annotations:
           message: |-
             {{ $labels.job }} {{ $labels.route }} is experiencing 429 errors.
           summary: "At any number of requests are responded with the rate limit error code."
         expr: sum(irate(loki_request_duration_seconds_count{status_code="429"}[1m])) by (job, namespace, route) / sum(irate(loki_request_duration_seconds_count[1m])) by (job, namespace, route) * 100 > 0
         for: 10s
         labels:
           severity: warning
   ```
3. Click **Create** to apply the configuration file to the cluster.

## Using the eBPF agent alert {#network-observability-netobserv-dashboard-ebpf-agent-alerts_network_observability}

Resolve the `NetObservAgentFlowsDropped` alert, which occurs when the eBPF agent hashmap is full, by increasing the `spec.agent.ebpf.cacheMaxFlows` value in the `FlowCollector` custom resource.

An alert, `NetObservAgentFlowsDropped`, is also triggered when the capacity limiter is triggered. If you see this alert, consider increasing the `cacheMaxFlows` in the `FlowCollector`, as shown in the following example.

> [!NOTE]
> Increasing the `cacheMaxFlows` might increase the memory usage of the eBPF agent.

**Procedure**

1. In the web console, navigate to **Ecosystem** → **Installed Operators**.
2. Under the **Provided APIs** heading for the **Network Observability Operator**, select **Flow Collector**.
3. Select **cluster**, and then select the **YAML** tab.
4. Increase the `spec.agent.ebpf.cacheMaxFlows` value, as shown in the following YAML sample:

   ```yaml
   apiVersion: flows.netobserv.io/v1beta2
   kind: FlowCollector
   metadata:
     name: cluster
   spec:
     namespace: netobserv
     deploymentModel: Service
     agent:
       type: eBPF
       ebpf:
         cacheMaxFlows: 200000
   ```

   where:

   `spec.agent.ebpf.cacheMaxFlows`
   :   Specifies the maximum number of flows to cache. If a `NetObservAgentFlowsDropped` alert occurs, increase this value from its current level.

**Additional resources**
{._additional-resources}

- [Creating alerting rules for user-defined projects](https://docs.redhat.com/en/documentation/monitoring_stack_for_red_hat_openshift/latest/html/managing_alerts/managing-alerts-as-a-developer#creating-alerting-rules-for-user-defined-projects_managing-alerts-as-a-developer)
