> For the complete documentation index, see [llms.txt](https://docs.roadrunner.dev/docs/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.roadrunner.dev/docs/logging-and-observability/metrics.md).

# Metrics

RoadRunner offers metrics plugin, which provides an embedded metrics server based on [Prometheus](https://prometheus.io/). The metrics plugin allows users to monitor the performance and health of their applications and servers by collecting and displaying various metrics.

## Observing Application Server Metrics

To observe application server metrics, you can enable the metrics plugin by adding a `metrics` section to your configuration file.

**Here is an example:**

{% code title=".rr.yaml" %}

```yaml
version: "3"

metrics:
  address: 127.0.0.1:2112
```

{% endcode %}

After enabling the metrics plugin, you can access the Prometheus metrics by visiting the <http://127.0.0.1:2112/metrics> URL. The metrics plugin provides several general metrics that apply to all plugins and specific metrics for each supported plugin.

### General Metrics

The general metrics provided by the metrics plugin include:

* `{{plugin}}_total_workers` - Total number of workers used by the plugin.
* `{{plugin}}_worker_memory_bytes` - Single worker's memory usage.
* `{{plugin}}_workers_memory_bytes` - Total worker's memory usage.
* `{{plugin}}_worker_state` - Worker's current state.
* `{{plugin}}_workers_ready` - Number of workers currently in ready state.
* `{{plugin}}_workers_working` - Number of workers currently in working state.
* `{{plugin}}_workers_invalid` - Number of workers currently in invalid, killing, destroyed, errored, inactive states.

{% hint style="info" %}
`{{plugin}}` is the concatenation of the rr + plugin name. For example, for the `jobs` plugin, the metric would be `rr_jobs_total_workers`.
{% endhint %}

Per-worker memory metrics have a `pid` label. Worker-state metrics have `pid` and `state` labels. Pool totals have no plugin-defined labels.

### HTTP Metrics

{% hint style="info" %}
To enable specific HTTP metrics, you need to add an HTTP middleware called `http_metrics` to the `http.middleware` section of your configuration file.

**Here is an example:**

{% code title=".rr.yaml" %}

```yaml
http:
  middleware: [ "http_metrics" ]
```

{% endcode %}
{% endhint %}

The HTTP metrics provided by the metrics plugin include:

* `rr_http_request_total` - Total number of handled HTTP requests after server restart.
* `rr_http_request_duration_seconds` - HTTP request duration.
* `rr_http_requests_queue` - Requests currently inside the metrics middleware, including requests waiting for a worker and requests being processed.
* `rr_http_uptime_seconds` - Plugin uptime in seconds.
* `rr_http_no_free_workers_total` - Total number of NoFreeWorkers errors.

The request counter and duration histogram have a `status` label. A response body written without an explicit status records `status="200"`. The duration histogram has finite buckets from `0.005` through `60` seconds, including `20`, `30`, and `60` for slow requests.

The duration histogram includes downstream request execution. The `http_metrics` and `http_metrics:post` spans described in [OpenTelemetry](/docs/logging-and-observability/otel.md) measure work before and after the downstream handler. Their duration does not change histogram duration.

### gRPC Metrics

The gRPC metrics provided by the metrics plugin include:

* `rr_grpc_requests_queue` - Requests currently inside the gRPC interceptor, including requests waiting for a worker and requests being processed.
* `rr_grpc_request_total` - Total number of handled gRPC requests after server restart.
* `rr_grpc_request_duration_seconds` - gRPC request duration.

The request counter has `grpc_method` and `status_code` labels. The duration histogram has a `grpc_method` label. The status value is the gRPC code name, such as `OK` or `Internal`.

### Redis Metrics

The Redis metrics provided by the redis plugin include:

* `rr_redis_pool_conn_idle_current` - Current number of idle connections in the pool.
* `rr_redis_pool_conn_stale_total` - Number of times a connection was removed from the pool because it was stale.
* `rr_redis_pool_conn_total_current` - Current number of connections in the pool.
* `rr_redis_pool_hit_total` - Number of times a connection was found in the pool.
* `rr_redis_pool_miss_total` - Number of times a connection was not found in the pool.
* `rr_redis_pool_timeout_total` - Number of times a timeout occurred when looking for a connection in the pool.

### JOBS Metrics

The JOBS metrics provided by the metrics plugin include:

* `rr_jobs_jobs_err` - Number of jobs that failed while processing in the worker.
* `rr_jobs_jobs_ok` - Number of successfully processed jobs.
* `rr_jobs_jobs_requeue` - Number of jobs requeued by worker responses.
* `rr_jobs_push_ok` - Number of successful job pushes.
* `rr_jobs_push_err` - Number of jobs that failed to push.
* `rr_jobs_requests_total` - Job push attempts for an existing pipeline.
* `rr_jobs_push_latency_bucket` - Histogram buckets for successful job push duration, in seconds.
* `rr_jobs_push_latency_sum` - Total duration of successful job pushes, in seconds.
* `rr_jobs_push_latency_count` - Number of successful job pushes observed by the histogram.

The push request counter and latency histogram have `driver`, `job`, and `source` labels. `job` is the pipeline name. `source` is `single` or `batch`. Each job in a batch is counted separately. The outcome counters have no plugin-defined labels.

The v6 Jobs plugin exports `rr_jobs_jobs_ok`, `rr_jobs_jobs_err`, `rr_jobs_push_ok`, and `rr_jobs_push_err` as counters instead of gauges. Their names do not change. Use `rate()` or `increase()` for dashboards that measure changes over time.

Failed and requeued worker responses no longer increment `rr_jobs_jobs_ok`. They increment `rr_jobs_jobs_err` or the new `rr_jobs_jobs_requeue` counter instead. Update success-rate queries to keep these outcomes separate. The job outcome counters measure processing attempts, not unique jobs.

For example, compare processing rates with these PromQL queries:

```promql
rate(rr_jobs_jobs_ok[5m])
rate(rr_jobs_jobs_err[5m])
rate(rr_jobs_jobs_requeue[5m])
```

### Temporal Metrics

In temporal each SDK, has its own metrics - RoadRunner retransmits Go SDK metrics to the metrics storage. For example, we can observe a workflow failing or a nondeterminism state.

Full list of metrics we can look at here [Sdk Metrics](https://docs.temporal.io/references/sdk-metrics)

```yaml
version: "3"

temporal:
  address: 127.0.0.1:7233
  namespace: default

  # Temporal metrics
  #
  # Optional section
  metrics:

    # Metrics driver to use
    # Optional, default: prometheus. Available values: prometheus, statsd
    driver: prometheus

    # ---- Prometheus
    prometheus:
      # Server metrics address
      # Required for the production. Default: 127.0.0.1:9091, for the metrics 127.0.0.1:9091/metrics
      address: 127.0.0.1:9091
      # Metrics type
      type: "summary"
      # Temporal metrics prefix
      # Default: (empty)
      prefix: "foobar"
```

## Application Metrics

The RoadRunner metrics plugin also allows you to publish application metrics to the server via RPC and collect them in Prometheus. To do this, you need to register collectors in your configuration file.

**Here is an example:**

{% code title=".rr.yaml" %}

```yaml
version: "3"

rpc:
  listen: tcp://127.0.0.1:6001

metrics:
  address: 127.0.0.1:2112
  collect:
    registered_users:
      type: counter
      help: "Total registered users."
```

{% endcode %}

{% hint style="info" %}
Supported types for collectors include `gauge`, `counter`, `summary`, and `histogram`.
{% endhint %}

### Tagged metrics

You can also use tagged (labels) metrics to group values:

{% code title=".rr.yaml" %}

```yaml
version: "3"

rpc:
  listen: tcp://127.0.0.1:6001

metrics:
  address: 127.0.0.1:2112
  collect:
    registered_users:
      type: counter
      help: "Total registered users."
      labels: [ "type", "is_admin" ]
```

{% endcode %}

Choose either the basic configuration or this labeled configuration for `registered_users`. Every update to the labeled collector must include two label values, in the order `type`, `is_admin`.

### PHP client

The RoadRunner metrics plugin comes with a convenient PHP package that simplifies the process of integrating the plugin with your PHP application.

#### Installation

To get started, you can install the package via Composer using the following command:

```bash
composer require spiral/roadrunner-metrics
```

#### Usage

After the installation, you can create an instance of the `Spiral\RoadRunner\Metrics\Metrics` class, which will allow you to use the available class methods.

Use the [basic configuration](#application-metrics) for this example. Its collector has no labels:

{% code title="metrics.php" %}

```php
use Spiral\RoadRunner\Metrics\Metrics;
use Spiral\Goridge\RPC\RPC;

$metrics = new Metrics(RPC::create('tcp://127.0.0.1:6001'));

$metrics->add('registered_users', 1);
```

{% endcode %}

#### Labels

Using labeled metrics allows you to attach additional metadata to your metrics, which can be useful for filtering, grouping, and aggregating the data.

**Some benefits of using labeled metrics include:**

* **Increased granularity:** You can attach multiple labels to a metric, allowing you to slice and dice the data in various ways.
* **Better organization:** Labels can help you group and organize your metrics, making it easier to find and understand the data you are looking for.
* **Simplified querying:** You can use labels to filter and aggregate your metric data, making it easier to extract meaningful insights from the data.

Use the [tagged configuration](#tagged-metrics) for this call. Pass one value for `type` and one for `is_admin`, in that order:

{% code title="metrics.php" %}

```php
// ...

$labels = ['guest', 'false'];

$metrics->add('registered_users', 1, $labels);
```

{% endcode %}

#### Declare metrics

In addition to sending metrics to the server, you can also declare metrics from your PHP application. This can be useful if you want to dynamically declare custom metrics in your application.

**Here is an example of how to declare a custom metric:**

{% code title="metrics.php" %}

```php
use Spiral\RoadRunner\Metrics\Collector;

$metrics->declare(
    'earned_money',
    Collector::counter()->withHelp('Total earned money.'),
);

$metrics->add('earned_money', 100_000_000);
```

{% endcode %}

You can also declare the labeled collector in PHP instead of YAML. For this example, remove `registered_users` from `metrics.collect` before starting RoadRunner. A declaration does not change the labels of an existing collector. Use the two-label `add()` call above after this declaration:

{% code title="metrics.php" %}

```php
$metrics->declare(
    'registered_users',
    Collector::counter()->withHelp('Total registered users counter.')
        ->withLabels('type', 'is_admin'),
);
```

{% endcode %}

## API

### RPC API

RoadRunner provides an RPC API, which allows you to manage metrics in your applications using remote procedure calls. The RPC API provides a set of methods that map to the available methods of the `Spiral\RoadRunner\Metrics\Metrics` class in PHP.

#### Add

Add a value to a declared gauge or counter. A negative counter increment returns an RPC error. Use a gauge for values that must decrease.

```go
func (r *rpc) Add(m *Metric, ok *bool) error {}
```

#### Sub

Method is used to subtract a value from a declared metric (for gauge metrics only).

```go
func (r *rpc) Sub(m *Metric, ok *bool) error {}
```

#### Set

Method is used to set the value of a metric (for gauge metrics only).

```go
func (r *rpc) Set(m *Metric, ok *bool) (err error) {}
```

#### Declare

Method is used to register a new collector in Prometheus.

```go
func (r *rpc) Declare(nc *NamedCollector, ok *bool) error {}
```

#### Unregister

Method is used to remove a collector from the Prometheus registry.

```go
func (r *rpc) Unregister(name string, ok *bool) error {}
```

#### Observe

Method is used to observe the value of a metric (for histogram and summary metrics only).

```go
func (r *rpc) Observe(m *Metric, ok *bool) error {}
```
