> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ocient.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Datadog Integration

> Monitor Ocient cluster health and query performance by collecting metrics with the official Datadog Ocient integration.

export const Prometheus = "Prometheus®";

export const Datadog = "Datadog®";

export const Ocient = "Ocient®";

export const OcientDataIntelligencePlatform = "OcientAIQ™ Unified Data Platform";

The {OcientDataIntelligencePlatform} provides an official {Datadog} integration that collects performance metrics from the {Prometheus}-compatible [`/metrics` endpoint](/prometheus-metrics-endpoint) and forwards them to the Datadog cloud platform. Use this integration to monitor cluster health, query performance, ingestion throughput, and storage capacity alongside your other infrastructure in Datadog.

For installation and general configuration instructions, see the [Datadog Ocient Integration](https://docs.datadoghq.com/integrations/ocient/) page. This page covers {Ocient}-specific configuration guidance that supplements the Datadog documentation.

## Prerequisites

You must have:

* Datadog Agent version 7 or later installed on a host that can reach the Ocient cluster.
* Network access from the Datadog Agent host to port 9090 on each Ocient node that you want to monitor.
* The Ocient integration installed on the Agent (`agent integration install -t datadog-ocient==1.0.0`).

## Filter Metrics by Level

The Ocient `/metrics` endpoint supports a `level` query parameter that controls which metrics the endpoint returns. Append `?level=info` to the `openmetrics_endpoint` URL to limit collection to production-level metrics, reducing cardinality and Datadog cost. The supported levels are `info` (production metrics), `debug` (all metrics including internal diagnostics), and `default` (baseline set).

This configuration collects only `info`-level metrics from a single SQL Node.

```yaml conf.d/ocient.d/conf.yaml theme={null}
instances:
  - use_openmetrics: true
    openmetrics_endpoint: http://oc1-sql0:9090/metrics?level=info
```

## Monitor Multiple Nodes

The integration scrapes one endpoint for each instance. To monitor an entire cluster, add one instance per node in the configuration file. Each Ocient node exposes its own metrics independently. SQL Nodes expose query lifecycle metrics (`cmdcomp_*`), while Foundation Nodes expose storage and ingestion metrics (e.g., `local_storage_service_*`, `stream_loader_*`).

This configuration monitors one SQL Node `oc1-sql0` and two Foundation Nodes, `oc1-lts0` and `oc1-lts1`, in the same cluster, collecting `info`-level metrics from each.

```yaml conf.d/ocient.d/conf.yaml theme={null}
instances:
  - use_openmetrics: true
    openmetrics_endpoint: http://oc1-sql0:9090/metrics?level=info
  - use_openmetrics: true
    openmetrics_endpoint: http://oc1-lts0:9090/metrics?level=info
  - use_openmetrics: true
    openmetrics_endpoint: http://oc1-lts1:9090/metrics?level=info
```

## Name Metrics in Datadog

Metrics appear in Datadog with the prefix `ocient.` followed by the metric name from the `/metrics` endpoint. For example, the endpoint metric `cmdcomp_queries` becomes `ocient.cmdcomp_queries` in Datadog.

Datadog removes the `_total` suffix for counter metrics and replaces it with a `.count` suffix for submission. For example, `stream_loader_api_push_rows_requests_total` becomes `ocient.stream_loader_api_push_rows_requests.count`.

For a complete reference of available metrics, their types, labels, and descriptions, see the [Prometheus Metrics Endpoint — Key Metrics Reference](/prometheus-metrics-endpoint#key-metrics-reference).

## Service Checks

The integration reports a service check named `ocient.openmetrics.health` that indicates the reachability of the configured `openmetrics_endpoint` endpoint. The `OK` value indicates successful scraping, whereas the `CRITICAL` value indicates the endpoint is unreachable.

## Troubleshooting

Troubleshoot common error messages when you work in Datadog:

* `Agent cannot reach the endpoint.` — Verify that the host running the Agent can connect to port 9090 on the target Ocient node. The `/metrics` endpoint does not require authentication. Execute this `curl` command on the Agent host to confirm that the endpoint is reachable and returns metric data.

  ```bash theme={null}
  curl -s http://oc1-sql0:9090/metrics | head -5
  ```

  If the command returns metric lines starting with `#` (comment headers) or metric name-value pairs, the endpoint is accessible. An empty response or connection error indicates a network or firewall issue.

* `No metrics appear in Datadog.` — Execute the `agent check ocient` command and inspect the output for errors. Common issues include incorrect endpoint URLs, typos in the hostname or port, and firewall rules blocking port 9090.

* `High metric cardinality.` — Use the `level=info` query parameter to limit collection to production-level metrics. Alternatively, use the `exclude_metrics` option in the `conf.d/ocient.d/conf.yaml` file to drop high-cardinality metric families, such as `jemalloc_*`.

* `Metric names differ from the endpoint.` — Datadog applies the `ocient.` namespace prefix and converts counter suffixes. For details, see [Name Metrics in Datadog](#name-metrics-in-datadog).

## Related Links

* [Datadog Ocient Integration](https://docs.datadoghq.com/integrations/ocient/)
* [Prometheus Metrics Endpoint](/prometheus-metrics-endpoint)
* [Statistics Monitoring](/statistics-monitoring)
* [System Information REST Endpoints](/system-information-rest-endpoints)
