Insights Metrics and Checks

Insights Advanced Monitoring shows metrics and checks (status information) on the various default dashboards provided by SAP. These metrics and checks provide you with in-depth information (statistics and states) collected from the event broker services in your estate. You can also use these metrics to build custom visualizations or customize existing dashboards. The information from the metrics and checks is useful to not only manage all aspects of your event broker services, but give you better insight to manage your event-driven architecture (EDA).

Customers using Disaster Recovery (DR) should be aware that Insights does not collect certain metrics for the backup site, including Queues, Client Username, Topic Endpoints, and REST Delivery Point (RDP) metrics. Collecting these metrics for both the primary and backup sites would cause Insights to generate duplicate alerts. Insights continues to collect other metrics from the back up site to ensure the health of your DR configuration, including Queue metrics for backup sites with ACK disabled.

Here's a brief summary of the metrics and status information available:

Service Checks: These metrics characterize the status of a service for monitoring within Datadog. For more information about service checks, see the Datadog documentation. These are the service checks that are available with Insights:

Collected and Derived Metrics: There are two types of metrics as follows:

  • Collected — These are basic metrics collected directly from the event broker services. They typically represent measurements, statistics, and operational numbers provided by the event broker services.
  • Derived — These metrics are calculated or determined by some means. The metric is usually a combination of information sources that are used to calculate a value or status.

The collected and derived metrics are organized into the following alphabetical groupings:

  • Insights Metrics: A-B — Appliance Environmental Sensor Statistics, Bridge Statistics, Broker Resource Utilization Statistics
  • Insights Metrics: C — Cache Instance Statistics, Certificate Expiry Monitoring Statistics, Client Statistics, Client Username Statistics, Config-Sync Statistics
  • Insights Metrics: D — Disk Detailed Statistics, Disk Statistics, DMR Cluster Statistics, DMR Link Statistics, DNS Statistics
  • Insights Metrics: H-L — Hardware HBA Link Statistics, Health Statistics, Interface Statistics, Kafka Bridge Statistics, Limit Usage
  • Insights Metrics: Ma-Mem (Message Spool) — Maximum Guaranteed Message Size Statistics, Memory Statistics, Message Spool Statistics
  • Insights Metrics: Message VPN — Message VPN Detailed Statistics, Message VPN Message Spool Statistics, Message VPN Replication Statistics, Multi-Node Routing Statistics
  • Insights Metrics: P-R — Partitioned Queue Statistics, Queue Rate Statistics, Queue Statistics, Redundancy Statistics, Replication Statistics, REST Delivery Point Statistics
  • Insights Metrics: S-T — Server Certificate Statistics, Storage Statistics, Subscription Statistics, Time Server Statistics, Topic Endpoint Statistics

Filtering Metrics

Each Insights metric has a variety of tags associated with it. The tags provide information about where the metric's data was sourced. For example, the env tag allows you to determine the environment the data in the metric was derived from. You can filter the Insights metrics list using the tags to get data sourced from a specific object, for example, from a specific environment or event broker service.

To filter the metrics list, follow these steps:

  1. In Datadog, go to MonitorsSummary to access the list of metrics.

  2. In the Tag field, either:

    • select a tag from the list of options

    • enter the tag you want to filter the list by. For example, enter env:MyFirstEnvironment to filter the metric list to show only metrics for the environment named MyFirstEnvironment.

Service Checks

These metrics characterize the status of a service for monitoring within Datadog. For more information about service checks, see the Datadog documentation.

Service Check Name Polling
Frequency
in Seconds
Description

Bridge

bridge.status

10

Whether the inbound and outbound Message VPN bridges are healthy and the associated queues are bound.

Cache Instance

cache_instance.status

10

Whether the cache instance is operational.

Config-Sync

config_sync.up_or_disabled

60

The Config-Sync status. A valid status is Shutdown or Up.

Disk

disk.status

10

The internal disk status. A valid status is disabled or up.

DNS

system.dns_status

60

The Domain Name System (DNS) reachability status.

DMR

dmr.cluster.status

10

Whether the Dynamic Message Routing (DMR) cluster is operational.

DMR link

dmr.link.status

10

Whether the DMR links and redundancy are operational.

Interface

interface.status

30

The network-interface status. A valid status is disabled or up.

Queue

queue.unbound_and_enabled_with_messages

30

Whether an enabled queue has messages, but no bound clients. If no clients are consuming from the enabled queue, the queue could fill.

Redundancy

derived_metrics.redundancy.up_or_disabled

10

The redundancy status. A valid status is Shutdown or Up.

derived_metrics.redundancy.adb_links_status

10

Whether the ADB Hello is Up.

derived_metrics.redundancy.is_service_active

10

Whether one event broker in the High-Availability (HA) group is active.

derived_metrics.message_spool_service_status

10

Whether the message spool disk on the active event broker is in the AD_ACTIVE state.

derived_metrics.message_spool_standby_status

10

Whether the message spool disk on the inactive node is in the AD_Standby state.

System POST

system.post_status

60

Whether the Power-On Self-Test (POST) is successful.

Time Server

ntp.in_sync

60

Whether the event broker clock is synchronized with the configured time server.

Topic Endpoint

topic_endpoint.unbound_and_enabled_with_messages

30

Whether an enabled topic endpoint has messages, but no bound clients. If no clients are consuming from the enabled topic endpoint, the topic endpoint could fill.

VPN Replication

vpn.replication_sync_eligible

10

Whether the replication service entered a degraded state.

Collected and Derived Metrics

The list of supported metrics and details that can be summarized as collected and derived metrics.

  • Collected: These are basic metrics collected directly from the event broker services. They typically represent measurements, statistics, and operational numbers provided by the event broker services.
  • Derived: These metrics are calculated or determined by some means. The metric is usually a combination of information sources that are used to calculate a value or status.

Common Metric Tags

These common tags are associated with every metric listed in the table below and are useful for filtering and aggregating metrics:

  • ha_role: The identity of the software event broker in a high-availability (HA) event broker service for which the metric is measuring. The value can backup, monitoring, or primary.

  • service_id: The unique identifier of the event broker service, such as s1h0z84b1ar.

  • service_name: The name of the event broker service, such as My-First-Service.

  • service_type: The type of event broker service. This can be developer or enterprise.

  • service_class: The service class the metric, such as developer or enterprise-kilo.

Depending on the metric, these tags are also useful for aggregating and filtering the metrics:

  • vpn_name: The name of the Message VPN.

  • remote_name: The name of the remote Message VPN name in a bridge.

  • topic_endpoint: The name of the topic endpoint.

  • bridge_name: The name of the bridge.

  • cache_cluster_name: The name of the cache cluster.

  • cache_instance_name: The name of the cache instance.

  • cache_name: The name of the cache.

  • queue_name:The name of the queue, such as my-first-queue.

  • rdp_name: The name of the REST Delivery Point.

For detailed metric reference information, see the following pages:

  • Insights Metrics: A-B — Appliance Environmental Sensor Statistics, Bridge Statistics, Broker Resource Utilization Statistics

  • Insights Metrics: C — Cache Instance Statistics, Certificate Expiry Monitoring Statistics, Client Statistics, Client Username Statistics, Config-Sync Statistics

  • Insights Metrics: D — Disk Detailed Statistics, Disk Statistics, DMR Cluster Statistics, DMR Link Statistics, DNS Statistics

  • Insights Metrics: H-L — Hardware HBA Link Statistics, Health Statistics, Interface Statistics, Kafka Bridge Statistics, Limit Usage

  • Insights Metrics: Ma-Mem (Message Spool) — Maximum Guaranteed Message Size Statistics, Memory Statistics, Message Spool Statistics

  • Insights Metrics: Message VPN — Message VPN Detailed Statistics, Message VPN Message Spool Statistics, Message VPN Replication Statistics, Multi-Node Routing Statistics

  • Insights Metrics: P-R — Partitioned Queue Statistics, Queue Rate Statistics, Queue Statistics, Redundancy Statistics, Replication Statistics, REST Delivery Point Statistics

  • Insights Metrics: S-T — Server Certificate Statistics, Storage Statistics, Subscription Statistics, Time Server Statistics, Topic Endpoint Statistics