Skip to main content

Custom Workload Metrics

A workload can report its own application metrics, such as a count of robot falls or a battery level, and Admiral charts them next to the system telemetry. The workload writes StatsD-style lines to a local unix socket. The device agent aggregates them and sends them to the same telemetry pipeline as CPU, memory and disk metrics, with the same device and fleet labels.


1. Enable it per fleet​

Custom metrics are off by default and are not billed.

  1. Open the fleet in the dashboard and go to Settings > Custom metrics.
  2. Turn on the switch on the Custom metrics card.

Devices open the metrics socket as soon as they receive the change. The socket directory is mounted into the workload container when the container is next created, so restart or redeploy the workload (or let the next configuration rollout recreate it) before it can write. Turning the setting off closes the socket on the device; the mount is removed the next time the container is recreated.

Changing the setting needs permission to update the fleet.


2. The metrics socket​

ItemValue
Socket path in the container/run/admrl/metrics/metrics.sock
Environment variableADMRL_METRICS_SOCKET (set to the path above)
Socket typeunix datagram (SOCK_DGRAM)

Use the environment variable rather than hard-coding the path. The socket is writable by any user in the container, so the workload does not need to run as root. If the file does not exist, the setting is off for the fleet or the container has not been recreated since it was enabled; treat that as "metrics unavailable" and carry on.


3. Line format​

Each datagram holds one or more lines separated by \n:

name:value|type
name:value|type|#tag1:value1,tag2:value2
TypeMeaning
cCounter. Each line adds value to a running total.
gGauge. The last value written wins.

Examples:

robot.fall:1|c
robot.fall:1|c|#site:warehouse3,model:r2
battery.level:87.5|g|#slot:0

Malformed lines are dropped.

Reserved labels​

Tags become labels on the series. These tag names are dropped because they are used by the platform to keep data separated per organisation, fleet and device: org_id, device_id, fleet_id, device, job, instance, __name__, and anything starting with __.

Limits​

LimitValue
Datagram size8 KiB
Distinct series per device500
Tags per line8
Metric name length128 characters
Tag value length128 characters

Lines or new series over a limit are dropped. The agent counts them in edge_custom_dropped_total{reason}. Keep tag values low-cardinality (a site name, not a request id), since every distinct tag combination is a separate series.


4. Naming in charts​

The agent turns the name into a Prometheus series name: dots and dashes become _, and the result is prefixed with edge_custom_.

You writeSeriesType
robot.fall:1|cedge_custom_robot_fall_totalcounter
battery.level:87|gedge_custom_battery_levelgauge

Counters get a _total suffix. The dashboard lists the series that reported in the last day and shows them with a best-effort name (robot.fall). Because dots and underscores both map to _, the dashboard cannot always recover the original spelling.

Charts are in the Custom tab of the fleet's and each device's Telemetry page. On a fleet, counters are charted as the sum of the increase per interval across all devices, and gauges as a median with a p5 to p95 band and the worst device. Both also get a Top devices chart. On a device, each metric gets a plain chart.

To query series yourself, use the Query tab on the telemetry page, for example sum(increase(edge_custom_robot_fall_total[1h])).


5. Examples​

Python​

import os
import socket

SOCK = os.environ.get("ADMRL_METRICS_SOCKET", "/run/admrl/metrics/metrics.sock")
_sock = socket.socket(socket.AF_UNIX, socket.SOCK_DGRAM)
_sock.setblocking(False)


def emit(line: str) -> None:
try:
_sock.sendto(line.encode(), SOCK)
except OSError:
pass # metrics disabled or the agent is not listening; never break the app


# A robot fall, tagged with the site it happened at
emit("robot.fall:1|c|#site:warehouse3")
emit("battery.level:87.5|g")

Go​

package main

import (
"net"
"os"
)

func main() {
path := os.Getenv("ADMRL_METRICS_SOCKET")
if path == "" {
path = "/run/admrl/metrics/metrics.sock"
}
conn, err := net.Dial("unixgram", path)
if err != nil {
return // metrics disabled; carry on
}
defer conn.Close()

conn.Write([]byte("robot.fall:1|c|#site:warehouse3\nbattery.level:87.5|g"))
}

Shell​

With socat:

printf 'robot.fall:1|c|#site:warehouse3' | socat - UNIX-SENDTO:"$ADMRL_METRICS_SOCKET"

With a netcat that supports unix datagram sockets (nc -uU):

printf 'robot.fall:1|c' | nc -uU -w1 "$ADMRL_METRICS_SOCKET"

6. Worked example: counting robot falls​

  1. Enable custom metrics on the fleet and recreate the workload.
  2. In the robot software, call emit("robot.fall:1|c|#site:warehouse3") each time a fall is detected.
  3. Open the fleet's Custom metrics page and pick robot.fall (counter). The first chart shows falls per interval across the fleet, and Top devices shows which robots fall most.

Troubleshooting​

  • The socket does not exist in the container. Check that the setting is on and that the container was recreated after enabling it.
  • Nothing appears in the dashboard. Allow a minute or two for the first values, and make sure the time range covers the writes. A datagram over 8 KiB or a malformed line is dropped without an error to the workload.
  • Some series are missing. The device may be over its 500-series limit; check edge_custom_dropped_total in the Query tab.