Memory Testing
Failing RAM rarely announces itself. It shows up as crashes with no pattern, a perception stack that returns a wrong answer once a week, files that come back corrupted, or a robot that restarts at random under load. Logs and metrics look normal because the fault is below the software.
Admiral can test a device's memory in place, from the dashboard, from the on-device console, or from the API. The result tells you whether the memory is healthy, which kind of fault it has if not, and where. The device also takes known-bad memory out of use so it can keep working until you replace the board.
Choosing a Test
Two choices decide how a test runs:
- Impact: does the workload keep running, or is it stopped for the test?
- Placement: does the device stay booted and connected (online), or does it reboot into a dedicated test boot (offline)?
| Online | Offline (test boot) | |
|---|---|---|
| Keep workload running | Live | Not possible |
| Stop workload | Full online | Test boot |
Offline with the workload running is impossible by design: a test boot reboots the device, and nothing is running after a reboot until the test is over. The dashboard greys the combination out and the API rejects it with 400 invalid.
| Mode | What it does | What it costs | What it covers |
|---|---|---|---|
| Live | Tests free memory beside the running workload. It backs off, shrinking or pausing, when the device comes under memory pressure. | No downtime. A little CPU and memory bandwidth. | Free memory only, leaving a safety margin. Memory the workload is using cannot be tested. One pass by default. |
| Live, Quick preset | A shorter, lighter Live test. | No downtime. | At most a quarter of the available memory, capped at 512 MiB, one pass, no idle-decay check. A fast sanity check, not a verdict on the whole module. |
| Full online | Stops the workload, frees caches, tests nearly all free RAM, then restarts the workload. The device stays connected and reports progress. | Workload downtime for the whole test. | Nearly all of RAM. Two passes by default, plus an idle-decay check. |
| Test boot | Reboots the device. The kernel sweeps memory before anything else loads, then the full pattern suite runs with the workload held. The device then returns to service on its own. | Reboot plus workload downtime for the whole test. | The most thorough option: memory the operating system would normally already be using is tested too. |
Pick Live to investigate while the robot is working, Full online when you can afford a maintenance stop, and Test boot when a device has been crashing and you want the most complete answer.
A Live test cannot see memory in use, so a clean Live result does not rule out a fault in the rest of RAM. If crashes continue after a clean Live test, run Full online or a test boot.
What happens during a test boot
- The device records the request and reboots. The reboot is logged as a
memory_testboot reason. - Before anything else loads, the kernel sweeps memory with a configurable number of passes (4 by default).
- The device comes up in maintenance with the workload held. It is connected, so the dashboard shows live progress.
- The full pattern suite runs, including the idle-decay check.
- The result is sent to Admiral Cloud, the device reboots again, and the workload starts normally.
A test boot is one-shot. If the device resets in the middle, it does not enter another test boot; it boots normally and reports the interrupted test.
Optionally choose Hold in maintenance if the test fails. After a failed test the device then stays in maintenance with the workload held, so a technician can inspect it, until someone cancels or 30 minutes pass; after 30 minutes with no response the hold is released and the device reboots to normal. Cancelling releases the hold and reboots to normal straight away. Without this option a failed device returns to service automatically.
Time budget
Every test has a time budget, 2 hours by default (configurable from 60 seconds to 12 hours through the API). If the budget runs out, the test stops cleanly and the result is reported as an error with the message time budget reached. The idle-decay check alone waits 5 minutes, so a full pass over a large amount of RAM can take a long time. The progress bar shows the estimated time remaining.
What the Test Looks For
The suite writes known data patterns across memory and reads them back. Different patterns expose different faults.
| Pattern | What it catches |
|---|---|
| Stuck address | Address lines that are shorted or stuck, so two locations alias to one. |
| Random data | General cell faults and data-dependent faults, with a seed that can be replayed. |
| XOR, subtract, multiply, OR, AND compare | Faults in cells, buses and the memory controller that simple patterns miss. |
| Solid bits | Cells stuck at 0 or 1. |
| Block sequential | Faults that depend on neighbouring blocks. |
| Checkerboard | Coupling between adjacent cells. |
| Bit spread | Interference between bits that are far apart. |
| Bit flip | Cells that flip when their neighbours change repeatedly. |
| Walking ones, walking zeros | A single data line that is stuck, shorted, or crosstalking. |
| Moving inversions | Pattern-sensitive and refresh-related faults, swept up and down through memory. |
| Idle decay | Cells that lose their value when left alone. The test writes memory, waits 5 minutes, then verifies. On by default for full online tests and test boots; the API can also enable it for Live (not Quick). |
The test tracks the real physical addresses it covers, so a reported fault points at a specific location rather than a program variable.
Reading the Results
Each run ends with a result, shown in the dashboard history and available from the API.
Outcome says how the run ended:
| Outcome | Meaning |
|---|---|
pass | The run finished with no confirmed errors. |
fail | At least one error was found. |
cancelled | Someone cancelled it. Results found so far are kept. |
interrupted | The device reset or crashed during the test. See below. |
error | The test could not complete, for example the time budget ran out. |
Verdict says what kind of fault the errors look like:
| Verdict | Meaning |
|---|---|
ok | No errors. |
weak_cells | Scattered single-bit errors in individual cells. Typical of ageing or marginal memory chips. |
data_line | The same bit fails across many addresses. Points at a data line, a connector, or the board, not at individual cells. |
address_line | Failures follow an address pattern. Points at an address line or the memory controller. |
thermal | Errors only appear when the device is hot. Check cooling and enclosure temperature before blaming the memory. |
unknown | Errors were found but do not fit a pattern. |
Coverage is how much distinct physical memory the run reached, against the device's total. Always read a pass together with its coverage: a Live run that covered 40% of RAM is a weaker statement than a test boot that covered nearly all of it.
Other fields in each result:
- Physical addresses of the failing locations, with the expected and actual values, the pattern that found them, whether the error reproduced, and the temperature at the time (up to 256 errors are kept per result; the full count is also reported).
- Retired pages: the memory pages taken out of use by this run (see below).
- Kernel sweep: for a test boot, the passes run and any bad ranges the kernel found at boot.
- Peak temperature during the run.
Passive memory health
Separately from tests, the device watches for memory trouble all the time and shows it in a Memory health block: whether the board supports each mode, the amount of memory the kernel has flagged as corrupted, error-correction counters on boards with ECC, the retired pages, and recent memory-error lines from the kernel. Memory trouble found this way is a reason to run a test.
Interrupted tests
If the device resets or crashes while a test is running (a power loss, a watchdog reset, a kernel panic), the next boot reports the run as interrupted, with the phase, pass, and pattern it last reached. This is not just bookkeeping. A device that reliably resets while a memory test is stressing it is telling you the memory or power path is faulty. Treat repeated interrupted runs as a strong sign of a hardware problem, and run a test boot to find out where.
An interrupted test shows as a warning in the device diagnosis.
Bad Pages and Memory Faults
When the test finds a failing location, it does not trust a single read. It re-tests that page in place with targeted patterns, three rounds (eight for pages whose errors look intermittent), and only a reproducible error is retired. One-off mismatches are counted but not acted on.
A retired page is taken out of use by the operating system. The data it held is moved, so nothing running is killed. The list of retired pages is saved on the device and re-applied on every boot before the workload starts, so a bad page does not come back after a reboot.
The device is marked with the MemoryFault condition (shown in the dashboard as a Memory fault badge, with the advice to replace the board or RAM) when any of these is true:
- More than 32 pages have been retired.
- More than 1 MiB of distinct memory is bad.
- The verdict is
data_lineoraddress_line.
A device with a memory fault keeps running: pages are still retired and the workload still starts. The fault is a clear signal to schedule a board or RAM replacement, not a shutdown. It appears as a critical issue in the diagnosis with the verdict, and as a badge on the device page.
Safety
- One test at a time per device, across the dashboard, the console, and the API. A second start is refused with
409 busyand the response includes the running test. - The test runs in a separate process. Under memory pressure the system sacrifices the test first, before the platform manager or the workload, and the test's own memory use is capped.
- The workload is always restored. For stopped modes the workload is held for the test and restarted when it ends, whether the test passes, fails, is cancelled, or its process dies. The hold does not count as a workload crash.
- Cancel at any time. Results found so far are kept. Cancelling a test boot releases the hold and reboots the device to normal.
- Rollouts wait. A device that is holding its workload for a memory test is not given a new release and is not counted as failed. It receives the rollout after the test ends.
- Rate limit. A device accepts one start every 5 seconds (
429 rate_limitedwithRetry-After).
Starting a Test
From the dashboard
Open the device and choose Observe > Diagnostics, then use the Memory test card.
- Pick Impact (Keep workload running or Stop workload) and Where (Online or Offline (reboots)). Keep workload running with Offline is disabled.
- For Live, optionally turn on Quick.
- Start the test. Stopped and offline modes ask you to confirm. A test boot also asks you to type the device name and offers Hold in maintenance if the test fails.
The card then shows the phase, pass, current pattern, coverage, rate, errors, and temperature as they happen, with a Cancel test button that asks you to confirm. When the test ends, the Results list below holds the history table: date, mode, outcome, verdict, coverage, errors, and retired pages, with a Report link per run. Each row expands to the failing addresses.
If the device runs firmware without memory testing, the card says Requires newer device firmware. If the board cannot do a test boot, the Offline option is disabled and the card shows the reason.
From the device console
On the on-device console, open Diagnostics and press M, or choose Memory test in the power menu. The console can start every mode, including a test boot, after a confirmation prompt for any mode that stops the workload or reboots. Live does not need one. Progress, errors, and the verdict show on screen, and Esc cancels after a confirmation. A person standing at the machine does not need network access to run a test.
From the API and MCP
Online modes (Live and Full online) can be started through the REST API and the MCP server. Test boots cannot.
Test Boots Are Operator-Only
A test boot reboots the device and holds its workload, so Admiral only starts one on an explicit human action:
- A signed-in person in the dashboard who has permission to reboot the device, or
- A person at the device's local console.
API tokens, service accounts, the CLI, the MCP server, schedules, update windows, fleet policy, and rollouts cannot start a test boot. A request for one from any of them is refused with 403 operator_required.
There is no fleet-wide or batch memory test. Tests are per device, started one at a time, so a mistake cannot take a whole fleet out of service.
Board Support
| Mode | Support |
|---|---|
| Live, Quick, Full online | All boards running current firmware. |
| Test boot | Boards whose bootloader supports it: currently x86 PCs and servers and NVIDIA Jetson boards, on an OS image that includes test boot support. When a board cannot run a test boot, the dashboard shows why. |
On other boards the online modes work as normal and the test boot option is disabled, with the reason shown: test boot is not supported by this board's bootloader yet. The same reason is returned by the API. See Installation for how each board boots.
A test boot also needs the device's current bootloader configuration to support it. A board that does not have it yet reports test boots as unavailable.
API
Memory tests are per device. Calls use the usual authentication and organisation header and the same errors as other calls that reach a device (503 device_offline, 504 device_timeout, 429 rate_limited). Starting and cancelling need the same permission as rebooting the device. Reading needs view access to the device.
| Method and path | Purpose |
|---|---|
POST /v1/devices/{id}/memory-test | Start a test. Returns 202 with the run. |
POST /v1/devices/{id}/memory-test/cancel | Cancel the running test. |
GET /v1/devices/{id}/memory-test | Current or last run, memory health, and recent results. |
GET /v1/devices/{id}/memory-test/results | Stored results, newest first. |
Start a test
curl -s -X POST "https://api.admrl.co/v1/devices/$DEVICE_ID/memory-test" \
-H "X-Organization-ID: $ORG_ID" \
-H "X-API-Token-ID: $TOKEN_ID" -H "X-API-Secret-Key: $TOKEN_SECRET" \
-H "Content-Type: application/json" \
-d '{"impact": "running", "placement": "online", "quick": true}'
| Field | Type | Meaning |
|---|---|---|
impact | running or stopped | Keep the workload running, or stop it for the test. |
placement | online or offline | Stay booted, or reboot into a test boot. offline requires stopped and a signed-in dashboard session. |
quick | boolean | Quick preset for Live. |
passes | integer, 1 to 20 | Passes over memory. Default 1 for Live, 2 for stopped modes. |
kernelPasses | integer, 1 to 17 | Kernel sweep passes, test boot only. Default 4. |
bitFade | boolean | Idle-decay check. On for stopped modes, off otherwise. |
holdOnFailure | boolean | Test boot only: stay in maintenance with the workload held if the test fails. |
timeBudgetSeconds | integer, 60 to 43200 | Time budget. Default 7200. |
maxBytes | integer | Optional cap on the amount of memory tested. |
Omitted fields use the mode default. Successful responses use the usual data envelope; data contains deviceId, runId, and status (the run's current state).
Cancel
POST /v1/devices/{id}/memory-test/cancel takes an optional body {"runId": "..."}. With no body it cancels whatever is running. The response has the same shape as a start.
Status
GET /v1/devices/{id}/memory-test returns:
| Field | Meaning |
|---|---|
source | live when the device answered just now, stored when it is offline or did not answer (last known health and stored results). |
supported | false when the device is online but its firmware has no memory test. |
status | The current or last run: runId, plan, phase, pass and passes, pattern, testedBytes, coverageBytes, totalBytes, rateBps, etaSeconds, errors, confirmedErrors, tempMilliC, startedAt. |
health | Capabilities (online, offline, offlineReason), hardwareCorruptedBytes, EDAC counters (-1 when unavailable), retiredPages, fault and faultReason, recent kernel memory-error lines, and the last result. |
history | Up to 10 recent results. |
healthReportedAt | When a stored health snapshot was taken (stored only). |
Phases: idle, preparing, stopping_workload, rebooting, kernel_sweep, testing, confirming, retiring, restoring, done, failed, cancelled, interrupted. Progress is also streamed on the device state stream as an operation of kind memory_test.
Results
GET /v1/devices/{id}/memory-test/results?limit=20 returns {"results": [...]}, newest first (limit 1 to 100, default 20). Each result has runId, plan, source, startedAt, finishedAt, outcome, verdict, summary, passesCompleted, patternsRun, testedBytes, coverageBytes, totalBytes, maxTempMilliC, kernel, errors (each with physAddr, expected, actual, mask, pattern, pass, tempMilliC, confirmed), errorCount, retired (physical page addresses), fault, and, for interrupted runs, interrupted. Results also carry a per-page error summary (pages) and the hardware and version details used by the run report.
Errors
Errors reported by the device or the memory-test checks use a bare body: {"error": "<code>", "message": "...", "status": {...}}. status is present for busy and carries the running test. 503, 504, and 429 use the same bodies as other calls that reach a device.
| Status | error | Meaning |
|---|---|---|
400 | invalid | The plan is invalid, for example offline with running, or a number out of range. message names the field. |
403 | operator_required | A test boot was requested by an API token, CLI token, or service account. |
409 | busy | A memory test is already running on this device. |
409 | not_running | Cancel was called with nothing running. |
422 | unsupported | The firmware or board does not support the requested test. |
502 | workload_stop_failed | The device could not stop the workload, so nothing was started. |
502 | internal | The device failed to handle the request. |
503 | device_offline | The device is not connected. |
504 | device_timeout | The device did not answer in time. |
429 | rate_limited | A start was accepted a moment ago; honour Retry-After. |
Every start and cancel is recorded in the organisation audit log with the requested plan.