Skip to content

Check redfish-storage

Overview

Checks the state of all physical drives, volumes and their storage controllers in a Redfish-compatible server via the Redfish API. Alerts when any drive, volume or storage controller reports a degraded or failed state. System-level health (processors, BIOS, power, temperature, indicator LED, etc.) is deliberately ignored by this check so that a system warning unrelated to storage does not mask the storage status; use redfish-systems for that.

Important Notes:

  • Tested on DELL iDRAC and DMTF Simulator
  • A check usually completes within a few seconds, but a slow or retried request can take longer. The bundled Director basket allows a 60 second runtime timeout.
  • This check runs with both HTTP and HTTPS. It uses GET requests only.
  • No additional Python Redfish modules need to be installed.

Data Collection:

  • Queries /redfish/v1/Systems to enumerate system members
  • For each member, follows the Storage link and queries every storage controller, its physical drives and, where present, its volumes (logical drives) for health status
  • Reads each collection in a single request via the Redfish $expand query where the controller supports it, otherwise falls back to one request per member
  • Uses HTTP Basic authentication if --username and --password are provided
  • Only evaluates systems, drives, volumes and storage controllers in "Enabled" or "Quiesced" state

Fact Sheet

Fact Value
Check Plugin Download https://github.com/Linuxfabrik/monitoring-plugins/tree/main/check-plugins/redfish-storage
Nagios/Icinga Check Name check_redfish_storage
Check Interval Recommendation Every 5 minutes
Can be called without parameters No (--url is required)
Runs on Cross-platform
Compiled for Windows No (runs with Python interpreter)
Uses State File $TEMP/linuxfabrik-monitoring-plugins-redfish.db

Help

usage: redfish-storage [-h] [-V] [--always-ok] [--brief]
                       [--cache-expire CACHE_EXPIRE] [--ignore IGNORE]
                       [--insecure] [--inventory] [--match MATCH]
                       [--no-insecure] [--no-perfdata] [--no-proxy]
                       [--password PASSWORD] [--proxy PROXY]
                       [--retries RETRIES] [--timeout TIMEOUT] --url URL
                       [--username USERNAME] [--verbose]

Checks the state of all physical drives, volumes and their storage controllers
in a Redfish-compatible server via the Redfish API. Alerts when any drive,
volume or storage controller reports a degraded or failed state. System-level
health (processors, BIOS, power, temperature, indicator LED, etc.) is
deliberately ignored by this check so that a system warning unrelated to
storage does not mask the storage status; use `redfish-systems` for that.

options:
  -h, --help            show this help message and exit
  -V, --version         show program's version number and exit
  --always-ok           Always returns OK.
  --brief               Hide items that are OK and show only those in
                        WARN/CRIT state. Alerting is unaffected: all items
                        still drive the overall check state.
  --cache-expire CACHE_EXPIRE
                        The amount of time after which the credential/data
                        cache expires, in minutes. Default: 5
  --ignore IGNORE       Ignore items whose name matches this Python regular
                        expression. Case-sensitive by default; use `(?i)` for
                        case-insensitive matching. Can be specified multiple
                        times.
  --insecure            This option explicitly allows insecure SSL
                        connections.
  --inventory           Output the parsed components as JSON on stdout and
                        exit OK, instead of running a health check. Use this
                        to collect a hardware inventory: the JSON is a single
                        object keyed by component type, so the output of
                        several Redfish checks can be merged into one
                        inventory document with `jq --slurp`. Ignores --brief,
                        --match and --ignore.
  --match MATCH         Only check items whose name matches this Python
                        regular expression. Case-sensitive by default; use
                        `(?i)` for case-insensitive matching. Can be specified
                        multiple times. If both `--match` and `--ignore` are
                        given, an item must match `--match` AND not match
                        `--ignore` to be reported (include first, exclude
                        second).
  --no-insecure         Verify the TLS certificate against the system trust
                        store, overriding the insecure default of this check.
                        Use it once the endpoint presents a publicly trusted
                        certificate, or once its CA has been added to the
                        system trust store.
  --no-perfdata         Suppress the performance data section from the output.
                        The status message and the exit code are unaffected,
                        so alerting keeps working while trending data is
                        dropped.
  --no-proxy            Do not use a proxy, not even one the environment
                        names. Overrides `--proxy`.
  --password PASSWORD   Redfish API password.
  --proxy PROXY         Proxy to reach the target through. The scheme defaults
                        to `http` when omitted. Overrides the proxy the
                        environment names (`http_proxy`, `https_proxy`,
                        `all_proxy`) together with the exceptions it lists in
                        `no_proxy`, and is itself overridden by `--no-proxy`.
                        Without either parameter the environment applies.
                        Credentials belong into the environment variable
                        rather than here, because a command-line argument is
                        visible to every user on the host. Example:
                        `--proxy=http://proxy.example.com:3128`.
  --retries RETRIES     Number of extra attempts if a request to the Redfish
                        API fails, before the check gives up. Helps against an
                        occasionally slow or flaky management controller.
                        Default: 3
  --timeout TIMEOUT     Network timeout in seconds. Default: 8 (seconds)
  --url URL             Redfish API URL.
  --username USERNAME   Redfish API username.
  --verbose             Makes this plugin verbose during the operation. Useful
                        for debugging and seeing what is going on under the
                        hood. For this check that also appends every Redfish
                        response it evaluated to its output, ready to be
                        attached to a bug report, and writes a trace of every
                        Redfish request, with timings, to linuxfabrik-
                        monitoring-plugins-redfish-trace.log below the
                        temporary directory. Unlike this check's output, the
                        trace survives a check that the monitoring server
                        terminates for exceeding its timeout, which is what
                        makes it useful against a slow management controller.
                        Passwords and session tokens are kept out of both. The
                        output grows with the responses, so keep this switched
                        on in a service definition only while chasing a
                        problem.

Documentation:
https://linuxfabrik.github.io/monitoring-plugins/check-plugins/redfish-storage/

Usage Examples

./redfish-storage --url=https://bmc --username=redfish-monitoring --password='linuxfabrik'

Output:

Everything is ok. Checked storage on 1 member.

Member: Contoso 3500, HostName: web483, SKU: 8675309, SerNo: 437XR1138R2

Disk  ! Type ! Proto ! Manufacturer ! Model         ! SerialNumber ! Size     ! LifeLeft % ! State
------+------+-------+--------------+---------------+--------------+----------+------------+------
SSD 0 ! SSD  ! SATA  ! MICRON       ! MTFDDAV480TDS ! 202629662A88 ! 447.1GiB ! 100        ! [OK]
SSD 1 ! SSD  ! SATA  ! MICRON       ! MTFDDAV480TDS ! 202629662A78 ! 447.1GiB ! 100        ! [OK]

Volume         ! RAID  ! Size     ! Encrypted ! State
---------------+-------+----------+-----------+------
Virtual Disk 0 ! RAID1 ! 931.3GiB ! False     ! [OK]

ID        ! Name            ! Description     ! Drives ! State
----------+-----------------+-----------------+--------+------
RAID.SL.1 ! PERC H740P Mini ! PERC H740P Mini ! 2      ! [OK]

States

  • OK if all enabled drives, volumes and storage controllers report a healthy state.
  • WARN if any of them reports a health or health rollup state of "Warning".
  • CRIT if any of them reports a health or health rollup state of "Critical".
  • System-level health is not considered. Use redfish-systems for that.
  • --always-ok suppresses all alerts and always returns OK.

Perfdata / Metrics

The per-drive metrics depend on your hardware: a drive is only reported when it exposes the corresponding value.

Name Type Description
\<drive>_media_life_left Percentage Predicted remaining life of a drive's media (0-100%). Example: SSD_0_media_life_left
\<drive>_power_on_hours Number Hours a drive has been powered on. Example: SSD_0_power_on_hours
\<drive>_temperature Temperature Drive temperature in degrees Celsius. Example: SSD_0_temperature
drives Number Number of enabled drives checked.
drives_not_ok Number Number of drives whose health is not OK.
storage_controllers Number Number of enabled storage controllers checked.
storage_controllers_not_ok Number Number of storage controllers whose health is not OK.
volumes Number Number of enabled volumes (logical drives) checked.
volumes_not_ok Number Number of volumes whose health is not OK.

For Maintainers

You don't need a physical server with a real BMC (the management controller that serves the Redfish API, e.g. HPE iLO or Dell iDRAC) to develop or test this plugin. The official DMTF Redfish mockup server serves a static, read-only Redfish tree over plain HTTP, which is exactly what this GET-only plugin needs. Note that the bundled public-rackmount1 mockup ships no storage subsystem, so the offline fixtures are the primary way to exercise the drive and volume paths.

Run the mockup server and point the plugin at it, from the repository root:

podman run \
    --detach --rm \
    --name lfmp-redfish-mock \
    --publish 5000:8000 \
    docker.io/dmtf/redfish-mockup-server:latest
sleep 3
check-plugins/redfish-storage/redfish-storage --url=http://127.0.0.1:5000 --no-proxy
podman stop lfmp-redfish-mock

Use http://127.0.0.1:5000 rather than http://localhost:5000, because localhost may resolve to IPv6 (::1) while the published container port is bound to IPv4.

The fixtures under unit-test/stdout/ are the output of --verbose runs, one file per scenario, with a ### GET <path> block for every Redfish response the plugin evaluated. The test suite replays them instead of calling a controller, so the output a user attaches to a bug report becomes a new scenario as it is, whatever the number of members. To simulate a fault, copy a healthy scenario and edit a drive's, volume's or controller's Status.Health to Critical or Warning. The offline test suite is run with ./run from the unit-test directory.

Credits, License