Check redfish-managers¶
Overview¶
Checks the health of all managers (BMCs such as iLO, iDRAC, or a generic BMC) of a Redfish-compatible server via the Redfish API. Reports every enabled manager with its identification (type, model, firmware version, power state, UUID) and rolled-up health status, and alerts whenever any manager's status leaves OK. Use redfish-logservices to inspect a manager's event log.
Important Notes:
- Tested on DELL iDRAC and DMTF Simulator
- A check usually completes within a few seconds, but a slow or retried request can take longer. The bundled Director basket allows a 60 second runtime timeout.
- This check runs with both HTTP and HTTPS. It uses GET requests only.
- No additional Python Redfish modules need to be installed.
Data Collection:
- Queries
/redfish/v1/Managersto enumerate manager members - For each member, reads the manager type, model, firmware version, power state, UUID and rolled-up health status
- Reads each collection in a single request via the Redfish
$expandquery where the controller supports it, otherwise falls back to one request per member - Uses HTTP Basic authentication if
--usernameand--passwordare provided - Only evaluates managers in "Enabled" or "Quiesced" state
Fact Sheet¶
| Fact | Value |
|---|---|
| Check Plugin Download | https://github.com/Linuxfabrik/monitoring-plugins/tree/main/check-plugins/redfish-managers |
| Nagios/Icinga Check Name | check_redfish_managers |
| Check Interval Recommendation | Every 5 minutes |
| Can be called without parameters | No (--url is required) |
| Runs on | Cross-platform |
| Compiled for Windows | No (runs with Python interpreter) |
| Uses State File | $TEMP/linuxfabrik-monitoring-plugins-redfish.db |
Help¶
usage: redfish-managers [-h] [-V] [--always-ok] [--cache-expire CACHE_EXPIRE]
[--ignore IGNORE] [--insecure] [--inventory]
[--match MATCH] [--no-insecure] [--no-perfdata]
[--no-proxy] [--password PASSWORD] [--proxy PROXY]
[--retries RETRIES] [--timeout TIMEOUT] --url URL
[--username USERNAME] [--verbose]
Checks the health of all managers (BMCs such as iLO, iDRAC, or a generic BMC)
of a Redfish-compatible server via the Redfish API. Reports every enabled
manager with its identification (type, model, firmware version, power state,
UUID) and rolled-up health status, and alerts whenever any manager's status
leaves `OK`. Use `redfish-logservices` to inspect a manager's event log.
options:
-h, --help show this help message and exit
-V, --version show program's version number and exit
--always-ok Always returns OK.
--cache-expire CACHE_EXPIRE
The amount of time after which the credential/data
cache expires, in minutes. Default: 5
--ignore IGNORE Ignore items whose name matches this Python regular
expression. Case-sensitive by default; use `(?i)` for
case-insensitive matching. Can be specified multiple
times.
--insecure This option explicitly allows insecure SSL
connections.
--inventory Output the parsed components as JSON on stdout and
exit OK, instead of running a health check. Use this
to collect a hardware inventory: the JSON is a single
object keyed by component type, so the output of
several Redfish checks can be merged into one
inventory document with `jq --slurp`. Ignores --brief,
--match and --ignore.
--match MATCH Only check items whose name matches this Python
regular expression. Case-sensitive by default; use
`(?i)` for case-insensitive matching. Can be specified
multiple times. If both `--match` and `--ignore` are
given, an item must match `--match` AND not match
`--ignore` to be reported (include first, exclude
second).
--no-insecure Verify the TLS certificate against the system trust
store, overriding the insecure default of this check.
Use it once the endpoint presents a publicly trusted
certificate, or once its CA has been added to the
system trust store.
--no-perfdata Suppress the performance data section from the output.
The status message and the exit code are unaffected,
so alerting keeps working while trending data is
dropped.
--no-proxy Do not use a proxy, not even one the environment
names. Overrides `--proxy`.
--password PASSWORD Redfish API password.
--proxy PROXY Proxy to reach the target through. The scheme defaults
to `http` when omitted. Overrides the proxy the
environment names (`http_proxy`, `https_proxy`,
`all_proxy`) together with the exceptions it lists in
`no_proxy`, and is itself overridden by `--no-proxy`.
Without either parameter the environment applies.
Credentials belong into the environment variable
rather than here, because a command-line argument is
visible to every user on the host. Example:
`--proxy=http://proxy.example.com:3128`.
--retries RETRIES Number of extra attempts if a request to the Redfish
API fails, before the check gives up. Helps against an
occasionally slow or flaky management controller.
Default: 3
--timeout TIMEOUT Network timeout in seconds. Default: 8 (seconds)
--url URL Redfish API URL.
--username USERNAME Redfish API username.
--verbose Makes this plugin verbose during the operation. Useful
for debugging and seeing what is going on under the
hood. For this check that also appends every Redfish
response it evaluated to its output, ready to be
attached to a bug report, and writes a trace of every
Redfish request, with timings, to linuxfabrik-
monitoring-plugins-redfish-trace.log below the
temporary directory. Unlike this check's output, the
trace survives a check that the monitoring server
terminates for exceeding its timeout, which is what
makes it useful against a slow management controller.
Passwords and session tokens are kept out of both. The
output grows with the responses, so keep this switched
on in a service definition only while chasing a
problem.
Documentation:
https://linuxfabrik.github.io/monitoring-plugins/check-plugins/redfish-managers/
Usage Examples¶
./redfish-managers --url=https://bmc --username=redfish-monitoring --password='linuxfabrik'
Output:
Everything is ok. Checked manager health on 1 member.
Manager: BMC Joo Janta 200, Firmware: 1.45.455b66-rev4, Power: On, UUID: 58893887-8974-2487-2389-841168418919
States¶
- OK if all enabled managers report a healthy rolled-up state.
- WARN if a manager health or health rollup state is "Warning".
- CRIT if a manager health or health rollup state is "Critical".
--always-oksuppresses all alerts and always returns OK.
Perfdata / Metrics¶
| Name | Type | Description |
|---|---|---|
| managers | Number | Number of enabled managers checked. |
| managers_not_ok | Number | Number of managers whose health is not OK. |
For Maintainers¶
You don't need a physical server with a real BMC (the management controller that serves the Redfish API, e.g. HPE iLO or Dell iDRAC) to develop or test this plugin. The official DMTF Redfish mockup server serves a static, read-only Redfish tree (including a Managers collection) over plain HTTP, which is exactly what this GET-only plugin needs.
Run the mockup server and point the plugin at it, from the repository root:
podman run \
--detach --rm \
--name lfmp-redfish-mock \
--publish 5000:8000 \
docker.io/dmtf/redfish-mockup-server:latest
sleep 3
check-plugins/redfish-managers/redfish-managers --url=http://127.0.0.1:5000 --no-proxy
podman stop lfmp-redfish-mock
Use http://127.0.0.1:5000 rather than http://localhost:5000, because localhost may resolve to IPv6 (::1) while the published container port is bound to IPv4.
The fixtures under unit-test/stdout/ are the output of --verbose runs, one file per scenario, with a ### GET <path> block for every Redfish response the plugin evaluated. The test suite replays them instead of calling a controller, so the output a user attaches to a bug report becomes a new scenario as it is, whatever the number of members. To simulate a fault, copy a healthy scenario and edit the manager's Status.Health to Critical or Warning. The offline test suite is run with ./run from the unit-test directory.
Credits, License¶
- Authors: Linuxfabrik GmbH, Zurich
- License: The Unlicense, see LICENSE file.