Skip to content

Check huawei-dorado-controller

Overview

Checks the health and running status of all controllers on a Huawei OceanStor Dorado storage system via the REST API (/controller endpoint). Alerts when any controller reports a non-normal health or running state.

Important Notes:

  • Tested on Huawei OceanStor Dorado 8000 V6 6.1.0 and Dorado 6000 V6 V700R001C10SPH128
  • CPU and memory usage have their own thresholds: --warning and --critical for the CPU, --warning-mem and --critical-mem for the memory. All of them are off by default. A controller keeps its memory well filled in normal operation (75 to 83% on a Dorado 6000 V6, and 89 and 96% in the vendor's own example of a healthy controller), so a memory threshold only makes sense once you know what your array normally sits at.
  • A controller board without a temperature sensor answers with -1 (the boards of a Dorado 6000 V6 do). It is shown as --, left out of the temperature check and out of the performance data.
  • Create a read-only API user that can perform queries only
  • The default session timeout period on the storage system is 20 minutes; --cache-expire defaults to 15 minutes to stay within that window

Data Collection:

  • Queries the Huawei OceanStor Dorado REST API at https://<ip>:<port>/deviceManager/rest/<deviceId>/controller
  • Authenticates via session tokens (iBaseToken + cookie), cached in a SQLite database to avoid repeated logins
  • If the appliance rejects a request, the check logs in again and retries, up to three attempts one second apart

Fact Sheet

Fact Value
Check Plugin Download https://github.com/Linuxfabrik/monitoring-plugins/tree/main/check-plugins/huawei-dorado-controller
Nagios/Icinga Check Name check_huawei_dorado_controller
Check Interval Recommendation Every 5 minutes
Can be called without parameters No (--device-id, --password, --url and --username are required)
Runs on Cross-platform
Compiled for Windows No (runs with Python interpreter)
Uses State File $TEMP/linuxfabrik-monitoring-plugins-huawei-dorado.db

Help

usage: huawei-dorado-controller [-h] [-V] [--always-ok]
                                [--cache-expire CACHE_EXPIRE] [-c CRIT]
                                [--critical-mem CRIT_MEM]
                                [--critical-temperature CRIT_TEMPERATURE]
                                [--device-id DEVICE_ID] [--ignore IGNORE]
                                [--insecure] [--lengthy] [--match MATCH]
                                [--no-insecure]
                                [--no-match-severity {ok,warn,crit,unknown}]
                                [--no-perfdata] [--no-proxy]
                                [--password PASSWORD]
                                [--password-file PASSWORD_FILE]
                                [--performance] [--proxy PROXY]
                                [--scope SCOPE] [--timeout TIMEOUT] -u URL
                                --username USERNAME [-v] [-w WARN]
                                [--warning-mem WARN_MEM]
                                [--warning-temperature WARN_TEMPERATURE]

Checks the health and running status of all controllers on a Huawei OceanStor
Dorado storage system via the REST API (/controller endpoint). Alerts when any
controller reports a non-normal health or running state, and optionally when
its CPU usage, memory usage or temperature exceeds the configured thresholds.
Supports extended reporting via --lengthy, and reporting the I/O and cache
counters via --performance.

options:
  -h, --help            show this help message and exit
  -V, --version         show program's version number and exit
  --always-ok           Always returns OK.
  --cache-expire CACHE_EXPIRE
                        The amount of time after which the credential/data
                        cache expires, in minutes. Default: 15
  -c, --critical CRIT   CRIT threshold for CPU usage, as a Nagios range in
                        percent. Off by default, because a controller under
                        load is doing its job; set it once you know what your
                        array normally sits at. Example: `--critical=90`
  --critical-mem CRIT_MEM
                        CRIT threshold for memory usage, as a Nagios range in
                        percent. Off by default, because a controller keeps
                        its memory well filled in normal operation; the
                        vendor's own example of a healthy controller reports
                        96%. Set it once you know what your array normally
                        sits at. Example: `--critical-mem=99`
  --critical-temperature CRIT_TEMPERATURE
                        CRIT threshold in degrees Celsius. Off by default,
                        because a healthy operating temperature depends on the
                        controller model and on where the array stands.
                        Example: `--critical-temperature=55`
  --device-id DEVICE_ID
                        Huawei OceanStor Dorado API device ID. Optional: the
                        appliance reports its own at login, so this is only
                        needed to override that answer.
  --ignore IGNORE       Skip controllers. Any item matching this Python regex
                        will be ignored. Can be specified multiple times.
                        Example: `(?i)linuxfabrik` for a case-insensitive
                        match. The regex is anchored at the start of the
                        string (Python `re.match`) and is matched against
                        `UUID`, `LOCATION`, so prefix with `.*` to match
                        anywhere.
  --insecure            This option explicitly allows insecure SSL
                        connections.
  --lengthy             Extended reporting.
  --match MATCH         Limit to controllers. Filter by this Python regular
                        expression. Case-sensitive by default; use `(?i)` for
                        case-insensitive matching. Can be specified multiple
                        times. If both `--match` and `--ignore` are given, an
                        item must match `--match` AND not match `--ignore` to
                        be reported (include first, exclude second). Examples:
                        `(?i)example` to match "example" regardless of case.
                        `^(?!.*example).*$` to match any string except
                        "example" (negative lookahead). The regex is anchored
                        at the start of the string (Python `re.match`) and is
                        matched against `UUID`, `LOCATION`, so prefix with
                        `.*` to match anywhere.
  --no-insecure         Verify the TLS certificate against the system trust
                        store, overriding the insecure default of this check.
                        Use it once the endpoint presents a publicly trusted
                        certificate, or once its CA has been added to the
                        system trust store.
  --no-match-severity {ok,warn,crit,unknown}
                        State to report when no item matches the filters and
                        nothing is checked. Default: ok
  --no-perfdata         Suppress the performance data section from the output.
                        The status message and the exit code are unaffected,
                        so alerting keeps working while trending data is
                        dropped.
  --no-proxy            Do not use a proxy, not even one the environment
                        names. Overrides `--proxy`.
  --password PASSWORD   Huawei OceanStor Dorado API password.
  --password-file PASSWORD_FILE
                        Path to a file holding the password, read from its
                        first line. Keeps the password out of the process
                        list, where a command-line argument is visible to
                        every user on the host. Takes precedence over
                        `--password`. Keep the file readable only by the
                        monitoring user. Example: `--password-
                        file=/etc/icinga2/secrets/storage`.
  --performance         Additionally report the I/O counters of every
                        controller. Costs one API request per object, so a
                        large appliance may need a higher --timeout.
  --proxy PROXY         Proxy to reach the target through. The scheme defaults
                        to `http` when omitted. Overrides the proxy the
                        environment names (`http_proxy`, `https_proxy`,
                        `all_proxy`) together with the exceptions it lists in
                        `no_proxy`, and is itself overridden by `--no-proxy`.
                        Without either parameter the environment applies.
                        Credentials belong into the environment variable
                        rather than here, because a command-line argument is
                        visible to every user on the host. Example:
                        `--proxy=http://proxy.example.com:3128`.
  --scope SCOPE         Huawei OceanStor Dorado API scope.
  --timeout TIMEOUT     Network timeout in seconds. Default: 3 (seconds)
  -u, --url URL         Huawei OceanStor Dorado API URL.
  --username USERNAME   Huawei OceanStor Dorado API username.
  -v, --verbose         Makes this plugin verbose during the operation. Useful
                        for debugging and seeing what is going on under the
                        hood. Appends what every API request returned, so the
                        appliance's own answers can be read while working out
                        how it reports something. Session tokens are redacted.
                        The output is as long as those answers are, so this is
                        a debugging aid rather than something to leave
                        switched on.
  -w, --warning WARN    WARN threshold for CPU usage, as a Nagios range in
                        percent. Off by default, because a controller under
                        load is doing its job; set it once you know what your
                        array normally sits at. Example: `--warning=80`
  --warning-mem WARN_MEM
                        WARN threshold for memory usage, as a Nagios range in
                        percent. Off by default, because a controller keeps
                        its memory well filled in normal operation; the
                        vendor's own example of a healthy controller reports
                        96%. Set it once you know what your array normally
                        sits at. Example: `--warning-mem=97`
  --warning-temperature WARN_TEMPERATURE
                        WARN threshold in degrees Celsius. Off by default,
                        because a healthy operating temperature depends on the
                        controller model and on where the array stands.
                        Example: `--warning-temperature=45`

Documentation:
https://linuxfabrik.github.io/monitoring-plugins/check-plugins/huawei-dorado-controller/

Usage Examples

./huawei-dorado-controller --url=https://oceanstor:8088 --device-id=123456789 --username=monitoring --password=linuxfabrik

Output:

There are critical errors.

UUID   ! Location ! Master ! CPU (%) ! Mem (%) ! Health     ! Running     ! State
-------+----------+--------+---------+---------+------------+-------------+-----------
207:0A ! CTE0.A   ! -      ! 3       ! 75      ! Faulty (2) ! Online (27) ! [CRITICAL]
207:0E ! CTE0.E   ! -      ! 33      ! 87      ! Normal (1) ! Online (27) ! [OK]
207:0B ! CTE0.B   ! x      ! 17      ! 87      ! Normal (1) ! Online (27) ! [OK]
207:0C ! CTE0.C   ! -      ! 33      ! 86      ! Normal (1) ! Online (27) ! [OK]

--lengthy adds the board model, its role in the pair and the board voltage:

./huawei-dorado-controller --url=https://oceanstor:8088 --device-id=123456789 --username=monitoring --password=linuxfabrik --lengthy

Output:

There are critical errors.

UUID   ! Location ! Model              ! Role      ! Master ! CPU (%) ! Mem (%) ! Volt ! Health     ! Running     ! State
-------+----------+--------------------+-----------+--------+---------+---------+------+------------+-------------+-----------
207:0A ! CTE0.A   ! Unknown            ! Primary   ! -      ! 3       ! 75      ! 12.0 ! Faulty (2) ! Online (27) ! [CRITICAL]
207:0E ! CTE0.E   ! 4U4C control board ! Secondary ! -      ! 33      ! 87      ! 12.0 ! Normal (1) ! Online (27) ! [OK]
207:0B ! CTE0.B   ! 4U4C control board ! Primary   ! x      ! 17      ! 87      ! 12.0 ! Normal (1) ! Online (27) ! [OK]
207:0C ! CTE0.C   ! 4U4C control board ! Secondary ! -      ! 33      ! 86      ! 12.0 ! Normal (1) ! Online (27) ! [OK]

States

  • OK if all controllers report normal health and running status.
  • WARN if any controller reports a degraded health status, or one this check does not know.
  • WARN if any controller's running status is not "Normal", "Running" or "Online", unless it reports an outright failure.
  • CRIT if any controller reports health status "Faulty", "No Input", "Invalid" or "Offline".
  • CRIT if any controller's running status reports a failure ("Not running", "Sleep in High Temperature", "Offline", "Invalid", "Migration fault", "Error/Faulty", "To be synchronized", "Power-on failed", "Abnormal" or "Rollback failure").
  • WARN or CRIT if a controller's CPU usage reaches --warning or --critical. Both are off by default.
  • WARN or CRIT if a controller's memory usage reaches --warning-mem or --critical-mem. Both are off by default.
  • WARN or CRIT if a controller's temperature reaches --warning-temperature or --critical-temperature. Both are off by default.
  • A CPU, memory or temperature threshold that fires is marked on the value that crossed it, and the row's State column always shows the worst state of that controller.
  • UNKNOWN if the appliance lists no controllers at all, which points at the query rather than at the hardware.
  • --match limits the check to the controllers whose identifier, location or name matches the regex; --no-match-severity sets what to report when nothing matches (default: OK).
  • UNKNOWN on invalid API responses or responses with error codes.
  • --always-ok suppresses all alerts and always returns OK.

Perfdata / Metrics

Name Type Description
\<UUID>_cpu_usage Percentage CPU utilization.
\<UUID>_memory_usage Percentage Memory utilization.
\<UUID>_temperature Number Temperature in degrees Celsius. A board without a sensor reports -1 and is left out.
\<UUID>_voltage Number Board voltage in volts. The appliance counts it in tenths of a volt.

The health, running and location indicator status codes stay out of the performance data. The state already carries them, and a code does not read as a curve.

Have a look at the API documentation for details.

Troubleshooting

No valuable response from the API

Got no valuable response from https://...

Check the --url, --device-id, --username and --password parameters. Verify that the API user has query permissions and that the storage system REST API is reachable.

This operation fails to be performed because of the unauthorized REST.

This is a known transient issue with the Huawei REST API. The check makes up to three attempts and forces a fresh login before the second one. If the error persists, verify the API credentials and session timeout settings.

Credits, License