❤️ This project now has a Sponsor button. It stays free and open source, and always will. If it has made your server(s) quieter, you can thank me by sponsoring me. Thank you 🙏
- Container console log example
- Requirements
- Supported architectures
- Download Docker image
- Usage
- Parameters
- Stopping the container
- Troubleshooting
- Contributors
- Contributing
- License
This container drives the fans with two IPMI raw commands Dell removed from its recent hardware. What decides whether yours takes them is the iDRAC firmware, not the model :
| iDRAC | Dell's IPMI raw fan control commands |
|---|---|
| iDRAC 6, 7 and 8 (generations 11 to 13) | Available |
iDRAC 9 up to firmware 3.30.30.30 |
Available, that being the newest version they are confirmed to work on |
iDRAC 9 from firmware 3.34.34.34 on |
Removed by Dell, deliberately and for good |
| iDRAC 10 (17th generation) | Never had them |
3.31.31.31 and 3.32.32.32 sit between those two versions and what either answers has never been reported. Blades and modular servers refuse them whatever their generation, their fans belonging to the enclosure.
Downgrading is no longer a way out : an iDRAC 9 that has installed Dell's June 2024 release or any later one cannot go back to 4.40.10.00 or older.
None of it has to be worked out in advance : the container asks the server, logs its iDRAC firmware version at startup and says what it was answered.
On the 11th generation the commands are there, but the fans are addressed differently. Dell's speed command normally carries 0xff, meaning "every fan at once". An iDRAC 6 has been reported to refuse that byte — answering completion code 0xcc, "invalid data field in request" — while accepting the very same command addressed to one fan at a time (#378, on an R510). Nothing has to be configured for it : the container tries 0xff first, and on that answer it asks the server which fans it will take, walking the whole range upwards from 0x00, then sets each of the accepted ones on every cycle. It says so once when it switches. The walk itself runs once : whichever answer it comes back with — a set of fans, or nothing at all — that answer is kept and reused, so a server that refuses every identifier is not re-asked thirty-two times a minute forever (#397). The single 0xff command is still sent on every check, so a server that starts accepting it is picked up on the very next one — and that is also what drops the answer, since a server taking 0xff is no longer the one the walk was asked about. Should it later refuse 0xff again, the range is walked afresh rather than answered from a verdict older than the machine.
The set is discovered rather than counted from ipmitool sdr type fan, because the two do not match : the R510 above exposes ten fan RPM sensors — five modules, an A and a B rotor each — plus a Fan Redundancy row that is not a fan at all, and accepts eight identifiers, 0x00 to 0x07. A count taken from that list would have addressed ten identifiers, or eleven, on a server that accepts eight ; the surplus are refused, so the profile would have been reported not applied on a server where it was. Had the mismatch gone the other way, fans would have been left running at whatever speed they had while the table said the profile was applied. Every identifier in the range is tried rather than stopping at the first refusal, for the same reason : nothing says the ones a BMC accepts form an unbroken run from 0x00.
The account this container uses must be an Administrator on the iDRAC : fan control is not read-only, and a lesser account is refused it with the very completion code (0xd4) a firmware that removed the commands returns.
- Log into your iDRAC web console
- In the left side menu, expand "iDRAC settings", click "Network" then click "IPMI Settings" link at the top of the web page.
- Check the "Enable IPMI over LAN" checkbox then click "Apply" button.
- Test access to IPMI over LAN running the following commands :
apt -y install ipmitool
ipmitool -I lanplus \
-H <iDRAC IP address> \
-U <iDRAC username> \
-P <iDRAC password> \
sdr elist allThis Docker container is currently built and available for the following CPU architectures :
- AMD64
- ARM64
- with local iDRAC:
docker run -d \
--name Dell_iDRAC_fan_controller \
--restart=unless-stopped \
-e IDRAC_HOST=local \
-e FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64> \
-e CPU_TEMPERATURE_THRESHOLD=<decimal temperature threshold in °C, from 20 to 125, or auto> \
-e HIGH_FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64, higher than FAN_SPEED ; setting it turns line interpolation on> \
-e CPU_TEMPERATURE_THRESHOLD_TO_START_LINE_INTERPOLATION=<decimal temperature in °C, from 20 to 125, lower than CPU_TEMPERATURE_THRESHOLD> \
-e CPU_TEMPERATURE_SOURCE=<auto, ipmi or lm-sensors> \
-e CHECK_INTERVAL=<seconds between each check, or a suffixed duration like 5m, up to 15 minutes> \
-e MAXIMUM_IPMI_UNREACHABLE_DURATION=<how long the iDRAC may stay unreachable before exiting, in seconds or suffixed like 5m, or empty> \
-e MAXIMUM_CONSECUTIVE_IPMI_FAILURES=<the same threshold in cycles instead, 1 or more, or empty> \
-e DISABLE_THIRD_PARTY_PCIE_CARD_DELL_DEFAULT_COOLING_RESPONSE=<true or false> \
-e KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXIT=<true or false> \
-e MONITORING_ONLY_MODE=<true or false> \
--device=/dev/ipmi0:/dev/ipmi0:rw \
tigerblue77/dell_idrac_fan_controller:latest- with LAN iDRAC:
docker run -d \
--name Dell_iDRAC_fan_controller \
--restart=unless-stopped \
-e IDRAC_HOST=<iDRAC IP address> \
-e IDRAC_USERNAME=<iDRAC username> \
-e IDRAC_PASSWORD=<iDRAC password> \
-e FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64> \
-e CPU_TEMPERATURE_THRESHOLD=<decimal temperature threshold in °C, from 20 to 125, or auto> \
-e HIGH_FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64, higher than FAN_SPEED ; setting it turns line interpolation on> \
-e CPU_TEMPERATURE_THRESHOLD_TO_START_LINE_INTERPOLATION=<decimal temperature in °C, from 20 to 125, lower than CPU_TEMPERATURE_THRESHOLD> \
-e CPU_TEMPERATURE_SOURCE=<auto, ipmi or lm-sensors> \
-e CHECK_INTERVAL=<seconds between each check, or a suffixed duration like 5m, up to 15 minutes> \
-e MAXIMUM_IPMI_UNREACHABLE_DURATION=<how long the iDRAC may stay unreachable before exiting, in seconds or suffixed like 5m, or empty> \
-e MAXIMUM_CONSECUTIVE_IPMI_FAILURES=<the same threshold in cycles instead, 1 or more, or empty> \
-e DISABLE_THIRD_PARTY_PCIE_CARD_DELL_DEFAULT_COOLING_RESPONSE=<true or false> \
-e KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXIT=<true or false> \
-e MONITORING_ONLY_MODE=<true or false> \
tigerblue77/dell_idrac_fan_controller:latestdocker-compose.yml examples:
- to use with local iDRAC:
version: '3.8'
services:
Dell_iDRAC_fan_controller:
image: tigerblue77/dell_idrac_fan_controller:latest
container_name: Dell_iDRAC_fan_controller
restart: unless-stopped
environment:
- IDRAC_HOST=local
- FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64>
- CPU_TEMPERATURE_THRESHOLD=<decimal temperature threshold in °C, from 20 to 125, or auto>
- HIGH_FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64, higher than FAN_SPEED ; setting it turns line interpolation on>
- CPU_TEMPERATURE_THRESHOLD_TO_START_LINE_INTERPOLATION=<decimal temperature in °C, from 20 to 125, lower than CPU_TEMPERATURE_THRESHOLD>
- CPU_TEMPERATURE_SOURCE=<auto, ipmi or lm-sensors>
- CHECK_INTERVAL=<seconds between each check, or a suffixed duration like 5m, up to 15 minutes>
- MAXIMUM_IPMI_UNREACHABLE_DURATION=<how long the iDRAC may stay unreachable before exiting, in seconds or suffixed like 5m, or empty>
- MAXIMUM_CONSECUTIVE_IPMI_FAILURES=<the same threshold in cycles instead, 1 or more, or empty>
- DISABLE_THIRD_PARTY_PCIE_CARD_DELL_DEFAULT_COOLING_RESPONSE=<true or false>
- KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXIT=<true or false>
- MONITORING_ONLY_MODE=<true or false>
devices:
- /dev/ipmi0:/dev/ipmi0:rw- to use with LAN iDRAC:
version: '3.8'
services:
Dell_iDRAC_fan_controller:
image: tigerblue77/dell_idrac_fan_controller:latest
container_name: Dell_iDRAC_fan_controller
restart: unless-stopped
environment:
- IDRAC_HOST=<iDRAC IP address>
- IDRAC_USERNAME=<iDRAC username>
- IDRAC_PASSWORD=<iDRAC password>
- FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64>
- CPU_TEMPERATURE_THRESHOLD=<decimal temperature threshold in °C, from 20 to 125, or auto>
- HIGH_FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64, higher than FAN_SPEED ; setting it turns line interpolation on>
- CPU_TEMPERATURE_THRESHOLD_TO_START_LINE_INTERPOLATION=<decimal temperature in °C, from 20 to 125, lower than CPU_TEMPERATURE_THRESHOLD>
- CPU_TEMPERATURE_SOURCE=<auto, ipmi or lm-sensors>
- CHECK_INTERVAL=<seconds between each check, or a suffixed duration like 5m, up to 15 minutes>
- MAXIMUM_IPMI_UNREACHABLE_DURATION=<how long the iDRAC may stay unreachable before exiting, in seconds or suffixed like 5m, or empty>
- MAXIMUM_CONSECUTIVE_IPMI_FAILURES=<the same threshold in cycles instead, 1 or more, or empty>
- DISABLE_THIRD_PARTY_PCIE_CARD_DELL_DEFAULT_COOLING_RESPONSE=<true or false>
- KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXIT=<true or false>
- MONITORING_ONLY_MODE=<true or false>For security reasons, it is recommended to store your credentials in a .env file instead of hardcoding them in a docker run command or a docker-compose.yml file. Copy .env.example to .env, fill in your values, then reference it:
- with
docker run:
docker run -d \
--name Dell_iDRAC_fan_controller \
--restart=unless-stopped \
--env-file .env \
tigerblue77/dell_idrac_fan_controller:latest- with
docker-compose.yml:
version: '3.8'
services:
Dell_iDRAC_fan_controller:
image: tigerblue77/dell_idrac_fan_controller:latest
container_name: Dell_iDRAC_fan_controller
restart: unless-stopped
env_file:
- .env(if using local iDRAC, add the devices: section shown in the examples above)
Every parameter has a default value except IDRAC_USERNAME and IDRAC_PASSWORD, which are credentials : the image ships none, and both have to be set whenever IDRAC_HOST is not local.
-
IDRAC_HOSTparameter can be set to "local" or to your distant iDRAC's IP address. Default value is "local". -
IDRAC_USERNAMEparameter is only necessary if you're adressing a distant iDRAC. It has no default value : the image ships none. -
IDRAC_PASSWORDparameter is only necessary if you're adressing a distant iDRAC. It has no default value : the image ships none. -
FAN_SPEEDparameter is the duty cycle the fans are held at while your fan control profile is applied. It can be set as a decimal percentage (from 0 to 100%) or as the same value in hexadecimal (from 0x00 to 0x64). Default value is 5(%).-
Anything outside that range stops the container at startup rather than being clamped or passed through to the fans,
200having once reachedipmitoolas0xc8. -
⚠️ The0xprefix is the only thing that tells the two notations apart, and both are accepted, so a value that lost its prefix is not refused — it just applies a different duty cycle than the one you meant.0x64is 100% while64is 64%;0x30is 48% while30is 30%. The startup log always states the one that was resolved:Fan speed objective: 64%That is the line to check first if the fans do not behave the way you expected.
-
-
CPU_TEMPERATURE_THRESHOLDparameter is the T°junction (junction temperature) threshold beyond which the Dell fan mode defined in your BIOS will become active again (to protect the server hardware against overheat). It can be set to a decimal number of degrees Celsius, from 20 to 125, or to "auto" to let the container read the threshold from the CPUs themselves. Default value is "auto".-
An explicit value has to be between 20°C and 125°C, and the container refuses to start on anything outside that window rather than run with it. No CPU throttles below the lower bound and none tolerates more than the upper one, so a value outside it is a typo (
500for50) or a Fahrenheit reading rather than a setting — and left in place it would be silent, every comparison against it reading as "not overheating" while the whole chassis stayed atFAN_SPEED. -
A value above the maximum is not a stricter setting but the absence of one, which is why it is refused rather than clamped. No PowerEdge CPU reaches 125°C — the server's own thermal protection powers the machine off first — so such a threshold could never be crossed, the overheat fallback it governs could never fire, and the container would print
CPU temperature threshold: 160°Cat startup while supervising nothing. That fallback cannot be switched off, by design, and no value of this parameter is meant to do it.MONITORING_ONLY_MODEis not that switch either: it never applies any profile, yours included. -
In "auto" mode, the threshold is the "high" temperature your CPU manufacturer defined, as reported by the
lm-sensorsutility (thehigh = +62.0°Cvalue below), which is far more relevant than a single fixed value shared by every CPU model :coretemp-isa-0000 Adapter: ISA adapter Package id 0: +45.0°C (high = +62.0°C, crit = +72.0°C) Core 0: +44.0°C (high = +62.0°C, crit = +72.0°C)"high" is the temperature at which your CPU expects cooling to be at full, and it always sits below "crit", the temperature at which the CPU throttles itself — so it is the one that leaves the fans time to act. How far below varies a lot by CPU model : 10°C on the example above, but only 2°C on a PowerEdge T630 reporting
high = +83.0°C, crit = +85.0°C. On a multi-socket server, the lowest "high" value of all detected CPUs is used as the threshold, and every detected CPU is compared against it. -
Automatic detection needs
lm-sensorsto be describing the controlled server. In"local"mode that is true by construction. In network mode it is true only when the container can be shown to be running on the serverIDRAC_HOSTnames — the same serial number comparison the CPU readings rest on (#465, extended to the threshold in #473) — and anything short of a proven match falls back to the default rather than guessing. It also requires your Docker host's kernel to expose CPU temperatures through/sys(thecoretempmodule). -
Automatic detection only works on Intel CPUs. AMD's
k10tempdriver publishes no "high" value at all on Zen parts (every EPYC server), and on older parts it publishes a fixed 70°C that is a Linux driver constant rather than an AMD specification, so it is deliberately ignored. AMD servers use the fallback value below. -
Whenever the threshold can't be detected, the container falls back to 50(°C) and logs why at startup.
-
⚠️ This default changed in v1.28. Versions up to v1.27 used a fixed 50°C. On Intel servers in "local" mode, "auto" typically resolves to a higher value (roughly 62 to 96°C depending on the CPU model — read the exact one from the startup log), so the fans stay atFAN_SPEEDlonger than they used to before Dell's profile takes over. This matches what your CPU actually asks for, but it also means the whole chassis runs atFAN_SPEEDfor longer, and the CPU is the only component this container watches. If you were relying on the old behaviour, setCPU_TEMPERATURE_THRESHOLD=50explicitly. -
⚠️ The accepted range is also new in v1.28. Versions up to v1.27 took any integer, so a value such as160was accepted and simply never reached — the container ran with its overheat fallback unable to fire, while printing that threshold at startup as though something were being supervised. From v1.28 it stops the container at startup instead, and arestart: unless-stoppedpolicy then restarts it straight into the same refusal, which reads as a container flapping rather than as a configuration mistake. Set the temperature your CPUs should not exceed, orauto.
-
-
HIGH_FAN_SPEEDparameter is the fan speed the ramp reaches once the hottest detected CPU reachesCPU_TEMPERATURE_THRESHOLD, in the same notation asFAN_SPEED(a decimal percentage from 0 to 100, or the same value in hexadecimal from 0x00 to 0x64), instead of jumping straight fromFAN_SPEEDto the Dell default fan control profile the momentCPU_TEMPERATURE_THRESHOLDis crossed (issue #44). It has no default value : the image ships none, and setting it is what turns line interpolation on — a container that never mentions it behaves exactly as one that predates this feature. It must be higher thanFAN_SPEED, and the container refuses to start otherwise.CPU_TEMPERATURE_THRESHOLDstill applies unchanged as the final safety fallback : this parameter only changes what happens below it.- The ramp is driven by the hottest of every CPU this container detects, whichever socket it is on and however many the server has.
- It is a memoryless function of the current temperature, not a state : it changes the speed below the threshold, but not what happens at it. A CPU whose load holds it exactly on the threshold still crosses it every
CHECK_INTERVAL, ramp or not. What decides whether that crossing is audible is the size of the step it makes at the top of the ramp — setHIGH_FAN_SPEEDas close as your server tolerates to the speed Dell's own profile actually reaches (read it from the "Active fan speed profile" column once it engages) : the ramp only ever removes the step below this speed, and a lowHIGH_FAN_SPEEDagainst a Dell profile that ramps far higher is still an audible jump at the threshold, just a smaller one.
-
CPU_TEMPERATURE_THRESHOLD_TO_START_LINE_INTERPOLATIONparameter is the CPU temperature (in degrees Celsius, same 20 to 125 window asCPU_TEMPERATURE_THRESHOLD) at which the ramp starts climbing away fromFAN_SPEED. Below it,FAN_SPEEDis used unchanged. It must be lower thanCPU_TEMPERATURE_THRESHOLD, and the container refuses to start otherwise. Default value is 30(°C).-
Only read, and only validated, once the ramp is on : shipping this one a default is safe, unlike the parameter above with none, precisely because it stays unread until that one turns the ramp on.
-
Example, with
FAN_SPEED=10,HIGH_FAN_SPEED=50,CPU_TEMPERATURE_THRESHOLD_TO_START_LINE_INTERPOLATION=30andCPU_TEMPERATURE_THRESHOLD=70:Hottest detected CPU Fan speed 15°C 10% 30°C 10% 35°C 15% 50°C 30% 69°C 49% 70°C 50%, or Dell's default profile the very next cycle a reading is still at or above the threshold The container prints the same shape as a small ASCII chart at startup, once for the values actually configured, so the ramp can be read at a glance without doing the arithmetic by hand.
-
-
CPU_TEMPERATURE_SOURCEparameter selects where the CPU temperatures the container supervises are read from. Default value is "auto".autoreads them from your iDRAC, and falls back tolm-sensorsonly if your iDRAC turns out to report no CPU temperature at all. Some older iDRACs accept Dell's raw fan control commands but answer nothing usable to a temperature query (issue #216) : on those, the container used to be able to do nothing but hand the fans back to Dell's own profile forever. The fallback is tried on the check that found no sensor, just before the container would otherwise refuse to run over it, and it is logged when it engages. One check is enough to conclude: an iDRAC that exposes no processor entity does so on every check rather than on that one, which is the same reason the container refuses instead of retrying. Once engaged it stays for the life of the container, that being a property of the firmware rather than a passing condition, and changing the meaning of the table's numbers mid-run would be worse than keeping a source that works.ipminever readslm-sensors, whatever your iDRAC reports. Set it if you want the source to be the iDRAC and nothing else.lm-sensorsreads them fromlm-sensorsfrom the start, without waiting for your iDRAC to prove it cannot report them. The container refuses to start iflm-sensorsreports no CPU temperature, rather than silently supervising nothing.- Whatever the source, fan control still goes through your iDRAC :
lm-sensorsreplaces the readings, not the IPMI commands. A server whose iDRAC rejectsraw 0x30 0x30cannot be cooled by this container at all, and this parameter changes nothing for it. lm-sensorsreads the CPUs of the machine the container runs on, so it can only answer for the controlled server when that machine is the controlled server. In "local" mode it is, by construction. In network mode the container now checks rather than assumes : it compares the serial numbers your host reports about itself with the ones your iDRAC reports for the server it manages, pair by pair —/sys/class/dmi/id/product_serialagainst the FRU'sProduct Serial(the service tag on a Dell), and/sys/class/dmi/id/board_serialagainst itsBoard Serial. A pair is only ever compared with itself and never crossed with the other one's value (#469) : a server whose FRU carries noProduct Serialat all — the R510 of #378 is one — is answered by the board pair instead. Both sides are canonicalised before being compared : whitespace removed, case folded, and the padding a firmware wraps the value in trimmed off the ends, one DMI table reporting..CN1374XXXXXXXX.for the very board its iDRAC callsCN1374XXXXXXXX(#479).autofalls back only when one whole pair is readable on both sides, holds real identifiers, and is equal (#465). Anything else keeps the old refusal — an unreadable/sys(both files are root-only on most distributions, and absent when/sysis not mounted into the container), a field nobody filled in, or two values that differ. Asking forlm-sensorsoutright in network mode is still refused whatever the machines are, that being an assertion whereautomakes a check.- A CPU with no package temperature sensor is read from its hottest core.
lm-sensorsnormally publishes one temperature for the whole chip — the package — alongside one per core, and the package is what gets read. Some older Xeons publish no package sensor at all (issue #378, on an R510 whose iDRAC6 reports no CPU temperature either, leavinglm-sensorsas the only source). There, the hottest core stands in for it — not the average of the cores, which on a partly loaded CPU reads far below both the package and the per-core limit it is compared against, and would keep the fans low while one core approached throttling. The startup line that names the source of each CPU column says which one it is —coretemp-isa-0000 (hottest core)rather thancoretemp-isa-0000— because the two readings look identical in the table and the core one runs slightly warmer. - It also only works on Intel CPUs, again like threshold detection : AMD's
k10tempdriver reportsTctl, a control value that is not the physical temperature your iDRAC reports for the same CPU. Onautoandipmi, an AMD server therefore keeps reading its CPUs through IPMI; asking explicitly forlm-sensorson one makes the container refuse to start rather than supervise nothing. - The value is read leniently : case is ignored, surrounding whitespace and quotes are stripped,
lm_sensorsandlmsensorsare accepted as spellings oflm-sensors, and an empty value meansauto. Anything that is still none of the three stops the container at startup rather than silently meaningauto. - Your inlet and exhaust temperatures keep coming from your iDRAC :
lm-sensorshas no equivalent for them, so only the CPU rows are replaced. On a server that reports neither, both columns show-, which is what they already do today.
-
CHECK_INTERVALparameter is the time between each temperature check and potential profile change, in seconds unless a unit suffix (s,m,hord) says otherwise, so90,90sand5mare all valid. Fractions of a second are not. The container refuses to start on a valuesleepcannot wait for, and on zero, as either would leave the monitoring loop unpaced and running at full speed against your iDRAC. Default value is 5(s). A short interval makes the controller react quickly to temperature spikes, at the cost of more IPMI traffic towards the iDRAC and more container log lines. If your iDRAC struggles to keep up (especially over LAN) or if you prefer quieter logs, increase this value.This interval is also the controller's reaction time, so it is bounded from above. While your fan control profile is applied, Dell's own dynamic fan control is disabled and the fans are pinned at
FAN_SPEED: nothing raises them until the next check reads a temperature aboveCPU_TEMPERATURE_THRESHOLD. The interval is therefore the longest your server can heat up with its cooling frozen at a speed you chose for an idle machine. Above 60 seconds the container starts and prints a warning saying so. Above 15 minutes it refuses to start, that delay being long enough that the controller is not really controlling anything anymore.Both limits are lifted by
MONITORING_ONLY_MODE, where no profile is ever applied, Dell's dynamic fan control keeps the fans and the interval is only how often temperatures are logged. If you want a slow polling cadence for logging purposes, that is the mode to use. -
MAXIMUM_IPMI_UNREACHABLE_DURATIONparameter is how long the iDRAC may stay completely unreachable before the container gives up and exits. Default value is 60s, in seconds unless a unit suffix (s,m,hord) says otherwise, exactly likeCHECK_INTERVAL. Empty disables the escalation. It counts only failures to reach the iDRAC: a server correctly reported as powered off is a state that was observed, not a failure, and never counts however long it stays off; any cycle that reaches the iDRAC resets the count. An unreachable iDRAC accepts no command, so exiting cannot and does not try to move the fans — the point is to obtain a fresh IPMI session, which is what clears an expired session, a rebooted iDRAC or an exhausted session limit, and to make the loss visible todocker psand to anything watching container state instead of it being buried in logs. Read the paragraph below before relying on it.This only helps if something restarts the container. With Docker's default
norestart policy, a container that exits stays dead:graceful_exittries to restore Dell's profile on the way out, but that command goes through the same unreachable iDRAC and fails too, so the fans keep the speed they were last set to with nothing watching them at all. Worse, a container that keeps retrying recovers on its own the moment the iDRAC answers again, whereas one that exited does not. Run with--restart unless-stopped(or a Composerestart:policy) if you leave this enabled, or set it empty to keep the previous retry-forever behaviour.The duration is counted in whole
CHECK_INTERVALcycles: it is divided by the interval and rounded up, never below one, nothing being concluded from less than one observed failure. So a duration at or below a singleCHECK_INTERVALmeans exiting on the first unreachable reading — which is what a0is refused for, reached without a refusal. With the defaults, 60s against a 5s interval, that is 12 cycles; if you shorten this parameter, keep it comfortably above yourCHECK_INTERVAL. The startup log states what it resolved to, and warns you whenever the escalation would fire on the first failure, whichever of the two parameters put it there:iDRAC unreachable escalation: After 12 checks (60s, rounded up to whole check intervals) -
MAXIMUM_CONSECUTIVE_IPMI_FAILURESparameter expresses that same threshold as a raw number of consecutive unreachable cycles instead of a duration. Default value is (empty), the duration above being used. When set it takes precedence, being the more specific of the two. Prefer the duration unless you need the count exactly: it keeps meaning the same thing whenCHECK_INTERVALchanges, where a cycle count silently would not.- The minimum is 1,
0being refused: nothing can be concluded from fewer than one observed failure. Use an empty value, not0, to disable the escalation. 1is accepted but warned about at startup, since it means exiting on the very first unreachable reading — on any transient glitch. It is legitimate on a rock-solid LAN, which is why it is a warning rather than a refusal; the same warning fires whenMAXIMUM_IPMI_UNREACHABLE_DURATIONresolves to a single check.
- The minimum is 1,
-
DISABLE_THIRD_PARTY_PCIE_CARD_DELL_DEFAULT_COOLING_RESPONSEparameter is a boolean that allows to disable third-party PCIe card Dell default cooling response. Default value is false.- From the 14th generation the IPMI command this parameter used is gone. Dell moved the setting at that generation, from one command covering the whole server to one attribute per PCIe slot reachable only over Redfish, so the command this container sends is answered "invalid command". A server older than the generations that ever had the setting answers exactly the same way, and the completion code does not tell the two apart — so a refusal in your log is not by itself proof that yours is a 14th generation machine, and on an older one there is nothing for Redfish or for network mode to reach (#481). Setting it yourself by hand takes a moment, should you ever want to : Configuration > System Settings > Hardware Settings, under Cooling Configuration (named Fans Configuration on older iDRAC 9 firmware), then in the PCIe Airflow Settings table set LFM Mode to
Disabledon the slot holding the card. - On those servers the container now applies it over Redfish, so the parameter works again. Only the slots actually holding a third-party card are written to — a slot with a Dell card or no card at all is left alone, its airflow being something Dell has real data for — and every slot goes in a single request, because a Redfish write creates a configuration job on the iDRAC. It is written once rather than on every cycle for the same reason, and a slot already in the wanted state is not written at all.
KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXITkeeps its meaning on this transport :falseputs Dell's default back on the way out, not whatever value was there before the container started, which is exactly what it has always done over IPMI.- The whole exchange is given ten seconds per cycle, shared between the requests it makes rather than granted to each of them, so that an iDRAC whose HTTPS interface hangs cannot hold up the temperature reading the next check depends on. A healthy one answers in a fraction of that.
- Reaching the setting at all is retried up to three times, one
CHECK_INTERVALapart, when what stopped it describes a moment the iDRAC was having — busy, its configuration job queue full, or a request that never completed. An answer about the request, the resource or the credentials is concluded on the first one instead, none of those changing while the container runs. Either way the log says which, and how long it went on. This is not tunable, and does not need to be. - An iDRAC that could not be reached is never reported as a server without the setting. Those are different statements, and only the second one names your hardware : a readable answer carrying no per-slot control says
Not supported by this server, while an unreachable HTTPS interface saysRedfish could not be reached (see the log)and local mode saysNot over IPMI (Redfish needs network mode). - The whole thing runs only in network mode : in local mode the container reaches the BMC through
/dev/ipmi0and is given no iDRAC address or credentials, so there is nothing to address an HTTPS request to. The iDRAC's certificate is not verified, because iDRACs ship self-signed ones and the request carries the same credentials to the same host as the IPMI session already does. - This half has never run against a real iDRAC. It was built from attribute dumps posted on #360 by owners of seven machines, and its test suite runs against a mock built from them — a good mock, and not a server. Nothing it does can pin your fans or change a fan speed : it writes one attribute, on the slots reporting a third-party card, and puts Dell's own value back when the container stops. If it misbehaves on your hardware, the log names the exact path to set it by hand, and reporting it on that issue is welcome.
- From the 14th generation the IPMI command this parameter used is gone. Dell moved the setting at that generation, from one command covering the whole server to one attribute per PCIe slot reachable only over Redfish, so the command this container sends is answered "invalid command". A server older than the generations that ever had the setting answers exactly the same way, and the completion code does not tell the two apart — so a refusal in your log is not by itself proof that yours is a 14th generation machine, and on an older one there is nothing for Redfish or for network mode to reach (#481). Setting it yourself by hand takes a moment, should you ever want to : Configuration > System Settings > Hardware Settings, under Cooling Configuration (named Fans Configuration on older iDRAC 9 firmware), then in the PCIe Airflow Settings table set LFM Mode to
-
KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXITparameter is a boolean that decides what becomes of the third-party PCIe card cooling response when the container stops :trueleaves it as the container set it,falserestores Dell's own. Default value is false. -
MONITORING_ONLY_MODEparameter is a boolean that allows to run the container in a read-only, monitoring-only mode: temperatures are still read and logged at eachCHECK_INTERVAL, but no fan control profile (neither the user-defined one nor Dell's default) and no third-party PCIe card cooling response change is ever sent to the server. Useful to observe temperatures and validate yourFAN_SPEED/CPU_TEMPERATURE_THRESHOLDvalues before letting the container actually take control of the fans. Default value is false.
The three boolean parameters above accept only the lowercase literals true and false, and the container refuses to start on anything else — True, TRUE, 1, on and yes included. This is not pickiness: these parameters are dispatched by running their value, so those spellings used to be taken as false without a word. MONITORING_ONLY_MODE=True would take control of the fans on a server you had explicitly asked the container to leave alone, while logging "Monitoring only mode: Disabled".
While the container runs, the fans are held at your FAN_SPEED and the server's own thermal regulation is switched off. Stopping the container is what gives it back, so that is not a detail of the shutdown — it is the moment the server stops depending on this container to stay cool.
On docker stop, docker restart, a host shutdown or a docker compose down, the container:
- applies Dell's default dynamic fan control profile, handing the fans back to the iDRAC;
- resets the third-party PCIe card cooling response to Dell's default, unless you set
KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXIT=true; - exits.
It says so in the log, and that line is the one to look for when you stop it:
/!\ Warning /!\ Container stopped, Dell default dynamic fan control profile applied for safety. Exiting.
In MONITORING_ONLY_MODE nothing was ever applied, so nothing is restored and no command is sent.
The container's PID 1 is a small supervisor whose only job is to make sure step 1 happens even if the monitoring process cannot do it itself. It forwards the stop signal, gives the monitoring process 3 seconds to bow out cleanly, kills it if it has not, and then applies Dell's profile on its behalf.
That needs Docker's stop timeout to be longer than those 3 seconds. The default is 10 seconds, so nothing has to be configured — but if you shorten it, you can cut the container off before it has handed the fans back:
| Where | Keep it above 3 seconds |
|---|---|
docker stop -t <seconds> |
the default is 10 |
docker run --stop-timeout <seconds> |
the default is 10 |
stop_grace_period: in docker-compose.yml |
the default is 10s |
terminationGracePeriodSeconds: on Kubernetes |
the default is 30 |
Past that timeout the container is SIGKILLed, which no program can catch or delay, and the fans stay at your FAN_SPEED with nothing left to raise them. The safety net is defeated precisely in the case it exists for, so if you have a reason to stop containers quickly, exclude this one.
Read its log rather than its restart count: a container that stops on a configuration mistake prints a block naming the parameter, its current value and what would have been accepted, then exits.
/!\ Error /!\ Invalid configuration, the container will not start.
Parameter : CPU_TEMPERATURE_THRESHOLD
Value : "160°C"
Expected : a temperature between 20°C and 125°C, [...]
Restarting will not help : this is read the same way on every start, so a container under
an "always", "unless-stopped" or "on-failure" restart policy stops here again on every
attempt, until the configuration itself is corrected.
Nothing about that can self-correct — the value is the same on every attempt — so the restart policy the Usage examples recommend, which is what protects you against a transient failure, turns a permanent refusal into an endless loop. docker logs --tail 40 Dell_iDRAC_fan_controller shows the block; fix what it names and the loop ends. The most common case is a CPU_TEMPERATURE_THRESHOLD above 125, which versions up to v1.27 accepted (see the Parameters section).
The other thing that restarts a container here is its healthcheck, and it now watches two things rather than one. It has always asked whether the temperatures can still be read — an iDRAC that stopped answering fails it, and the restart gets a fresh IPMI session. Since #440 it also asks whether the monitoring loop is still completing its cycles, because an iDRAC that keeps answering while the loop has stopped used to leave the container reported healthy with the fans held at your FAN_SPEED and nothing evaluating the temperature threshold any more. That check is deliberately slack — three check intervals, and never less than a minute, on top of Docker's own three retries — and a container that has not completed its first cycle yet is never called unhealthy, a server that is switched off being something this container waits out by design.
A refusal without that "Restarting will not help" line is the opposite case, and the only one worth waiting out: an iDRAC that did not answer this time may answer the next, so the restart policy is what recovers it without you doing anything.
That server does not let its fans be driven over IPMI. The container says so once, in full, and then stops asking :
/!\ Warning /!\ This server refused fan control, so the container will stop asking and keep reading temperatures instead.
Its fans are not left in an unknown state : the command that was refused is the one that takes them away from Dell's own dynamic fan control profile, so they never left it.
[...]
Three things are worth knowing about it.
Nothing is left half-done. The refused command is the one that takes the fans from Dell's own dynamic profile, so they never left it : the server cools itself exactly as it did before the container started, and stopping the container changes nothing either.
It is a verdict about this server, not about this cycle. An iDRAC that answers with an IPMI completion code answers with the same one every time, so re-sending the command only fills the log. An iDRAC that could not be reached prints no completion code at all, and that case is never treated as a refusal — the commands keep being sent for as long as the outage lasts.
Two things produce this answer, and the completion code does not say which : a firmware that removed the commands, or an iDRAC account that is not an Administrator. The iDRAC version and iDRAC user privileges sections cover both. Check the account first if this server used to work — that is the half you can fix.
The container keeps running, keeps reading the temperatures and keeps logging them, which is what MONITORING_ONLY_MODE does on purpose.
If it is the firmware, downgrading will not get you out of it. Dell removed the commands on purpose — "going forward access is not going to be allowed as it affects the thermal algorithms and cooling the system", on its own forum — and since the June 2024 release (7.00.00.172 on the 14th generation, 7.10.50.00 on the 15th) an iDRAC 9 cannot be downgraded to 4.40.10.00 or older, the bootloader having changed in a way older firmware cannot boot. The restriction carries forward to every version since, and Dell documents that "while the firmware downgrade job may reflect a successful status, the Lifecycle Log records a RAC0181 informational event on iDRAC9 reboot" : check the version the iDRAC reports afterwards rather than the job's outcome.
What the server still lets you do, without this container. The iDRAC keeps a coarser set of cooling controls of its own, in Configuration > System Settings > Hardware Settings, under Cooling Configuration — named Fans Configuration on older iDRAC 9 firmware, so look for either. None of them is a substitute for FAN_SPEED — Dell's algorithm keeps deciding, and its own documentation is explicit that "the server does not allow fan speeds to drop below the threshold that is required to cool the server" — but two of the three are worth knowing about, and the third is worth knowing about precisely so you do not reach for it :
- Thermal profile — Minimum Power or Sound Cap instead of Default. This is the only one of the three that can make a server quieter, and it selects a profile rather than a speed.
- Fan speed offset — Low, Medium, High and Max add +25%, +50%, +75% and +100% over the baseline the algorithm computed. It only ever raises, so it answers a server that runs too hot and never one that runs too loud, whatever its name suggests.
- Minimum fan speed — a floor, not a setpoint : "System fans can run higher than the fan speed that the MFS option sets [...] but not lower". It cannot take the fans below what the algorithm asks for, which is the one thing this container does and the one thing these servers no longer allow. It is also bounded from below by a read-only limit that varies with the machine, so a low value may not even be accepted.
Dell's own documents disagree on which licence this needs — KB 000257346 says "Enterprise or Datacenter" for the web interface, the iDRAC 9 RACADM guide says "iDRAC Express or iDRAC Enterprise" for the same attributes, and the iDRAC 10 attribute registry marks them as needing none — so look rather than assume. Whether this container could ever drive any of it is #360, which is waiting on hardware and is not a promise.
- Check
Tcase(case temperature) of your CPU on Intel Ark website and then setCPU_TEMPERATURE_THRESHOLDto a slightly lower value. Example with my CPUs (Intel Xeon E5-2630L v2) : Tcase = 63°C, I setCPU_TEMPERATURE_THRESHOLDto 60(°C). Note that the default "auto" value does not do this for you : Tcase is a case (heat spreader) temperature, while "auto" uses the junction-scale "high" value, which is usually well above Tcase (on a Xeon Gold 5122 for instance, Tcase is 71°C but "high" is around 94°C). If you want a Tcase-derived threshold, set it explicitly. The startup log always states which threshold was picked and where it comes from. - If it's already good, adapt your
FAN_SPEEDvalue to increase the airflow and thus further decrease the temperature of your CPU(s) - If neither increasing the fan speed nor increasing the threshold solves your problem, then it may be time to replace your thermal paste
There is no switch to flip while the container runs, and there does not need to be: stopping a container that drives the fans already hands them back to Dell's own dynamic profile, which is precisely the transition you are asking for. Recreating it with the other value of MONITORING_ONLY_MODE is therefore the whole mechanism, and it goes through the handover Stopping the container describes — the supervisor's safety net included, which is what makes it safe rather than merely convenient (see #407).
Mind which way round the two modes are, because they read backwards from what you are after:
| What drives the fans | How loud | |
|---|---|---|
MONITORING_ONLY_MODE=false |
this container, holding them at your FAN_SPEED |
as quiet as that speed |
MONITORING_ONLY_MODE=true |
the iDRAC's own dynamic profile | whatever Dell's algorithm decides, which is loud under load |
So "quiet at night" is false, and "let the server look after itself during the day" is true.
Two containers, one configuration each, only one of them running at a time. Take whichever Usage recipe you already run and create it twice, with docker create in place of docker run so that neither starts before it is asked to: the same parameters both times, two --names, and the one value that differs — MONITORING_ONLY_MODE=false on the quiet one, MONITORING_ONLY_MODE=true on the other. Name them fans_quiet and fans_loud, and give both --restart=unless-stopped.
Then, from the host's cron:
0 22 * * * docker stop fans_loud && docker start fans_quiet
0 8 * * * docker stop fans_quiet && docker start fans_loudThat restart policy is deliberate: after a host reboot only the one that was actually running comes back, a container explicitly stopped not being started again.
With docker compose, one service is enough. Interpolate the value in your docker-compose.yml:
- MONITORING_ONLY_MODE=${MONITORING_ONLY_MODE:-false}and drive it the same way:
0 22 * * * cd /path/to/compose && MONITORING_ONLY_MODE=false docker compose up -d
0 8 * * * cd /path/to/compose && MONITORING_ONLY_MODE=true docker compose up -ddocker compose up -d recreates the container when the configuration it resolves has changed, and leaves it alone when it has not.
Either way there is no moment when the fans are held at a static speed with nothing watching them: the stop hands them back before the next container starts.
Each mode is also validated for what it is actually going to do, which one container switching in place could not be. CHECK_INTERVAL is bounded to 15 minutes when the container drives the fans and unbounded when it only logs, so the monitoring container may poll every hour while the controlling one keeps a reaction time worth having. The monitoring one is also allowed to run with no readable CPU temperature sensor at all, where a controlling container refuses to start.
If what you want is two speeds rather than "this container or Dell's algorithm" — quiet at night, less quiet by day — it is the same recipe with MONITORING_ONLY_MODE=false on both containers and two different FAN_SPEED values. Worth weighing: Dell's dynamic profile under load is considerably louder than a static 30%, so "let Dell decide by day" and "run at 30% by day" are quite different outcomes.
- Run the image using usual
docker runcommand instead of UnRAID Community Apps or Docker UI. More informations here.
At startup, the container logs the CPU temperature sensors it found, with the IPMI entities they were read from (4 CPU temperature sensors detected (entities 3.1, 3.2, 3.3 and 3.4).), and prints one column per detected CPU. There is no built-in limit on that number, so 4-socket servers (R930, R830, R920, R940...) get all of their CPUs monitored.
If fewer sensors are listed than the number of CPUs installed, your iDRAC isn't reporting the missing sockets as readable IPMI processor entities. Check what it does report with :
ipmitool -I lanplus \
-H <iDRAC IP address> \
-U <iDRAC username> \
-P <iDRAC password> \
sdr type temperatureCPUs are the lines whose 4th column is an entity 3.<something> and whose reading ends in degrees C. They need not be contiguous nor in order: a socket that is empty or unreadable is usually still listed, but as Disabled or No Reading instead of a temperature, and is therefore not monitored. Please open an issue with your server model and that output if a CPU that does report a temperature is missing from the table.
The container follows those sensors while it runs, so there is no need to restart it after changing the CPUs of the target server. A CPU that starts reporting a temperature is picked up and monitored on the next check.
A CPU that stops reporting one keeps its column, reading -, and the Dell default fan control profile is applied meanwhile, since its temperature is unknown. It is only dropped from the table if it is still silent on several consecutive checks after the server has been switched off and back on, that being the only way a CPU can physically leave the machine: a sensor going quiet on a running server is a fault, not a missing socket, and dropping it would silently stop watching a CPU that is still installed. The conclusion is logged as such :
CPU 3 and CPU 4 are considered removed from the server: their temperature sensors (entities 3.3 and 3.4) reported nothing on the 5 readings that followed the server powering back on. 2 CPU temperature sensors detected (entities 3.1 3.2).
Several agreeing readings are required because a populated socket can still be unreadable for a few checks after a reboot, while its iDRAC reports it exactly like a socket that is gone. Following the CPUs this way costs no extra IPMI command: it reuses the sensor data each cycle already reads.
Note that on chassis products (VRTX, FX2, M1000e, MX7000) each server node has its own iDRAC with its own address: point the container at a node's iDRAC, not at the chassis CMC, which doesn't answer IPMI at all.
If your iDRAC reports no readable CPU temperature sensor at all, the container has nothing to supervise. Every PowerEdge has at least one CPU, so rather than sit and wait it hands the fans back to Dell's own dynamic profile and refuses to run, naming what to check:
/!\ Error /!\ No CPU temperature sensor could be read from DELL PowerEdge R730xd, and every PowerEdge has at least one CPU.
MONITORING_ONLY_MODE=true is the exception: it drives no fan, so a CPU it cannot read costs it a column and nothing else. It keeps running and logs the chassis temperatures it can read, with no CPU column in the table:
No CPU temperature sensor detected, only the chassis temperatures will be monitored.
Some older iDRACs are in that state permanently: they accept Dell's raw fan control commands but answer nothing usable to a temperature query. In "local" mode, the machine running the container is the server, so its CPUs can be read directly instead, through lm-sensors. That is what CPU_TEMPERATURE_SOURCE=auto (the default) does, on the check that found no sensor and before the container would otherwise refuse to run:
08-08-2026 15:04:31 The iDRAC reports no readable CPU temperature sensor, reading the CPUs from lm-sensors instead. Fan control keeps going through the iDRAC.
2 CPU temperature sensors detected (lm-sensors chips coretemp-isa-0000 and coretemp-isa-0001).
If that line never appears, the fallback couldn't engage. In order:
- You're in network mode and the container could not prove it runs on the controlled server.
lm-sensorsreads the machine the container runs on, so it may only answer for a server it can show is that same machine. The check compares the serial numbers your host reports about itself with the ones your iDRAC reports, pair by pair —/sys/class/dmi/id/product_serialagainst the FRU'sProduct Serial,/sys/class/dmi/id/board_serialagainst itsBoard Serial— and anything short of a match on one whole pair refuses. The likeliest reason is that neither file is readable from inside the container : both are root-only on most distributions, and absent altogether when/sysis not mounted. Check both before concluding, and check the FRU side withipmitool -I open fru: a server whose FRU carries noProduct Serialis answered by the board pair alone. Comparing the two by eye can also mislead — the container trims the padding a firmware wraps the value in, so a host reporting..CN1374XXXXXXXX.for a board the iDRAC callsCN1374XXXXXXXXis a match. SettingIDRAC_HOST=localand exposing/dev/ipmi0sidesteps the question entirely, and is the simpler answer when the container runs on the server anyway: the CPUs are then read locally while every fan control command still goes to the very same BMC. - Your Docker host doesn't expose its CPU temperatures. Check with
docker exec <container name> sensors -u, which must print acoretemp-chip carrying at least onetemp*_input:line. APackage id 0:block is not required — a Xeon that publishes no package sensor is read from its hottest core instead, and a host whose/etc/sensors.drenames the feature is read by sub-feature number rather than by label (#378). If it prints nothing at all, load thecoretempkernel module on the host (modprobe coretemp) and make sure/sysis readable from the container. - Your CPUs are AMD.
k10tempreportsTctl, a control value that is not the physical temperature your iDRAC reports, so it is deliberately not read. Please open an issue with your server model,sensors -uandipmitool -I open sdr elist allif you have such a server: hardware output is what's missing to support it.
Whatever the source, fan control itself always goes through your iDRAC. If the raw commands are rejected too:
/!\ Error /!\ Failed to enable manual fan control. ipmitool said: Unable to send RAW command (channel=0x0 netfn=0x30 lun=0x0 cmd=0x30 rsp=0xc1): Invalid command.
then no temperature source changes anything: your server's fans cannot be driven through this container. That is expected on blades and sleds, whose fans belong to their enclosure and are driven by its CMC.
Thanks to everyone who already has :
Contributions are what make the open source community such an amazing place to learn, inspire, and create. Any contributions you make are greatly appreciated.
If you have a suggestion that would make this better, please fork the repo and create a pull request. You can also simply open an issue with the tag "enhancement". Don't forget to give the project a star! Thanks again!
- Fork the Project
- Create your Feature Branch (
git checkout -b feature/AmazingFeature) - Commit your Changes, signed off (
git commit -s -m 'Add some AmazingFeature') - Push to the Branch (
git push origin feature/AmazingFeature) - Open a Pull Request
Please read CONTRIBUTING.md before you do : it explains the sign-off in step 3, and the terms your contribution arrives under in a project that is dual-licensed.
To test locally, use either :
docker build -t tigerblue77/dell_idrac_fan_controller:dev .
docker run -d ...or
export IDRAC_HOST=<iDRAC IP address>
export IDRAC_USERNAME=<iDRAC username>
export IDRAC_PASSWORD=<iDRAC password>
export FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64>
export CPU_TEMPERATURE_THRESHOLD=<decimal temperature threshold in °C, from 20 to 125, or auto>
export HIGH_FAN_SPEED=<fan speed in %, from 0 to 100, or hexadecimal from 0x00 to 0x64, higher than FAN_SPEED ; setting it turns line interpolation on>
export CPU_TEMPERATURE_THRESHOLD_TO_START_LINE_INTERPOLATION=<decimal temperature in °C, from 20 to 125, lower than CPU_TEMPERATURE_THRESHOLD>
export CPU_TEMPERATURE_SOURCE=<auto, ipmi or lm-sensors>
export CHECK_INTERVAL=<seconds between each check, or a suffixed duration like 5m, up to 15 minutes>
export MAXIMUM_IPMI_UNREACHABLE_DURATION=<how long the iDRAC may stay unreachable before exiting, in seconds or suffixed like 5m, or empty>
export MAXIMUM_CONSECUTIVE_IPMI_FAILURES=<the same threshold in cycles instead, 1 or more, or empty>
export DISABLE_THIRD_PARTY_PCIE_CARD_DELL_DEFAULT_COOLING_RESPONSE=<true or false>
export KEEP_THIRD_PARTY_PCIE_CARD_COOLING_RESPONSE_STATE_ON_EXIT=<true or false>
export MONITORING_ONLY_MODE=<true or false>
chmod +x Dell_iDRAC_fan_controller.sh
./Dell_iDRAC_fan_controller.shThe repository ships an automated test suite that runs the controller against a mocked ipmitool, so it needs no Dell hardware, no iDRAC and no network : bash, coreutils, GNU grep and awk are enough.
./tests/run_tests.sh # run everything
./tests/run_tests.sh --list # list the test cases without running them
./tests/run_tests.sh -f temperature # only run the cases whose name, or whose case file, matches
./tests/run_tests.sh --tap # emit TAP output for a CI parserIt covers every PowerEdge generation from the 9th (2006) to the 17th (2024) — including the recent ones whose firmware no longer accepts Dell's IPMI raw fan control commands — in their single, dual and quad socket variants, plus the sensor layouts they report (missing exhaust sensor, empty second socket, unreadable reading, two-digit sensor IDs...).
Blades and modular servers are covered too : the M1000e and VRTX blades, the FX2 and MX7000 sleds and the nodes of a C-series chassis. They carry no fan of their own — the enclosure does, driven by its CMC — so this container cannot cool them, and the suite pins what it does instead : identify the server, report that the fan control commands were rejected, and keep monitoring.
The suite also runs on every pull request, and on every push to master, through the Tests workflow, both directly and inside the built Docker image. Each run publishes a report on the pull request : the test count compared against the base commit, and, behind the check run, every test case that ran with what a failing one expected and what it obtained. See tests/README.md for the layout and for how to add a test case.
This project is dual-licensed.
By default, it is free software under the GNU Affero General Public License version 3 (AGPL-3.0-only). You may use it, study it, modify it and redistribute it, at no cost and with no formality. The one thing asked in return is reciprocity : if you distribute the program — as-is or modified, as scripts, as an image, or inside a product — the people who receive it must get the corresponding source under those same terms. The full text is in LICENSE.
Running the container is never restricted. On a homelab, on a company's own servers, in production, at any scale : the AGPL asks nothing of you for that, and no permission is needed.
A separate commercial licence is available for the parties who cannot meet those obligations — typically a vendor embedding the controller in a product whose source cannot be published, or anyone needing a warranty, an indemnity or a support commitment, none of which the AGPL provides. The choice between the two is yours ; see LICENSE-COMMERCIAL.md.
Using it
Under AGPL-3.0-only |
|
|---|---|
| Run it in a homelab | ✅ |
| Run it at work, in production, at any scale | ✅ no permission needed |
| Modify it for your own use, without distributing it | ✅ |
| Redistribute it, modified or not, commercially or not | ✅ provided the corresponding source goes with it |
| Combine it with GPL / AGPL code | ✅ |
| Patent licence | ✅ granted (§11) |
Putting it inside something you ship
| What you need | |
|---|---|
| Ship it in your product, publishing your modified source | ✅ nothing — the AGPL covers it |
| Ship it in your product, keeping your source closed | 💼 commercial licence |
| Build a service on it and decline the §13 source obligation | 💼 commercial licence |
| Get a warranty, an indemnity or a support commitment | 💼 commercial licence |
The short version : using it never requires a commercial licence, and never requires permission. Only conveying it while withholding the corresponding source does.
Copyright and attribution notices, the licence history and the third-party terms that apply to the published Docker image are recorded in NOTICE.
Already running a version from before the change ? Nothing is withdrawn from you. This project was under CC BY-NC-SA 4.0 until the relicensing tracked in #304, and that licence is irrevocable (§2(a)(1)) : the copies obtained under it keep those terms for good. Everything released from that point on comes under the AGPL, which permits more than the old licence did — commercial use included, for everybody — at the cost of the one obligation in the first table, which never triggers unless you redistribute the program.



