How cpuidle works on Linux and why it affects resource consumption so much

Last update: 11/05/2026
Author Isaac
  • cpuidle separates policy and mechanism by using governors and drivers to manage CPU sleep states.
  • State selection is based on target residence, output latency, inactivity history, and upcoming timers.
  • PM QoS and the tick scheduler condition which deep states can be used without breaking latency requirements.
  • In ARM and other modern SoCs, cpuidle is integrated with firmware via PSCI, being key to real power consumption and battery life.

cpuidle linux idle time management

If you use Linux on a laptop, desktop, or ARM motherboard and you're worried about battery, heat, or why your CPU doesn't go to sleep "as it should", understand how the subsystem works cpuidle It's key. Behind something seemingly as simple as "the CPU is idle" lies a rather sophisticated mechanism that decides what state of rest to use, how long to sleep, and how long it takes to wake up.

Furthermore, if you come from projects like Asahi Linux on Mac with M1/M2It's normal to be confused: we're talking about drivers of cpuidle...from immature or completely absent resting states, and that this is precisely what prevents the system from being used as a daily driver. In this article, we will calmly break down... What exactly is cpuidle, how does it work internally, what role do governors and drivers play, how is it configured, what kernel options affect it, and how does it integrate into modern ARM platforms? such as those that use PSCI or TF-A.

What is the cpuidle subsystem and why does it exist?

Decades ago, the kernel's "idle" state was an empty loop : when there was nothing to execute, the idle loop ran, essentially an infinite loop waiting for the next interrupt. Simply by not executing complex code, some energy was saved: the cache, FPU, etc., weren't used as much.

With the evolution of hardware, processors began offering multiple idle states (C-states, idle states) , each with different levels of energy savings and penalties: entering a deep state can save a lot of energy, but it takes more time and energy to enter and exit it. If you enter too deep a state and are woken up too quickly, you've lost.

This is where cpuidle comes into play , the kernel subsystem dedicated to managing CPU idle time . Its purpose is to decide, whenever a CPU runs out of tasks (only the idle task remains), the best sleep state it can use to save power without breaking response latency.

Conceptually, cpuidle separates two pieces: on one hand the mechanism (drivers) , which knows how to talk to the hardware and enumerate the idle states; and on the other hand, the policy (governors) , which decides which specific state will be used at any given time, based on inactivity history, upcoming timers, and latency constraints.

All this happens in the idle loop : when the scheduler sees that a CPU has no more runnable tasks, it executes the special "idle" task, whose code first calls the governor to choose state and then the driver to enter it.

cpuidle linux governors and drivers

Logical CPUs, idle task, and what it means to be idle

The cpuidle subsystem always works in terms of logical CPUs , that is, the entities that the scheduler sees: they can be physical cores, hardware threads (hyper-threads) or combinations, depending on the architecture and implementation.

From the kernel's perspective, a logical CPU is "idle" when it has no executable tasks associated with it other than the idle task itself . The scheduler manages processes and threads as "tasks," and these can be in different states; when a task becomes runnable, it is assigned to a CPU. If a CPU is left with only the idle runnable task, the kernel considers it inactive.

The idle task executes the call idle loopThis loop, in each iteration, calls a cpuidle governor to decide on a sleep state and then invokes the cpuidle driver to request the hardware to enter that state. If no sleep states are available, there is not enough time before the next event, or the latency constraints are too strict, the CPU either executes a relatively useless loop or uses the most basic wait instruction (like `wait`). hlt or similar).

In multi-core or SMT processors, idle decisions affect unit hierarchies : a thread-level sleep request can cause its core to enter a deeper state if all threads are idle, and in turn, a cluster of cores can go to sleep if all its members allow it. Cpuidle needs to model this through states that represent combinations of hierarchical levels , with latencies and residencies that reflect the deepest possible state.

When a state represented by an object of type is requested struct cpuidle_stateThe driver can authorize the hardware to go as deep as the design allows. Therefore, The exit latency and target residency must be aligned with the real worst-case state combination within the hierarchy (core, cluster, package, etc.).

CPUIDLE governors: how they decide which state to use

CPU Idle Governors are policy modules that run whenever a CPU enters the idle loop. Their purpose is to use available information to select the idle state that saves the most energy without violating latency constraints.

Each governor is defined as a structure struct cpuidle_governorwith callbacks enable, disable, select y reflect, a priority field (rating) and a name. Once registered with cpuidle_register_governor(), can be chosen automatically by the kernel (based on rating, default configuration or parameter cpuidle.governor=) or manually from user space via sysfs.

When a governor is activated for a CPU via its callback enable(), receives a struct cpuidle_device that represents that CPU and a struct cpuidle_driver with the list of available states (struct cpuidle_state). Yes enable() It fails; the kernel uses a default idle code. architecture-specific instead of cpuidle for that CPU.

  Right Mouse Button Not Working. Causes and Solutions

The heart of the governor is the callback select()which receives the cpuidle device, the driver, and a pointer to a boolean stop_tickThis callback returns the index of the chosen state within the array of states, or a negative error code. Additionally, it can decide whether to stop the tick scheduler on that CPU (clearing the boolean or not).

When the CPU wakes up, the governor receives the call to reflect() with information about which state was chosen and how long the downtime actually lasted. This allows you to refine their predictions based on historical dataIt is also required to respect PM QoS latency restrictions: by cpuidle_governor_latency_req() obtains the effective latency limit and should never select a state whose exit_latency exceed that value.

Main governors: ladder, haltpoll, menu and teo

In Linux, there are several cpuidle governors, each with its own strategy and target audience. The four main ones are ladder, haltpoll, menu, and teo , and which one is chosen by default largely depends on whether the kernel is tickless (capable of stopping the tick scheduler) or not.

Ladder is designed for systems with active periodic ticking. It uses a simple approach based almost exclusively on idle duration history : it moves up and down a "ladder" of states, promoting to deeper states when it detects sufficiently long inactivity and retreating when wake-ups occur too soon.

haltpoll is a specialized governor for virtual machines. Instead of entering deep hardware idle states, it relies heavily on polling (waiting loops) to reduce apparent latency in environments where physical idle states may not contribute much or are not well modeled by the hypervisor.

menu y teo These are the governors used in tickless systems (CONFIG_NO_HZ_IDLE o CONFIG_NO_HZ_FULLBoth combine the History of idle periods with information on the next timerand they try to save energy without spending more time computing the decision than is saved by making it.

The governor menu attempts to explicitly predict how long the CPU will be idle . It uses the time until the next timer as an upper limit and applies a correction factor based on recent idle history to approximate a "typical" duration. With this predicted value, it consults the state table to select the deepest state whose target duration matches.

The TEO (Timer-Event Oriented) governor tackles the problem differently: instead of trying to predict the exact downtime, it quantifies historical data in "bins" or intervals associated with each state . Each bin corresponds to the range of times where that state is typically optimal. TEO maintains metrics for hits (wake-ups where the actual duration closely matches the target residence time) and intercepts (wake-ups due to untimed events that disrupt the prediction) and, with this information, directly infers which specific state is most likely to be the correct one.

Only when it's advantageous does TEO check the time until the next timer to avoid entering a state whose target residency is greater than the actual available window . Furthermore, its design attempts to reduce decision costs: in scenarios with very short idle times, it's better to quickly choose a shallow state than to expend more energy thinking than is saved.

cpuidle drivers: bridge between the kernel and the hardware

While the governors are busy with politics, the cpuidle drivers implement the actual state input/output mechanismEach driver represents the list of states supported by the processor (or by a set of CPUs) through a struct cpuidle_driver that contains the array of struct cpuidle_state.

Each struct cpuidle_state defines, among other fields, the target residence (target_residency in microseconds), the maximum output latency (exit_latency), flags such as CPUIDLE_FLAG_POLLING and, very importantly, a callback enter() which is the one that does the delicate part: executing the instructions or calls necessary to ask the hardware to enter that state.

The entries in the state array must be sorted by ascending target_residency , so that index 0 usually corresponds to the shallowest (and cheapest to use) state, and subsequent indices are associated with increasingly deeper states. Governors assume this sorting for their calculations.

The callback enter() receives the cpuidle device, driver, and state index to useFor cases of suspension type suspend-to-idle, it is used instead enter_s2idle()which must comply with stricter restrictions: it cannot reactivate interrupts or manipulate timing devices during its execution, something that enter() Yes, it could be done depending on the platform.

In addition to describing the states, the driver must record which CPUs are under its control: each CPU has its own struct cpuidle_device, which is normally registered with cpuidle_register_device()If there are no "coupled" states (which require coordination between multiple CPUs), the register is done with cpuidle_register_driver()If there are any, they are used cpuidle_register()which is also responsible for registering the devices.

Modern platforms try to reduce the number of specific drivers: for example, ARM typically uses a generic driver that delegate to standard interfaces such as PSCI (Power State Coordination Interface)And in RISC-V something similar via SBI (Supervisor Binary Interface). Even so, drivers like these still exist in x86. intel_idle (with the state table "burned" into the driver) and acpi_idle (which obtains the states from the ACPI tables).

Inactivity states: parameters, sysfs and metrics

Each resting state that cpuidle exposes to governors is characterized by several main parameters. The two most important are the target residence (target_residency) and output latency (exit_latency), both in microseconds.

La target_residency In practice, it marks the energetic depth of the state.This is the minimum time the hardware must remain in that state (including the entry cost) to be worthwhile compared to shallower states. If the system wakes up before reaching that residence time, it has likely expended more energy entering that state than it has saved.

  Find out how to get the most out of Termux on Android

La exit_latency It sets the worst-case time that elapses between the CPU requesting to exit that state and actually executing the next useful instruction.It also includes the case where a wakeup event arrives while the hardware is still entering the state, because in general the internal transition must be completed before it can exit in an orderly fashion.

In addition, there are flags that describe additional properties of the state: for example, CPUIDLE_FLAG_POLLING indicates that this “state” is not actually a hardware restbut a polling loop that is used as a special mechanism to avoid touching the real states when convenient (for example, in certain virtualized or debugging environments).

The kernel exposes very detailed information about CPU idle states through sysfs, in /sys/devices/system/cpu/cpu<N>/cpuidle/There are directories there. state0, state1etc., one for each input of the driver array, and in each one we can find attributes such as name, desc, latency, residency, usage, time, power, above, below y rejected.

These attributes allow us to see how many times has each state been requested (usage), how much total time has been spent on it according to the kernel (time)and to what extent the choice was good or bad (above y below They count cases where the actual idle duration was clearly too short or too long relative to the target residence). rejected It counts the times when the request was rejected, typically because an interruption occurred right at the time of the transition.

There is one particularly useful attribute, disableThis allows you to activate or deactivate that specific state for a CPU from user space (by typing 1 or 0). If disabled for a CPU, the governor will not consider it when selecting; if you want to completely remove a system state, you must disable it on all CPUs. Attribute default_status indicates whether the state is enabled or not by default.

The tick scheduler and tickless systems

The famous scheduler tick It is a periodic timer (100, 250 or 1000 Hz, depending on CONFIG_HZ) that the kernel uses, among other things, to allocate CPU time among tasks, update counters, and trigger timer expirations.

From cpuidle's point of view, the periodic tick is a nuisance: while it is active in a CPU idle, that CPU cannot sleep longer than the tick period , and furthermore, each wake-up per tick involves entering and exiting a sleep state, wasting energy if too deep a sleep state is chosen.

By definition, on a CPU that is only running the idle loop It is not strictly necessary to keep the tick For CPU sharing: there are no more runnable tasks. Therefore, Linux can be configured as tickless in idle (CONFIG_NO_HZ_IDLE) or in “full” mode when there is only one isolated task on the CPU (CONFIG_NO_HZ_FULL), deactivating the tick under those conditions.

The decision to stop the tick or not is made by the governor using the parameter stop_tick from your callback select()If you expect an interruption (timer or otherwise) in the short term (within what would be a tick period), There's no point in deactivating the tick.Time would be spent reprogramming it, and the downtime period would possibly be spent in too superficial a state if it turns out that nobody wakes up the CPU.

Conversely, if the governor believes the CPU will be idle for longer than the tick and the chosen state is deep, It's best to stop the ticking so as not to ruin the savingsSome kernel configurations (parameter nohz=off or disable CONFIG_NO_HZ_IDLE) force the tick to never stop, in which case the system is not tickless and the default governor is usually ladder instead of menu or teo.

PM QoS: How to control resting latencies

The Power Management Quality of Service (PM QoS) framework allows drivers and user space processes to express constraints on the system's energy behavior, particularly on input/output latencies of sleep states.

For cpuidle there are two main types of restrictions: a global CPU latency limit and restrictions of CPU resume latency (pm_qos_resume_latency_us)Internally, requests are stored in priority lists and the effective value is, in this case, the minimum of all those requested.

From user space, the global limit can be modified by opening /dev/cpu_dma_latency and writing in that descriptor a 32-bit integer with the maximum tolerated latency in microseconds. Each open descriptor represents an independent requestWhen it closes, that request disappears and the system recalculates the effective value with the rest.

For CPU restrictions, there is a file power/pm_qos_resume_latency_us en /sys/devices/system/cpu/cpu<N>/Writing a value there changes the request associated with that specific CPU (shared by the entire user space, so it's advisable to arbitrate who accesses it). Kernel drivers can also register their own requests via the internal PM QoS APIs.

The governors of cpuidle must, in each state selection, respect the minimum between the effective global latency and that of the affected CPUThey cannot choose states whose exit_latency exceed that limit. This is extremely important in cases of software with soft real-time requirements, such as audio or video, where a resumption that is too slow from a deep state could cause underruns or glitches.

There is also another QoS, cpu_wakeup_latencywhich affects the choice of idle states during the mode suspend-to-idle (s2idle) of the system. Its handling, from user space, is similar to that of cpu_dma_latency and it is also expressed in microseconds.

cpuidle control via kernel parameters

Linux allows you to adjust the behavior of cpuidle and idle drivers from the kernel command line. The most drastic parameter is cpuidle.off=1This completely disables the subsystem: the idle loop still exists, but cpuidle governors and drivers are no longer invoked, and instead the architecture's "default" mechanism is used, which is usually much simpler and less efficient.

  What Is DriverMax. Uses, Features, Reviews, Prices

Parameter cpuidle.governor=<nombre> It allows you to force the governor to use, for example cpuidle.governor=menu o cpuidle.governor=teo, instead of the one that would be automatically chosen. This is useful for experimenting with power consumption and latency on the same hardware without recompiling the kernel.

In x86 architectures, there are also specific parameters related to how idle mode is entered. For example, idle=halt y idle=poll They disable the drivers. intel_idle y acpi_idleforcing the system to use the instruction hlt or a pure polling loop for rest, respectively. This simplifies the behavior but at the cost of efficiency: idle=pollIn particular, it can prevent the use of P-states that require CPUs to be idle, worsening power consumption and single-threaded performance.

Parameter idle=nomwait prohibits the use of the MWAIT instruction to enter sleep states, forcing acpi_idle to use hlt and deactivating intel_idle In Intel processors, only ACPI manages the states. Furthermore, the drivers intel_idle y processor (the latter includes) acpi_idle) accept options such as intel_idle.max_cstate=<n> y processor.max_cstate=<n> to narrow down the list of available states in the driver and discard all those that are deeper than the specified index.

In the case of intel_idle.max_cstate=0This is equivalent to disabling the Intel driver specifically and allowing the ACPI driver to take over, while processor.max_cstate=0 is interpreted as processor.max_cstate=1These are useful options for to diagnose stability problems, abnormal latencies, or unusual power consumption, at the cost of limiting the system's ability to save energy at idle.

Integration with ARM, PSCI and standby platforms

On modern ARM platforms (such as many SoCs from TI, NXP, Rockchip, Apple via Asahi, etc.), cpuidle is typically integrated with low-level firmware via PSCI and is common in IoT devices with intelligent IoT service management . The generic ARM driver for cpuidle communicates with PSCI using Secure Monitor Calls (SMC) and delegates the actual state transitions to Arm Trusted Firmware (TF-A) or equivalent firmware.

A typical example is a SoC like the AM62x, where the standby state is implemented as a CPUIdle state based on the WFI (Wait For Interrupt) instruction . From the user's perspective, the system enters and exits standby continuously, many times per second, without requiring any interaction: this is the default "light" sleep state, with entry and exit times on the order of microseconds.

The execution path when the system enters idle mode on these platforms is approximately: The idle loop detects that there are no tasks, and the governor chooses the corresponding state (for example, one called stby)The generic ARM driver calls PSCI through the layer drivers/firmware/psci.c, and on the TF-A side the handler is invoked cpu_standby() defined in the structure plat_psci_opsThat's where it really happens. WFI.

When an interrupt arrives, the processor automatically exits WFI, TF-A returns control to the kernel, and the CPU continues execution from where it left off. All of this is orchestrated transparently, provided the device tree accurately describes the idle states (node idle-states, references from each CPU, and properties such as entry-latency-us, exit-latency-us y min-residency-us).

This lightweight CPU standby should not be confused with full system-level deep sleep modes , where entire power blocks are shut down, peripherals are reconfigured, and input/output times are in the millisecond or second range. CPUidle states are more geared towards micromanaging short-term sleep , while deep sleep modes require additional coordination from runtime PM, driver suspend/resume, and often specific support in the bootloader or firmware.

Once the kernel has cpuidle enabled and a working driver (for example, the generic ARM + PSCI driver correctly described in the DT), The user does not need to "activate" anything.The governors handle it automatically. However, you can consult statistics in /sys, change the current governor in /sys/devices/system/cpu/cpuidle/current_governor or adjust the properties of the states exposed by the driver.

On relatively new platforms or in porting processes like Asahi Linux, many battery life, overheating, or lack of functional sleep mode issues are linked to the fact that cpuidle's full integration with the hardware is not yet complete (or mature) : well-described states are missing from the development documentation, there's no firmware that implements reasonable states, or the driver is still under development. Until this stabilizes, it's common to only use a very superficial Wi-Fi state or, even worse, to disable cpuidle altogether, resulting in a rather rudimentary sleep mode.

Ultimately, understanding how cpuidle models logical CPUs, how governors make decisions regarding target residencies, latencies, and PM QoS , how drivers connect with ACPI, PSCI, or proprietary firmware, and how the tick scheduler conditions idle periods, allows you to see in a new light why your laptop lasts longer (or shorter) on battery life, why a "tickless" kernel usually consumes less power, and why, in certain ports to exotic hardware, sleep and idle are literally the difference between a test system and a usable daily driver.

Configure advanced power policies with Power Profiles and Modern Standby
Related articles:
Configure advanced power policies with Power Profiles and Modern Standby