Architecting Near-Optimal Leakage Efficiency: The Power of Application-Driven SPM Subbanking
Architectural Leakage Power Minimization of Scratchpad Memories by Application-Driven Sub-Banking
The paper proposes an application-driven subbanking technique for Scratchpad Memories (SPMs) to minimize leakage power in embedded systems. It introduces an optimal partitioning algorithm that exploits the idleness of memory blocks by placing them into state-preserving low-leakage sleep modes.
TL;DR
As technology scales below 65nm, leakage power has overtaken dynamic power as the primary energy bottleneck in embedded systems. This paper introduces a breakthrough architectural technique to partition Scratchpad Memories (SPMs) into subbanks that can independently enter low-leakage sleep states. By refining the search space through a proven "v-range" property and employing memoization, the authors achieve up to 89% energy reduction with virtually zero performance impact.
The Leakage Crisis in Deep Submicron Design
In the world of embedded systems, Scratchpad Memories (SPMs) are the preferred alternative to caches due to their deterministic behavior and lack of complex mapping logic. However, being the largest on-chip structures, they are also the biggest "leakers."
The Flaw in Prior Work
Most earlier subbanking efforts targets Dynamic Power. Their logic was simple: split memory so that frequent accesses stay in small, low-capacitance banks. But in modern 65nm+ nodes, memory consumes power even when it is not being accessed. If we ignore static power during partitioning, we risk creating "optimized" structures that actually consume more total energy because inactive banks stay in a high-leakage active state.
Research Insight: Why Most Barriers are Useless
The search space for partitioning a memory of words into banks is . For a standard SPM, this is astronomically large.
The authors' core contribution is the v-range property. They mathematically prove that the optimal energy boundary between two subbanks must coincide with a change in the access pattern (idleness distribution). They call these segments "v-ranges." Instead of checking every address, the algorithm only checks the boundaries where an address range's "idle profile" changes.
Fig 1: The state machine for SPMs - Active vs. Sleep. The "Break-even" point determines if the idle period is long enough to offset the Sleep-to-Active energy cost.
Methodology: The Optimal Search Algorithm
The solution involves two distinct phases:
- Discretization and v-range Identification: The address map is scanned to identify intervals with identical temporal behavior. This reduces the number of candidate boundaries from millions to a few hundred.
- Memoized Exhaustive Search: To avoid redundant calculations, the algorithm stores the energy cost of every sub-range in a lookup table.
- Complexity: Shifted from to where is the number of v-ranges and is the trace length.
Fig 2: Visualization of partitioning logic. By grouping highly accessed addresses into smaller banks and idle ones into larger banks, both dynamic and leakage power are minimized.
Experimental Results
Using the MiBench suite and an industrial 65nm STMicroelectronics library, the authors demonstrated:
- Energy Savings: Average of 60% total energy reduction.
- Performance Overhead: Approximately 0.01%, as the 1-cycle reactivation delay is rare relative to the total execution time.
- Scalability: The speedup provided by the memoized search was up to 85,554x (for the
ispellbenchmark) compared to standard exhaustive search.
Fig 3: Technology characterization showing linear-exponential scaling of leakage based on Vdd and memory size.
Final Thoughts: A Blueprint for Future SPMs
The beauty of this work lies in its transparency. It does not require modifying the internal SRAM bitcells or changing the compiler toolchain. It is a purely architectural optimization that leverages mathematical proof to solve what was previously an intractable search problem.
While modern designers are shifting toward more dynamic power management (DPM), this "static partitioning" approach provides a foundational limit on how efficient a memory hierarchy can truly be when matched perfectly to its workload.
