Fine-Grained Trace & Scan: Redefining Observability in Post-Silicon Debug
14380_Efficient Selection of Trace and Scan Signals for Post-Silicon Debug.
This paper introduces a fine-grained debug architecture for post-silicon validation that intelligently combines trace and scan signals using multiple scan chains with varying dumping periods. The proposed method utilizes two novel algorithms, Constrained Signal Selection (CSS) and Flexible Signal Selection (FSS), to maximize the restoration ratio (RR) of internal chip states.
TL;DR
Post-silicon validation is the "last line of defense" in chip design, yet it is hampered by the "black box" nature of manufactured silicon. This paper proposes a fine-grained hybrid architecture that breaks the binary choice between trace (fast, expensive) and scan (slow, cheap) signals. By multiplexing the trace buffer into multiple scan chains with different "dumping periods," the authors achieve over 125% improvement in state restoration with almost zero impact on chip area.
The Observability Crisis in Modern Chips
As integrated circuits scale, the gap between internal complexity and external observability widens. Pre-silicon simulation can only do so much; bugs often emerge only under real operating conditions. Standard solutions like Trace Buffers store a few signals every cycle, while Scan Chains dump thousands of signals over many cycles.
The Problem with SOTA: Existing hybrid methods are too "coarse." They treat signals as either top-tier (traced every cycle) or bottom-tier (scanned infrequently). This ignores the reality of logic networks, where many signals have intermediate importance—not critical enough to trace every cycle, but too important to wait 100 cycles to see.
Methodology: The Power of Fine-Grained Dumping
The core innovation is a hardware-software co-design.
1. Fine-Grained Architecture
Instead of one big scan chain, the authors divide the trace buffer width () into separate partitions. Each partition can have a different length, creating a spectrum of Dumping Periods ().
- : Standard trace signals.
- : Signals scanned at high to low frequencies.

2. Selection Metrics: RP and RI
How do you decide which signal goes into which "speed" lane? The authors propose two metrics:
- Restoration Power (RP): Used for constrained hardware. It balances the number of states a signal can help restore against the resources () it consumes.
- Restoration Impact (RI): Used for flexible architectures. It measures the absolute gain in restored states during a mock simulation window.
The intuition is simple: if a signal's value allows us to deduce 10 other signals through forward/backward logic implication, it deserves a "faster" lane (lower ).
Experimental Validation
Using the ISCAS’89 benchmark suite, the researchers compared their methods—Constrained Signal Selection (CSS) and Flexible Signal Selection (FSS)—against established baselines.
Performance Gains
The results are transformative. By shifting from a coarse-grained to a fine-grained approach, the Restoration Ratio skyrocketed. In some benchmarks like s38584, the improvement reached 125% compared to prior trace+scan combinations.

Efficiency & Overhead
One might fear such complexity requires massive hardware. However, the logic synthesis using a 45-nm technology node showed:
- Area Overhead: Average of 0.57%.
- Power Overhead: Average of <1%. The "controller" is essentially a small counter and some shift-enable logic, which is negligible compared to the massive logic gates of an industrial SoC.
Critical Insight: Why This Works
The success of this method lies in its alignment with logic topology. In a circuit, some signals are "hubs" with high fan-out/fan-in. Tracing these hubs at a medium frequency is often more efficient than tracing a few signals at 100% frequency or many signals at 1% frequency. This "middle-ground" captures the spatial and temporal correlations of the netlist more effectively.
Conclusion and Future Outlook
This work demonstrates that the efficiency of post-silicon debug is not just about how much data we collect, but the frequency at which we collect specific pieces of it.
Limitations: The current algorithms rely on random input simulations to calculate probability and . In real-world applications with proprietary IP, selecting representative input vectors remains a significant challenge.
Future Work: Integrating these selection algorithms into High-Level Synthesis (HLS) could allow for "visibility-aware" chip design, where the debug hardware is optimized simultaneously with the logic.
