Case Study on Hardware Reliability Failures in Memory Systems
As DRAM The craftsmanship has been miniaturized to the Fly Method ( fF )‑level capacitors, the physical reliability of memory is facing unprecedented challenges. This paper focuses on… 2025 To 2026 Some key failure cases disclosed over the years have been systematically reviewed, starting from “line interference ( RowHammer ) to “column interference ( ColumnDisturb The evolution of the read‑interference threat model has been thoroughly analyzed. The aim is to provide engineers and manufacturers with a reference guide on the latest hardware failure modes, root‑cause analyses, and mitigation strategies.
I. Read-disturb failure: from Row-to-Column Dimension Expansion
1. ColumnDisturb: A Paradigm Shift from “Row Interference” to “Column Interference”
(1) Phenomenon
In October 2025, at the 58th IEEE/ACM International Symposium on Microarchitecture (MICRO), researchers conducted the first experimental demonstration of a novel, widespread read‑disturb phenomenon—ColumnDisturb. Unlike RowHammer and RowPress, which affect only a few rows adjacent to the target row, ColumnDisturb exploits repeated activation or retention of a specific row (the target row) to induce disturbances across DRAM columns (i.e., bitlines), thereby impacting memory cells sharing the same column as the target row—and it can propagate across multiple DRAM subarrays.
Specifically, activating a single row can simultaneously disrupt DRAM cells in up to three subarrays—on the DDR4 chips tested, this could affect as many as 3,072 rows. The research team conducted a comprehensive characterization using 216 DDR4 chips and four HBM2 chips sourced from the world’s three largest DRAM manufacturers.
Key experimental findings include:
ColumnDisturb affects all test chips from the world’s three major DRAM manufacturers and becomes more pronounced as DRAM technology nodes shrink.
Even in existing DRAM chips, ColumnDisturb can induce multiple bit flips within the standard DDR4 refresh window (63.6 ms);
Beyond the standard refresh window, the number of rows experiencing bit flips induced by ColumnDisturb is up to 198 times greater than that of rows where the hold‑off failure occurs.

(2) Analysis
The interference paths of RowHammer and RowPress are row-to-row—inducing disturbances in adjacent rows via capacitive coupling between wordlines. ColumnDisturb, on the other hand, reveals another physical pathway: row-to-column.
When an attack row is repeatedly accessed or kept open for an extended period, voltage fluctuations on the bitline propagate through the column‑select path, affecting the memory cells in other rows of the same column. Because a single column in a subarray spans a large number of rows, the scope of the attack expands sharply from “a few adjacent rows” to “hundreds or even thousands of rows.”
(3) Summary
The threat model for read disturbances needs to be completely rewritten. The assumption from the RowHammer era that “only adjacent rows are affected” has been invalidated. Manufacturers must assess the effectiveness of existing protective mechanisms, such as PRAC and RFM, against ColumnDisturb.
The perceptual refresh mechanism requires a reassessment. Such mechanisms sacrifice some of the refresh margin in favor of performance and power efficiency, and ColumnDisturb precisely exploits this “margin.”
It is recommended to incorporate the ColumnDisturb test into the standardized quality‑inspection procedures for subsequent DDR6 and DDR5 steppings.
2. LeakyHammer: The RowHammer protection mechanism itself becomes a side channel.
(1) Phenomenon
The LeakyHammer study, submitted in March 2025 and formally published at MICRO 2025 in October, for the first time systematically analyzes the timing‑based covert channels and side‑channel vulnerabilities introduced by RowHammer mitigation mechanisms.
RowHammer mitigation mechanisms—such as PRAC and RFM, introduced in DDR5—temporarily block access to DRAM when preventive actions are triggered, resulting in a measurable and significant increase in memory access latency. Attackers can exploit this by crafting specific memory-access patterns to invoke these preventive measures on demand, then use high‑precision timers to measure the resulting latency changes for encoding and transmitting information.
Specifically, the researchers constructed covert-channel attacks targeting two mainstream protection mechanisms—PRAC and RFM—separately:
The covert channel capacity of PRAC reaches 39.0 Kbps;
The covert channel capacity of RFM reaches 48.7 kbps.
As a proof of concept, the researchers demonstrated a website fingerprinting attack—identifying which websites a user has visited by analyzing differences in the RowHammer mitigation response patterns triggered when accessing different sites.

(2) Analysis
LeakyHammer is rooted in a classic problem in computer architecture: timing side channels. The RowHammer mitigation mechanism exhibits two exploitable characteristics:
Bandwidth Preemption: After PRAC/RFM is triggered, the memory controller issues additional refresh commands, forcing the attacker’s subsequent requests to queue and wait.
Inducibility: The effectiveness of the defensive measure is highly dependent on the attacker’s controllable address-access patterns.
(3) Summary
“Protective measures” themselves must be incorporated into security threat modeling. Manufacturers should not merely assess their resistance to interference; they must also quantitatively evaluate the risk of information leakage.
“Covert protection” is a core challenge in next-generation design. Ideally, protective operations should be time‑invariant, but this would impose substantial performance overhead. Manufacturers must clearly communicate to customers the engineering trade-offs between performance and security.
It is recommended to introduce randomized jitter. Adding a small random delay when the protection is triggered can effectively reduce channel capacity.
3. A Novel RowHammer Attack Framework—Bypassing the Defenses of Next-Generation Intel Platforms
(1) Phenomenon
In October 2025, a domestic research team presented a paper at MICRO 2025, proposing a novel attack framework to address the failure of existing RowHammer attacks on next-generation Intel CPU architectures.
This study systematically addresses the three core challenges associated with launching attacks on new platforms:
DRAM Address Mapping Reverse Engineering: A highly efficient, general-purpose method for reverse-engineering DRAM address mappings is proposed, capable of deriving increasingly complex, complete mapping relationships within seconds.
Attack Paradigm Innovation: For the first time, an attack paradigm based on x86 software prefetch instructions has been proposed, breaking through the activation‑rate bottleneck of conventional attacks.
Platform Compatibility: The memory controller scheduling policy has been specifically optimized for the latest-generation Intel CPUs.
(2) Analysis
Although the new-generation Intel platform has strengthened its RowHammer protection, its protective logic still relies on an assumption about the physical address mapping of DRAM—that is, the assumption that an attacker cannot precisely determine which physical rows are adjacent. This study dismantled that assumption through high-speed reverse engineering.
The introduction of software prefetch instructions addresses another bottleneck in traditional attacks: the activation rate generated solely by load instructions is limited, whereas prefetch instructions can more efficiently populate the memory controller’s request queue.
(3) Summary
The “covertness” of DRAM address mapping should not be regarded as a security mechanism. This study demonstrates that, regardless of how complex the mapping algorithm is, an attacker can reverse it within seconds.
The attack surface at the instruction set architecture (ISA) level requires reevaluation. The x86 prefetch instruction, originally designed for performance optimization, has now been weaponized as an enhancement tool in the RowHammer attack.
II. DDR5 power management failure (PMIC-related fault)
1. HPE Gen11 Server PMIC PWR_GOOD Error — Logging Failed
(1) Phenomenon
HPE’s official defect database documents a known issue with Gen11 servers: when the PMIC on a DDR5 DIMM malfunctions, the PWR_GOOD signal is pulled low (open-drain output—floating/high‑impedance under normal conditions, pulled low in case of failure), causing the system to shut down immediately to prevent potential damage.
The PMIC will attempt to write the error status to non-volatile memory (the error status register). However, the VIN_Bulk discharge time may vary depending on the system configuration or the load of the DDR5 DIMM itself, and it cannot be guaranteed that PMIC errors will be successfully recorded during a power-off event.
Operations personnel can only see “Server Critical Fault” entries in the iLO (Integrated Lights-Out) IML (Integrated Management Log).

(2) Analysis
The DDR5 PMIC monitors the VIN_Bulk input voltage and all output regulators (VOUT_A, VOUT_B, VOUT_C, VOUT_D, VOUT_1.8V, VOUT_1.0V, and VBias). Any abnormality will trigger a low‑level PWR_GOOD signal. The issue is that, during a power‑off event, the PMIC may not have sufficient time to complete the nonvolatile write of the error registers.
(3) Summary
The PMIC shall ensure that, after the PWR_GOOD signal goes low, there is sufficient hold time to complete error log writing.
For deployed systems, it is recommended to provide a standardized mapping table that correlates “Server Critical Fault” events to specific DIMMs.
It is recommended that PMIC manufacturers consider using a supercapacitor to provide an independent power supply for writing error logs at the moment of power loss.
2. DDR5 voltage telemetry all zero—PMIC communication compatibility issue
(1) Phenomenon
In January 2026, a user on the HWiNFO forum reported that, among four Kingston DDR5‑5200 32GB memory modules, DIMM #0 showed all voltage readings at 0.000V—VDD, VDDQ, VPP, 1.8V, 1.0V, and VIN were all zero. However, the SPD Hub temperature remained normal (38–39°C), and the module functioned perfectly—MemTest86 ran overnight without any errors.
The author of HWiNFO notes that this may be due to the simultaneous operation of other monitoring software, which can cause bus conflicts among multiple tools.
(2) Analysis
A zero voltage reading with normal functionality indicates a problem in the PMIC’s internal ADC read path or in I²C/I³C communication. Specifically, the issue could be:
Bus conflict: Multiple monitoring software applications simultaneously access the PMIC via SMBus/I²C.
Firmware Compatibility: Certain PMIC implementations are incompatible with the access patterns of some monitoring software.
Software bug: The PMIC register parsing logic in a specific version of the monitoring software contains a defect.
(3) Summary
The reliability of telemetry data is the foundation of system health monitoring. Manufacturers should ensure that the PMIC’s voltage and temperature telemetry channels continue to function properly even when multiple monitoring software applications access them concurrently.
It is recommended to clearly specify the access rules for telemetry registers in the PMIC datasheet.
Manufacturers should establish communication channels with monitoring software developers, such as HWiNFO.
III. Serial Presence Detect ( SPD) and firmware failure
1. SPD read failure — HWiNFO detected loss of SPD information
(1) Phenomenon
In 2025, a Chiphell community user reported that two stock‑style green‑striped SK Hynix DDR5 modules stopped reading SPD information within 10 minutes of booting the system. The user found that only HWiNFO triggers this issue. After the SPD data is lost, only Zentiming can still retrieve memory information.

(2) Analysis
The SPD EEPROM communicates with the system via the SMBus (based on I²C). Non‑compliant software that accesses SMBus devices without adhering to the proper read/write protocols may inadvertently overwrite or corrupt data stored in the SPD EEPROM. Furthermore, if the module does not have JEDEC‑defined Reversible Software Write Protection (RSWP) enabled, the SPD data remains vulnerable to arbitrary writes.
(3) Summary
SPD write protection (RSWP) shall be included as standard equipment from the factory.
It is recommended that, when selecting an SPD EEPROM, priority be given to devices that support the JEDEC reversible write‑protect specification.
For shipped modules that have not yet had their SPD locked, it is recommended that manufacturers provide a firmware update utility to allow users to manually enable write protection.
2. Startup Failure Caused by SPD Data Anomalies — In-Depth Review of Nine Representative Cases
(1) Phenomenon
In September 2025, a comprehensive research report based on the SPD 4.1.2.M-2 standard was published. Through empirical re‑examinations of nine typical failure cases, the report identified the root causes of startup failures, including SPD data anomalies, timing incompatibilities, and hardware discrepancies.
Typical scenarios include:
SPD data corruption: The EEPROM contents have been corrupted due to abnormal writes or physical damage.
Timing incompatibility: The timing parameter set stored in the SPD exceeds the memory controller’s supported range on the motherboard.
Hardware differences: Memory SPD parameters may vary slightly between different production batches of the same model.
(2) Analysis
The SPD stores critical memory parameters such as timing, frequency, and voltage. The motherboard BIOS reads the SPD to configure the memory controller. If the SPD data is corrupted or incorrect, the motherboard may apply improper settings, resulting in a POST failure or system instability.
(3) Summary
It is recommended to establish an enterprise‑level compatibility validation platform that encompasses automated testing, knowledge graph construction, and intelligent diagnostics.
Add cross-platform compatibility verification to the factory testing process.
It is recommended to add a redundancy check field (such as CRC) to the SPD.
IV. Single-event upset ( SEU)
1. The exponential impact of SRAM cell scaling in 12nm FinFET technology on SEU sensitivity
(1) Phenomenon
A study published in 2025, based on commercial 12nm FinFET SRAMs, conducted experiments using heavy ions with linear energy transfer (LET) values ranging from 0.476 MeV·cm²/mg to 85.59 MeV·cm²/mg, systematically investigating the influence of supply voltage on single-event upset (SEU) sensitivity.
Key finding: Under advanced FinFET processes, the single-event upset cross section increases exponentially as the supply voltage decreases.

(2) Analysis
The physical mechanism underlying SEUs is as follows: when a high-energy particle traverses a semiconductor material, it deposits energy along its trajectory, generating electron–hole pairs. These charge carriers are collected by the PN junction of the storage node, and when the accumulated charge exceeds the critical charge (Qcrit), the logical state of the storage node flips.
The critical charge Qcrit is proportional to the supply voltage Vdd (Qcrit ≈ C_load × Vdd). As Vdd decreases, Qcrit also decreases, making it easier for particles with the same energy to trigger a flip.
(3) Summary
While low-power design enhances energy efficiency, it also significantly increases the risk of SEUs. Manufacturers must make clear engineering trade-offs between low power consumption and high reliability.
For high-reliability applications such as automotive, aerospace, medical, and data centers, it is recommended that manufacturers provide “hardened” product versions that have been certified through radiation‑testing.
V. Supply Chain Security and Emerging Threats
1. Upgraded DDR5 Memory Fraud: The “Hollow Chip” Incident
(1) Phenomenon
In June 2026, as DRAM prices continued to rise, a wave of memory‑fraud incidents came to light. A memory module advertised as “SK Hynix DDR5‑5600 16GB” was being sold for just 179 yuan, but testing revealed that it failed to pass the system’s POST check entirely.
This fake memory boasts an extremely convincing disguise. However, a closer look at its hardware components quickly gives it away:
The surfaces of the eight DRAM chips bear no laser‑etched markings.
Capacitors on the circuit board come loose with just a light touch.
Bulging has been observed in certain areas of the PCB, and the inductor exhibits chipped edges as well as blackened discoloration resulting from high‑temperature soldering.
The most outrageous finding was that, upon cutting open the DRAM die, the interior was completely empty—there were no silicon chips (dies) whatsoever, only a magnetic plate affixed to the bottom of the PCB.


(2) Analysis
From 2025 to 2026, global DRAM prices continued to rise, with the three major manufacturers—Samsung, SK Hynix, and Micron—accounting for roughly 90% of total production capacity. These high prices have prompted illicit actors to exploit various methods for profit:
PCB recycling and refurbishment;
Hollow-shell encapsulation;
DDR4 masquerading as DDR5.
(3) Summary
It is recommended that manufacturers strengthen product anti-counterfeiting measures, such as QR‑code traceability systems and special laser‑etched markings.
Manufacturers should establish a rapid reporting and removal channel with e-commerce platforms for counterfeit products.
It is recommended to include a genuine‑product verification guide on the product packaging and in the user manual.
2. The failure rate of a specific batch of Cisco DIMMs is unusually high.
(1) Phenomenon
In April 2026, Cisco issued vulnerability advisory CSCwb98743 (FN72464): Certain DIMMs from a specific manufacturing batch (with particular date codes) exhibit a higher-than-expected failure rate. The most common symptom is an “Uncorrectable DRAM ECC” error.
(2) Analysis
An abnormally high failure rate in a specific batch of DIMMs typically stems from localized quality issues in the manufacturing process—such as deviations in the die‑sorting stage, process variability during module assembly, or incoming‑material quality problems with components in that batch.
(3) Summary
It is recommended that manufacturers establish a comprehensive batch traceability system.
When an abnormal batch‑level failure rate is detected, a product recall or replacement program should be initiated promptly.
It is recommended to incorporate statistical process control (SPC) into the factory acceptance testing.
VI. On-site failure rate versus industry benchmark
1. Large-Scale On-Site Study of DDR5 DRAM Failures—Failure Rate Lower Than That of DDR4
(1) Phenomenon
In June 2025, a large-scale study on field‑level failures of DDR5 DRAM was presented at the IEEE International Reliability Physics Symposium. The research was conducted within a large‑scale commercial server cluster comprising 64‑GB dual‑rank DDR5 DIMMs.
Key finding: The observed error rate of DDR5’s error-correction capability is lower than that of DDR4.
However, the researchers note that part of the reduction reflects data‑capture bias—DDR5’s on‑die ECC (OD‑ECC) limits the visibility of single‑bit errors and other fault modes. Nevertheless, the rate of uncorrectable errors with OD‑ECC remains lower than that observed in DDR4.
Time-to-failure (TTF) analysis employs the Weibull distribution and aligns with the bathtub curve—two out of three manufacturers are in the early failure phase (infant mortality phase).
(2) Analysis
DDR5 introduces several reliability enhancements at the architectural level:
More granular refresh management (RFM);
On-chip ECC (OD-ECC);
More robust PMIC protection mechanisms.
(3) Summary
The reliability improvements of DDR5 have been validated in real-world deployments.
The identification of early‑failure failures in both manufacturers suggests that they should strengthen their factory burn-in testing.
It is recommended that the manufacturer continuously monitor on-site failure rate data.
VII. Conclusion
Looking back at the evolution of memory hardware reliability from 2025 to 2026, three distinct trends run consistently throughout:
First, the threat model of read interference is expanding from the “row level” to the “column level.” ColumnDisturb’s discovery expands the scope of read‑disturb effects from a few adjacent rows to thousands of rows. LeakyHammer further reveals that the behavioral characteristics of the protection mechanisms themselves can be weaponized as side channels.
Second, the PMIC has become the critical boundary for DDR5 reliability and security. From HPE’s PWR_GOOD log‑loss issue to the unsecured PMIC interface vulnerability tracked as CVE‑2025‑48516, the PMIC is no longer just a power‑management device.
Third, real-world deployment data is becoming the gold standard for reliability assessment. Field failure studies of DDR5 have demonstrated that architecture-level reliability improvements indeed translate into lower failure rates.
For manufacturers, the 15 cases examined in this paper collectively point to four core recommendations for action:
Redefining the Threat Model of Read Disturbance — Expanded from the “row level” to the “column level”;
Design the PMIC as a security boundary. — Enable interface access protection by default;
SPD write protection should be standard equipment out of the factory.
Establish a closed-loop process from on-site failures to product improvements.
For engineers, it is recommended to pay particular attention during system deployment to the following: configuring RFM security mode in the BIOS, verifying the SPD write-protection status, performing consistency checks on PMIC telemetry readings, and monitoring power‑input quality.
The next phase of memory reliability will focus on achieving a systematic balance among physical limits, security threats, and supply-chain risks.
Hot News
2026 MCU Industry Outlook: Market Expansion, Technological Upgrades, and Supply Chain Restructuring
In 2026, the global microcontroller (MCU) industry is expected to maintain steady growth, driven by robust demand across multiple sectors, including automotive electrification, industrial automation, and IoT and AI infrastructure. At the same time, supply-chain constraints and successive price hikes will persist throughout the year. International industry leaders are accelerating their localization strategies, while domestic manufacturers continue to make breakthroughs in high-end offerings and product differentiation.