Embedded Memory Testing: MBIST, Fault Coverage, Diagnosis, and Repair - Yenra

Separate embedded-memory test, fault coverage, diagnosis, repair and ECC, using a small worked MBIST example and a practical test-plan checklist.

Conceptual embedded memory grid with one amber cell and separate spare cells
A conceptual memory test highlights a fault; usable diagnosis and repair depend on the implemented architecture.

Memory built-in self-test, or MBIST, lets on-chip logic exercise embedded memories with controlled operations and compare the results. A useful test plan specifies the faults it targets, the operating conditions, how failures are diagnosed and what repair resources exist. “Test passed” is meaningful within that defined scope.

Separate test, diagnosis, repair and correction

Siemens describes Tessent MemoryBIST as supporting embedded-memory test, diagnosis and repair. These are related stages with different outputs. A test detects behavior inconsistent with expectations; diagnostic information helps locate and classify it; a repair flow uses resources designed into the memory architecture.

On narrow screens, focus the table and use the arrow keys or swipe to see every column.

What each mechanism contributes
MechanismTypical purposeQuestion to resolve
MBISTGenerate operations and compare responsesWhich faults and conditions does this test target?
DiagnosisRetain information about failing operations or locationsIs there enough detail for analysis and repair allocation?
Built-in self-repairApply a supported repair mapping using available resourcesCan the actual fault pattern be covered by those resources?
ECCDetect or correct errors according to the code's capabilityHow does correction interact with the test's observations?

Siemens' discussion of ECC in memory test and repair highlights the interaction between these mechanisms. Define whether a test observes raw failures, corrected data or error status. Otherwise an apparently successful read may answer a different question from the one the test engineer intended.

A small example shows both detection and limits

Four one-bit cells, one stuck-at-zero fault

Consider addresses 0 through 3. Write zero to every address and read each back; all four return zero. Then write one to every address and read each back. If address 2 is stuck at zero, its second read returns zero when one is expected, so the comparison flags a failure.

This toy sequence performs 4 + 4 + 4 + 4 = 16 operations. Assuming exactly one operation per 100 MHz clock cycle, its idealized execution time is 160 nanoseconds. Setup, access latency, controller overhead and result collection are excluded.

The example illustrates a particular observable fault. It does not establish coverage for interactions between cells, address-decoding defects, retention loss or timing-dependent behavior. Each additional fault model needs an appropriate sequence and a reason that the sequence makes the fault observable.

Real memories add width, depth, ports, masks, latency and access restrictions. Translate the algorithm into legal operations for each memory instance. Record any disabled operations or constrained conditions because they alter what a passing result establishes.

Build a test contract

  1. Inventory: list the memory instances, dimensions, ports, clocks, power domains and repair features.
  2. Coverage objective: name the fault models and algorithms, with supporting qualification evidence.
  3. Execution conditions: define clocks, voltage and temperature conditions, allowed concurrency and power limits.
  4. Observability: specify pass/fail signals, diagnostic detail, ECC treatment and how results leave the device.
  5. Repair and retest: define resource allocation, mapping persistence and the tests that validate the repaired configuration.

Treat test patterns as destructive unless the specific implementation and procedure explicitly preserve application data. Manufacturing test, boot-time checks and in-service diagnostics have different opportunities to interrupt normal use. The firmware and system owners need to know when contents may change and how normal state is restored.

Report what was actually exercised

A useful report identifies the design revision, memory instances, algorithm configuration, conditions, diagnostic result and any repair applied. Record excluded memories and incomplete tests explicitly. Separate an unrepaired failure, an exhausted repair budget and a test-access problem; they call for different engineering responses.

After a repair mapping is applied, rerun the required checks and retain both pre-repair and post-repair results. For a field diagnostic, define the response to failure before deployment: logging, degraded operation, reset or service escalation should follow the product's own reliability requirements.

The MRAM guide adds the application-level problem of persistent-state recovery. Hardware test results and a sound update protocol address complementary parts of a dependable memory system.