PCBA failure analysis is a controlled process for moving from an observed symptom to a supported cause and an effective prevention decision. It is not simply a repair activity. Replacing a part, reheating a joint or rerunning a test may restore a board, but the result does not explain why the failure occurred, how many units may share the risk, or whether the proposed action prevents recurrence.
The analysis must protect original evidence, define the affected configuration, reproduce the symptom under controlled conditions, isolate the fault, test competing hypotheses and verify the causal mechanism. Manufacturing, design, component, firmware, fixture, test method, environment and handling may interact. The investigation should therefore avoid deciding “process problem” or “design problem” before the evidence supports that boundary.
This guide is for OEM quality, hardware, NPI, supplier-quality and procurement teams working with an EMS provider. It focuses on the handoff, evidence and decision structure needed to contain risk and reach a useful root cause. It does not replace product-safety procedures, laboratory-specific methods or qualified engineering judgment.
Begin PCBA failure analysis by preserving the first failure
When a unit fails, record its serial or board identity, product and BOM revision, PCB lot, firmware and configuration, work order, manufacturing route, station, fixture, program and limit revision, operator or automated identity, time, test conditions, exact failure code, measurements, waveform or image, and every attempt. Preserve customer or field context such as installation, duty cycle, power source, connected equipment, temperature, vibration, moisture, event history and packaging condition.
Do not overwrite the first failure with a final PASS. A reseat, restart or repeat can be diagnostically useful, but every action changes the evidence. Record who acted, what changed and what happened. Avoid cleaning, touching suspect areas, reflowing, probing aggressively, updating firmware or removing parts until photographs, visual observations and non-destructive data are captured.
Containment should follow the shared exposure, not an arbitrary lot label. Identify units that share the relevant PCB, component lot, approved source, paste or flux, line, stencil, oven recipe, time window, operator, fixture, software, rework route, transport or field condition. Hold the smallest defensible population while the hypothesis develops, and expand it when evidence points to a wider cause.
Containment protects evidence and limits exposure before the failure mechanism is known.
Create a chain of custody
Label each sample and control, seal or protect it against unintended handling, and record transfers between production, quality, supplier, laboratory and customer. Note whether the board was powered after the event, whether conformal coating or residue was removed, and whether earlier repair occurred. Keep packaging and mating items when connector damage, contamination, mechanical stress or electrical transients may matter.
Choose known-good and, where possible, known-failed controls from the same and different populations. A comparison board helps determine whether an observed mark, residue, void, measurement or material structure is unusual. Without controls, a visually striking feature can be mistaken for the cause simply because it is easy to see.
Write a testable problem statement
A useful problem statement describes the observable gap without embedding an assumed cause. Write what failed, where, when, under which conditions, how often, in which population and compared with what expected behavior. “U12 bad” is a diagnosis claim. “Serial 418 loses 3.3 V output after 20 minutes at a 1.8 A load while three controls remain within limit” is a testable observation.
Separate confirmed facts, reported observations, assumptions and unknowns. Customer language such as “intermittent,” “dead” or “overheating” needs measurable definition. Translate it into power rails, timing, communication, output accuracy, current, temperature, mechanical connection or another observable response. Define the expected requirement and tolerance source.
Build a timeline from manufacturing through the failure event. Include configuration changes, transport, storage, installation, firmware updates, calibration, power cycles, environmental exposures, repairs and diagnostic steps. Time order can distinguish an original assembly defect from later damage, but correlation alone still does not prove causality.
Reproduce the symptom without destroying evidence
Reproduction confirms that the team is investigating the same failure. Use the documented supply, load, firmware, interfaces, sequence, warm-up, enclosure, orientation and environmental conditions. Instrument the relevant nodes while protecting the assembly. Repeat enough to understand stability, but stop if additional operation could propagate damage or erase a transient signature.
When the failure does not reproduce, do not automatically close it as “no fault found.” Compare the laboratory and original conditions, including connectors, cables, grounding, loads, software versions, timing, vibration, temperature, humidity and installation. Review retained logs and fixture behavior. Define escalation rules for recurring intermittent units and link them across lots or field returns.
Use non-destructive methods first when they can answer the question: visual and microscopic inspection, AOI record review, X-Ray, electrical measurements, curve tracing, boundary scan, ICT/FCT logs, thermal imaging, acoustic microscopy or computed tomography as appropriate. Each method has detection limits. An X-Ray image can reveal geometry or density contrast but may not prove metallurgical integrity, crack propagation or electrical behavior.
A reproduced symptom is useful only when configuration, stimulus, limits and attempts are recorded.
Check the measurement system
A test station can create or hide a failure. Verify instrument calibration and range, fixture contacts, connector wear, software version, timing, reference units, known-failure challenges, ground paths and operator loading. Compare the suspect unit on an independent setup when practical. If moving the board between stations changes the result, preserve that fact instead of choosing the preferred reading.
Review raw values near limits, not only PASS/FAIL. An unstable contact, marginal guard band or software timeout can show as repeated retries. The original result, each attempt and any limit change must remain linked. The GNS article on auditing PCBA inspection and test evidence explains the record fields a buyer should request.
Isolate the fault and test competing hypotheses
Fault isolation narrows the failing function to a circuit, connection, component, solder joint, firmware state, mechanical interface or interaction. Use the schematic, layout, BOM, signal flow, power tree and product specification. Divide the system at logical boundaries and compare good and bad units under the same conditions. Measure upstream and downstream of the suspected point rather than replacing the most visible component immediately.
Create hypotheses from plausible mechanisms. For an intermittent reset, candidates may include input drop, regulator protection, cracked connection, connector contact, firmware watchdog, noise, thermal drift or fixture interruption. For a communication failure, candidates may include physical layer, termination, clock, power, configuration, firmware, cable or external device. Rank them by fit to the evidence, not by convenience.
For each hypothesis, predict an observation that should occur if it is true and one that would challenge it. Then choose the least destructive discriminating test. Record both positive and negative findings. Eliminating a mechanism is valuable because it prevents the same unsupported explanation from returning later.
Plan destructive analysis around a specific question
Destructive analysis may include component removal, cross-sectioning, dye-and-pry, decapsulation, microsection, chemical analysis or material characterization. Use it after documenting the original state and preserving controls. Write the question, selected location, orientation, preparation method, expected feature and interpretation before cutting or removing the sample.
Sampling matters. A section through one location can miss a crack elsewhere, and preparation can introduce smearing, pull-out or deformation. Keep pre-section images and X-Ray orientation so the laboratory can target the suspected feature. Compare with a control prepared by the same method. Have qualified personnel interpret material or metallurgical findings; apparent morphology alone may not establish when or why damage occurred.
NASA failure-analysis resources emphasize systematic evidence and root-cause discipline for reliable electronics. Their public material can inform method planning, but NASA program requirements and conclusions should not be applied automatically to a commercial project. The project must define its own safety, acceptance and contractual authority.
Begin with non-destructive evidence so later cross-sectioning or component removal has a defensible target.
Verify root cause and corrective action
Root cause should explain the failure mechanism, why the affected unit or population was exposed, and why existing controls did not prevent or detect it. There may be a physical cause, an escape cause and a system cause. A solder crack may be the physical mechanism; insufficient support or excessive strain may create it; an inspection/test gap may allow escape; an uncontrolled change or incomplete review may explain why the risk entered production.
Strengthen the conclusion by challenging it. Where safe and practical, recreate the mechanism in a representative sample and remove or control it in another. Confirm that the predicted symptom follows. Reconcile contradictory evidence. State remaining uncertainty and the population boundary rather than overstating confidence.
Corrective action should act on the cause. Options can include design, land pattern, component, support, material, process window, tooling, work instruction, maintenance, supplier control, software, test coverage, limit, training or handling changes. Additional inspection may contain risk, but it is not always prevention. Review the action for new risks, validate it under representative conditions and approve its effectivity.
Verify effectiveness across time and population. Track recurrence, related failure modes, process metrics, repair/retest, field returns and audit evidence. Define the number of builds, units or time period required before closure. Ensure temporary containment is removed only after the permanent control is released.
Corrective action closes only after the proposed cause, changed control and validation evidence agree.
Build a buyer-ready failure analysis report
The final report should include product and sample identity, revisions, complaint and expected behavior, timeline, containment and population, evidence preservation, reproduction, measurement-system checks, inspections and tests, hypotheses, discriminating results, destructive methods, findings, causal chain, corrective and preventive actions, validation, effectivity, remaining risk and effectiveness plan. Attach images and raw records with scale, orientation, sample identity, method and reviewer.
Distinguish observation from interpretation. “Dark region visible at one interface in image 12” is an observation. “Thermal fatigue caused an open circuit” is an interpretation that needs electrical, structural, history and mechanism evidence. Explain method limitations and unresolved questions. A transparent evidence boundary is more useful than a confident but unsupported conclusion.
Link the report back to manufacturing traceability so the OEM can identify all potentially affected units and confirm the corrected configuration. The GNS PCBA traceability guide provides a buyer-side check for retrievability before a failure occurs.
Manage external laboratories and supplier handoffs
When analysis moves to a component supplier or independent laboratory, send a written question rather than only a failed board. Include product and sample identity, handling restrictions, symptom and reproduction conditions, relevant schematic region, known measurements, hypotheses, requested methods, prohibited actions, evidence format and return disposition. Ask the laboratory to approve any destructive step that differs from the plan.
Request calibrated images with sample identity, scale, orientation, location and acquisition method. For measurements, retain units, instrument, setup, uncertainty where relevant and raw data. For cross-sections, document target selection, cut plane, preparation and controls. A polished report without traceable sample-to-image linkage cannot support an affected-population decision.
Component-supplier analysis should be reconciled with board-level evidence. A supplier may report that the returned part meets its electrical test, yet the system symptom can still involve board interaction, intermittent connection, application stress or a condition absent from the supplier test. Conversely, a failed part does not prove assembly caused the failure. Compare the component result, manufacturing history, application conditions and known-good controls before assigning responsibility.
Use decision checkpoints to prevent open-ended analysis. After containment, confirm whether the population is defensible. After reproduction, decide whether the original symptom is understood. Before destruction, confirm the planned test can distinguish hypotheses. Before closure, verify causal evidence, corrective action, effectivity and recurrence monitoring. If the evidence remains insufficient, state the uncertainty and control the risk rather than forcing a root-cause label.
Commercial responsibility can be reviewed after the technical record is stable. Separating fact finding from immediate blame improves sample preservation and data sharing. The final report can then map verified causes to design, material, process, test or handling ownership using the contract and change history.
Keep rejected hypotheses in the case file with the evidence that challenged them. This prevents repeated work during a later recurrence and shows why the team selected one mechanism over another.
Preserve the case record after technical closure
The closed case should remain searchable by product, revision, serial or lot, symptom, failure mechanism, component, process and corrective action. Retain original images, raw measurements, fixture and software versions, rejected hypotheses, destructive analysis records, approvals and effectivity. A later recurrence may look different at first, and the earlier negative findings can save time when engineers compare the two populations.
Define who can reopen the case and which evidence triggers that decision. A related field return, supplier notice, process drift or corrective action failure may require renewed containment and a wider population review. Keep the previous conclusion visible while recording the new evidence. This history lets buyers judge whether the action controlled the mechanism across time instead of treating every closed report as permanent proof.
Conclusion
Effective PCBA failure analysis preserves evidence, defines the population, reproduces a measurable symptom, checks the measurement system, tests competing mechanisms and uses destructive methods only for a specific question. It closes with a causal chain and an action proven across the affected product and process—not with a repaired sample.
For supplier review, send the original failure record, released design and firmware, unit and lot identity, production/test history, field conditions, prior actions and available samples. Agree containment, decision authority, evidence format and escalation before analysis begins.
Request a PCBA Failure Evidence Review
FAQ
What is the first step in PCBA failure analysis?
Protect the original evidence while containing the affected population. Record the unit, configuration, symptom, conditions, first failure, handling history and related production records before repair, retest, cleaning or destructive work changes the evidence.
Is replacing a failed component the same as finding root cause?
No. Replacement can localize a fault or restore function, but it does not by itself explain why the component, solder joint, connection, process or design failed. Root cause needs a supported causal mechanism and evidence that distinguishes competing explanations.
When should destructive analysis begin?
Begin after non-destructive observations, hypotheses and decision needs are documented and representative controls are preserved. The planned section or removal should test a specific question because destructive work can erase evidence and introduce preparation artifacts.
How is a PCBA corrective action verified?
Verify that the action changes the suspected causal mechanism, passes affected functional and manufacturing checks, does not create new risks, remains effective under representative conditions and reduces recurrence across the defined population and review period.