Senior Failure Analysis Manager
The Senior Failure Analysis Manager will lead a specialized engineering team and manage the state-of-the-art diagnostic laboratory. They are responsible for the forensic isolation of electrical, mechanical, and thermal root causes underlying hardware failures occurring during pilot builds, large-scale production testing, and field deployments of next-generation data center switches.
The ideal candidate will possess extensive diagnostic experience, ranging from silicon-level component anomalies to system-level signal and power integrity issues. Acting as a fundamental technical pillar, this leader will drive Root Cause Corrective Actions (RCCA), coordinate closely with chip designers and material suppliers, and ensure strict adherence to "Copy Exact" diagnostic baselines across all laboratories to prevent defect recurrence.
Responsibilities
- Laboratory Management and Diagnostic Leadership
- Technical Team Responsibility: Train, manage, and technically direct a high-level team composed of failure analysis engineers and laboratory technicians specializing in semiconductor, PCBA, and interconnect diagnostics.
- Advanced Equipment Strategy: Oversee capital expenditure (CapEx) planning, operational calibration, and maintenance of advanced laboratory equipment, including scanning electron microscopes (SEM), energy dispersive X-ray spectroscopy (EDS), Fourier-transform infrared spectroscopy (FTIR), 3D X-ray/CT systems, and microsectioning stations.
- Standardized Failure Analysis Flowcharts: Establish and enforce rigorous analytical testing workflows, from non-destructive to destructive, to ensure component evidence is preserved during classification.
- Root Cause Isolation and Cross-functional Assessment
- Multi-level Printed Circuit Board (PCBA) Diagnostics: Drive root cause investigation of complex board defects, including latent via cracks, pad craters, "Head-in-Pillow" (HiP) solder joints, electrochemical migration (ECM), and signal integrity degradation in high-frequency laminates.
- Silicon and Package FA: Collaborate closely with leading ASIC, GPU, and memory suppliers to evaluate chip-level anomalies, thermal interface material (TIM) loss, and integration failures between the Ball Grid Array (BGA) package and the board.
- Production Line Alignment: Act as the primary technical liaison with the production plant. When in-line test failures exceed statistical thresholds, provide rapid classification feedback to Process, Quality, and Automation engineering groups to contain the deviation.
- Closed-Loop Corrective Actions and Supplier Management
- Rigorous Execution of Closed-Loop Corrective Actions (RCCA): Develop and report comprehensive engineering failure analysis reports. Utilize structured methodologies (8D, 5 Whys, fault tree analysis) to convert laboratory findings into permanent changes in engineering and manufacturing processes.
- Technical Escalations with Suppliers: Present irrefutable evidence to component and raw material suppliers when failures point to material impurities or manufacturing defects, thereby driving the implementation of corrective measures at the supplier level.
Requirements:
- Academic Background: Bachelor's degree in Materials Science, Electrical Engineering, Solid State Physics, or a closely related engineering discipline. Master's (MS) or PhD will be highly valued.
- Experience: 8 to 10 or more years of specific experience in failure analysis or hardware reliability engineering in enterprise-class equipment (e.g., core network switches, high-performance servers, telecommunications equipment), with at least 3 years in a direct engineering team management role or as a laboratory head.
- Technical Domain Expertise:
- Deep knowledge of multi-layer printed circuit board (PCB) manufacturing, high-speed signal constraints, and advanced solder metallurgy in Surface Mount Technology (SMT) (SAC305, low-temperature alloys).
- Expert proficiency in using non-destructive localization techniques, such as Time-Domain Reflectometry (TDR), thermography (InGaAs/lock-in thermography), and high-resolution 3D X-ray.
- Extensive knowledge of physical material characterization techniques (decapsulation and separation, cross-sectioning, polishing, and chemical etching).
- Methodologies: Deep knowledge of IPC standards (IPC-A-610, IPC-TM-650 diagnostic methods) and JEDEC microelectronics reliability standards.
Key Performance Indicators (KPI) for Success
- Mean Time to Root Cause Identification (MTTRC): Accelerate the time between initial failure logging and definitive confirmation of the physical or electrical root cause.
- First-Time Correct Detection Rate: Ensure initial diagnostic strategies correctly identify failures without destroying critical component evidence or compromising data integrity.
- Closed-Loop Action Verification: Ensure 100% verification and closure of design or manufacturing-related engineering changes arising from laboratory findings.
Requirements
- Minimum Education: Higher education - Bachelor's 5 years of experience
- Languages: English
- Age: between 35 and 65 years
- Keywords: lider, jefe, gerente, manager, director, chief, lead, jefatura, regente, senior, sr