System-Level Reliability Researcher

Interuniversitair Micro-Electronica Centrum VZW· Arr. Leuven· EURES· publicada el 14/07/2026
Imprescindible:AI

Join imec’s center of excellence for hardware-software-technology co-design to define the future of system-level reliability in compute systems. Apply

Compute System Architecture (CSA) is a center of excellence at imec for hardware-software-technology co-design for future compute systems. We work in close collaboration with other expertise centers in imec specializing in applications, technology, circuits and design to innovate and pathfind next-generation compute system architectures across multiple domains – AI, HPC, Automotive, Space, and more. CSA has presence in 6 centers of IMEC, in Belgium, Netherlands, Germany, UK, USA and Qatar. This position is primarily for Leuven, Belgium.

What you will do

We are looking for an experienced System-Level Reliability Researcher to join our team. In this position, you will play a key role in developing methodologies to evaluate and improve the reliability of advanced compute architectures. You will focus on how hardware faults, degradation effects, and technology-level reliability observations translate into system-level behavior, application correctness, availability, and product lifetime.

This R&D role focuses on compute system architecture, reliability modeling, and resilience evaluation. You will collaborate with experts from the technology, device, circuit, and design domains to translate lower-level reliability information into architectural fault models, simulation inputs, and system-level reliability metrics.

Your responsibilities will include:

  • Developing system-level reliability assessment methods for processors, accelerators, memory subsystems, SoCs, and chiplet-based architectures.
  • Creating architectural fault models and abstractions based on reliability information provided by lower layers of the technology stack.
  • Building and using simulation, emulation, and fault-injection frameworks to study fault propagation, error masking, and application-level impact.
  • Quantifying reliability outcomes such as Silent Data Corruption, detected errors, service interruptions, availability loss, and lifetime-related degradation at the system level.
  • Evaluating Reliability, Availability, and Serviceability (RAS) mechanisms, including error detection, containment, recovery, redundancy, and graceful degradation techniques.
  • Studying workload-dependent reliability behavior under realistic execution conditions and operating profiles.
  • Connecting technology-informed fault characteristics with system architecture models to support reliability-aware design decisions.
  • Collaborating closely with colleagues within and outside CSA to interface with workload models, architectural simulators, circuit-level reliability data, and technology-level observations.
  • Contributing to research publications, partner discussions, etc.

Who you are

  • You hold a Master’s or PhD in Electrical/Electronics/Computer Engineering/Science or relevant domains with strong R&D experience in system-level reliability.
  • You have 6+ years of relevant experience in industry and/or academia (we encourage you to apply if you have slightly less experience but strong alignment with the role).
  • You have a strong background in compute system architecture and microarchitecture.
  • You have a good understanding of core concepts in system-level reliability, resilience, and dependability, including:

-

  • Reliability, Availability, and Serviceability (RAS) architecture

-

  • Fault tolerance and resilience techniques

-

  • Error detection, containment, recovery, and graceful degradation

-

  • Fault propagation and architectural error masking
  • You have experience developing architectural models, quantitative analysis frameworks, or system-level evaluation methodologies.
  • You possess strong analytical and quantitative skills, including statistical analysis, data interpretation, and uncertainty-aware reasoning.
  • You have hands-on experience designing and executing experiments to derive insights from complex system behavior.
  • You are able to collaborate across disciplines and abstraction layers, translating reliability information from technology, circuit, and design teams into meaningful architectural models and system-level insights.
  • You have the knowledge and experience with good software and research engineering practices (e.g., version control, reproducibility, automation, testing, documentation).
  • You enjoy being hands-on while working with your colleagues and while mentoring students/interns, promoting a team culture of creativity, collaboration, and excellence.
  • You are a team player and flexible in accommodating changing priorities to support business needs; you are self-motivated, responsible and willing to take ownership.
  • You feel at home and know how to integrate into a multicultural environment; you are open-minded, you seek and embrace differences and accept constructive challenges.
  • You have effective English communication skills for discussions, documentations, and dissemination of ideas and results within CSA/IMEC and in wider ecosystems.