Senior Backend Engineer (SRE / Reliability & Performance)

Appier· Taipei, Taiwan· greenhouse· publicerad 2026-07-06
Krav:PythonJavaGraphQLGitAWSAzureGoogle CloudDockerKubernetesCloudBackendDevOpsAgileCI/CDAISeniorLead

About Appier

Appier (TSE: 4180) is an AI-native Agentic AI as a Service (AaaS) company that empowers businesses to create value through cutting-edge AdTech and MarTech solutions. Founded in 2012 with the vision of "Making AI Easy by Making Software Intelligent," Appier helps businesses turn AI into ROI through its Ad Cloud, Personalization Cloud, and Data Cloud—each powered by Agentic AI that enables autonomous, adaptive, and real-time decision-making. Today, Appier operates 17 offices across APAC, the US, and EMEA, and is listed on the Tokyo Stock Exchange. Learn more at www.appier.com .

About the role

Engineers at Appier build a wide range of platforms and services that interconnect data and AI with our customers and users—operating at scale where every millisecond and every nine of availability matters. As a Senior Backend Engineer (SRE / Reliability & Performance) , you will sit at the intersection of backend development and site reliability engineering: designing and building scalable, performant backend services while owning the reliability, observability, and performance characteristics of large-scale, high-traffic systems (QPS > 10k). You will profile and tune systems end to end, lead the response to production incidents, and drive the engineering practices that keep our services fast and resilient as traffic grows. Seniority and title are determined by job-related skills, experience, and evaluation following the interview. We welcome international candidates to our teams. This position is ideally to be based in Taiwan.

Responsibilities

Design and build scalable, reliable, and maintainable backend services and the components that support them

Own system reliability and performance for high-traffic services, including SLO/SLA definition, error budgets, and capacity planning

Profile, benchmark, and tune critical components to resolve latency, throughput, and scalability bottlenecks in medium-to-large systems (QPS > 10k)

Diagnose and solve production traffic problems—load spikes, hot paths, resource contention, and cascading failures

Lead system design and provide technical guidance on reliability and performance trade-offs

Continuously improve observability (logging, metrics, tracing), incident management, DevOps, and production operational SOPs

Lead incident response, troubleshooting, and blameless post-mortems, then drive the follow-up engineering work

Build and optimize CI/CD pipelines and deployment automation to ship safely and frequently

Lead code reviews to ensure high quality coding and operational standards

Mentor engineers and facilitate agile collaboration across cross-functional teams

Participate in on-call rotation to ensure product reliability and scalability

About you

[Minimum qualifications]

5+ years of experience in backend software development, with hands-on SRE / reliability / infrastructure responsibilities

Proven experience tuning system reliability and performance for production services

Demonstrated experience solving traffic and scalability problems in medium-to-large systems

Ability to build and operate web services on Linux

Proficient in one or more of the following languages: Go / Python / Java / Scala / C++

Good knowledge of Network API design (e.g. REST or GraphQL)

Good understanding of SQL/NoSQL databases (MySQL / PostgreSQL / MongoDB / Redis / etc.)

Hands-on experience with observability tooling (e.g. Prometheus, Grafana, distributed tracing)

Familiar with AWS, GCP, or Azure

Familiar with Git

Proactive, with strong interpersonal and problem-solving skills

[Preferred qualifications]

BS/MS degree in Computer Science or related field

Technical leadership experience, such as mentoring engineers and facilitating agile processes

Strong skills with profiling and debugging tools, and building high-performance network services on Linux

Experience designing and architecting large-scale distributed systems and implementing distributed algorithms and data structures

Experience with container orchestration (Kubernetes, Docker) and infrastructure-as-code / configuration management (Terraform, Ansible)

Hands-on experience with CI/CD platforms (Jenkins, GitLab CI, GitHub Actions, ArgoCD)

Expert in some of the following CS domains:

Nginx / HAProxy and load balancing

Caching strategies and data-intensive application design

Monitoring and alerting systems (Prometheus / Nagios)

Capacity planning and chaos / resilience engineering

Operation automation

Continuous integration / continuous deployment

#LI-TC1