AI Foundation Architect Lead
Univers combines global operating scale, deep industrial intelligence, and recognition from the institutions shaping energy, technology, and sustainability. The platforms built for the last decade were designed to monitor. Univers was built to act — not just report. Univers operates at global scale: 1,070 GW+ Energy assets under AI management 450M+ Connected sensors and devices 800+ Enterprise customers Leader Gartner Magic Quadrant Leader Home - Univers Job Summary We are seeking a forward-thinking and hands-on AI Platform Foundation Architect to spearhead the evolution of our core platform foundation into an AI-native, highly scalable tech ecosystem. In this role, you will be responsible for defining and executing the architecture across Infrastructure as a Service (IaaS), Middleware, Big Data platforms, and AI Infrastructure (AI Infra). Beyond core infrastructure management (HA, auto-scaling, backup, and performance), you will act as a key transformation catalyst—building enterprise AI Gateway capabilities for LLM/SLM token and cost optimization, hosting small models, and embedding AI capabilities into automated operations, deployment, and root-cause troubleshooting. You will bridge traditional infrastructure engineering with modern AI agentic operations to empower all upper-level platform and product teams. Key Responsibilities: Oversee the evaluation, selection, architecture standards, and operational frameworks for enterprise middleware (e.g., distributed messaging, caches, databases, API gateways). Architect and optimize high-throughput Big Data platforms, time-series databases (TSDB), and log/metric storage engines handling high-concurrency real-time telemetry. Provision scalable cloud-native and hybrid IaaS solutions (compute, networking, storage, K8s) using Infrastructure as Code (IaC) principles. Design automated horizontal and vertical resource scaling (expansion/shrinkage) to ensure performance under peak traffic while minimizing infrastructure spend. Establish full-stack telemetry and monitoring to track critical hardware and software metrics (IO, CPU, RAM, Storage, Network). Architect multi-region, resilient fault-tolerant systems to guarantee non-disruptive platform availability. Design unified backup workflows, automated snapshots, and rapid DR execution to safeguard core data assets. Integrate AI agents and LLM/SLM capabilities directly into CI/CD pipelines, automated deployments, anomaly detection, real-time log analysis, and root-cause analysis (RCA) to accelerate incident resolution. Architect and deploy centralized AI Gateways (e.g., LiteLLM, Portkey, Kong AI) to manage, route, and optimize token usage, semantic prompt caching, rate-limiting, and cost allocation across foundational LLMs/SLMs. Design low-latency, scalable hosting and inference infrastructure for small/edge models (SLMs) and open-source models to support lightweight, specialized internal tasks. Build a unified AI Infra layer providing model serving, GPU/NPU resource orchestration, and model context protocol (MCP) routing for upper-level teams. Basic Qualifications: AI Infra & Model Serving: Hands-on experience with AI Gateways, token optimization strategies (prompt caching, model routing), and hosting small language models (SLMs) or open-source models using tools like vLLM, Ollama, or Triton. AIOps & Smart O&M: Familiarity with leveraging LLMs/AI Agents for log analysis, automated deployment checks, root cause analysis (RCA), or AI-driven alerting. Middleware & Data Engines: Deep architecture expertise in messaging queues (Kafka, RabbitMQ), caches (Redis), and modern data/time-series platforms (IoTDB, HBase,InfluxDB, Flink Lakehouse architectures). IaaS & Cloud-Native: Strong experience in public/hybrid clouds (AWS, GCP, Azure), Kubernetes orchestration, and Infrastructure as Code (Terraform, Ansible). Observability & Resource Optimization: Expert-level knowledge of observability tools (Prometheus, Grafana, Loki/ELK) for monitoring CPU, IO, Storage, and system bottlenecks. Solid understanding of distributed systems; hands-on experience in designing and developing platforms for traffic risk control, transaction fraud detection, or account security is highly preferred. Ability to thrive in a fast-paced, dynamic environment and drive projects forward with efficiency. Experience integrating heterogeneous enterprise systems and working with both IT data and OT/IoT data from industrial or physical environments. Excellent communication, with the ability to align executives, engineers, product teams, field teams and customers around difficult technical decisions. Bachelor’s or Master’s degree in Computer Science, Software Engineering or equivalent practical experience. Preferred: Experience with agentic AI, RAG, model and agent governance, tool permissioning, human-in-the-loop workflows or governed write-back. Experience establishing developer platforms, APIs, SDKs, connector ecosystems or open-source and partner programs. Experience leading globally distributed teams and delivering platforms across multiple regulatory and deployment environments.