Data Engineer
WHO WE ARE As Singapore’s longest‑established bank, OCBC has supported individuals and businesses in achieving their aspirations since 1932. We are transforming into a future‑ready learning organization – leveraging technology and innovation while staying true to our ambition to be Asia’s leading financial services partner for a sustainable future. Join us to build the bank of the future, work in collaborative teams, and create lasting value for our customers and communities. ROLE We are seeking a Data Engineer (VP) to design, build, and scale enterprise‑grade data pipelines and platforms within a banking environment. This role owns the end‑to‑end architecture of batch and real‑time data pipelines, AI knowledge base, sets engineering standards for the team, and works closely with the data leaders to turn data into scalable, reusable, and AI‑ready products. You will play a hands‑on technical leadership role — architecting solutions, writing production‑grade code, and mentoring other engineers — across use cases such as risk management, customer engagement, fraud detection, and intelligent automation. This role reports to Head of Data Product, Group data office. KEY RESPONSIBILITIES Data Platform Architecture Own and continuously optimize scalable data warehouse / lakehouse architectures across data platforms (e.g., Cloudera, AWS or GCP)
Design and evolve modern architecture patterns for batch and streaming data pipelines
Define and enforce data modeling, partitioning, and performance‑optimization standards and best practices
Batch & Streaming Data Processing and Orchestration Build and optimize large‑scale batch pipelines using Spark, SQL, Python, Map/Reduce.
Design and implement real‑time streaming pipelines using Kafka, Flink, or equivalent engines
Architect and maintain CDC pipelines using Debezium, Confluent, or Fivetran for real‑time data synchronization
Design and optimize complex Airflow DAGs for end‑to‑end orchestration
Drive standardization of orchestration patterns and reusable components across teams
Cloud Infrastructure & DevOps Architect design and deliver on cloud‑native data infrastructure i.e., Cloudera, AWS, GCP (Docker, Kubernetes, Cloud Run or equivalent)
Own CI/CD pipelines for data engineering workloads and use infrastructure‑as‑code practices, i.e., Terraform.
Design and build REST APIs and backend services using Python / Flask to serve data products
Design and implement caching strategies using Redis to support low‑latency, high‑throughput access
AI Knowledge Base Design solutions for vector databases and embedding pipelines to power semantic search and knowledge bases for AI agents
Architect design experience, including chunking, embedding generation, indexing, and retrieval strategies
Design and build tool‑ready, contextual data layers that LLMs and AI agents can query and reason over
Ensure online/offline consistency and freshness of knowledge base content feeding AI applications
Cross‑functional Collaboration & Mentorship Partner with the data team leaders, AI teams, Infra/SRE team, and business stakeholders
Translate business needs into scalable, production‑ready data products
Mentor mid‑level and junior data engineers; review code and uphold engineering best practices
Drive continuous improvement of data engineering standards, tooling, and processes
REQUIREMENTS Bachelor’s or Master’s degree in computer science or a related field
At least 10 years of experience in data engineering, data platforms, or related roles, including experience leading pipeline design and delivery
Strong understanding of modern data architectures including Data Warehouse, Data Lake, Lakehouse, and batch/streaming systems
Experience building tool‑ready APIs and contextual data layers for LLM / AI agent consumption is preferred
Hands‑on experience owning production data platforms end‑to‑end, including on‑call/reliability ownership
Exposure to LLM applications, RAG architectures, vector databases, or AI agent / tool‑calling frameworks is a plus
Strong product mindset: ability to treat data as a product, not just a project
Ability to abstract complex data problems into scalable solutions
Excellent communication skills across technical and business stakeholders, with demonstrated ability to mentor others
Experience in banking or financial services is preferred
Technical Stack Data Warehouse / Platform: Cloudera, BigQuery, Redshift, Teradata, or similar
Batch Processing: Spark, SQL, ETL, Python, Map/Reduce
Streaming: Flink or other real‑time data processing engines
CDC: Debezium, Confluent, Fivetran, or similar
Data Ingestion: APIs, GA4, Pub/Sub, Kafka, Python pipelines
Orchestration: Airflow or equivalents
Cloud & Infrastructure: GCP (Docker, Kubernetes, Cloud Run) or AWS equivalents
DevOps / DataOps: CI/CD pipelines, Terraform or equivalents
Backend & Serving: Python, Flask, REST APIs, Redis
AI Knowledge Base: RAG pipelines end‑to‑end (OCR, chunking, embedding, indexing, tuning, etc)
Experience building tool‑ready APIs and contextual data layers for LLM / AI agent consumption is strongly preferred
WHAT WE OFFER Competitive base salary and comprehensive benefits.
Strong learning and development opportunities.
Exposure to impactful, enterprise‑scale data and AI initiatives across the OCBC Group.
A collaborative environment that values innovation, craftsmanship, and continuous improvement.
Your wellbeing, growth, and aspirations matter to us as much as delivering value to our customers.