AVP - Site Reliability Engineer
2 days ago
Singapur, singapore
Singapore Exchange Limited
Full-time
SGD 210,000/year
Free with email or Google
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Free with email or Google
By continuing, you agree to our Terms & Privacy Policy.
- 5-10+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or a related technical discipline, with experience leading reliability initiatives in complex enterprise environments.
- Proven experience defining and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and reliability governance frameworks.
- End-to-End Service Ownership to Support the full operational lifecycle of assigned applications and platforms, including onboarding, release readiness, production monitoring, incident response, problem management, recovery validation, and continuous service improvement.
- Hands-on expertise with observability and monitoring platforms such as Prometheus, Grafana, ELK, AWS CloudWatch, DataDog, Coralogix, or similar technologies.
- Strong experience operating and supporting cloud-native platforms, Kubernetes/EKS environments, and distributed systems at scale.
- Proficiency in automation and infrastructure-as-code using tools such as Terraform, Ansible AAP, Python, or equivalent technologies.
- Experience leading major incident management, problem management, root-cause analysis, and post-incident review processes.
- Demonstrated capability in capacity planning, performance engineering, resilience testing, and operational risk management within high-availability environments.
- Strong stakeholder management and communication skills, with the ability to collaborate effectively across engineering, infrastructure, security, product, and business teams.
- Experience coaching and mentoring engineers while driving operational excellence, reliability culture, and continuous improvement initiatives.
Preferred Qualifications
- Experience or exposure within financial services, financial market infrastructure, exchanges, index calculation platforms, reference data systems, market data services, trading, clearing, settlement, or other highly regulated environments.
- Professional certifications in cloud platforms (AWS, Azure, GCP), Kubernetes, Site Reliability Engineering, DevOps, or related disciplines.
- Experience implementing AIOps, observability analytics, predictive monitoring, and intelligent automation solutions.
- Hands-on experience leveraging AI or agentic AI technologies for alert triage, incident investigation, operational automation, runbook optimisation, and knowledge management.
- Exposure to FinOps practices, cloud cost optimisation, and platform engineering operating models.
- Familiarity with security, resilience, and regulatory requirements applicable to critical financial systems.
- Experience driving large-scale technology transformation, platform modernisation, or cloud adoption initiatives.
- Advanced knowledge of software engineering practices, CI/CD pipelines, and reliability-focused architecture patterns.