Site Reliability Engineer

1 week ago

Singapur, singapore Momcozy Full-time SGD 180,000 Contract

Join our overseas backend team to own the stability, observability, and automation of the Singapore production environment. Our smart hardware and mobile applications run across multiple AWS regions, and this role exists to ensure the region has independent release, monitoring, and emergency-response capability. You will hold operational access to the regional production environment, handle incidents alongside the regional tech lead, and collaborate with the HQ infrastructure team on containerization, unified gateway, and middleware governance initiatives.

Who You’ll Work With

  • Reports to: Regional Tech Lead (based in Singapore)
  • Collaborates with: HQ backend engineering team (China), HQ infrastructure team, regional backend engineers, and product/operations teams
  • Team context: Part of a growing regional engineering team; works across time zones with China HQ

Responsibilities

  • Operate and optimize regional AWS resources and Kubernetes clusters, including capacity planning and cost optimization, ensuring environment consistency and configuration traceability
  • Build and maintain monitoring dashboards, alert rules, and log platforms for applications, middleware (MySQL / Redis / Kafka / RocketMQ), and infrastructure, ensuring alerts are accurate and actionable
  • Participate in regional incident response, executing and automating mitigation actions (rollback, scaling, degradation switches); run degradation and failure drills and maintain incident runbooks
  • Build and maintain CI/CD pipelines supporting canary releases and fast rollback; enforce change review and checklists; advance infrastructure-as-code practices
  • Own backup, recovery, tuning, and capacity assessment for regional databases and middleware; partner with developers on slow-query and performance issue resolution
  • Manage regional IAM accounts and access under company policy; support rollout of WAF, gateway, and other security capabilities; enforce least-privilege access and audit trails for data operations

Requirements

Must-Have

  • Bachelor's degree or above in Computer Science, Software Engineering, or a related field, or equivalent industry experience
  • 5+ years of experience in operations, SRE, or platform engineering
  • Proficient with AWS (EC2, ALB, EKS, RDS, ElastiCache, S3, IAM, VPC, CloudWatch), with multi-region operations experience
  • Strong Kubernetes skills: cluster operations, troubleshooting, and resource governance
  • Experience with Terraform or similar infrastructure-as-code tools, and CI/CD tools such as Jenkins / GitLab CI
  • Solid Linux and networking fundamentals; scripting proficiency in at least one of Python / Shell / Go
  • Experience deploying, monitoring, and troubleshooting MySQL, Redis, and Kafka / RocketMQ
  • Familiar with Prometheus / Grafana / ELK or comparable observability stacks
  • Experience handling production incidents and postmortems; able to execute mitigation calmly and methodically under pressure
  • Legally authorized to work in Singapore; this role does not provide visa sponsorship

Nice-to-Have

  • Able to read Java / Spring Boot code and understand application-side issues
  • Hands-on experience with API gateways (Higress / Shenyu / Kong / Nginx) and canary releases
  • Experience with traffic replay, chaos engineering, or failure drills
  • Experience with multi-region, cross-time-zone collaboration

Language

  • Professional working proficiency in English (CEFR B2 or above)
  • Mandarin Chinese is highly advantageous for collaboration with China HQ, but not required

What We Offer

  • Opportunity to build a regional engineering function from the ground up
  • Cross-regional collaboration with a mature HQ engineering team
  • Exposure to large-scale e-commerce and IoT infrastructure
  • Competitive compensation and benefits package