Current jobs related to MD, Head of Site Reliability Engineering, Group Technology - Singapore - DBS Bank Limited


  • Singapore GIC Full time

    GIC, one of the world's largest sovereign wealth funds, is seeking a seasoned leader to helm its Total Portfolio Risk Performance Technology Group as the SVP MD Engineering Head.The ideal candidate will have extensive experience in technology delivery in financial organizations, particularly in risk and performance, and possess a proven track record in...


  • Singapore Aptitude Asia Full time

    At Aptitude Asia, we're seeking a skilled Site Reliability Engineer to join our team. This role is crucial in ensuring the high reliability, availability, and performance of our applications throughout their lifecycle.Key Responsibilities:Develop and implement automation scripts to streamline repetitive tasks and address recurring issues.Collaborate with...


  • Singapore Aptitude Asia Full time

    Job SummaryAptitude Asia seeks a skilled Site Reliability Engineer to ensure the high reliability, availability, and performance of applications throughout their lifecycle.Key ResponsibilitiesReliability and Performance: Ensure applications operate with high reliability, availability, and performance.Automation and Innovation: Automate repetitive tasks and...


  • Singapore Mapletree Full time

    The Role The Head, IS&T will lead the IS&T teams in driving IT strategy, digital transformation & building a reliable and scalable IT system to support the rapid business growth across the Group. To be successful, you should be a strong leader with extensive team management experience and have an in-depth knowledge of the current and up-and-coming trends in...


  • Singapore AIA Singapore Private Limited Full time

    About the RoleWe are seeking a highly skilled Senior Site Reliability Engineer to join our team at AIA Singapore Private Limited. As a key member of our operations team, you will be responsible for ensuring the reliability and stability of our critical production services and applications.Key ResponsibilitiesLead complex system and champion services...


  • Singapore BYTEPLUS PTE. LTD. Full time

    Role OverviewAt ByteDance, we're seeking a skilled Site Reliability Engineer to join our team. As a Site Reliability Engineer, you'll be responsible for ensuring the reliability and normal operation of multiple core systems for big data and online computing. This includes building automated operation solutions for large-scale systems, cooperating with the...


  • Singapore Vortexa Full time

    Vortexa is a cutting-edge company that leverages satellite data and AI to provide real-time insights into global energy flows. We're looking for a skilled Site Reliability Engineer to join our Data Services Team, responsible for the developer platform and Amazon AWS estate.The ChallengeOur platform processes massive amounts of data from various sources,...


  • Singapore Singtel Group Full time

    At Singtel, our mission is to Empower Every Generation. We are dedicated to fostering an equitable and forward-thinking work environment where our employees experience a strong sense of Belonging, to make meaningful Impact and Grow both personally and professionally. By joining Singtel, you will be part of a caring, inclusive and diverse workforce...


  • Singapore Ripple Labs Singapore Full time

    As a Senior Site Reliability Engineer at Ripple Labs Singapore, you will be responsible for ensuring the high availability and scalability of our systems. Your primary goal will be to design, implement, and maintain a robust and efficient infrastructure that can handle high traffic and complex distributed systems.Key Responsibilities:Design and implement...


  • Singapore NodeFlair Full time

    Senior Site Reliability EngineerWe are working with NodeFlair, a leading pioneer in the Cryptocurrency space, to search for a Senior Site Reliability Engineer to join their Singapore/Remote team.Summary:Our client, a top player in cryptocurrency data monitoring, tracks over 10,000 tokens on 400+ exchanges with 300 million page views from 100+...


  • Singapore Oxford Knight Full time

    Job Title: Senior Site Reliability EngineerOxford Knight is seeking a highly skilled Senior Site Reliability Engineer to join our team. As a Senior Site Reliability Engineer, you will be responsible for designing, developing, and maintaining our Linux trading infrastructure on a day-to-day basis.Key Responsibilities:Lead the design and development of major...


  • Singapore Sea Full time

    About SeaSea is a cutting-edge technology company with a hyper-growing business scale, transforming complex problems into technical challenges. Our team of passionate engineers is dedicated to delivering world-class experiences for our users.The Games Site Reliability Engineer (SRE) team at Sea Labs Indonesia plays a crucial role in ensuring the stability...


  • Singapore Helius Full time

    Job Title: Site Reliability EngineerJob Summary: Helius is seeking a skilled Site Reliability Engineer to join our team. As a Site Reliability Engineer, you will be responsible for designing, implementing, and operating highly scalable and reliable systems. Your main focus will be on ensuring the smooth operation of our services, resolving technical issues,...


  • Singapore BYTEDANCE PTE. LTD. Full time

    About the JobAt ByteDance, we are looking for a talented Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability and normal operation of multiple core systems for big data and online computing, while paying attention to system capacity and stability.Key Responsibilities Ensure the reliability and normal...


  • Singapore ITCAN PTE. LIMITED Full time

    Roles & ResponsibilitiesRoles & Responsibilities:The Site Reliability Engineer (SRE) combines software development and system engineering to build and run distributed solutions in a secured multi-tier heterogeneous environment to safeguard, provide and continuously improve the software and systems behind the organization’s cloud platform solutions.The Job:...


  • Singapore LANDI INTERNATIONAL (SINGAPORE) PTE. LTD. Full time

    Landi International (Singapore) PTE. LTD.As a Site Reliability Engineer at Landi International (Singapore) PTE. LTD., you will play a crucial role in ensuring the availability, reliability, and scalability of our platforms. Your primary responsibilities will include:· Building, operating, and maintaining our platform infrastructures across various...


  • Singapore Ripple Labs Singapore Full time

    Job Title: Senior Site Reliability EngineerWe are seeking a highly skilled Senior Site Reliability Engineer to join our team at Ripple Labs Singapore. As a key member of our infrastructure team, you will be responsible for ensuring the high availability and scalability of our systems.Key Responsibilities:Design and Implement High Availability Solutions:...


  • Singapore ACCESS PEOPLE (SINGAPORE) PTE. LTD. Full time

    Roles & ResponsibilitiesA global energy trading firm is transitioning to a data-centric platform and is seeking a Site Reliability Engineer to support this multi-year program. The role will focus on enhancing the reliability, scalability, and stability of the company's evolving platform. The successful candidate will work on integrating a new event-based,...


  • Singapore Snaphunt Full time

    The OpportunityWe're seeking an experienced Site Reliability Engineer to empower users with a rich feature set, high availability, and stellar performance at First Digital Finance Corp.As we expand customer deployments, the ideal candidate will deliver insights from massive-scale data in real-time, collaborating with a cross-functional team to develop...


  • Singapore Tower Research Capital Full time

    Tower Research Capital Job DescriptionJob Title: Site Reliability EngineerJob Summary:We are seeking a highly skilled Site Reliability Engineer to join our team at Tower Research Capital. The successful candidate will be responsible for ensuring the continuous operation of our Linux-based trading infrastructure and addressing day-to-day operational needs.Key...

MD, Head of Site Reliability Engineering, Group Technology

2 months ago


Singapore DBS Bank Limited Full time
Business Function

Group Technology enables and empowers the bank with an efficient, nimble and resilient infrastructure through a strategic focus on productivity, quality & control, technology, people capability and innovation. In Group Technology, we manage the majority of the Bank's operational processes and inspire to delight our business partners through our multiple banking delivery channels.

Job Overview:

As the Head of Site Reliability Engineering (SRE) & Governance, will be responsible for leading and overseeing the SRE & Governance team to ensure the stability, reliability, and scalability of our banking platforms and services. This role requires a visionary leader with a deep understanding of SRE principles, a strategic mindset, and a hands-on approach to problem-solving. The ideal candidate will have a strong background in software engineering, infrastructure management, and a proven track record in driving reliability improvements in complex technical environments.


Key Responsibilities:


Leadership & Strategy:

o Develop and execute the overall SRE strategy in alignment with the bank's business goals and technological vision.

o Lead, mentor, and manage a team of SRE engineers, fostering a culture of collaboration, innovation, and continuous improvement.

o Consolidate and oversee the bankwide production support team to ensure consistent and reliable service across all platforms.

o Define and implement the SRE framework and policy, standardizing reliability practices across the organization.

o Establish and lead a Centre of Excellence (COE) for SRE, promoting best practices and continuous learning.

Reliability & Performance:

o Design, implement, and maintain robust monitoring, alerting, and incident response systems to ensure high availability and performance of banking services.

o Drive the adoption of best practices in system design, capacity planning, and performance optimization.

o Identify and mitigate potential risks to system reliability, proactively addressing issues before they impact customers.

Monitoring & Analysis:

o Develop and implement monitoring & analysis strategies to proactively identify and address potential issues within the technology infrastructure.

o Utilize data-driven insights to optimize system performance, reliability, and scalability.

o Collaborate with cross-functional teams to establish monitoring tools and metrics, ensuring alignment with business objectives and goals.

Automation & Tooling:

o Champion automation efforts to streamline operational processes, reduce manual intervention, and increase system efficiency.

o Develop and maintain tools and scripts for infrastructure management, deployment, and monitoring.

Incident Management & Troubleshooting:

o Lead incident response efforts, ensuring timely resolution of issues and effective communication with stakeholders.

o Conduct root cause analysis and implement corrective actions to prevent recurrence of incidents.

o Maintain comprehensive documentation of incidents, resolutions, and preventive measures.

Capacity Planning & Scalability

o Lead capacity planning efforts to ensure that technology infrastructure can support current and future business needs.

o Develop scalability strategies to accommodate growth and fluctuations in demand.

o Collaborate with cross-functional teams to implement scalable solutions and optimize resource allocation for efficient and effective operations.

Quality Assurance & Resiliency:

o Build and maintain robust quality assurance processes to ensure the reliability and performance of all banking services.

o Lead efforts to enhance the resiliency of the bank's infrastructure, ensuring systems are robust against potential failures and disasters.

o Oversee Level 1.5 technical risk management, identifying and mitigating risks that could impact the bank's technology operations.

Tech Risk:

o Establish and enforce robust technology governance frameworks, policies, and controls to mitigate operational, cyber, and compliance risks.

o Ensure adherence to all regulatory and legal requirements.

Processes for Build and Operate:

o Develop and optimize processes for building and operating technology infrastructure to ensure reliability, performance, and scalability.

o Implement best practices and standards for configuration management, deployment, and monitoring of systems.

Software Development Life Cycle (SDLC)

o Implement and oversee best practices for SDLC.

o Collaborate with development teams to ensure that SDLC processes are aligned with reliability and performance goals.

o Drive continuous improvement in SDLC processes to enhance the quality, efficiency, and reliability of technology solutions.

DevOps

o Champion the adoption and implementation of DevOps principles and practices.

o Collaborate with development and operations teams to streamline processes, automate tasks, and improve collaboration.

o Drive continuous integration, continuous delivery (CI/CD) pipelines, and infrastructure as code initiatives to enhance efficiency and reliability.

Collaboration & Communication:

o Work closely with the application & infrastructure teams to ensure that reliability is built into the architecture and design of new features and services.

o Communicate reliability goals, progress, and challenges to executive leadership and other stakeholders.

o Promote a culture of transparency and accountability within the SRE team and across the organization.

Vendor Management:

o Manage relationships with external vendors and service providers relevant to SRE, ensuring their services align with the bank's reliability and performance standards.

o Negotiate contracts and service level agreements (SLAs) with SRE-related vendors to secure favourable terms and ensure accountability.

o Continuously evaluate vendor performance in the context of SRE, addressing any issues and exploring new opportunities to optimize service delivery.

o Collaborate with vendors to stay abreast of the latest SRE technologies and tools that can enhance the bank's infrastructure and reliability.

Qualifications & Requirements:

Education & Experience:

o Bachelor's or Master's degree in Computer Science, Engineering, or a related field.

o Minimum 20 years of experience in software engineering, infrastructure management, or a related technical field.

o Minimum 5 years of experience in a leadership role within an SRE or DevOps team, preferably in the banking or financial services industry.

Technical Skills:

o Proficiency in programming languages such as Python, Go, Java, or similar.

o Strong understanding of cloud platforms (AWS, Azure, GCP) and containerization technologies (Docker, Kubernetes).

o Experience with infrastructure-as-code tools (Terraform, Ansible, etc.) and CI/CD pipelines.

o Deep knowledge of monitoring and observability tools (Prometheus, Grafana, ELK stack, etc.)

Soft Skills:

o Excellent leadership, mentoring, and team-building skills.

o Strong problem-solving and analytical abilities.

o Effective communication and interpersonal skills, with the ability to convey complex technical concepts to non-technical stakeholders.

o Strategic thinking and a proactive approach to identifying and addressing potential issues.

Apply Now

We offer a competitive salary and benefits package and the professional advantages of a dynamic environment that supports your development and recognizes your achievements.