Site Reliability Engineer

1 day ago


Singapore BYTEPLUS PTE. LTD. Full time
Roles & Responsibilities

ByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.


Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.


Why Join Us

Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible.

Together, we inspire creativity and enrich life - a mission we aim towards achieving every day.

To us, every challenge, no matter how ambiguous, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.

At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve.

Join us.


About the Team

The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.


In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.


Responsibilities

Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world.


- Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices.

- Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance.

- Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems.


Qualifications

Minimum Qualifications

- B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions.

- Minimum of 5 years of R&D experience in the fields of cloud computing or large-scale model systems.

- Proficiency in cloud-native technologies and understanding of the relevant technology stack.

- Expertise in one of the following programming languages: Golang, Python, or Java, with the ability to use it proficiently in a professional setting.

- Familiarity with cloud-native technologies for log collection, monitoring, and alerting.


Preferred Qualifications:

- Prior experience in the construction and maintenance of stability systems for large-scale infrastructures.

- Experience in operating and maintaining large-scale systems.

- Experience with infrastructure as code, particularly Terraform, is highly desirable.


ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.


Tell employers what skills you have

Machine Learning
Troubleshooting
Construction
Kubernetes
Cloud Computing
Scripting
Reliability
Networking
Python
Docker
Ansible
Java
Scheduling
Linux

  • Singapore Aptitude Asia Full time

    At Aptitude Asia, we're seeking a skilled Site Reliability Engineer to join our team. This role is crucial in ensuring the high reliability, availability, and performance of our applications throughout their lifecycle.Key Responsibilities:Develop and implement automation scripts to streamline repetitive tasks and address recurring issues.Collaborate with...


  • Singapore BYTEPLUS PTE. LTD. Full time

    Role OverviewAt ByteDance, we're seeking a skilled Site Reliability Engineer to join our team. As a Site Reliability Engineer, you'll be responsible for ensuring the reliability and normal operation of multiple core systems for big data and online computing. This includes building automated operation solutions for large-scale systems, cooperating with the...


  • Singapore Aptitude Asia Full time

    Job SummaryAptitude Asia seeks a skilled Site Reliability Engineer to ensure the high reliability, availability, and performance of applications throughout their lifecycle.Key ResponsibilitiesReliability and Performance: Ensure applications operate with high reliability, availability, and performance.Automation and Innovation: Automate repetitive tasks and...


  • Singapore Oxford Knight Full time

    Job Title: Senior Site Reliability EngineerOxford Knight is seeking a highly skilled Senior Site Reliability Engineer to join our team. As a Senior Site Reliability Engineer, you will be responsible for designing, developing, and maintaining our Linux trading infrastructure on a day-to-day basis.Key Responsibilities:Lead the design and development of major...


  • Singapore BYTEDANCE PTE. LTD. Full time

    About the JobAt ByteDance, we are looking for a talented Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability and normal operation of multiple core systems for big data and online computing, while paying attention to system capacity and stability.Key Responsibilities Ensure the reliability and normal...


  • Singapore LANDI INTERNATIONAL (SINGAPORE) PTE. LTD. Full time

    Landi International (Singapore) PTE. LTD.As a Site Reliability Engineer at Landi International (Singapore) PTE. LTD., you will play a crucial role in ensuring the availability, reliability, and scalability of our platforms. Your primary responsibilities will include:· Building, operating, and maintaining our platform infrastructures across various...


  • Singapore AIA Singapore Private Limited Full time

    About the RoleWe are seeking a highly skilled Senior Site Reliability Engineer to join our team at AIA Singapore Private Limited. As a key member of our operations team, you will be responsible for ensuring the reliability and stability of our critical production services and applications.Key ResponsibilitiesLead complex system and champion services...


  • Singapore ACCESS PEOPLE (SINGAPORE) PTE. LTD. Full time

    Roles & ResponsibilitiesA global energy trading firm is transitioning to a data-centric platform and is seeking a Site Reliability Engineer to support this multi-year program. The role will focus on enhancing the reliability, scalability, and stability of the company's evolving platform. The successful candidate will work on integrating a new event-based,...


  • Singapore Vortexa Full time

    Vortexa is a cutting-edge company that leverages satellite data and AI to provide real-time insights into global energy flows. We're looking for a skilled Site Reliability Engineer to join our Data Services Team, responsible for the developer platform and Amazon AWS estate.The ChallengeOur platform processes massive amounts of data from various sources,...


  • Singapore ASIA GULF CLOUD PTE. LTD. Full time

    Roles & ResponsibilitiesAbout SGB:SGB is a new digital bank that will offer a secure and integrated platform to access andmanage conventional and digital assets and financial solutions, including round-the-clock realtime settlement, trading connectivity, custody and asset management. It serves globalinvestors, innovators and institutions looking for a...


  • Singapore Ripple Labs Singapore Full time

    As a Senior Site Reliability Engineer at Ripple Labs Singapore, you will be responsible for ensuring the high availability and scalability of our systems. Your primary goal will be to design, implement, and maintain a robust and efficient infrastructure that can handle high traffic and complex distributed systems.Key Responsibilities:Design and implement...


  • Singapore NodeFlair Full time

    Senior Site Reliability EngineerWe are working with NodeFlair, a leading pioneer in the Cryptocurrency space, to search for a Senior Site Reliability Engineer to join their Singapore/Remote team.Summary:Our client, a top player in cryptocurrency data monitoring, tracks over 10,000 tokens on 400+ exchanges with 300 million page views from 100+...


  • Singapore ITCAN PTE. LIMITED Full time

    Roles & ResponsibilitiesRoles & Responsibilities:The Site Reliability Engineer (SRE) combines software development and system engineering to build and run distributed solutions in a secured multi-tier heterogeneous environment to safeguard, provide and continuously improve the software and systems behind the organization’s cloud platform solutions.The Job:...


  • Singapore Ripple Labs Singapore Full time

    Job Title: Senior Site Reliability EngineerWe are seeking a highly skilled Senior Site Reliability Engineer to join our team at Ripple Labs Singapore. As a key member of our infrastructure team, you will be responsible for ensuring the high availability and scalability of our systems.Key Responsibilities:Design and Implement High Availability Solutions:...


  • Singapore DBS Bank Limited Full time

    Job Title: Associate, Site Reliability EngineeringDBS Bank Limited is seeking a highly skilled Associate, Site Reliability Engineering to join our team. As a key member of our Group Technology and Operations (T&O) team, you will play a critical role in enabling and empowering the bank with an efficient, nimble, and resilient...


  • Singapore ADDVALUE INNOVATION PTE LTD Full time

    Roles & ResponsibilitiesReliability EngineerResponsibilities Work with product development teams to develop relaibility requirements, establish a reliability / test program and perform appropriate analyse to ensure that new products meet all the relaibility targets. Perform risk / reliabilty analysis (FMEA, FMECA, MTBF) for existing and new products Able...


  • Singapore Hireio, Inc. Full time

    Job SummaryWe are seeking a highly skilled Site Reliability Engineer to join our Machine Learning Systems team. As a Site Reliability Engineer, you will be responsible for ensuring the stability and efficiency of our ML systems, including large model deployment, training, evaluation, and inference.You will work closely with our global team to develop and...


  • Singapore GARENA ONLINE PRIVATE LIMITED Full time

    Job OverviewWe are seeking a skilled Site Reliability Specialist to join our team at GARENA ONLINE PRIVATE LIMITED. The ideal candidate will have a strong background in Linux operating systems, computer networks, and programming languages such as Bash, Python, and Go.Key Responsibilities:Deep dive into development lines, learning and understanding the...


  • Singapore Ripple Labs Singapore Full time

      WHAT YOU’LL DO: Keeping your assigned site or service up and running or rapid recovery from failures Actively troubleshoot any issues that arise during testing and production, catching and solving issues before launch, Automating work including infrastructure needs, testing, failover solutions, failure mitigation, and much more, Monitor and...


  • Singapore JOHN CRANE SINGAPORE PTE LTD Full time

    Roles & ResponsibilitiesPurpose of roleTo provide on-site support and manage reliability contract of mechanical seals for customers in Singapore.Roles and Responsibilities Assist and guide seal installation / removal on site On site pump /seal initial assessment and inspection Commissioning of seal system Perform 5-point checks on the equipment’s prior...