Site Reliability Engineer
1 week ago
Join us as we work together to inspire creativity and enrich life around the globe. Location: Team: Employment Type: Regular Job Code: A67178 Share this listing: Responsibilities ByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa. About the Team The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services. In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities. Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world. Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices. Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance. Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems. Qualifications Minimum Qualifications B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions. Minimum of 5 years of R&D experience in the fields of cloud computing or large-scale model systems. Proficiency in cloud-native technologies and understanding of the relevant technology stack. Expertise in one of the following programming languages: Golang, Python, or Java, with the ability to use it proficiently in a professional setting. Familiarity with cloud-native technologies for log collection, monitoring, and alerting. Preferred Qualifications Prior experience in the construction and maintenance of stability systems for large-scale infrastructures. Experience in operating and maintaining large-scale systems. Experience with infrastructure as code, particularly Terraform, is highly desirable. Job Information About Us Why Join ByteDance Inspiring creativity is at the core of ByteDance's mission. Our innovative products are built to help people authentically express themselves, discover and connect – and our global, diverse teams make that possible. Together, we create value for our communities, inspire creativity and enrich life - a mission we work towards every day. As ByteDancers, we strive to do great things with great people. We lead with curiosity, humility, and a desire to make impact in a rapidly growing tech company. By constantly iterating and fostering an "Always Day 1" mindset, we achieve meaningful breakthroughs for ourselves, our Company, and our users. When we create and grow together, the possibilities are limitless. Join us. Diversity & Inclusion ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too. #J-18808-Ljbffr
-
Site Reliability Engineer
2 weeks ago
Singapore NTT Data Singapore Full timeAs a Site Reliability Engineer you will be filling a mission-critical role ensuring that our systems are healthy, monitored, automated, fault tolerant and designed to scale. You will collaborate and work closely with engineering teams to continually improve our production services, facilitating fast delivery of new products, and reducing downtime. Key...
-
Site Reliability Engineer
1 week ago
Singapore RigNet Full timeAbout us One team. Global challenges. Infinite opportunities. At Viasat, we’re on a mission to deliver connections with the capacity to change the world. For more than 35 years, Viasat has helped shape how consumers, businesses, governments and militaries around the globe communicate. We’re looking for people who think big, act fearlessly, and create an...
-
Site Reliability Engineer
1 week ago
Singapore Viasat Full timeAbout us One team. Global challenges. Infinite opportunities. At Viasat, we’re on a mission to deliver connections with the capacity to change the world. For more than 35 years, Viasat has helped shape how consumers, businesses, governments and militaries around the globe communicate. We’re looking for people who think big, act fearlessly, and create an...
-
Site Reliability Engineer
2 weeks ago
Singapore NodeFlair Full time**Job Summary**: **Salary** S$11,500 - S$16,500 / Monthly **Job Type** **Seniority** Senior **Years of Experience** At least 7 years **Tech Stacks** Microsoft Puppet Java Ansible Python **This is Adyen** Adyen provides payments, data, and financial products in a single solution for customers like Meta, Uber, H&M, and Microsoft - making us the...
-
Site Reliability Engineer
1 week ago
Singapore Rapsys Technologies Full timeDrive the Site Reliability Engineering agenda forward at an Enterprise Level to improve availability, reliability, and performance of services. - Drive cross-team efforts in resiliency assessment exercises and reporting - Draft and/or contribute to internal SRE training materials - Support services before they go live through activities such as Chaos testing...
-
Site Reliability Engineer
6 days ago
Singapore Imperva Full time**Site Reliability Engineer**:** About the role** Imperva’s Infrastructure and Cloud team is looking for a highly technical Site Reliability Engineer to drive innovation, scale, and create operational excellence for the Imperva globally distributed network. As an SRE in the ICO organization, you approach solving, supporting, and optimizing the...
-
Site Reliability Engineer
1 week ago
Singapore Point72 Full timeJoin to apply for the Site Reliability Engineer role at Point72 About the role As part of Point72’s Technology Team, you will focus on developing and maintaining complex, distributed, real-time systems that support our Global Macro business. Your responsibilities will include optimizing operations through automation, building foundational SRE components,...
-
Site Reliability Engineer
6 days ago
Singapore DT One Full timeAbout DT One DT One was founded to provide mobile carriers with the infrastructure and services they need to help migrant workers stay in touch with their family and friends back home. Today we operate a leading global network for mobile top‑up solutions, innovative mobile rewards, and Phone‑to‑Phone solutions. Our global network delivers better...
-
Site Reliability Engineer
2 weeks ago
Singapore Pan Asia Group Resources Full time**Key Responsibilities**: - Drive Site Reliability Engineering agenda to improve availability, reliability, and performance of services - Drive optimise-operate initiative, example, reduction of operation toil - Work with enterprise team in deploying SRE enablers/initiatives. - Strong background in machine learning and deep learning algorithms. -...
-
Site Reliability Engineer
8 hours ago
Singapore IFUN GAMES Full time**Responsibilities** - Design, implement, and maintain tools and processes for monitoring, alerting, and incident response - Collaborate with developers to improve the design and operation of systems, with a focus on reliability, performance, and scalability - Participate in on-call rotations to respond to incidents and handle escalations - Analyze system...