Site Reliability Engineer, ARK Large Model Platform
3 weeks ago
ByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.
Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.
Why Join Us
Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible.
Together, we inspire creativity and enrich life - a mission we aim towards achieving every day.
To us, every challenge, no matter how ambiguous, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.
At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve.
Join us.
About the Team
The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.
In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.
Responsibilities
Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world.
- Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices.
- Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance.
- Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems.
Qualifications
Minimum Qualifications
- B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions with R&D experience in the fields of cloud computing or large-scale model systems.
- Proficiency in cloud-native technologies and understanding of the relevant technology stack.
- Expertise in one of the following programming languages: Golang, Python, or Java, with the ability to use it proficiently in a professional setting.
- Familiarity with cloud-native technologies for log collection, monitoring, and alerting.
Preferred Qualifications:
- Prior experience in the construction and maintenance of stability systems for large-scale infrastructures.
- Experience in operating and maintaining large-scale systems.
- Experience with infrastructure as code, particularly Terraform, is highly desirable.
ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.
Tell employers what skills you have
Machine Learning
Troubleshooting
Construction
Kubernetes
Cloud Computing
Scripting
Reliability
Networking
Python
Docker
Ansible
Java
Scheduling
Linux
-
Site Reliability Engineer
3 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeRoles & ResponsibilitiesByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the...
-
Large Language Model Algorithm Engineer
2 months ago
Singapore BYTEDANCE PTE. LTD. Full timeRoles & ResponsibilitiesFounded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create...
-
Cloud Native Engineer, ARK Large Model Platform
3 weeks ago
Singapore BYTEDANCE PTE. LTD. Full timeRoles & ResponsibilitiesByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the...
-
Cloud Native Engineer
3 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeRoles & ResponsibilitiesByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the...
-
Backend Engineer
3 weeks ago
Singapore BYTEDANCE PTE. LTD. Full timeRoles & ResponsibilitiesByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the...
-
Backend Engineer, ARK Large Model Platform
3 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeRoles & ResponsibilitiesByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the...
-
Large Language Model Algorithm Engineer
3 weeks ago
Singapore BYTEDANCE PTE. LTD. Full timeAbout the RoleByteDance PTE. LTD. is seeking a talented AI researcher to join our Machine Learning Platform team. As a key member of our team, you will contribute to the advancement of next-generation artificial intelligence technologies, including large models, multimodal capabilities, text comprehension, generation algorithms, and reinforcement learning...
-
Singapore BYTEDANCE PTE. LTD. Full timeAbout the RoleAs a Senior Software Engineer, Large Model Development at ByteDance PTE. LTD., you will be responsible for the development of the Ark Large Model Platform on Volcano Engine. This involves researching and implementing systematic solutions for large model applications in various industries, with a focus on reducing the IT cost of large models and...
-
Product Solution Architect
3 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeRoles & ResponsibilitiesByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the...
-
Product Solution Architect, Volcano ARK
3 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeRoles & ResponsibilitiesByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the...
-
Singapore Jobscentral Full timeAssume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.As a Lead Site Reliability Engineer at JPMorgan Chase within the Chief Technology Office, you hold a leadership role in your team, demonstrate strong knowledge across multiple...
-
Cloud Native Engineer for Large Model Platform
3 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeRoles and ResponsibilitiesAt ByteDance PTE. LTD., we are seeking an exceptional Cloud Native Engineer for Large Model Platform to join our team. The ideal candidate will have a strong background in cloud computing and large-scale model systems, with expertise in Golang, C++, Cuda, and Linux systems. Key Responsibilities:Develop and maintain large-scale AI...
-
Site Reliability Engineer
3 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeRole OverviewAt ByteDance, we're seeking a skilled Site Reliability Engineer to join our team. As a Site Reliability Engineer, you'll be responsible for ensuring the reliability and normal operation of multiple core systems for big data and online computing. This includes building automated operation solutions for large-scale systems, cooperating with the...
-
Site Reliability Engineer
3 weeks ago
Singapore BYTEDANCE PTE. LTD. Full timeAbout the JobAt ByteDance, we are looking for a talented Site Reliability Engineer to join our team. In this role, you will be responsible for ensuring the reliability and normal operation of multiple core systems for big data and online computing, while paying attention to system capacity and stability.Key Responsibilities Ensure the reliability and normal...
-
Singapore Goldman Sachs Group, Inc. Full timeYour Impact Site Reliability Engineering (SRE) is an engineering discipline that combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. At Goldman Sachs, SRE is responsible for improving the availability and reliability of some of the firm’s most critical platform services, and ensures they...
-
Site Reliability Engineer
5 days ago
Singapore LANDI INTERNATIONAL (SINGAPORE) PTE. LTD. Full timeJob DescriptionWe are seeking a highly skilled Site Reliability Engineer to join our team at LANDI INTERNATIONAL (SINGAPORE) PTE. LTD. in Singapore.Role Summary:The successful candidate will be responsible for the operation and maintenance of our infrastructures, ensuring their reliability and performance while learning from senior engineers.Main...
-
Site Reliability Engineer
3 weeks ago
Singapore Hireio, Inc. Full timeJob SummaryWe are seeking a highly skilled Site Reliability Engineer to join our Machine Learning Systems team. As a Site Reliability Engineer, you will be responsible for ensuring the stability and efficiency of our ML systems, including large model deployment, training, evaluation, and inference.You will work closely with our global team to develop and...
-
Site Reliability Engineer
3 weeks ago
Singapore LANDI INTERNATIONAL (SINGAPORE) PTE. LTD. Full timeLandi International (Singapore) PTE. LTD.As a Site Reliability Engineer at Landi International (Singapore) PTE. LTD., you will play a crucial role in ensuring the availability, reliability, and scalability of our platforms. Your primary responsibilities will include:· Building, operating, and maintaining our platform infrastructures across various...
-
Senior Site Reliability Engineer
3 weeks ago
Singapore ACCESS PEOPLE (SINGAPORE) PTE. LTD. Full timeRoles & ResponsibilitiesA global energy trading firm is transitioning to a data-centric platform and is seeking a Senior Site Reliability Engineer to support this multi-year program. The role will focus on enhancing the reliability, scalability, and stability of the company's evolving platform. Key Responsibilities: Establish and track SLOs/SLIs, ensuring...
-
Site Reliability Engineer
3 weeks ago
Singapore ACCESS PEOPLE (SINGAPORE) PTE. LTD. Full timeRoles & ResponsibilitiesA global energy trading firm is transitioning to a data-centric platform and is seeking a Site Reliability Engineer to support this multi-year program. The role will focus on enhancing the reliability, scalability, and stability of the company's evolving platform. The successful candidate will work on integrating a new event-based,...