
Site Reliability Engineer, Ark Large Model Platform
2 weeks ago
ByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa.
Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content.
Why Join Us
Creation is the core of ByteDance's purpose. Our products are built to help imaginations thrive. This is doubly true of the teams that make our innovations possible.
Together, we inspire creativity and enrich life - a mission we aim towards achieving every day.
To us, every challenge, no matter how ambiguous, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.
At ByteDance, we create together and grow together. That's how we drive impact - for ourselves, our company, and the users we serve.
Join us.
About the Team
The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.
In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.
**Responsibilities**:
- Manage and oversee the stability of both control and data aspects of large-scale model systems through effective DevOps practices.
- Develop and enhance observability systems for monitoring the stability of large model systems, ensuring high reliability and performance.
- Handle super large-scale cluster management and ensure efficient operation and maintenance of large model systems.
**Qualifications**:
Minimum Qualifications
- B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions with R&D experience in the fields of cloud computing or large-scale model systems.
- Proficiency in cloud-native technologies and understanding of the relevant technology stack.
- Expertise in one of the following programming languages: Golang, Python, or Java, with the ability to use it proficiently in a professional setting.
- Familiarity with cloud-native technologies for log collection, monitoring, and alerting.
Preferred Qualifications:
- Prior experience in the construction and maintenance of stability systems for large-scale infrastructures.
- Experience in operating and maintaining large-scale systems.
- Experience with infrastructure as code, particularly Terraform, is highly desirable.
ByteDance is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At ByteDance, our mission is to inspire creativity and enrich life. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.
-
Singapore NodeFlair Full time**Job Summary**: **Salary** S$6,500 - S$13,000 / Monthly **Job Type** **Seniority** Junior **Years of Experience** At least 1 year **Tech Stacks** Terraform Go Java Python TikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa. TikTok is the leading destination for...
-
Site Reliability Engineer
7 days ago
Singapore NodeFlair Full time**Job Summary**: **Salary** S$11,250 - S$22,500 / Monthly **Job Type** **Seniority** Mid **Years of Experience** At least 5 years **Tech Stacks** Terraform Go Java Python TikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa. TikTok is the leading destination for...
-
Site Reliability Engineer
2 weeks ago
Singapore BYTEPLUS PTE. LTD. Full timeByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa. Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market,...
-
Site Reliability Engineer
7 days ago
Singapore ByteDance Full time**Location**: Singapore **Team**: Backend **Employment Type**: Regular **Job Code**: A118803 **Responsibilities**: About the Team The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with...
-
Site Reliability Engineer
7 days ago
Singapore ByteDance Full time**Location**: Singapore **Team**: Backend **Employment Type**: Regular **Job Code**: A67178 **Responsibilities**: ByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa. About the Team The Applied Machine Learning (AML) - Enterprise team provides machine learning platform...
-
Large Language Model Algorithm Engineer
7 days ago
Singapore ByteDance Full timeFounded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content. Why Join...
-
Large Language Model Algorithm Engineer
1 week ago
Singapore BYTEPLUS PTE. LTD. Full timeFounded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market, including Toutiao, Douyin, and Xigua, ByteDance has made it easier and more fun for people to connect with, consume, and create content. **Why Join...
-
Cloud Native Engineer, Ark Large Model Platform
9 hours ago
Singapore ByteDance Full timeByteDance will be prioritizing applicants who have a current right to work in Singapore, and do not require ByteDance's sponsorship of a visa. Founded in 2012, ByteDance's mission is to inspire creativity and enrich life. With a suite of more than a dozen products, including TikTok, Helo, and Resso, as well as platforms specific to the China market,...
-
Cloud Native Engineer
1 week ago
Singapore NodeFlair Full time**Job Summary**: **Salary** S$11,250 - S$22,500 / Monthly **Job Type** **Seniority** Mid **Years of Experience** At least 5 years **Tech Stacks** AWS C++ Go LoRa Azure Linux NoSQL Kubernetes SQL TikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa. TikTok is the leading...
-
Singapore NodeFlair Full time**Job Summary**: **Salary** S$6,500 - S$13,000 / Monthly **Job Type** **Seniority** Mid **Years of Experience** At least 3 years **Tech Stacks** AWS C++ Go Azure Linux NoSQL Kubernetes SQL TikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa. TikTok is the leading...