Cloud Native Engineer

3 weeks ago


Singapore TIKTOK PTE. LTD. Full time
Roles & Responsibilities

TikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa.


TikTok is the leading destination for short-form mobile video. Our mission is to inspire creativity and bring joy. TikTok has global offices including Los Angeles, New York, London, Paris, Berlin, Dubai, Singapore, Jakarta, Seoul and Tokyo.


Why Join Us

Creation is the core of TikTok's purpose. Our platform is built to help imaginations thrive. This is doubly true of the teams that make TikTok possible.

Together, we inspire creativity and bring joy - a mission we all believe in and aim towards achieving every day.

To us, every challenge, no matter how difficult, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.

At TikTok, we create together and grow together. That's how we drive impact - for ourselves, our company, and the communities we serve.

Join us.


About the Team

The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.


In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.


Responsibilities

Responsible for Ark Large Model Platform development on Volcano Engine, researching systematic solutions on large model solution implementations and applications in various industries, striving to reduce the IT cost of large model applications, meeting the users' ever-growing demand for intelligent interaction and improving the lifestyle and communications of users in the future world.


- Maintain a large-scale AI cluster and develop state-of-the-art machine learning platforms to support a diverse group of stakeholders.

- Tackle extremely challenging tasks which include, but are not limited to, delivering highly efficient training and inference for large language models, managing extremely effective distributed training jobs across clusters with over 10,000 nodes and GPU chips, and constructing highly reliable ML systems with unparalleled scalability.

- The work encompasses various aspects of LLMOps (Large Language Model Operations), such as resource scheduling, task orchestration, model training, model inference, model management, dataset management, and workflow orchestration.

- Investigate cutting-edge technologies related to large language models, AI, and machine learning at large, such as state-of-the-art distributed training systems with heterogeneous hardware, GPU utilization optimization, and the latest in hardware architecture.

- Employ a variety of technological and mathematical analyses to enhance cluster efficiency and performance.


Qualifications

Minimum Qualifications

- B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions with 5 years of R&D experience in the fields of cloud computing or large-scale model systems.

- Experience in Golang/C++/Cuda development with a solid understanding of Linux systems and popular cloud platforms such as Volcano Engine Cloud, AWS, and Azure Cloud.

- Profound knowledge of cloud-native orchestration technologies like Kubernetes, coupled with experience in large-scale cluster maintenance, job scheduling optimization, and cluster efficiency enhancement with a strong grasp on various foundational areas of computer science, including computer networking, the Linux file system, object storage services, and SQL as well as NoSQL databases.

- Experience in developing ML platforms or MLOps platforms. Experience in distributed machine learning model training, ML model fine-tuning, and deployment.

- Self-motivated, thirst for innovation, collaborative working aptitude, and consistently uphold high standards in coding and documentation quality.


Preferred Qualifications:

- Familiar with High-Performance Computing (HPC) stacks, which may include but is not limited to computing with Cuda/OpenCL, networking with NCCL/MPI/RDMA/DPDK, and model compiling with MLIR/TVM/Trition/LLVM.

- Experience in developing ML platforms or MLOps platforms. Experience in distributed machine learning model training, ML model fine-tuning, and deployment.


TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse voices and to creating an environment that reflects the many communities we reach. We are passionate about this and hope you are too.


Tell employers what skills you have

Machine Learning
Sponsorship
Hardware Architecture
Lifestyle
Scalability
Kubernetes
Azure
Hardware
Cloud Computing
SQL
Networking
Orchestration
Scheduling
Databases
Linux
  • School of Infocomm

    3 weeks ago


    Singapore GVT Government Technology Agency Full time

    [What the role is]School of Infocomm - Lecturer (Cloud Native Infrastructure)[What you will be working on]The school is looking for dynamic individuals with a high degree of self-motivation and the ability to work effectively in a team environment. You will be involved in the following to support our full time and part-time course offerings in areas related...


  • Singapore Dell Full time

    There is a paradigm shift towards cloud-native technologies and methodologies calling for new capabilities. We are looking for interns who are passionate about cloud native architecture to join us in our Cloud Native Architecture Centre of Excellence to co-develop and architect cloud native products, uplift the Cloud Native Ecosystem in Singapore and...

  • Cloud Engineer

    2 weeks ago


    Singapore First Wave Technology Pte. Ltd. Full time

    Job Description:We are looking for an experienced and detail-oriented Cloud Engineer with a minimum of 2 years of hands-on experience in the cloud industry, particularly in Microsoft Azure. The ideal candidate will be responsible for designing, implementing, and maintaining Azure-based cloud solutions, ensuring the highest levels of reliability, security,...

  • Cloud Engineer

    2 weeks ago


    Singapore Evolution Recruitment Solutions Pte. Ltd. Full time

    Cloud EngineerWest, OnsiteCloud Native + Kubernetes Expert!Roles & ResponsibilityInstall, Configure, Maintain and Monitor Kubernetes ClustersAdministration and troubleshooting of Kubernetes environmentMaintaining Kafka, Grafana, EFK related cloud applicationsDeploy and Improve Kubernetes-based solutions and infrastructureCollaborate with Developer, DBA,...

  • Cloud Engineer

    2 weeks ago


    Singapore FIRST WAVE TECHNOLOGY PTE. LTD. Full time

    Roles & ResponsibilitiesJob Description:We are looking for an experienced and detail-oriented Cloud Engineer with a minimum of 2 years of hands-on experience in the cloud industry, particularly in Microsoft Azure. The ideal candidate will be responsible for designing, implementing, and maintaining Azure-based cloud solutions, ensuring the highest levels of...

  • Cloud Engineer

    2 weeks ago


    Singapore EVOLUTION RECRUITMENT SOLUTIONS PTE. LTD. Full time

    Roles & ResponsibilitiesCloud EngineerWest, OnsiteCloud Native + Kubernetes Expert!Roles & Responsibility Install, Configure, Maintain and Monitor Kubernetes Clusters Administration and troubleshooting of Kubernetes environment Maintaining Kafka, Grafana, EFK related cloud applications Deploy and Improve Kubernetes-based solutions and infrastructure ...


  • Singapore TIKTOK PTE. LTD. Full time

    Roles & ResponsibilitiesTikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa. TikTok is the leading destination for short-form mobile video. Our mission is to inspire creativity and bring joy. TikTok has global offices including Los Angeles, New York, London, Paris, Berlin,...

  • Cloud Engineer

    1 day ago


    Singapore FIRST WAVE TECHNOLOGY PTE. LTD. Full time

    Roles & ResponsibilitiesWe are looking for an experienced and detail-oriented Cloud Engineer with a minimum of 2 years of hands-on experience in the cloud industry, particularly in Microsoft Azure. The ideal candidate will be responsible for designing, implementing, and maintaining Azure-based cloud solutions, ensuring the highest levels of reliability,...

  • Cloud Engineer

    1 week ago


    Singapore Matrix Process Automation Pte. Ltd. Full time

    Cloud Engineer (AWS/Azure)Develop, deploy, and manage secure, scalable, and reliable cloud infrastructure on AWS and Azure platforms.Collaborate with the solutions architecture team to implement cloud solutions that meet project and business requirements.Provide technical guidance and support to project teams in the deployment and integration of applications...

  • Cloud Engineer

    7 days ago


    Singapore MATRIX PROCESS AUTOMATION PTE. LTD. Full time

    Roles & ResponsibilitiesCloud Engineer (AWS/Azure) Develop, deploy, and manage secure, scalable, and reliable cloud infrastructure on AWS and Azure platforms. Collaborate with the solutions architecture team to implement cloud solutions that meet project and business requirements. Provide technical guidance and support to project teams in the deployment...


  • Singapore IHS GLOBAL PTE. LTD. Full time

    Roles & ResponsibilitiesWe are seeking an experienced and visionary Director of Cloud Solutions Engineering to lead our talented team of Cloud Engineers. The ideal candidate will drive the development and delivery of innovative cloud solutions, leveraging their expertise in cloud automation, governance, and financial operations (FinOps) controls.Key...


  • Singapore S&P GLOBAL ASIAN HOLDINGS PTE. LTD. Full time

    Roles & ResponsibilitiesWe are seeking an experienced and visionary Director of Cloud Solutions Engineering to lead our talented team of Cloud Engineers. The ideal candidate will drive the development and delivery of innovative cloud solutions, leveraging their expertise in cloud automation, governance, and financial operations (FinOps) controls.Key...

  • Cloud Architect

    2 weeks ago


    Singapore Enggsol Pte. Ltd. Full time

    Apply technical knowledge and customer insights to create a migration and modernization roadmap with customers using the Cloud Adoption Framework (CAF). Architect solutions to meet business and IT needs, ensuring technical viability of new projects and successful deployments, while orchestrating key resources and infusing key Infrastructure technologies...

  • Cloud Architect

    2 weeks ago


    Singapore ENGGSOL PTE. LTD. Full time

    Roles & Responsibilities Apply technical knowledge and customer insights to create a migration and modernization roadmap with customers using the Cloud Adoption Framework (CAF). Architect solutions to meet business and IT needs, ensuring technical viability of new projects and successful deployments, while orchestrating key resources and infusing key...

  • Cloud Architect

    1 week ago


    Singapore Atomic Group Full time

    Cloud Architect The Company: Leading IT systems integrator specializing in designing and implementing innovative technology solutions. With a focus on cloud computing, cybersecurity, and digital transformation, they help organizations leverage the latest technologies to achieve their business objectives. The Position: They are currently hiring for a highly...

  • Lead Cloud Architect

    3 weeks ago


    Singapore Ethos BeathChapman Full time

    Our client is a prestigious financial services company headquartered in Singapore, with global investments in various asset classes and businesses. This Lead Cloud Architect role has been established to bolster the organization\'s cloud strategy and investments, and you will be a key member of the Infrastructure and Platform Division. As a Lead Cloud...

  • Lead Cloud Architect

    3 weeks ago


    Singapore Ethos BeathChapman Full time

    Our client is a prestigious financial services company headquartered in Singapore, with global investments in various asset classes and businesses. This Lead Cloud Architect role has been established to bolster the organization\'s cloud strategy and investments, and you will be a key member of the Infrastructure and Platform Division. As a Lead Cloud...

  • Cloud Devops Engineer

    2 weeks ago


    Singapore Petroseraya Pte. Ltd. Full time

    Location: Singapore, SGCompany: ST Engineering GroupJob Req ID: 2196As a Cloud DevOps Engineer, you will be responsible to deploy and configure solutions in the cloud (Public, Hybrid, Private). You automate cloud operations, develop infrastructure automation scripts and participate in the continuous improvement of cloud solutions. You participate in the...

  • Cloud Devops Engineer

    3 weeks ago


    Singapore PETROS-CONSULTING PTE. LTD. Full time

    Roles & ResponsibilitiesLocation: Singapore, SGCompany: ST Engineering GroupJob Req ID: 2196As a Cloud DevOps Engineer, you will be responsible to deploy and configure solutions in the cloud (Public, Hybrid, Private). You automate cloud operations, develop infrastructure automation scripts and participate in the continuous improvement of cloud solutions. You...


  • Singapore Prudential Assurance Company Singapore (pte) Limited Full time

    As a Cloud Engineering Lead in the software engineering team at Prudential Singapore, you are an integral part of an agile team that works to enhance, build, and deliver trusted market-leading technology products in a secure, stable, and scalable way. Drive significant business impact through your capabilities and contributions, and apply deep technical...