Cloud Native Engineer

1 week ago


Singapore NodeFlair Full time

**Job Summary**:
**Salary**
S$11,250 - S$22,500 / Monthly

**Job Type**

**Seniority**

Mid

**Years of Experience**
At least 5 years

**Tech Stacks**
AWS C++ Go LoRa Azure Linux NoSQL Kubernetes SQL

TikTok will be prioritizing applicants who have a current right to work in Singapore, and do not require TikTok's sponsorship of a visa.

TikTok is the leading destination for short-form mobile video. Our mission is to inspire creativity and bring joy. TikTok has global offices including Los Angeles, New York, London, Paris, Berlin, Dubai, Singapore, Jakarta, Seoul and Tokyo.

**Why Join Us**

Creation is the core of TikTok's purpose. Our platform is built to help imaginations thrive. This is doubly true of the teams that make TikTok possible.

Together, we inspire creativity and bring joy - a mission we all believe in and aim towards achieving every day.

To us, every challenge, no matter how difficult, is an opportunity; to learn, to innovate, and to grow as one team. Status quo? Never. Courage? Always.

At TikTok, we create together and grow together. That's how we drive impact - for ourselves, our company, and the communities we serve.

Join us.

**About the Team**

The Applied Machine Learning (AML) - Enterprise team provides machine learning platform products on VolcanoEngine with cloud native resource scheduling system which intelligently orchestrates different tasks and jobs with minimised costs of every experiment and maximised resource utilisation, rich modelling tools including customised machine learning tasks and web IDE, and multi-framework high performance model inference services.

In 2021, through VolcanoEngine, we released this machine learning infrastructure to the public, to provide more enterprises with reduced costs of computation power, lower barriers to machine learning engineering and deeper developments in AI capabilities.

**Responsibilities**
- Maintain a large-scale AI cluster and develop state-of-the-art machine learning platforms to support a diverse group of stakeholders.
- Tackle extremely challenging tasks which include, but are not limited to, delivering highly efficient training and inference for large language models, managing extremely effective distributed training jobs across clusters with over 10,000 nodes and GPU chips, and constructing highly reliable ML systems with unparalleled scalability.
- The work encompasses various aspects of LLMOps (Large Language Model Operations), such as resource scheduling, task orchestration, model training, model inference, model management, dataset management, and workflow orchestration.
- Investigate cutting-edge technologies related to large language models, AI, and machine learning at large, such as state-of-the-art distributed training systems with heterogeneous hardware, GPU utilization optimization, and the latest in hardware architecture.
- Employ a variety of technological and mathematical analyses to enhance cluster efficiency and performance.

**Qualifications**

**Minimum Qualifications**
- B. Sc or higher degree in Computer Science or related fields from accredited and reputable institutions with 5 years of R&D experience in the fields of cloud computing or large-scale model systems.
- Experience in Golang/C++/Cuda development with a solid understanding of Linux systems and popular cloud platforms such as Volcano Engine Cloud, AWS, and Azure Cloud.
- Profound knowledge of cloud-native orchestration technologies like Kubernetes, coupled with experience in large-scale cluster maintenance, job scheduling optimization, and cluster efficiency enhancement with a strong grasp on various foundational areas of computer science, including computer networking, the Linux file system, object storage services, and SQL as well as NoSQL databases.
- Experience in developing ML platforms or MLOps platforms. Experience in distributed machine learning model training, ML model fine-tuning, and deployment.
- Self-motivated, thirst for innovation, collaborative working aptitude, and consistently uphold high standards in coding and documentation quality.

**Preferred Qualifications**:

- Familiar with High-Performance Computing (HPC) stacks, which may include but is not limited to computing with Cuda/OpenCL, networking with NCCL/MPI/RDMA/DPDK, and model compiling with MLIR/TVM/Trition/LLVM.
- Experience in large language model (LLM) training and development, including large-scale foundational model training aligned with Scaling Laws, efficient fine-tuning techniques such as Lora/P-Tuning/RLHF, model inference optimization, and transformations of model structures for optimizations like Sparse/MoE/LongContext.

TikTok is committed to creating an inclusive space where employees are valued for their skills, experiences, and unique perspectives. Our platform connects people from across the globe and so does our workplace. At TikTok, our mission is to inspire creativity and bring joy. To achieve that goal, we are committed to celebrating our diverse vo



  • Singapore TENCENT CLOUD INTERNATIONAL PTE. LTD. Full time

    **Responsibilities** - Analyze client technical architectures and deliver optimal containerization solutions tailored to client scenarios. - Act as a liaison between clients and product research teams, ensuring prompt resolution of client issues to enhance satisfaction levels. - Design and implement scalable, reliable, and secure cloud native architectures...


  • Singapore beBeeEngineering Full time $9,500

    Job Title: Cloud Native Engineering LeadDescriptionWe are seeking an experienced software professional to design, develop and optimize scalable cloud-based software solutions while leading a team of engineers. Our organization offers a hybrid working arrangement that balances work and personal life.This is a unique opportunity for you to make an impact in...


  • Singapore Luxoft Full time

    **Project description**: We are looking for a talented Cloud Native Developer proficient with building microservices on the back of Pivotal Cloud Foundry. We are looking for someone with strong experience in this area. The person would be a part of an experienced and established team working on a major rebuild of our platform. **Responsibilities**: Develop...


  • Singapore beBeeEngineering Full time

    Cloud Native Software Engineering LeaderWe are seeking a seasoned Cloud Native Software Engineering Leader to spearhead the design, development, and optimization of scalable cloud-based software solutions while mentoring a team of engineers.A hybrid working arrangement offers the perfect blend of work-life balance and collaboration. This is an exceptional...


  • Singapore LION CITY INNOVATIONS PTE. LTD. Full time

    **Job Responsibilities**: **1. Kubernetes Lifecycle Management**: - Oversee the full lifecycle management of Kubernetes clusters across development, testing, and production environments in a multi-cloud setting. - Handle the creation, upgrades, scaling, and decommissioning of Kubernetes clusters with efficiency. **2. User Support**: - Provide expert-level...


  • Singapore beBeeCloud Full time $100,000 - $120,000

    Cloud-Native Software EngineerWe are seeking a highly skilled software engineer with hands-on experience building cloud-native applications. In this role, you will design and implement scalable microservices, containerized applications, and modern deployment pipelines on public cloud platforms.Design, develop, and maintain Java-based microservices using...


  • Singapore beBeeSolutions Full time $120,000 - $250,000

    This is a challenging opportunity to work in AI and cloud native technologies.">">As a Solutions Architect, you will be responsible for designing and developing large-scale cloud native applications, leveraging your expertise in AI and machine learning.Your duties will include identifying new opportunities for our technology and solutions, optimally...


  • Singapore NEWBRIDGE ALLIANCE PTE. LTD. Full time

    Key Responsibilities: - Define and enforce robust cloud-native security practices, including identity and access management, data encryption, and threat prevention, catering to the financial services domain. - Provide guidance and mentorship to development teams, ensuring adherence to cloud-native best practices, design patterns, and development...


  • Singapore Lenovo Full time

    Why Work at Lenovo We are Lenovo. We do what we say. We own what we do. We WOW our customers. Lenovo is a US$69 billion revenue global technology powerhouse, ranked #196 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the...


  • Singapore beBeeSolution Full time $120,000 - $160,000

    Solution Architect RoleAbout the Position:As a seasoned Solution Architect, you will collaborate with our esteemed client to design and develop cloud-native applications from start to finish.You will be part of an agile team that delivers tailored products for the financial sector. Your key responsibility will involve working closely with business teams and...