LLM Evaluation Engineer – Benchmarking

10 hours ago

singapore UMELIFE (SINGAPORE) PTE. LTD. Full-time

UMELIFE (SINGAPORE) PTE. LTD. seeks an AI evaluation engineer to build and maintain an automated LLM evaluation pipeline covering general capabilities, agent capabilities, and persona/role-playing evaluations with one-click assessment and regression testing.

You will execute benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, manage results across training runs, and prepare checkpoint reports to guide model development.