LLM Evaluation Engineer – Benchmarking
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
UMELIFE (SINGAPORE) PTE. LTD. seeks an AI evaluation engineer to build and maintain an automated LLM evaluation pipeline covering general capabilities, agent capabilities, and persona/role-playing evaluations with one-click assessment and regression testing.
You will execute benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, manage results across training runs, and prepare checkpoint reports to guide model development.