Senior AI Harness Engineer
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
By continuing, you agree to our Terms & Privacy Policy.
Avensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation and managed services. Given our decade of success we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail, and supply chain.
We are currently looking to hire Senior AI Harness Engineer.
This is an exciting opportunity to expand your skill set, achieve job satisfaction and work-life balance. More details as below.
Job Role:
Job Description
Senior AI Harness Engineer
Position Overview
We are looking for a Senior AI Harness Engineer to design, build, and maintain the engineering frameworks, evaluation systems, testing infrastructure, and tooling required to develop reliable, scalable, and production-ready AI/LLM applications. The role will focus on building AI harnesses that enable systematic testing, evaluation, benchmarking, observability, and continuous improvement of Generative AI and Agentic AI systems.
The ideal candidate will have strong experience in Python, LLMs, Generative AI, AI agents, evaluation frameworks, RAG, prompt engineering, API integration, cloud platforms, and MLOps/LLMOps, with the ability to build robust engineering solutions around AI models.
Key Responsibilities
. Design and develop AI/LLM evaluation and testing harnesses for Generative AI and Agentic AI applications.
. Build reusable frameworks for model evaluation, prompt testing, regression testing, benchmarking, and performance validation.
. Develop automated test suites to evaluate accuracy, relevance, groundedness, hallucination, toxicity, safety, latency, cost, and response quality.
. Create testing frameworks for LLM-based agents, multi-agent workflows, RAG pipelines, and tool-calling applications.
. Develop and maintain Python-based AI engineering frameworks, utilities, SDKs, and automation tools.
. Implement LLMOps/MLOps pipelines for model, prompt, dataset, and evaluation lifecycle management.
. Build automated CI/CD pipelines for AI applications, including model and prompt regression testing.
. Integrate AI evaluation tools and frameworks such as MLflow, LangSmith, DeepEval, Ragas, Azure AI evaluation, OpenAI evaluation frameworks, or equivalent technologies.
. Design evaluation datasets, test cases, golden datasets, benchmark datasets, and synthetic test data.
. Develop automated mechanisms for prompt/version management and experiment tracking.
. Implement observability and monitoring for AI applications, including LLM traces, token usage, latency, errors, model performance, and cost.
. Develop harnesses for testing RAG systems, including retrieval quality, chunking strategies, embeddings, vector search, reranking, and grounded responses.
. Build evaluation capabilities for function calling, API/tool invocation, MCP-based integrations, and agentic workflows.
. Work with AI/ML engineers and application teams to identify failure patterns and improve model/application performance.
. Establish engineering standards for AI quality, reliability, security, scalability, and responsible AI.
. Troubleshoot complex issues across application code, LLM APIs, vector databases, cloud services, and AI infrastructure.
. Build dashboards and reports to communicate AI evaluation and quality metrics to technical stakeholders.
. Mentor junior engineers and contribute to architecture and technical design decisions.
Required Technical Skills
Programming
. Strong hands-on experience with Python.
. Good knowledge of REST APIs, JSON, asynchronous programming, SDK development, and microservices.
. Experience with Git, GitHub/GitLab/Bitbucket, code reviews, and software engineering best practices.
Generative AI & LLM
. Strong understanding of LLMs and Generative AI application architecture.
. Experience with models such as OpenAI, Azure OpenAI, Anthropic, Gemini, Llama, Mistral, or equivalent.
. Strong knowledge of:
o Prompt Engineering
o RAG
o Embeddings
o Vector databases
o Function/Tool Calling
o Structured outputs
o Agentic AI
o LLM APIs
o Context management
o Model evaluation
AI/Agent Frameworks
Hands-on experience with one or more:
. LangChain
. LangGraph
. LlamaIndex
. Semantic Kernel
. AutoGen
. CrewAI
. OpenAI Agents SDK
. MCP (Model Context Protocol)
AI Evaluation & Testing
Experience with A