AI Evaluation Engineer at Confidential Client (C2C)

Job Title: : Al Evaluation Engineer Location: Tampa Florida 12+ Months Contract Opportunity Hybrid Opportunity Skills : Al Evaluation, LangSmith. Langfuss, Ragas, DeepEval, MLflow. Job Description: We are seeking an experienced Al Evaluation Engineer to define and implement evaluation strategies and success metrics for Al agents. In this role, you will be responsible for assessing task completion, correctness, groundedness, hallucination, tool usage, safety, consistency, and overall user experience of Al-driven systems. You will build comprehensive test scenarios, including edge and adversarial cases, and develop automated evaluation pipelines leveraging deterministic checks, LLM-as-a-Judge techniques, and custom evaluators. Your analysis will involve detailed examination of agent traces, prompts, retrievals, tool calls, and outputs to identify failure modes and performance bottlenecks. You will also conduct regression testing spanning changes in models, prompts, tools, and workflows. Establishing quality thresholds, monitoring dashboards, and release gates will be key to maintaining robust Al agent deployments. This role requires close collaboration with Al engineering, product, business, and risk teams to ensure continuous improvement in agent functionality and safety. Key Responsibilities: Define evaluation strategies and success metrics for Al agents. Evaluate task completion, correctness, grounded Ness, hallucination, tool usage, safety, consistency, and user experience. Build test scenarios, edge cases, and adversarial cases. Develop automated eval pipelines using deterministic checks, LLM-as-a-Judge, and custom evaluators. Analyze agent traces, prompts, retrieval, tool calls, and outputs to identify failure patterns. Perform regression testing across model, prompt, tool, and workflow changes. Establish quality thresholds, dashboards, and release gates. Partner with Al engineering, product, business, and risk teams to continuously improve agent performance. Skills – understanding of GenAI, LLMS, RAG, prompts, tool calling, and Al agents.. Python and experience working with APIs, JSON, datasets, and evaluation frameworks.. Ability to translate business expectations into measurable evaluation criteria. Experience with Al evaluation / observability tools such as LangSmith, Langfuse, Ragas, DeepEval, MLflow, or similar is preferred.. Strong analytical, problem-solving, and experimentation mindset. Education: At least a bachelor’s degree (or equivalent experience) in Computer Science, Software/Electronics Engineering, Information Systems, or a closely related field is required for the project Contact Information Email: manasa.s@itechus.net Click the email address to contact the job poster directly. Related