We are looking for experienced AI Benchmark Quality Reviewers to support an advanced AI evaluation project focused on improving the quality and reliability of AI-generated solutions.
The role involves reviewing AI evaluation tasks, analyzing model outputs, validating grading logic, and providing evidence-based feedback to improve benchmark quality. ID 107627 _ AI Benchmark Qualit…
What You’ll Do
• Review task instructions, source materials, reference solutions, and evaluation criteria for accuracy and completeness
• Analyze AI agent execution traces, tool calls, and generated outputs
• Identify issues in grading logic, expected answers, and evaluation criteria
• Investigate whether failures are caused by model limitations, task issues, grader errors, or environment problems
• Provide clear,
structured feedback and document findings with supporting evidence ID 107627 _ AI Benchmark Qualit…
Requirements
• 5+ years of relevant professional experience
• Comfortable reading and understanding:
• Python
• SQL
• Shell scripts
• Structured data
• Execution logs
• Strong analytical and problem-solving skills
• Ability to verify calculations, reconcile conflicting information, and assess technical deliverables
• Strong written English communication skills
• High attention to detail and ability to provide clear, reproducible feedback
Skills: data analysis,technical quality assurance,ai evaluation
📌 AI Benchmark Quality Reviewer – AI Evaluation Project (Colombia)
🏢 eDataBae
📍 Colombia
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.