- Create complex, domain-specific prompts to test LLM capabilities.
- Evaluate AI-generated responses for factual accuracy, reasoning, and quality.
- Identify factual errors, logical inconsistencies, and model weaknesses.
- Develop challenging test cases that expose model limitations.
- Provide clear, evidence-based feedback to improve model performance.
Requirements
- Hold a Master's degree or higher in a relevant domain.
- Have a minimum of of professional, research, or teaching experience.
- Possess excellent written English communication skills.
- Demonstrate strong analytical skills and exceptional attention to detail.