Help evaluate how AI systems solve real mathematical problems. In this short-term project, you will turn research and computational workflows into challenging, reproducible tasks that an AI agent must complete in a Linux terminal.
What You'll Do
• Design multi-step tasks in areas such as numerical analysis, optimization, statistics, probability or mathematical modeling.
• Build working solutions, prepare inputs and expected outputs, and create automated tests that check mathematical correctness.
• Debug numerical issues and document assumptions, tolerances and reproducibility requirements.
What You'll Bring
• A PhD, postdoctoral experience or equivalent advanced technical experience in mathematics, statistics or a closely related field.
• Strong programming skills in Python, R, Julia, C/C++, Bash or another relevant language,
plus experience working in Linux or terminal environments.
• Experience independently implementing and validating computational algorithms, including numerical stability and error analysis.
• Benchmark design, research software, publications and open-source contributions are welcome but not required.
Project details
• Expected project length: five weeks.
• Availability: 40 hours per week, including four hours of overlap with Pacific working hours each day.
• Selection includes completing a short interest form, a 60-minute assessment and a final review/interview.
Compensation
• $300 per approved task.
• Tasks are estimated to take six to eight hours on average.