Anyone AI Labs is seeking a Research Scientist focused on LLM evaluations and benchmarking. You will design frontier-grade evaluation packages across reasoning, coding, agents, and multi-modal capabilities, grounded in expert-verified truth and validated against multiple models. You’ll own evaluation methodology, collaborate with labs, and contribute to public benchmarks and papers, with a strong emphasis on rigor and QC in a remote, LatAm/US‑centric setup.