Anyone AI Labs is seeking a Research Scientist focused on LLM evaluations and benchmarking. You will design frontier-grade evaluation packages across reasoning, coding, agents, and multi-modal capabilities, grounded in expert-verified truth and validated against multiple models.
You’ll own evaluation methodology, collaborate with labs, and contribute to public benchmarks and papers, with a strong emphasis on rigor and QC in a remote, LatAm/US‑centric setup.
📌 LLM Evaluation (Colombia)
🏢 Anyone AI
📍 Colombia
Postulate a este anuncio
Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.