Caseware is seeking an experienced applied scientist to design and run experiments for evaluating and improving LLM‑based applications. You will own the evaluation framework, build AI tools for the end‑to‑end agent builder, and develop synthetic data and eval builders.
You will shape agentic memory and ensure regulatory compliance while mentoring other engineers. You will work on cutting‑edge analytics inside a fully remote role based in Colombia, collaborating with Caseware teams and clients