Description:
Requirements:
What You Will Do
- Define and maintain an evaluation strategy for AI-powered workflows, including agent
behavior, output quality, accuracy, completeness, safety, and business-rule compliance.
- Build automated AI evaluations using tools such as DeepEval, Ragas, Promptfoo,
LangSmith, Langfuse, or similar frameworks.
- Create and maintain representative evaluation datasets, golden examples, regression
suites, synthetic test data, and edge-case scenarios.
- Combine deterministic checks, business rules, structured assertions, LLM-as-a-judge
techniques, human-review workflows, and production evidence where appropriate.
- Integrate evaluations into CI/CD pipelines so meaningful AI regressions are detected
before release.
- Establish practical quality thresholds, release signals, and investigation workflows for
changes to prompts, models, retrieval, tools, agent logic, data, and integrations.
- Test APIs, backend services, data flows, third-party integrations, asynchronous
workflows, retries, failure handling, and downstream system outcomes.
- Validate billing and collections workflows end to end, including workflow state, data
transformations, permissions, external-system behavior, and customer-impacting
outcomes.
- Use Langfuse, CloudWatch, logs, traces, dashboards, and data evidence to investigate
failures and distinguish product defects from model behavior, environment issues, test
defects, or accepted risk.
- Partner with Product, Engineering, Operations, and AI platform teams to clarify expected
behavior, identify risks early, and define meaningful validation coverage.
- Build and maintain targeted automated API, integration, and end-to-end tests. Use
Playwright where browser-level coverage is the right way to protect an important
workflow.
- Turn escaped defects, operational issues, and production learnings into stronger
evaluations, regression coverage, observability, or delivery guardrails.
- Provide clear release-readines
📌 Quality Engineer (Bogotá)
🏢 Cafeto
📍 Bogotá