Company Description
Technology is our how. And people are our why. For over two decades, we have been harnessing technology to drive meaningful change.
By combining world-class engineering, industry expertise and a people-centric mindset, we consult and partner with leading brands from various industries to create dynamic platforms and intelligent digital experiences that drive innovation and transform businesses.
From prototype to real-world impact - be part of a integral shift by doing work that matters.
We are seeking a hands-on Site Reliability Engineer (SRE) / AI Platform DevOps Engineer to own infrastructure provisioning, CI/CD automation, telemetry pipelines, and production deployment for AI-powered services, agents, and orchestration systems.
This is an SRE-heavy, infrastructure-first role, focused on ensuring AI systems operating in production are:
- Reliable
- Observable
- Scalable
- Secure
- Cost-efficient
- Safe to deploy and operate
You will play a critical role in building and maintaining the platform foundation that enables AI services to run safely and efficiently at scale.
Key Responsibilities
1. Infrastructure Provisioning & Automation
- Design and manage cloud infrastructure using Infrastructure as Code (Terraform or similar)
- Provision and maintain Kubernetes clusters and supporting services
- Automate environment setup across development, staging, and production
- Manage networking, IAM, secrets, storage, and compute scaling
- Ensure high availability, resilience, and disaster recovery readiness
2. CI/CD & Deployment Engineering
- Build and maintain CI/CD pipelines for:
- AI services
- Agent frameworks
- Orchestrators
- Model artifacts
- Implement automated testing and reliability validation gates
- Enable blue/green and canary deployments
- Build safe rollback mechanisms for services and models
- Integrate reliability and health checks into deployment workflows
3. Model & Agent Deployment Governance
- Package, version, and deploy models
📌 DevOps Engineer (Cali)
🏢 Endava
📍 Cali