03 sep
|
AgileEngine
|
Valle del Cauca
03 sep
AgileEngine
Valle del Cauca
DevOps / Site Reliability Engineer ID*****
Full time | AgileEngine | Colombia
Posted On 08/29/2026
Job Information
City Cali
State/Province Valle del Cauca
******
IT Services
Job Description
AgileEngine is an Inc. **** company that creates award-winning software for Fortune 500 brands and trailblazing startups across 17+ industries.
We rank among the leaders in areas like application development and AI/ML, and our people-first culture has earned us multiple Best Place to Work awards.
WHY JOIN US
If you're looking for a place to grow, make an impact, and work with people who care, we'd love to meet you!
ABOUT THE ROLE
We are looking for a
DevOps / Site Reliability Engineer
to maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment.
This role blends platform engineering with incident command, using Terraform, CI/CD pipelines, and CSPM tools like Wiz.
You will lead major-incident calls, own remediation follow-through, and build the playbooks that guide response.
WHAT YOU WILL DO
- Scale and maintain the ability to drive operational stability across multi-cloud environments (Azure, AWS, GCP).
- Engineer unified security policies and configuration baselines using IaC (Terraform) to prevent misconfigurations.
- Design, maintain, and optimize enterprise CI/CD pipelines to support continuous ASPM ingestion and deployment.
- Act on continuous monitoring alerts, utilizing Cloud Security Posture Management (CSPM) tools like Wiz to secure workloads.
- Serve as Incident Commander on major and critical incidents — running the bridge, directing technical workstreams,
making time-critical decisions, and coordinating cross-functional responders under pressure.
- Own the post-incident loop — track remediation items to closure, hold owning teams accountable to timelines, and drive systemic fixes and preventative actions across groups.
- Draft and send clear, accurate, audience-appropriate incident notifications and status updates to technical teams, management, and stakeholders throughout the incident lifecycle.
- Develop, maintain, and socialize divisional / group-level incident-management playbooks, runbooks, and escalation procedures that standardize response and reduce time-to-resolution.
MUST HAVES
-
5+ years of experience
.
- In-depth architectural expertise in
multi-cloud defense, federated IAM, and zero-trust principles
.
- Strong practical experience with
Kubernetes, Terraform, CI/CD orchestration, and Python/Go scripting
.
- Senior-level, hands-on
incident-command experience
driving major/critical incident calls to resolution in a
24x7 production environment
.
- Proven track record of
remediation follow-up
— coordinating with teams and holding owners accountable until issues are fully closed.
- Demonstrated skill drafting and issuing
incident notification communications
to both technical and executive audiences.
- Direct experience authoring
divisional/group incident-management playbooks
and escalation procedures.
- Fully autonomous.
- Drives the architecture of
complex automated runbooks
and mentors Middle-level SREs.
- Extensive experience deploying and tuning APIs from modern
CNAPP/CSPM platforms, ideally Wiz
.
- Prior experience building platforms subject to strict financial compliance standards (
PCI-DSS, SOC2
).
NICE TO HAVES
- PagerDuty — hands-on experience with on-call scheduling, alert routing, and incident orchestration.
- ServiceNow — familiarity with incident, problem, and change management workflows and reporting.
PERKS AND BENEFITS
-
Growth without limits
: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
-
Competitive compensation
: get recognition that reflects your skills and impact, with regular performance and compensation reviews
-
Flexibility
: work 100% remotely with versátil hours that support focus, autonomy, and a healthy work rhythm
-
Meaningful, modern projects
: build impactful products using modern technologies alongside global teams and leading brands
-
Collaborative culture
: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
-
Well-being & support
: access local well-being programs and people-focused support tailored to your location
#J-*****-Ljbffr
📌 Devops / Site Reliability Engineer Id70127 (Valle del Cauca)
🏢 AgileEngine
📍 Valle del Cauca