22 ago
|
Jobtailor
|
Bogotá
- Support and maintain business-critical production environments, ensuring high availability and system reliability.
- Monitor infrastructure and applications, proactively identifying and resolving issues before they impact users.
- Participate in incident response activities, troubleshooting production outages and coordinating recovery efforts.
- Perform RCA and contribute to postmortems, corrective actions, and continuous improvement initiatives.
- Manage and optimize Kubernetes clusters and cloud infrastructure.
- Develop and maintain monitoring dashboards, alerts, and observability solutions.
- Automate operational processes and infrastructure deployments using IaC and scripting.
- Collaborate with engineering and product teams to improve scalability, performance, and operational excellence.
- Support and enhance CI/CD pipelines to ensure reliable and efficient software delivery.
Requirements
- 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering.
- Strong hands-on experience with Linux administration, troubleshooting, and production support.
- Experience managing and supporting Kubernetes and containerized workloads (Docker/OpenShift is a plus).
- Solid knowledge of AWS, Azure, or GCP cloud environments.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, ELK, or CloudWatch.
- Experience with Infrastructure as Code (Terraform preferred) and CI/CD pipelines.
- Ability to troubleshoot complex production issues, perform Root Cause Analysis (RCA), and drive preventive improvements.
- Working knowledge of automation and scripting using Bash, Python, or Go.
- Intermediate to advanced English (B2+).
Core Competencies
Demonstrates expertise in Site Reliability Engineering and DevOps practices, with a strong focus on managing Kubernetes clusters, cloud infrastructure, and automation through Infrastructure as Code. Proficient in monitoring and observability tools to ensure high availability and system reliability.
Highest-signal resume keywords
- Site Reliability Engineering
- Kubernetes Management
- Cloud Infrastructure (AWS, Azure, GCP)
- Infrastructure as Code (Terraform)
- Monitoring and Observability Tools (Prometheus, Grafana, Datadog)
ATS Optimization Keywords
Hard Skills
- Linux Administration
- Troubleshooting
- Root Cause Analysis (RCA)
- Automation and Scripting (Bash, Python, Go)
- CI/CD Pipelines
Industry Keywords
- DevOps
- Cloud Operations
- Infrastructure Engineering
- Production Support
- Continuous Improvement
Tools & Technologies
- Kubernetes
- Docker
- OpenShift
- Prometheus
- Grafana
- Datadog
- Splunk
- ELK
- CloudWatch
#J-18808-Ljbffr
📌 SRE Software Engineer (Bogotá)
🏢 Jobtailor
📍 Bogotá