Service Reliability Engineer (Colombia)

Service Reliability Engineer (Colombia)

01 ago
|
1083 Amadeus IT Group Colombia
|
Colombia

01 ago

1083 Amadeus IT Group Colombia

Colombia

Service Reliability Engineer Location: Bogotá Open to SRE's with different technical backgrounds and all levels of experience, including cloud-native, platform, and application-focused reliability engineering.

Summary of the role We are looking for Site Reliability Engineers (SREs) to join our global teams supporting mission‑critical airline platforms and systems. In this role, you will focus on ensuring system reliability, availability, scalability, and performance, while driving automation and operational excellence across distributed environments. You will collaborate with globally distributed teams in a follow‑the‑sun model, supporting production systems and continuously improving reliability and operational processes.

Key Responsibilities - Reliability & Production Operations: Ensure high availability, scalability, and resilience of production systems

- Conduct incident, problem, and change management following ITIL practices
- Perform root cause analysis (RCA) and drive resolution of production issues
- Support systems across multiple environments (test, staging, production)
- Participate in on‑call / follow‑the‑sun rotations.
- Automation & Reliability Engineering: Automate repetitive operational tasks, deployments, and recovery processes
- Improve system reliability through engineering solutions
- Contribute to continuous improvement of operational processes and efficiency
- Support implementation of deployment strategies (e.g., blue/green, canary).
- Observability & Performance: Build and enhance monitoring, alerting, and observability frameworks
- Improve visibility across metrics, logs, and traces
- Track and improve SLOs/SLAs and system performance




- Perform proactive reliability analysis and capacity planning.
- Infrastructure & Platform Reliability: Operate and support systems across cloud and on‑prem environments
- Work with containerized and distributed systems
- Support Linux and/or Windows‑based environments
- Contribute to system architecture improvements, resilience, and scalability.
- Application Reliability (Multi‑stack): Troubleshoot distributed systems, microservices, and APIs
- Diagnose issues using logs, monitoring tools, and profiling techniques
- Support applications across different stacks, such as .NET/C# applications and other backend technologies
- Manage runtime configuration, dependencies, and system health.
- Collaboration & Continuous Improvement: Partner closely with engineering, platform, and product teams
- Contribute to post‑incident reviews and reliability improvements
- Document processes, playbooks, and troubleshooting guides
- Support knowledge sharing and mentoring of junior engineers.

Required Skills & Experience - Core SRE Capabilities: Experience as a Site Reliability Engineer or in a similar production engineering role (typically 3+ years, adaptable based on seniority)

- Strong experience in production operations and incident management
- Solid understanding of reliability concepts (availability,



latency, scalability, resilience);

Experience working in mission‑critical environments.

- Technical Skills: Operating systems: Linux and/or Windows Server
- Observability tools: Grafana, Prometheus, ELK, Splunk or similar
- CI/CD and automation tooling (Jenkins, GitHub Actions, etc.)
- Cloud platforms: Azure, AWS, or GCP
- Scripting/programming: Python, Shell, Go, or similar
- Understanding of distributed systems, microservices, and databases
- Familiarity with ITIL processes
- Container orchestration: Kubernetes, OpenShift; .NET/C# application debugging and IIS administration
- Automation tools: Ansible or similar
- Event streaming (Kafka), caching, or messaging systems
- Database technologies (SQL/NoSQL).
- Soft Skills: Good communication skills in English
- Strong ownership and reliability mindset
- Ability to perform under pressure in production environments
- Strong collaboration and communication skills across general teams
- Continuous improvement and problem‑solving mindset.

Benefits - Competitive remuneration and individual and company annual bonus.

- Vacation and holiday paid time off.
- Health insurances and other competitive benefits.
- Hybrid work at our Bogotá office.
- Professional development with online learning hubs covering technical and soft skills.
- Diverse and inclusive workplace.

Equal Opportunity Amadeus is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to gender, race, ethnicity, sexual orientation, age, beliefs, disability or any other characteristics protected by law.

📌 Service Reliability Engineer (Colombia)
🏢 1083 Amadeus IT Group Colombia
📍 Colombia

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: service reliability engineer (colombia) / colombia

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: service reliability engineer (colombia) / colombia