Job Title: Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure)
Key Skills: Azure IaaS, Azure Monitor, Application Insights, New Relic, Databricks, DBT, SQL, Azure DevOps, GitHub Actions, Site Reliability Engineering (SRE), Application Performance Monitoring (APM), Log Analytics (KQL), Observability, Distributed Tracing, Structured Logging
Experience: 5+ years of experience in Site Reliability Engineering, Cloud Operations, or related roles. Mandatory minimum of 1 year of hands-on experience with DBT, Databricks, and SQL.
Location: Legal residents of Peru, Colombia, Bolivia, Costa Rica, Mexico, and Brazil.
Work Mode: Remote
At Coforge, we are looking for a Senior Site Reliability Engineer (SRE) – Application Observability & Readiness (Azure) (#21013-18-1) with the following profile.
Main Responsibilities
• Collaborate with development teams to design and implement monitoring, alerting, dashboards, and APM instrumentation across applications and services.
• Lead the implementation, configuration,
and optimization of Application Performance Monitoring (APM) solutions.
• Apply observability best practices using tools such as Azure Monitor, Application Insights, New Relic, and Log Analytics (KQL).
• Enable code-level instrumentation, distributed tracing, and structured logging to improve application visibility and reliability.
• Design and maintain application-level monitoring dashboards and operational health metrics.
• Define and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and effective alerting strategies based on latency, error rates, traffic, and resource saturation.
• Continuously improve monitoring and alerting mechanisms through production insights and incident learnings.
• Participate in production readiness reviews, identifying operational risks, observability gaps, and potential failure scenarios before deployment.
• Support incident analysis and post-incident improvements throug
📌 Senior Site Reliability Engineer (Cali)
🏢 Encora
📍 Cali