09 sep
|
OpsGenius
|
Colombia
09 sep
OpsGenius
Colombia
Senior Site Reliability Engineer (Azure) Location: Remote, LATAM Job Type: Full-Time, with benefits OpsGenius is a boutique US-based cloud operations firm. We take over the operational side of large production platforms so engineering teams can get back to building. We are hiring a Senior SRE to join a small, senior team standing up the operational foundation for a large multi-tenant SaaS platform on Azure, running at significant scale across many environments and a high-volume API surface.
The platform has scaled quickly, and you will be one of the people building the operational foundation alongside it: observability, on-call, and deployment practices. This is hands-on production ownership, not a ticket queue.
WHAT YOU WILL DO - Build out observability: SLO definitions, alert rationalization, synthetic checks, and dashboards using Azure Monitor, Application Insights, and potentially Datadog - Stand up and participate in an on-call rotation, including paging tooling, escalation paths, and incident documentation - Modernize deployments: blue/green or canary patterns with validated rollback, automated smoke tests, and Infrastructure as Code with Terraform - Perform Azure SQL performance work: query optimization, indexing strategy, and execution plan analysis across a large multi-tenant estate - Own backup and restore strategy for Azure SQL, including point-in-time recovery testing and periodic restore drills - Tune elastic pool sizing and evaluate DTU versus vCore tradeoffs for cost and performance - Optimize Cosmos DB RU consumption and partition key design for high-throughput workloads - Operate containerized workloads on Azure Container Apps,
alongside App Service and Service Bus - Write operational runbooks (incident triage, rollback, backup restore, secret rotation, certificate renewal) that any on-call engineer can follow - Contribute to Azure cost optimization: right-sizing, autoscaling tuning, tagging, and cost reporting - Support security and reliability hardening over time, including IAM reviews, backup restore drills, and DR exercises - Collaborate daily with a US-based team during overlapping hours WHAT WE ARE LOOKING FOR Required - 7+ years in SRE, DevOps, or cloud infrastructure roles, including direct ownership of production systems under an on-call rotation - Deep, hands-on Azure experience. AWS-primary backgrounds with light Azure exposure will not be a fit - Hands-on Azure SQL DBA experience: indexing, query tuning, HA/DR (failover groups, geo-replication), and backup/restore. This is not a generalist SRE role; real database ownership is expected - Comfort reading query execution plans and diagnosing performance regressions at the database level, not just infrastructure monitoring - Terraform or equivalent IaC in production - Experience with observability tooling: Application Insights, Azure Monitor, Datadog,
or similar - Strong written and spoken English; daily communication with a US team and occasional client stakeholders - Work schedule overlapping US Eastern hours (roughly 9 to 5 Eastern is adecuado) Nice to Have - Experience in HIPAA, PHI, or other regulated environments.
A background check to healthcare-industry standards is required prior to production access - Cosmos DB RU optimization and partition key design at high throughput - Prior work in a multi-tenant SaaS environment ON-CALL EXPECTATIONS This role includes a pager-based on-call rotation covering SEV-1 and SEV-2 incidents, shared with the rest of the SRE team. On-call is a core part of the role. Expect it to be light in the first month and ramp as the team takes over production responsibility.
COMPLIANCE AND ACCESS - All personnel are named and approved by the client before any access is provisioned - Background checks to healthcare-industry standards are completed prior to production access - Production access is provisioned through the client identity provider with MFA and time-bound elevation BENEFITS - Paid time off - Supplemental health insurance - Learning credits for training and certification - Performance incentives - Regular one on ones and ongoing career support - Fully remote WHY THIS ROLE - You are building the observability, on-call, and deployment practices, not inheriting someone else’s - Small senior team, direct access to senior engineers, no layers of process - Interesting scale problems: multi-tenant data at volume and real performance challenges, in an environment that has to stay up while you harden it
📌 Senior Site Reliability Engineer (Azure) (Colombia)
🏢 OpsGenius
📍 Colombia