31 ago
|
OpsGenius
|
Colombia
31 ago
OpsGenius
Colombia
Senior Site Reliability Engineer (Azure)
Location: Remote, LATAM
Job Type: Full-Time, with benefits
OpsGenius is a boutique US-based cloud operations firm. We take over the operational side of large production platforms so engineering teams can get back to building.
We are hiring a Senior SRE to join a small, senior team standing up the operational foundation for a large multi-tenant SaaS platform on Azure, running at significant scale across many environments and a high-volume API surface.
The platform has scaled quickly, and you will be one of the people building the operational foundation alongside it: observability, on-call, and deployment practices. This is hands-on production ownership, not a ticket queue.
WHAT YOU WILL DO
- Build out observability: SLO definitions, alert rationalization, synthetic checks, and dashboards using Azure Monitor, Application Insights, and potentially Datadog
- Stand up and participate in an on-call rotation, including paging tooling, escalation paths, and incident documentation
- Modernize deployments: blue/green or canary patterns with validated rollback, automated smoke tests, and Infrastructure as Code with Terraform
- Perform Azure SQL performance work: query optimization, indexing strategy, and execution plan analysis across a large multi-tenant estate
- Own backup and restore strategy for Azure SQL, including point-in-time recovery testing and periodic restore drills
- Tune elastic pool sizing and evaluate DTU versus vCore tradeoffs for cost and performance
- Optimize Cosmos DB RU consumption and partition key design for high-throughput workloads
- Operate containerized workloads on Azure Container Apps,
alongside App Service and Service Bus
- Write operational runbooks (incident triage, rollback, backup restore, secret rotation, certificate renewal) that any on-call engineer can follow
- Contribute to Azure cost optimization: right-sizing, autoscaling tuning, tagging, and cost reporting
- Support security and reliability hardening over time, including IAM reviews, backup restore drills, and DR exercises
- Collaborate daily with a US-based team during overlapping hours
WHAT WE ARE LOOKING FOR
Required
- 7+ years in SRE, DevOps, or cloud infrastructure roles, including direct ownership of production systems under an on-call rotation
- Deep, hands-on Azure experience. AWS-primary backgrounds with light Azure exposure will not be a fit
- Hands-on Azure SQL DBA experience: indexing, query tuning, HA/DR (failover groups, geo-replication), and backup/restore. This is not a generalist SRE role; real database ownership is expected
- Comfort reading query execution plans and diagnosing performance regressions at the database level, not just infrastructure monitoring
- Terraform or equivalent IaC in production
- Experience with observability tooling: Application Insights, Azure Monitor, Datadog, or similar
- Strong written and spoken English; daily communication with a US team and occasional client stakeholders
- Work schedule overlapping US Eastern hours (roughly 9 to 5 Eastern is idóneo)
Nice to Have
- Experience in HIPAA, PHI, or other regulated environments. A background check to healthcare-industry standards is required prior to production access
- Cosmos DB RU optimization and partition key design at high throughput
- Kubernetes experience
- Prior work in a multi-tenant SaaS environment
ON-CALL EXPECTATIONS This role includes a pager-based on-call rotation covering SEV-1 and SEV-2 incidents, shared with the rest of the SRE team. On-call is a core part of the role. Expect it to be light in the first month and ramp as the team takes over production responsibility.
COMPLIANCE AND ACCESS
- All personnel are named and approved by the client before any access is provisioned
- Background checks to healthcare-industry standards are completed prior to production access
- Production access is provisioned through the client identity provider with MFA and time-bound elevation
BENEFITS
- Full-time employment
- Paid time off
- Supplemental health insurance
- Learning credits for training and certification
- Performance incentives
- Regular one on ones and ongoing career support
- Fully remote
WHY THIS ROLE
- You are building the observability, on-call, and deployment practices, not inheriting someone else's
- Small senior team, direct access to senior engineers, no layers of process
- Interesting scale problems: multi-tenant data at volume and real performance challenges, in an environment that has to stay up while you harden it
📌 Senior Site Reliability Engineer (Azure) (Colombia)
🏢 OpsGenius
📍 Colombia