AgileEngine is seeking a Senior Site Reliability Engineer to provide core system administration and operational stability for enterprise on-premise and SaaS-hosted systems, with a strong focus on Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry.
You will participate in on-call rotations, incident response, RCA, and automate infrastructure tasks using Bash, Python, or Go, ensuring system health across ESM and ECP platforms. Remote-friendly role.