AgileEngine is seeking a Senior Site Reliability Engineer to ensure operational stability across enterprise on-premise and SaaS-hosted systems. You will manage Kubernetes clusters, implement monitoring using OpenTelemetry, and automate infrastructure tasks with Python, Bash, or Go.
Expect on-call rotations, incident response, and RCA participation, with a strong emphasis on Linux/Unix administration and Snowflake/OpenTelemetry ecosystems. Remote role with integral collaboration.