AgileEngine is seeking a Senior Site Reliability Engineer to drive core system administration and operational stability for both on-premise and SaaS-hosted systems. The role emphasizes Kubernetes cluster management, monitoring, and observability with Snowflake and OpenTelemetry.
You will participate in on-call rotations, incident response, and RCA, and automate infrastructure tasks using Python, Bash, or Go, ensuring health across multiple platforms and environments.