AgileEngine is seeking a Senior Site Reliability Engineer to ensure core system administration and operational stability across enterprise on-premise and SaaS environments. The role emphasizes Kubernetes cluster management, monitoring, and observability with Snowflake and OpenTelemetry.
You will participate in on-call rotations, incident response, RCA, and automate infrastructure tasks using Python, Bash, or Go, ensuring health across ESM and ECP platforms.