AgileEngine is seeking a Senior Site Reliability Engineer to provide core system administration and operational stability for on-premise and SaaS-hosted environments. You will manage Kubernetes clusters, monitor systems, and advance observability using Snowflake and OpenTelemetry.
Responsibilities include on-call rotations, incident response, RCAs, and IaC automation using Bash, Python, or Go, with a strong emphasis on Linux/Unix health across diverse platforms.