AgileEngine is seeking a Senior Site Reliability Engineer to bolster on-premise and cloud systems, with a strong focus on Kubernetes cluster management, monitoring, and observability. You will participate in on-call rotations, incident response, and root-cause analysis, and automate infrastructure tasks using Python, Bash, or Go across ESM and ECP platforms.
Experience with service mesh architectures is valued for ECP roles, as are Snowflake and OpenTelemetry ecosystems.