AgileEngine, LLC.
is seeking a Senior Site Reliability Engineer to ensure enterprise on-premise and SaaS systems are reliable and scalable.
The role emphasizes Kubernetes cluster management, monitoring, and observability using Snowflake and OpenTelemetry.
You will contribute to on-call rotations, incident response, and root-cause analysis while automating tasks with Python, Bash, or Go.
You will stabilize health across ESM and ECP platforms, with a strong focus on automation and IaC to drive #J-*****-Ljbffr