Agile Engine, LLC. seeks a Senior Site Reliability Engineer to provide core system administration and operational stability for enterprise on-premise and SaaS-hosted systems, with a strong focus on Kubernetes cluster management, monitoring, and observability using Snowflake and Open Telemetry.
You will participate in on-call rotations, incident response, and root-cause analysis, and automate infrastructure tasks using Python, Bash, or Go, ensuring system health across ESM and ECP platform