Agile Engine is seeking a Senior Site Reliability Engineer to ensure core system administration and operational stability across on‑premise and SaaS environments. You will manage Kubernetes clusters, monitor systems using Open Telemetry, and automate tasks with Python, Bash, or Go.
The role requires 4+ years in infrastructure management, strong Linux expertise, and experience with Snowflake/Open Telemetry ecosystems. Remote-friendly with on-call rotations and incident response.