01 ago
|
EPAM Systems
|
Colombia
01 ago
EPAM Systems
Colombia
EPAM is a leading global provider of digital platform engineering and development services. We are committed to having a positive impact on our customers, our employees, and our communities. We embrace a dynamic and inclusive culture. Here you will collaborate with multi-national teams, contribute to a myriad of innovative projects that deliver the most creative and cutting-edge solutions, and have an opportunity to continuously learn and grow. No matter where you are located, you will join a dedicated, creative, and diverse community that will help you discover your fullest potential.Senior Site Reliability EngineerIn this role, you will serve as the connection point between software development and systems operations, applying engineering principles to automate operational work, expand infrastructure capacity, and keep our systems reliable, resilient, and performing well. Your objective is to build, operate, and safeguard the production environments that support our applications, reducing downtime while enabling fast and secure software releases.ResponsibilitiesArchitect, construct, and maintain cloud infrastructure through modern Infrastructure as Code (IaC) approaches such as Terraform or CloudFormationDevelop and refine CI/CD pipelines to streamline software releases, configuration management, and routine operational activitiesEstablish comprehensive logging, monitoring, and alerting frameworks using platforms such as Prometheus, Grafana, or DatadogDefine clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs)
to gauge system reliabilityAddress production incidents by leading troubleshooting efforts aimed at restoring service as quickly as possibleFacilitate blameless post-mortem reviews to pinpoint root causes and prevent future occurrencesWork alongside software developers to fine-tune system performance and forecast capacity requirementsGuarantee that services scale properly to accommodate growth and sudden spikes in trafficRequirementsAt least 3 years of relevant professional experienceSolid foundation in systems administration, DevOps, or systems-oriented software engineeringCompetency in at least one scripting or programming language, such as Python, Bash, Go, or RustPractical experience with public cloud platforms, including AWS, Azure, or GCPDirect experience using containerization technologies such as Docker and KubernetesThorough knowledge of Linux/Unix system administration and essential networking concepts, including TCP/IP, DNS, HTTP, and SSL/TLSStrong enthusiasm for automation, minimizing repetitive manual work, and designing systems that degrade gracefully under failureBackground working within Financial Services, Insurance, or Retail sectorsStrong written and spoken proficiency in English at a C1 level or higherWe offerInternational projects with top brandsWork with integral teams of highly skilled, diverse peersHealthcare benefitsEmployee financial programsPaid time off and sick leaveUpskilling, reskilling and certification coursesUnlimited access to the LinkedIn Learning library and 22,000+ coursesGlobal career opportunitiesVolunteer and community involvement opportunitiesEPAM Employee GroupsAward-winning culture recognized by Glassdoor, Newsweek and LinkedInEPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.#J-18808-Ljbffr
📌 Senior Site Reliability Engineer (Colombia)
🏢 EPAM Systems
📍 Colombia