AI Platform Operations Manager (Colombia)

AI Platform Operations Manager (Colombia)

08 oct
|
NextGenEnergyJobs
|
Colombia

08 oct

NextGenEnergyJobs

Colombia

STACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies.

Responsibilities

- The DevOps Engineer, AI Platform is responsible for automating, deploying, and operating the infrastructure and delivery pipelines that support STACK’s enterprise AI platform on Azure. This is a hands-on engineering role focused on build and run — not oversight.
- Reporting to Head of AI, Enterprise AI & Data Strategy org this individual owns the infrastructure-as-code, CI/CD, containerization, observability, and release automation that allow AI engineers and enterprise application teams to ship agentic AI solutions, RAG pipelines, and integration services reliably and repeatably. The role sits at the intersection of cloud infrastructure, platform engineering, and MLOps — turning platform architecture into automated, governed, observable, and cost-efficient environments that teams across the organization build on.

Requirements

- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent practical experience.
- 5+ years of hands-on experience in DevOps, site reliability engineering, or platform engineering roles with a strong delivery track record.
- Strong proficiency in Infrastructure as Code — Terraform, Bicep, or ARM — including module design, state management, and reusable patterns.
- Proven experience building and operating CI/CD pipelines in Azure DevOps or GitHub Actions.
- Hands-on experience with containerization and orchestration — Docker, Azure Kubernetes Service (AKS), and Azure Container Apps or equivalent.




- Solid working knowledge of Azure core services — compute, networking (VNet, NSG, Private Endpoints), storage, identity (Entra ID), and Key Vault.
- Strong scripting and automation skills in Python, PowerShell, or Bash.
- Experience with monitoring and observability tooling — Azure Monitor, Log Analytics, Application Insights, Prometheus, or Grafana.
- Working knowledge of Git-based workflows, code review practices, and artifact/registry management.
- Demonstrated ability to troubleshoot production issues across infrastructure, network, and application layers.
- Microsoft Certified: DevOps Engineer Expert (AZ-400), Azure Administrator (AZ-104), or Certified Kubernetes Administrator (CKA).
- Experience deploying and operating AI/ML workloads — model endpoints, RAG pipelines, vector databases, or agentic services in production.
- Familiarity with MLOps tooling and practices — Azure Machine Learning, MLflow, Databricks, or equivalent model lifecycle platforms.
- Experience deploying MCP (Model Context Protocol) servers or similar integration services connecting AI agents to enterprise systems.
- Exposure to agentic AI frameworks such as Semantic Kernel, LangGraph, or AutoGen from a deployment and operations perspective.
- Experience with GPU compute provisioning, quota management, and inference cost optimization.
- Knowledge of FinOps frameworks and Azure cost optimization practices.
- Experience integrating with enterprise systems such as Microsoft 365, Freshworks ITSM, Workday, NetSuite, or Procore.
- Experience in data center, hyperscale, or infrastructure-intensive industry environments.

#J-18808-Ljbffr

📌 AI Platform Operations Manager (Colombia)
🏢 NextGenEnergyJobs
📍 Colombia

Postulate a este anuncio

Muestra tus habilidades a la empresa, rellenar el formulario y deja un toque personal en la carta, ayudará el reclutador en la elección del candidato.

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai platform operations manager (colombia) / colombia

Suscribete a esta alerta:

Recibe por email las nuevas ofertas de trabajo para: ai platform operations manager (colombia) / colombia