23 sep
|
Brightgrove
|
Medellín
23 sep
Brightgrove
Medellín
Junior Data Engineer
Global energy and technology company
Location
Medellin, Colombia, Remote
Area
Data
Tech Level
Junior
Tech Stack
Azure Databricks, PySpark, Azure Data Factory, Azure Data Lake Storage Gen2, SQL & Data Engineering
About the Client
Our customer is a integral energy and technology company operating across multiple markets worldwide. With a strong focus on digital transformation, engineering, and innovation, the company uses advanced data and cloud technologies to improve business operations and decision-making at scale.
Project details
The project focuses on building and scaling a modern cloud-based data platform on Microsoft Azure. The platform brings together data from multiple sources and uses Azure Databricks, Data Factory, Data Lake Storage Gen2, PySpark, Spark SQL, and Delta Lake to create reliable, scalable data pipelines and transformation processes.
The solution is designed for large-scale data processing, with a strong focus on performance, data quality, security, governance, and automation.
Your Team
You will work as part of a cross-functional technology and client team, interacting with technical consultants, application specialists, senior IT professionals, engineers, account managers, business stakeholders, and customer leadership.
What's in it for you
- Interview process that respects people and their time
- Professional and open IT community
- Internal meet-ups and resources for knowledge sharing
- Time for recovery and relaxation
- Bright online and offline events
- Opportunity to become part of our internal volunteer community
Responsibilities
Design develop and maintain scalable data pipelines using Azure Databricks and Azure Data Factory. Develop ETLELT solutions for ingesting transforming and loading data from multiple sources.
Build and optimize PySpark applications and Spark SQL transformations for largescale data processing.
Create and manage Databricks notebooks workflows jobs and clusters.
Implement Delta Lakebased data solutions ensuring data quality reliability and performance.
Develop data ingestion frameworks using Azure Data Lake Storage ADLS Gen2.
Monitor troubleshoot and optimize data pipelines and Databricks workloads.
Implement CICD pipelines for Databricks artifacts using Azure DevOps or GitHub.
Collaborate with data architects business analysts and stakeholders to deliver enterprise data solutions.
Ensure security governance and compliance using Unity Catalog Azure Key Vault and related Azure services.
Skills
Must-Have Skills
- Python — 1–3 years of hands-on experience
- Strong PySpark and Spark SQL
- Hands-on Azure Databricks
- Azure Data Factory
- Azure Data Lake Storage Gen2
- Strong SQL and data modeling
- ETL/ELT development and optimization
- Databricks Workflows, Notebooks, and Delta Lake
- Git/GitHub
- Spark performance tuning and troubleshooting
Nice to Have
- Azure Synapse Analytics
- Unity Catalog
- Azure DevOps / CI/CD
- Azure Key Vault
- Snowflake
- Data governance and data quality
- Azure monitoring and logging
- CDC and real-time data processing
Certifications
- Microsoft Azure Data Engineer Associate (DP-203)
- Databricks Data Engineer Associate
Your personal recruiter
Valentina Brysina
📌 Junior Data Engineer (Medellín)
🏢 Brightgrove
📍 Medellín