02 sep
|
Aditi Consulting
|
Bogotá
02 sep
Aditi Consulting
Bogotá
Summary
Design and own data platforms and AI-enabled pipelines that turn raw, messy, real-world data into trustworthy products. You work end to end -- from the instruments that capture data, through declarative pipelines and quality gates, to the RAG and agentic-AI workloads that sit on top -- holding the line on data quality and governance so that everything built on this data can be trusted.
This is a hands-on architect role: you make high-level decisions, define reference patterns, and still write code.
Our operating belief -- AI moves the data. Quality earns the trust. The hardest part of AI at enterprise scale is not the model -- it is the data discipline underneath it.
Responsibilities:
- Architect scalable, secure, and observable data platforms across Lakehouse (Bronze to Silver to Gold) and serving layers aligned to business goals
- Design and integrate data collection instruments and enforce validation at the point of data capture
- Build declarative, production-grade pipelines using Databricks Lakeflow Declarative Pipelines and orchestrate them reliably
- Stand up and enforce auditable data-quality gates that control promotions and releases
- Design and deliver AI and GenAI workloads, including RAG systems and agentic pipelines with human-in-the-loop controls
- Define and enforce data and AI governance including lineage, access control, and responsible AI guardrails
- Make high-level architectural decisions, lead design reviews and proofs of concept, and define reusable best practices
- Own staffing, interviewing, mentoring, and teaching through pairing, reviews, and documentation
- Design and manage data ingestion across multiple capture mechanisms including APIs, telemetry, CDC, surveys, and regulated systems
- Implement validation at source including constraints, logic, vocabularies, and consistency checks
- Define data dictionaries, schemas, metadata,
and conformed models aligned with industry standards
- Own end-to-end data quality as an engineering discipline including incident detection, root cause analysis, and durable fixes
- Build and maintain code-first validation frameworks embedded in CI/CD and pipelines
- Implement ML-driven observability for data health across freshness, volume, schema, and distribution
- Ensure data quality supports trustworthy analytics and AI workloads
Required Qualifications:
- 8 to 12+ years of experience in data engineering with progression into architecture roles
- Strong Python expertise including vectorized data processing, profiling, and clean engineering practices
- Advanced SQL and solid NoSQL knowledge including optimization, indexing, and data modeling
- Experience with Databricks, Lakeflow Declarative Pipelines, Spark or PySpark, Delta Lake, and lakehouse architectures
- Experience with at least one major cloud platform (AWS or Azure)
- Experience with orchestration tools such as Apache Airflow, Dagster, or Prefect
- Proven hands-on experience with data quality engineering using code-first approaches such as dbt tests, Elementary, Great Expectations, Soda, or Deequ
- Experience designing and integrating data collection instruments with validation at the source
- Experience designing RAG systems and working with at least one vector database such as Pinecone, Weaviate, FAISS, or Milvus
- Strong understanding of distributed systems concepts including CAP theorem, ACID vs BASE,
and batch vs stream processing
- Experience implementing CI/CD practices for data pipelines
- Knowledge of data and AI governance including lineage and access control tools such as Unity Catalog, Snowflake Horizon, or Microsoft Purview
- Proven experience hiring, mentoring, and developing engineering talent
Preferred Qualifications:
- Experience with Snowflake and Cortex
- Experience with LangChain, LangGraph, MCP, or agentic pipeline patterns
- Familiarity with AI governance frameworks such as EU AI Act, NIST AI RMF, or ISO/IEC 42001 and LLM guardrails
- Experience with performance and load testing tools such as k6, Locust, or JMeter
- Understanding of linear algebra applied to embeddings and similarity search
- Experience with clinical or regulated data standards such as CDISC or CDASH and EDC systems
- Strong BI experience including Power BI, DAX, row-level security, and performance optimization
- Relevant certifications including AWS Certified Solutions Architect, Databricks, or Google Professional Data Engineer
- Bilingual English and Spanish
Soft Skills:
- Strong architectural thinking with a focus on trade-offs rather than perfect solutions
- Comfort working under uncertainty and adapting to changing requirements and systems
- Critical thinking and ability to challenge assumptions and validate conclusions
- Strong problem framing skills to define constraints and success criteria before executing
- Strong communication and teaching mindset with the ability to mentor and grow teams
- Ownership mindset with a focus on quality, governance, and long-term reliability
Must Have Skill:
- System-level data thinking -- the ability to design, govern, and ensure quality across the entire data lifecycle, from data capture to AI consumption in production systems
#AditiConsulting #26-03850
📌 Data Engineer (Bogotá)
🏢 Aditi Consulting
📍 Bogotá