04 sep
|
Icebergdata
|
Bogotá
04 sep
Icebergdata
Bogotá
Key Responsibilities
- Translate business requirements into maintainable technical solutions
- Design resilient and cost-efficient architectures on cloud platforms
- Develop robust scrapers (proxy rotation, CAPTCHA handling, authentication, DOM change tolerance)
- Design and implement REST APIs with authentication, pagination, and security best practices
- Orchestrate and monitor data pipelines (scheduling, retries, alerts, logging, metrics)
- Normalize and transform data (ETL/ELT) and design storage schemas
- Reduce downtime when target sources change (self-healing capabilities)
- Elevate quality and performance standards (profiling, optimization, automation)
- Ensure compliance with Terms of Service, copyright, robots.txt, and data protection regulations
- Promote best practices (security, versioning, CI/CD, documentation)
- Conduct unit and integration testing; validate data quality
- Implement observability: structured logging, metrics, and traces
- Create technical documentation and participate in code reviews
- Collaborate with stakeholders and clients (Spanish/English)
Requirements Must Have:
- ✓Advanced Python (async, error handling, typing, packaging)
- ✓Web scraping tools: Requests/HTTPX, Playwright/Selenium, BeautifulSoup/lxml, CSS/XPath selectors
- ✓API development with FastAPI (or equivalent), authentication (tokens, OAuth), pagination
- ✓Git (PRs, code review) and CI/CD (GitHub Actions/GitLab CI)
- ✓SQL (PostgreSQL/MySQL) and NoSQL (MongoDB/Redis); modeling and performance optimization
- ✓Docker; orchestration and deployment concepts (K8s desirable)
- ✓Observability (structured logging, metrics, traces, and alerts)
- ✓Cloud fundamentals (AWS/Azure/GCP)
- ✓Advanced English (B2+/C1, oral and written)
Nice To Have
- +Messaging/queues (Kafka/RabbitMQ),
scheduled tasks (Celery/Arq)
- +Data lakes/warehouses (S3/BigQuery/Snowflake) and dbt/Airbyte/Prefect
- +Vector DBs/embeddings (FAISS/Pinecone) and LLM ops
- +Go/Rust/Node.js for high-performance components
- +Security (secrets management, rate limiting, OWASP)
Benefits
- ★Competitive salary
- ★Lunch bonus: $500,000 COP
- ★Uber between office and home (specific hours)
- ★$2,000 USD for training after the first year
- ★Continuous learning opportunities
- ★Google Cloud certifications paid by the company
- ★Corporate travel opportunities
- ★Office snacks
- ★Laptop and tools provided
- ★Adaptable schedule with team coordination
What We're Looking For ◆Proactivity and sense of urgency; extreme ownership
◆Analytical thinking and creative problem-solving
◆Adaptability to frequent changes and effective prioritization
◆Clear communication with technical and non-technical audiences
◆Collaborative work and willingness to share knowledge
◆Customer orientation and quality focus
◆Continuous learning and curiosity about new tools/technologies
Success Metrics (First 90 Days)
Stabilize critical scrapers
Success rate / availability (%)
≥ 99% in 60 days
Accelerate deliveries
Lead time of changes / throughput
−30% in 90 days
Data quality
Validation errors per million records
≤ 0.5 ppm
API reliability
SLA/SLO (e.g., 99.5%)
Monthly compliance
30-60-90 Day Plan
30 hours (1 week)
Technical onboarding (repos, pipelines, standards). Take ownership of 1 scraper and 1 API. Map failure points and quick wins.
Success: 1 stable release; 5+ PRs; basic documentation
60 hours (2 weeks)
Fix active scrapers. Performance improvements (profiling/optimization). Automate key tests and observability.
Success: Availability ≥95%; −20% incidents; CI tests
90 hours (3 weeks)
Deliver architecture/orchestration improvements. Reduce change lead time. Propose quarterly roadmap.
Success: −30% lead time; SLOs met; roadmap approved In Scope
- ✓Development of robust scrapers (proxy rotation, CAPTCHA handling, authentication, DOM change tolerance)
- ✓Design/implementation of REST APIs with authentication, pagination, and security best practices
- ✓Pipeline orchestration and monitoring (scheduling, retries, alerts, logging, metrics)
- ✓Data normalization and transformation (ETL/ELT) and storage schema design
- ✓Unit and integration testing; data quality validation
- ✓Observability: structured logging, metrics, and traces
- ✓Technical documentation and code reviews; collaboration with stakeholders (ES/EN)
Out of Scope
- ✗UI/UX/front-end design beyond API endpoints and contracts
- ✗Direct sales/commercial negotiation (technical support when required)
- ✗On-premise physical infrastructure operation (cloud-based work)
Data Ethics & Compliance Responsibility to comply with Terms of Service, copyright, robots.txt policies, and applicable data protection regulations (e.g., Law 1581 of 2012 in Colombia, GDPR if applicable). Escalate legal doubts before proceeding with new data sources.
📌 Data Solutions & Automation Engineer (Bogotá)
🏢 Icebergdata
📍 Bogotá