05 ago
|
Svitla Systems
|
Colombia
05 ago
Svitla Systems
Colombia
Svitla Systems Inc. is looking for a Senior Data Engineer for a full-time position (40 hours per week) in Colombia, Mexico.
You'll take full ownership of a Python/Scrapy-based legal/regulatory data ingestion pipeline. This role combines large-scale data engineering, advanced web scraping against hostile sources, complex legal document processing, and cost-efficient cloud infrastructure design. It also includes maintenance of a React frontend application.
Requirements
- 5+ years of experience in data engineering or backend development, with a focus on production pipelines.
- Strong experience in Python, including production web scraping (Scrapy or equivalent) and structured-data processing.
- Practical understanding of defeating or working around anti-bot measures (rate limiting, Cloudflare, paywalled/gated legal databases) within legal and ethical bounds.
- Understanding of relational and non-relational databases; schema design for hierarchical data.
- Solid understanding of SQL and data modeling, with the judgment to design schemas that serve both change detection and human-readable display.
- Experience with cloud infrastructure (AWS, GCP, or Azure), IaC (Terraform/CloudFormation), and containers (Docker).
- Knowledge of Microsoft Azure, including VMs, Blob storage, App Service plans, and container services, with the ability to provision and tear down resources programmatically.
- Knowledge of search architectures (full-text, semantic search — Elasticsearch, OpenSearch, or vector DBs).
- Familiarity with messaging/notification systems (email APIs, Slack/Teams webhooks).
- Understanding of Infrastructure-as-code for repeatable, automated deployments.
- Be comfortable owning a system end-to-end with minimal hand-holding, including reading and improving inherited code and documentation.
- The ability to take ownership of complex legacy systems without extensive documentation.
- Meticulous attention to detail:
this work handles legal text where structural precision is critical.
- Cost- and efficiency-oriented mindset.
- Be comfortable working with ambiguity across heterogeneous, changing data sources.
Nice to have
- Experience with React and Vite for maintaining and extending the display layer.
- Prior experience in legal tech, regtech, or compliance software.
- Familiarity with anti-detection scraping techniques (proxy rotation, fingerprinting, simulated human behavior).
- Familiarity with deterministic hashing and content versioning systems.
- Experience designing systems for horizontal scale (clustering, multi-VM, multi-threaded, or containerized workloads).
- Experience AWS alongside Azure (the concepts transfer; multi-cloud).
- Familiarity with semantic / cross-jurisdictional search techniques.
- Exposure to LLM or vision-AI integration — e.g., using AI to parse document structure, generate summaries and action statements, or pre-classify regulatory changes as impacting vs. non-impacting for human review.
Responsibilities
- Take over and fully understand an existing Python/Scrapy ingestion pipeline, its codebase, and its documented database schemas, then make it your own.
- Build and maintain web scrapers across 50+ jurisdictions, each with its own structure, format, and obstacles — XML feeds, HTML pages, DOCX-only sources (e.g., West Virginia), and content behind Westlaw, LexisNexis, Cloudflare, and similar barriers.
- Consolidate tooling where possible: prefer one or two robust, broadly capable scraping tools over a sprawl of point solutions,
falling back to specialized tooling only for genuinely esoteric sources.
- Normalize ingested content into a structured, plain-text format for deterministic hashing and change detection, while preserving the original, formatted HTML so documents can be displayed exactly as the issuing agency intended.
- Faithfully preserve each source's original hierarchy (nested clauses, sub-paragraphs, lettered and numbered subsections, Roman numerals, Unicode markers, appendices, and tables), so stored content remains a true representation of the statute.
- Maintain and extend a multi-destination architecture, including the nested-hierarchy schema used by the SaaS compliance platform (white-labeled in some deployments) and a flattened schema supporting full-text, semantic, and cross-jurisdictional search.
- Operate and evolve change detection and notification (email, Slack, Teams), with a path toward more granular, paragraph- and clause-level change tracking and side-by-side visual diffs.
- Design the infrastructure for cost efficiency: model how many VMs/containers are needed and for how long, automate spin-up and shutdown so nothing runs idle, and right-size compute based on real scraping timings rather than assumptions.
- Maintain the React and Vite single-page application that displays ingested regulations, including a searchable table-of-contents navigation pattern.
We offer
- US and EU projects based on advanced technologies.
- Competitive compensation based on skills and experience.
- Remote-friendly culture and no micromanagement.
- Christmas Bonus in the amount of 50% of the monthly payment.
- Bonuses for article writing, public talks, other activities.
- Personalized learning program tailored to your interests and skill development.
- Free tech webinars and meetups organized by Svitla.
- Fun corporate online/offline celebrations and activities.
- Awesome team, friendly and supportive community!
#J-18808-Ljbffr
📌 SENIOR DATA ENGINEER (Colombia)
🏢 Svitla Systems
📍 Colombia