Sas to Databricks Migration Engineer AI
Improving
· Remote
Tiempo completo
$3400 - $4500
🤖 Machine Learning
Hoy
Atención: eliminamos este aviso de nuestra base de datos al cabo de un mes (limpieza mensual), así que el enlace puede dejar de funcionar. Lo guardamos en tu navegador para que puedas volver a encontrarlo.
Description
We’re looking for a senior-minded Nearshore engineer who can turn SAS-based logic into robust, governed Databricks pipelines using Python, SQL, and modern CI/CD practices—while maintaining strict delivery discipline.
Required experience
- 4+ years of professional data or software engineering experience.
- Strong Python and SQL skills.
- Hands-on Databricks experience with Unity Catalog, Workflows, and Databricks Asset Bundles.
- Proficiency with GitLab CI/CD (pipelines, merge request workflows, and automated testing) and disciplined Git branching and code reviews.
- Clear written and spoken English for client-facing collaboration.
- Demonstrated ownership: scoping work, delivering outcomes, and proactively flagging risks without being prompted.
Focus areas
- Pipeline reliability, validation, and delivery of converted code.
- Parity checks between SAS and Databricks outputs.
- Deployment and governance using DABs + GitLab CI/CD + Unity Catalog.
Additional experience
- Spark and Delta Lake performance tuning.
- Data validation and reconciliation experience.
- Infrastructure-as-code or DAB-based deployment experience.
- SAS reading ability and exposure to healthcare data (plus experience with Azure) are valued.
How we work: We value clarity, accountability, and continuous improvement. We’ll expect you to communicate trade-offs, confirm assumptions early, and build trust through predictable delivery, thoughtful reviews, and transparent risk management.
Functions
We’ll rely on you to drive the end-to-end conversion delivery from SAS inventories to validated Python/SQL outputs on Databricks, with a strong focus on pipeline reliability and data validation.
- Pipeline engineering: Build and run pipelines that process SAS inventories and produce converted outputs.
- Quality and parity validation: Validate converted code for parity against SAS outputs (e.g., row counts, checksums, schema, and data types).
- Deployment ownership: Own deployments through Databricks Asset Bundles (DABs) and GitLab CI/CD, ensuring repeatable releases.
- Databricks governance: Manage Unity Catalog objects, permissions, and promotion across environments.
- Operational excellence: Troubleshoot job failures and performance issues, and take preventive actions to improve pipeline stability.
- Performance tuning: Apply tuning techniques for Apache Spark and Delta Lake to meet reliability and execution-time expectations.
- Data reconciliation: Use data validation and reconciliation practices to ensure correctness and consistency.
- Infrastructure-as-code mindset: Implement DAB-based deployment patterns and support automated, testable delivery workflows.
- Client collaboration: Communicate progress, risks, and technical decisions clearly with client and partner stakeholders.
Desirable
- Experience converting SAS workflows to Python/SQL in production environments.
- Healthcare data exposure and familiarity with typical data quality and privacy expectations.
- Azure exposure and understanding of how cloud services fit into end-to-end delivery and operations.
- Deep experience optimizing Spark jobs (partitioning, caching strategies, skew handling) and Delta Lake (file sizing, compaction patterns).
Benefits
- Contrato a largo plazo.
- 100% Remoto.
- Vacaciones y PTOs
- Posibilidad de recibir 2 bonos al año.
- 2 revisiones salariales al año.
- Clases de inglés.
- Equipamiento Apple.
- Plataforma de cursos en linea
- Budget para compra de libros.
- Budget para compra de materiales de trabajo
- mucho mas..
PythonSQLDatabricks
Fuente: GetOnBoard