Improving

Sas to Databricks Migration Engineer AI

Improving · Remote
Full time $3400 - $4500 🤖 Machine Learning Today
Heads up: we remove this job from our database within a month (monthly cleanup), so the link may stop working. We've saved it in your browser so you can find it again.

Description

We’re looking for a senior-minded Nearshore engineer who can turn SAS-based logic into robust, governed Databricks pipelines using Python, SQL, and modern CI/CD practices—while maintaining strict delivery discipline.

Required experience

  • 4+ years of professional data or software engineering experience.
  • Strong Python and SQL skills.
  • Hands-on Databricks experience with Unity Catalog, Workflows, and Databricks Asset Bundles.
  • Proficiency with GitLab CI/CD (pipelines, merge request workflows, and automated testing) and disciplined Git branching and code reviews.
  • Clear written and spoken English for client-facing collaboration.
  • Demonstrated ownership: scoping work, delivering outcomes, and proactively flagging risks without being prompted.

Focus areas

  • Pipeline reliability, validation, and delivery of converted code.
  • Parity checks between SAS and Databricks outputs.
  • Deployment and governance using DABs + GitLab CI/CD + Unity Catalog.

Additional experience

  • Spark and Delta Lake performance tuning.
  • Data validation and reconciliation experience.
  • Infrastructure-as-code or DAB-based deployment experience.
  • SAS reading ability and exposure to healthcare data (plus experience with Azure) are valued.

How we work: We value clarity, accountability, and continuous improvement. We’ll expect you to communicate trade-offs, confirm assumptions early, and build trust through predictable delivery, thoughtful reviews, and transparent risk management.

Functions

We’ll rely on you to drive the end-to-end conversion delivery from SAS inventories to validated Python/SQL outputs on Databricks, with a strong focus on pipeline reliability and data validation.

  • Pipeline engineering: Build and run pipelines that process SAS inventories and produce converted outputs.
  • Quality and parity validation: Validate converted code for parity against SAS outputs (e.g., row counts, checksums, schema, and data types).
  • Deployment ownership: Own deployments through Databricks Asset Bundles (DABs) and GitLab CI/CD, ensuring repeatable releases.
  • Databricks governance: Manage Unity Catalog objects, permissions, and promotion across environments.
  • Operational excellence: Troubleshoot job failures and performance issues, and take preventive actions to improve pipeline stability.
  • Performance tuning: Apply tuning techniques for Apache Spark and Delta Lake to meet reliability and execution-time expectations.
  • Data reconciliation: Use data validation and reconciliation practices to ensure correctness and consistency.
  • Infrastructure-as-code mindset: Implement DAB-based deployment patterns and support automated, testable delivery workflows.
  • Client collaboration: Communicate progress, risks, and technical decisions clearly with client and partner stakeholders.

Desirable


  • Experience converting SAS workflows to Python/SQL in production environments.
  • Healthcare data exposure and familiarity with typical data quality and privacy expectations.
  • Azure exposure and understanding of how cloud services fit into end-to-end delivery and operations.
  • Deep experience optimizing Spark jobs (partitioning, caching strategies, skew handling) and Delta Lake (file sizing, compaction patterns).

Benefits


  • Contrato a largo plazo.
  • 100% Remoto.
  • Vacaciones y PTOs
  • Posibilidad de recibir 2 bonos al año.
  • 2 revisiones salariales al año.
  • Clases de inglés.
  • Equipamiento Apple.
  • Plataforma de cursos en linea
  • Budget para compra de libros.
  • Budget para compra de materiales de trabajo
  • mucho mas..
PythonSQLDatabricks
Source: GetOnBoard
Report this job
📬 Looking for remote work? Leave your email and we'll let you know when we launch job alerts.