← All jobs
Z
ZifoChennai, India

Senior Data Engineer - Real World Data & Healthcare Analytics

Data Engineer

We are looking for a highly skilled Data Engineer with Real World Data (RWD) experience to build and manage end-to-end data pipelines for multimodal healthcare datasets. The ideal candidate will work at the intersection of data engineering, analytics, and healthcare research, transforming complex healthcare data into analysis-ready assets that support advanced analytics, AI/ML initiatives, and evidence-generation studies. This role requires expertise in large-scale healthcare data processing, data harmonization, cloud platforms, and modern data engineering practices.

Responsibilities

  • Data Engineering & Pipeline Development • Design, develop, and maintain scalable ETL/ELT pipelines for large healthcare and real world datasets.
  • • Build and manage data ingestion, transformation, harmonization, and analytics layers.
  • • Implement data quality frameworks, governance controls, lineage tracking, and monitoring.
  • • Manage data lifecycle processes across raw, curated, and analytics-ready environments.
  • • Work with structured and unstructured healthcare datasets from multiple sources.

Healthcare Data Harmonization

• Harmonize heterogeneous healthcare data sources and coding systems into standardized formats.

• Map and transform clinical terminologies including o SNOMED CT o ICD-10 o LOINC o RxNorm o CPT/HCPCS • Support implementation of common data models such as OMOP and FHIR.

Analytics & Study Support

  • • Support data feasibility assessments and data quality evaluations.
  • • Collaborate with epidemiologists, biostatisticians, data scientists, and business stakeholders.
  • • Develop reusable data assets, cohorts, and model-ready datasets.
  • • Enable advanced analytics and AI/ML use cases through reliable data engineering practices.

Application & Platform Development

  • • Contribute to analyst-facing applications, dashboards, and self-service data products.
  • • Support development of data products using modern workflow automation and AI assisted engineering approaches.
  • • Provide guidance on efficient querying and optimization of large longitudinal datasets.

Requirements

Data Engineering • Strong experience with Python, SQL, Spark / PySpark • Experience building production-grade ETL/ELT pipelines.

• Strong understanding of data modelling concepts - Star schema, Snowflake schema, Normalization and denormalization • Experience with metadata management, lineage, monitoring, and data governance.

Platforms & Technologies

  • Experience in one or more of the following - Palantir Foundry, Databricks, Snowflake, AWS or equivalent cloud platforms, HPC environments • Containerized workloads Software Engineering Practices • Git • CI/CD pipelines • Unit testing and automation • Performance monitoring and optimization
  • Domain Expertise
  • Candidates should have working knowledge of Healthcare Real World Data (RWD), Claims data, Electronic Health Records (EHR), Registries, Patient-reported outcomes, Wearables and digital health datasets
  • Understanding of study feasibility, observational research, and healthcare analytics workflows is highly desirable.

AI & Automation Experience

  • Preferred experience with Large Language Models (LLMs), AI-assisted data engineering, Agentic workflows, Data profiling and automated data quality assessments, integration of ML outputs into production data pipelines
  • Qualification / Requirement • 4-8 years of experience in data engineering, healthcare analytics, or real-world data platforms.
  • • Experience working with large-scale healthcare datasets in regulated environments.
  • • Strong problem-solving and analytical skills.
  • • Excellent stakeholder communication capabilities.
  • • Formal educational qualifications are flexible; relevant experience and expertise are valued.

Nice to Have

  • • Experience with multimodal healthcare datasets (clinical, omics, imaging, genomics, proteomics, microbiome, etc.).
  • • Hands-on experience implementing OMOP/FHIR at scale.
  • • Experience building self-service applications and data products for business users.
  • • Familiarity with federated data networks and data quality frameworks.
  • What We're Looking For • Systems thinker who can work with complex and evolving datasets.
  • • Strong collaboration skills across technical and business teams.
  • • Agile mindset with a focus on delivery.
  • • Commitment to data privacy and ethics.