Principal Data Engineer (RWE)
Jobgether
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
7,274 open data roles across 649 companies are on ApplySarthi right now, most of them in Bengaluru (415), Hyderabad (313), Mumbai (155).
- Data Architect – Data Products & Data AnalysisCallista Group AG
- RE/RS, Data Understanding - FoundationsOpenAI
- Junior Data Engineer & MarTech Specialist (Mobile Apps) (m/w/d)Trg
- Senior Software Engineer, Backend - Data LayerCamunda
- Senior AI Data Expert (f/m/x)exmox
What data roles keep asking for: AWS (24%), SQL (23%), Python (22%) — counted across their open postings here.
Data Engineer jobs in India · Remote Data Engineer jobs · Azure jobs · Data modelling jobs · Databricks jobs · ETL jobs
Jobgether has 3,935 open roles listed here.
- AI Researcher — Distillation
- AI Researcher — Distillation
- Accounting & Regulatory Reporting
- Accounts Receivable Coordinator
- AI Security Analyst
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for data roles keep coming back to AWS, SQL, Python. Practise those questions before you sit with Jobgether.
Questions you are likely to be asked
- Why do you want to join Jobgether?
- What is your experience with ETL? Tell me one thing you learned the hard way.
- Tell me about a time the data was messy or wrong. What did you do?
- How would you explain your model's result to someone who is not technical?
- What would you check first if a model's accuracy dropped after going live?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Principal Data Engineer (RWE) at Jobgether interview free →Accountabilities:: Develop automated data processes for the ongoing generation of patient-level data products supporting dashboards, reports, studies, and other business needs. Transform raw datasets received from data vendors and partners into usable, structured data products for RWE studies, dashboards, and analytical outputs. Convert heterogeneous healthcare datasets into reusable data models that support observational research and epidemiology studies. Convert bespoke datasets, such as biomarkers and mutation data, into OMOP format where appropriate, while identifying residual data that cannot be standardized and determining how it can still support analysis. Build FAIR (Findable, Accessible, Interoperable, Reusable) data pipelines and semantic data engineering frameworks that improve healthcare data discoverability and interoperability. Develop AI-ready datasets and data products capable of supporting generative AI and other advanced analytics use cases. Engage with epidemiologists, statisticians, market access specialists, health economists, and other stakeholders to understand, scope, document, and translate business requirements into actionable technical data structures. Collaborate with the RWE programming team to develop data structures required for study outputs and provide technical support where data engineering expertise adds value. Liaise with IT teams to ensure inbound datasets from data partners are fit for their intended analytical purposes. Collaborate with technical teams from data and analytics software vendors, including Databricks, when required. Maintain clear documentation covering data flows, schemas, pipelines, and processes to support onboarding, troubleshooting, and auditing. Design and implement comprehensive testing, validation, and monitoring approaches to ensure the accuracy, reliability, and quality of data products. Troubleshoot issues related to data loading, extraction, transformation, and ETL processes. Collaborate with other members of the Data Engineering team, providing support and taking on additional workload when needed. Requirements Strong understanding of Real World Data (RWD) and Real World Evidence (RWE) concepts and their application to healthcare analytics. Ability to assess business requirements and recommend appropriate real-world healthcare datasets for analytical use cases. Deep understanding of healthcare data models, healthcare data ecosystems, and patient-level datasets. Strong expertise in OMOP Common Data Model v5.4 and v6, including extensions. Knowledge of healthcare standards and terminologies such as SNOMED CT, RxNorm, ICD-10, LOINC, and HCPCS/CPT. Strong experience building scalable ETL/ELT pipelines using Databricks, PySpark, Spark SQL, SQL, and Delta Lake. Experience working with large-scale healthcare and patient-level datasets and distributed data processing frameworks. Strong understanding of Semantic Data Engineering principles and experience developing FAIR-compliant data pipelines. Experience with cloud-based data platforms and large-scale distributed processing environments. Strong Power BI development and data modeling capabilities, with the ability to create reusable analytical datasets for dashboards and studies. Experience designing AI-ready datasets and analytics data products. Strong data profiling, validation, monitoring, and automated data quality framework experience. Understanding of healthcare data quality assessment methodologies. Excellent stakeholder management, communication, and collaboration skills, with the ability to translate complex business needs into effective technical solutions. Experience working with cross-functional and globally distributed teams. Exposure to one or more therapeutic areas such as Oncology, Respiratory, Immunology and Inflammation, or Infectious Diseases. Nice-to-have: working knowledge of R, sparklyR, R Shiny, observational research methodologies, OHDSI tools, and Azure Data Platform services. Benefits India-based opportunity with a flexible working environment. Opportunity to work with large-scale healthcare and patient-level datasets. Exposure to Real World Data, Real World Evidence, observational research, epidemiology, and healthcare analytics. Opportunity to contribute to FAIR data engineering and AI-ready data initiatives. Collaborative environment with cross-functional and globally distributed teams. Opportunities to work across data engineering, analytics, healthcare standards, and emerging AI use cases. Professional development through collaboration with experienced data engineering, programming, analytics, and healthcare specialists. Inclusive workplace culture that values diversity, integrity, honesty, and respect.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- .Net Software DeveloperJobgether
- Account DirectorJobgether
- Account Director, Renewals & GrowthJobgether
- Advisor, BMO SmartFolio WFHJobgether
- Agentic AI DeveloperJobgether
- AI Graphic Designer + Video EditorJobgether
- AI/ML Data ScientistJobgether
- Analista de Automação e IA com N8NJobgether
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on lever · posted 2026-09-22. ApplySarthi collects openings and links to application pages; the role is advertised by Jobgether, not by us.