Data Curation Intern
Karya
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
28 open curation roles across 7 companies are on ApplySarthi right now, most of them in Chennai (1), Bengaluru (1).
- Market Curation Editor (Sports)Jobgether
Hugging Face jobs · Machine learning jobs · NLP jobs · Python jobs
Karya has 15 open roles listed here.
- Research Intern - AI Evaluationsbengaluru
- Strategy & Operations Analystbengaluru
- Lead - Public Policydelhi ncr
- Technical Support & Solutions Associate - Platform-as-a-Service (PaaS)bengaluru
- Community Coordinator - Work Managementbengaluru
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
28 open curation roles are hiring right now across 7 companies, mostly in Chennai — so the questions repeat. Practise them before you sit with Karya.
Questions you are likely to be asked
- Why do you want to join Karya?
- What is your experience with NLP? Tell me one thing you learned the hard way.
- Tell me about yourself, and why this role is the right next step.
- Tell me about a problem you solved at work that you are proud of.
- Tell me about a time you disagreed with your manager. What happened?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Data Curation Intern at Karya interview free →About Karya:
Why was Karya on the cover of the Time Magazine , highlighted by Satya Nadella , and invited to present its work to Sundar Pichai one on one?
In part, because Karya is on a mission to provide AI enabled earning and learning opportunities to communities with high talent, but low access to opportunities. Karya achieves this while also delivering high quality, timely, and price competitive data to its clients.
Karya builds high quality datasets for large companies like Google and Microsoft, while providing ethical work opportunities and fair wages to its workforce.
Karya’s workers make nearly 20 times the Indian minimum wage and through our one-of-a-kind digital work platform, we have delivered over 40 million digital tasks and have positively impacted over 100 thousand workers. In the coming years, our goal is to rapidly scale our impact by bringing economic opportunities to millions of underserved users in India. With a rapidly growing global presence, we are also looking to expand our client base in the Indian market by partnering with leading Indian enterprises.
About the Role
We are looking for a detail-oriented and curious Data Curation Intern to help build high-quality datasets for training AI/ML models with a specific focus on Indian language and multilingual data. You will work with large open-source datasets (e.g., Sangraha by AI4Bharat) that require significant cleaning, structuring, and enrichment before they can be used effectively in model training pipelines.
This is a hands-on, high-impact role at the intersection of data engineering, linguistics, and AI. You will start with text data pipelines and progressively move toward preparing data for read-speech and voice model training.
What You'll Do
Phase 1: Text Data Curation
Audit and profile open-source datasets (Sangraha, Common Crawl, IndicCorp, etc.) to assess quality, coverage, and noise levels
Design and implement data cleaning pipelines: deduplication, script normalisation, encoding fixes, noise removal, sentence boundary detection
Create and apply metadata tagging schemas labelling text by domain (news, legal, literature, health, etc.), subdomain, language, register, and quality tier
Build validation checklists and quality scorecards to benchmark dataset readiness for model training
Document data provenance, licensing, and processing steps for reproducibility
Phase 2: Speech & Voice Data Preparation
Curate high-quality, phonetically diverse text passages suitable for read-speech recording
Ensure text selection covers domain, prosodic, and phonemic variety required for TTS/ASR model training
Assist in defining metadata standards for audio datasets (speaker demographics, recording conditions, transcription format)
Support the pipeline transition from text corpus to aligned speech dataset
What We're Looking For
Must Have
Strong attention to detail — you notice inconsistencies others miss
Comfort with Python for data processing (pandas, regex, basic NLP libraries like spaCy or NLTK)
Familiarity with text data formats: CSV, JSONL, Parquet, plain text corpora
Curiosity about AI/ML, language technology, or computational linguistics
Ability to work independently, document work clearly, and communicate blockers early
Good to Have
Prior exposure to NLP datasets or open-source language resources (IndicNLP, AI4Bharat, Hugging Face datasets)
Knowledge of one or more Indian languages beyond English
Experience with data versioning tools (DVC, Git-LFS) or dataset platforms (Hugging Face Hub)
Basic understanding of how language models or speech models are trained
Why This Role
Work directly on real data pipelines that feed AI model training — not toy projects
Gain hands-on experience with large-scale multilingual and Indic language datasets
Build skills that are in high demand across AI labs, speech companies, and NLP startups
Clear progression path: text → read speech → voice data, with increasing responsibility
Mentorship from people who have built data and AI systems at scale
Karya celebrates diversity and is an equal opportunity employer. All applicants will be considered without regard to race, religion, gender identity, sexual orientation, disability, or any other protected status.
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- Analytics Lead - Data, Measurement & Impact Karya · bengaluru
- Associate - Impact ProgramKarya · bengaluru
- Community Coordinator - Work ManagementKarya · bengaluru
- Content Developer - Learning and CurriculumKarya · bengaluru
- Customer Success LeadKarya · bengaluru
- Entrepreneur in Residence - Karya XKarya · bengaluru
- Head of FinanceKarya · bengaluru
- Lead Product DesignerKarya · bengaluru
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on greenhouse · posted 2026-05-11. ApplySarthi collects openings and links to application pages; the role is advertised by Karya, not by us.