Senior AI Researcher- Pre-training (f/m/d)
Alephalpha
Tailor my CV for this job, freeView job and applyYour CV rewritten for this role, from your real experience. Sign in with Google, nothing to install.
Got this interview? Our apps help you get the job.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
272 open researcher roles across 89 companies are on ApplySarthi right now, most of them in Bengaluru (11), Mumbai (2), Delhi NCR (2).
- Postdoctoral Researcher for Longitudinal Surveys (DRS-57)GESIS – Leibniz-Institut für Sozialwissenschaften
- Blockchain Security ResearcherJobgether
- Principal macOS Security ResearcherHuntress
- Experience Researcher, PLGAsana
- Senior Market Researcher, Insurance Product & MarketingOscar Health
What researcher roles keep asking for: Python (42%), Machine learning (35%), C++ (21%), LLMs (17%), Deep learning (15%), PyTorch (13%) — counted across their open postings here.
LLMs jobs · Machine learning jobs · PyTorch jobs · Python jobs
Alephalpha has 3 open roles listed here.
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for researcher roles keep coming back to Python, Machine learning, C++, LLMs. Practise those questions before you sit with Alephalpha.
Questions you are likely to be asked
- Why do you want to join Alephalpha?
- What is your experience with PyTorch? Tell me one thing you learned the hard way.
- What would you check first if a model's accuracy dropped after going live?
- When would you not use machine learning for a problem?
- Walk me through a model you built, from the data to how it was used.
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior AI Researcher- Pre-training (f/m/d) at Alephalpha interview free →Our Mission Aleph Alpha is one of the few companies in Europe doing serious foundation model pre-training. Our customers — in finance, manufacturing, and public administration — need models that understand German, meet European regulatory requirements, and work reliably in high-stakes settings. We’re building that in Heidelberg. We are hiring a Senior AI Researcher to join our Pre-training team and to advance the architecture and training of our next generation of foundation models. If you are excited about designing inference-efficient architectures, optimising training recipes that scale reliably, and training models on a large scale cluster (thousands of NVIDIA Blackwell GPUs), we would love to hear from you. Team Culture We foster a culture built on ownership, autonomy, and empowerment. Teams and individual contributors are trusted to take responsibility for their work and drive meaningful impact. We maintain a flat organisational structure with efficient, supportive management that enables quick decision-making, open communication, and a strong sense of shared purpose. We collaborate closely on complex technical problems, working in pairs or using mob programming to resolve challenging issues. About the Role As a Senior AI Researcher in Pre-training (f/m/d) , you will own the critical technical levers that determine the success of our next-generation models: architecture, optimization, stability, and scaling. Working at the high-leverage intersection of research and engineering, you will translate mathematical reasoning and empirical observations into principled training decisions - from small-scale proxy experiments to multi-thousand-GPU runs. We are looking for an expert who can combine rigorous experimental design with high-quality production code, directly influencing model quality, run reliability, and the efficiency of the models we ship. Your Responsibilities Recipe & Architecture Optimization: Own core elements of the training recipe (optimizers, schedules, initialization) and design PyTorch-based architectural improvements to maximize convergence, stability, and training efficiency. Scaling Strategy & Predictability: Develop hyperparameter scaling laws and scale-up methodologies, using small-scale proxy experiments to reliably predict multi-thousand-GPU behavior and de-risk major training decisions. Stability, Diagnostics & Debugging: Investigate complex convergence issues (loss spikes, divergence) and resolve hard-to-reproduce distributed system failures like communication bottlenecks, race conditions, and synchronization errors. System-Model Co-Design: Partner with Compute Performance, Data, Evaluation, and Post-Training teams to align the model lifecycle with hardware constraints, memory bandwidth, and communication topologies. Core Qualifications You are proficient in Python and deeply familiar with PyTorch-based training workflows. You have a strong track record in machine learning research and software engineering, demonstrated through shipped models, impactful open-source contributions, or published research. You have a strong mathematical foundation and are comfortable reasoning formally about optimisation, scaling behaviour, and training dynamics. You deeply understand transformer training dynamics, optimisation, and the behaviour of large distributed training jobs. You can design rigorous experiments, reason clearly from noisy results, and translate empirical observations into robust training decisions. Hands-on experience pre-training large models (e.g., 7B+ parameters) on substantial infrastructure (e.g., 100+ GPU clusters). You apply strong software engineering practices, including writing maintainable, well-tested code and supporting reproducible experimentation workflows. You are able to implement complex model architectures efficiently and reliably and to debug complex issues across model code, training dynamics, and distributed systems. You collaborate effectively within a research and engineering team and communicate clearly about your work across Pre-training and the broader AAR/AA organization. You are able to work in Germany and collaborate regularly on site in Heidelberg as part of the Pre-training team. Preferred Qualifications Large-Scale Training: Hands-on experience training LLMs or multimodal models on large GPU clusters using distributed frameworks (e.g., Megatron-LM, DeepSpeed, torchtitan). Predictive Scaling: Familiarity with scaling laws, hyperparameter transfer, or methods for predicting large-scale training behavior from smaller proxy runs. Stability & Performance: Experience profiling distributed jobs and diagnosing training anomalies like loss spikes, numerical instability, or optimizer pathologies. Advanced Architectures: Exposure to sparse training approaches (e.g., Mixture-of-Experts) and an understanding of their routing and systems trade-offs. Track Record of Impact: Demonstrated research excellence through top-tier publications (NeurIPS, ICML, ICLR), impactful open-source contributions, or significant shipped technical work. Systems Curiosity: Low-level kernel optimization is not required, but we highly value a strong curiosity about the hardware and systems constraints that shape scale. What we offer Become part of an AI revolution! 30 days of paid vacation Access to a variety of fitness & wellness offerings via Wellhub Mental health support through nilo.health Substantially subsidized company pension plan for your future security Subsidized Germany-wide transportation ticket Budget for additional technical equipment Flexible working hours for better work-life balance and hybrid working model JobRad® Bike Lease Find Jobs in Germany on Arbeitnow
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on arbeitnow · posted 2026-10-10. ApplySarthi collects openings and links to application pages; the role is advertised by Alephalpha, not by us.