Senior Software Engineer - AI Compute, Together Cloud
Together AI
Make my CV for this job, freeView job and applyYour CV, rewritten for this role using only your real experience. Sign in with Google and upload your CV. Nothing to install.
Skills named in this job
Read from the description itself, not inferred.
This role on the market
11,752 open software roles across 741 companies are on ApplySarthi right now, most of them in Bengaluru (808), Hyderabad (309), Pune (183).
- Software Development Engineer, AWS SecurityAmazon Development Centre (London) Limited
- Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - FederalServiceNow
- ETIC, Software Testing, ManagerPwc
- Staff Software Engineer, Web Application ServicesMozilla
- Software Student for x86 Validation ToolsIntel
What software roles keep asking for: AWS (24%), Python (23%), Java (23%), System design (20%), Kubernetes (18%), C++ (17%), Observability (16%), CI/CD (13%) — counted across their open postings here.
Software Engineer jobs in the Netherlands · Software Engineer jobs in Amsterdam · Remote Software Engineer jobs · AWS jobs · Ansible jobs · Azure jobs · CI/CD jobs
Together AI has 78 open roles listed here.
- Senior Recruiter, GTM & Business
- Staff Software Engineer - AI Compute, Together Cloud
- Research Intern, Frontier Agents (Summer 2027)
- Research Intern, Frontier Agents (Winter 2027)
- Research Intern, Inference (Summer 2027)
Counted across 14 company job boards, updated as roles open and close.
Preparing for this interview
Interviews for software roles keep coming back to AWS, Python, Java, System design. Practise those questions before you sit with Together AI.
Questions you are likely to be asked
- Why do you want to join Together AI?
- What is your experience with Kubernetes? Tell me one thing you learned the hard way.
- How would you explain your model's result to someone who is not technical?
- What would you check first if a model's accuracy dropped after going live?
- When would you not use machine learning for a problem?
Prep Sarthi gives you a free mock interview: an AI interviewer asks you questions like these out loud, from your own CV and this job, and shows your score and your weakest answer.
Practise the Senior Software Engineer - AI Compute, Together Cloud at Together AI interview free →About the Role
Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products.
As a Senior Software Engineer focusing on AI Compute in the Together Cloud org, you will build and own major components of the next generation AI cloud platform – a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products – inference, RL, and fine-tuning – and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs.
The hard problems in rapidly scaling heterogeneous GPU fleets are software problems: fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes. We solve them by building global and in-DC services, Kubernetes operators, and high-performance SDN libraries, forking hypervisors and Linux kernels, and building infra tailored to inference and fine-tuning. You'll own massive greenfield projects across their full lifecycle – scoping the problem and writing PRDs with our PMs, designing the system, and building it through to GA launch.
Responsibilities
- Build the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that makes GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware.
- Build and maintain our in-DC IaaS layer: the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers, including VMs, parallel filesystems, VPCs, and InfiniBand partitions. Implement and harden the bring-up path for a new Vera Rubin data center with thousands of GPUs.
- Scale the distributed GPU scheduling and global management plane: the control-plane services behind on-demand and reserved clusters, including the automation that onboards new capacity and raises per-cluster limits.
- Harden the monitoring and automated remediation layer for fault tolerance: automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference fault-tolerant.
- Own your components end-to-end: write the design docs, break the work into milestones that ship incrementally, and improve the reliability of what's already in production.
- Raise the bar around you: code review, design feedback, and mentoring junior engineers.
- Build the tooling other teams rely on: testing frameworks, developer tools, and documentation that make our systems robust and usable across teams, plus contributions to the core, open-source Together AI platform.
To be successful, you'll need to be deeply technical, ready to own projects truly end-to-end, and an excellent communicator — strong software development fundamentals, strong systems knowledge and troubleshooting instincts, and the collaboration and diplomacy skills to work across teams.
Requirements
- 5+ years of professional software development experience, with strong proficiency in at least one backend programming language (Golang desired).
- Demonstrated ownership of large-scale projects driven end-to-end to completion, writing high-performance, well-tested, production-quality code.
- Demonstrated experience building and operating high-performance and/or globally distributed micro-service architectures across one or more cloud providers (AWS, Azure, GCP).
- Strong systems knowledge across compute, networking, and storage, including concurrency, memory management, performant I/O, and scale.
- Excellent communication skills — able to write clear design docs and work effectively with both technical and non-technical team members.
- Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy.
Preferred Qualifications:
- Proficiency with Kubernetes internals, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself
- Proficiency with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV
- Proficiency with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN
- Experience with Cluster API or similar
- Experience working on high-performance compute, networking, and/or storage
- Experience virtualizing GPUs and/or InfiniBand
- Experience building IaaS or PaaS systems at scale
- Experience with DPUs/SmartNICs
- GPU programming, NCCL, CUDA knowledge
About Together AI
Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at https://www.together.ai/privacy
Match this job to your CV
ApplySarthi scores your CV against this role, shows the skills you are missing, and writes a tailored version for the application.
Check my match →Similar open roles
- AI Infrastructure System Engineer Bangalore Together AI · bengaluru
- Junior/Senior or Staff Software Engineer, Inference / Compute Infrastructure EngineeringTogether AI
- Senior Program Manager, Data Center DeliveryTogether AI
- Senior Software Engineer — Infra Agent Systems Remote IndiaTogether AI
- Technical Support Engineer (GPU Cluster), India Together AI · bengaluru
- Technical Support Engineer (Inference) - India WeekendsTogether AI · bengaluru
- Staff Engineer, Distributed Storage and HPC & AI InfrastructureTogether AI · bengaluru
- Technical Support Engineer - IndiaTogether AI · bengaluru
Need answers during your interview? Try Live Sarthi.
Live Sarthi, an Interview Sarthi app, shows answer suggestions during the call.
- Hidden from supported screen sharingThe overlay stays out of supported Windows screen captures.
- Answers start in about 1.5 secondsResponse time varies with your connection and model.
- From your own CVYour projects and your experience, not a generic script.
- 30 minutes freeThen ₹99 for a 2-day pass with unlimited calls — you pay for the days you are interviewing, not a subscription.
A Windows app, from the same team as ApplySarthi.
Listed on greenhouse · posted 2026-01-20. ApplySarthi collects openings and links to application pages; the role is advertised by Together AI, not by us.