Member of Technical Staff Research Software Engineer
NY, London, CA
On-site
Permanent
27 Applications
Job description
You will architect and optimize training infrastructure for reinforcement-learning loops, distributed GPU systems, and large-scale data pipelines. You will turn research ideas into reliable production systems, build experiment tooling, diagnose training bottlenecks, and improve numerical stability, performance, and reproducibility.
Responsibilities
- Design and optimize large-scale training loops and data pipelines.
- Implement state-of-the-art techniques with numerical stability and computational efficiency.
- Build tooling to launch, monitor, and reproduce complex experiments.
- Diagnose training-stack bottlenecks, including GPU memory, communication, and dataloader issues.
- Translate research prototypes into reusable production-grade infrastructure.
Requirements
- Software engineering skills and machine-learning knowledge.
- Experience implementing research papers.
- Deep experience in distributed training and inference or data infrastructure.
- Knowledge of machine-learning algorithms, distributed systems, and high-performance computing.
- Performance, numerical stability, and reproducibility expertise.
Benefits
- Stock options.
- Comprehensive medical, dental, vision, and life insurance.
- Annual wellness allowance.
- Daily in-office lunch and dinner.
- 22 weeks of paid parental leave.
- Unlimited paid time off in the U.S. and 30 days in the U.K.
- Visa sponsorship and long-term immigration-pathway support where applicable.
- Regular off-sites, happy hours, and team celebrations.
Personalized Google Job Alerts
Never miss high-paying UK jobs — get instant alerts on Google
See newly verified UK job openings and transparent salary benchmarks on Google before other candidates apply.
Get Instant Job Alerts on Google
1-Click on Google
•
No Sign-up
•
100% Free
Is there something wrong with this job listing? Let us know.