Senior Machine Learning Research Engineer
Senior / Staff Machine Learning Research Engineer
Location: San Francisco Bay Area (Hybrid)
Level: Senior or Staff
Reports to: Head of Machine Learning
The Opportunity
A well-resourced, research-driven technology organization is building out a dedicated Machine Learning Research Engineering function to support the development of cutting-edge machine learning systems across large-scale, high-dimensional datasets.
The team has reached a point where research ambition is no longer constrained by ideas, but by the speed, scalability, and reliability of the underlying engineering systems. This hire will help remove that constraint.
You will operate at the intersection of machine learning research and systems engineering, owning the infrastructure, tooling, and technical foundations that allow researchers to move efficiently from hypothesis to trained model to reproducible result.
This is a high-leverage role. The systems, libraries, and engineering standards you establish will directly shape how machine learning research is conducted across the organization for years to come.
What You'll Own:
Research Infrastructure
- Design and build the core platform that supports large-scale machine learning development, including distributed training orchestration, asynchronous execution, experiment management, and high-throughput data pipelines. You will identify bottlenecks before they become problems and create systems that enable researchers to move faster.
Model Engineering at Scale
- Partner with researchers to transform novel model architectures into efficient, production-grade implementations. You will optimize training and inference workloads, improve accelerator utilization, implement distributed training strategies, and solve performance bottlenecks across large-scale compute environments.
ML Systems & Framework Development
- Build and maintain critical tooling, libraries, and workflows that support model development in PyTorch and JAX. You will improve reproducibility, scalability, experimentation velocity, and overall developer experience for a growing ML organization.
Engineering Standards
- Establish best practices around testing, reproducibility, model versioning, evaluation, code quality, and experiment tracking. Success means improving reliability and engineering rigor without slowing research velocity.
Technical Partnership
- Work closely with machine learning researchers and data engineers to translate cutting-edge ideas into practical systems. You should be comfortable discussing model architectures, training methodologies, and systems constraints while helping bridge the gap between research and engineering.
What We're Looking For:
- MS, or PhD in Computer Science, Machine Learning, Engineering, or a related technical discipline
- Proven experience building machine learning systems that were widely adopted by researchers, engineers, or end users
- Expert-level Python development skills
- Deep hands-on experience with both PyTorch and JAX
- Experience scaling machine learning workloads using distributed training techniques across GPU and/or TPU environments
- Strong understanding of model optimization, accelerator performance, memory management, training throughput, and parallelization strategies
- Excellent software engineering fundamentals, including system design, API design, testing, maintainability, and developer tooling
- Demonstrated ability to independently drive ambiguous, multi-quarter technical initiatives from concept through production
- Comfort operating in a fast-moving environment where priorities evolve alongside scientific and technical discovery
Strong Signals
- Experience building internal ML platforms, frameworks, or tooling that researchers actively adopted
- Deep understanding of distributed training systems, including model parallelism, data parallelism, and large-scale experimentation workflows
- Prior work involving large-scale foundation models, multimodal models, sequence models, or other compute-intensive ML systems
- Experience with high-performance computing environments and large-scale datasets
- Meaningful open-source contributions, technical publications, or presentations in the machine learning community
- Experience operating within research-first organizations where processes, platforms, and standards were built from the ground up rather than inherited
Why This Role
This is an opportunity to work on technically ambitious machine learning problems while having outsized influence on the direction of an organization's ML infrastructure.
The ideal candidate enjoys building systems that accelerate research, thrives in highly technical environments, and wants to have a direct impact on how next-generation machine learning capabilities are developed at scale.
FAQs
Congratulations, we understand that taking the time to apply is a big step. When you apply, your details go directly to the consultant who is sourcing talent. Due to demand, we may not get back to all applicants that have applied. However, we always keep your resume and details on file so when we see similar roles or see skillsets that drive growth in organizations, we will always reach out to discuss opportunities.
Yes. Even if this role isn’t a perfect match, applying allows us to understand your expertise and ambitions, ensuring you're on our radar for the right opportunity when it arises.
We also work in several ways, firstly we advertise our roles available on our site, however, often due to confidentiality we may not post all. We also work with clients who are more focused on skills and understanding what is required to future-proof their business.Â
That's why we recommend registering your resume so you can be considered for roles that have yet to be created.Â
Yes, we help with CV and interview preparation. From customized support on how to optimize your CV to interview preparation and compensation negotiations, we advocate for you throughout your next career move.
