Senior Machine Learning Engineer | Foundation Models & CUDA Optimization
Lead the development and production deployment of a novel foundation model for automated software delivery in embedded systems. Design custom CUDA kernels, optimize distributed ML systems, and architect scalable infrastructure. Requires deep expertise in large-scale foundation models, CUDA, and production ML frameworks like PyTorch.
Key Highlights
Key Responsibilities
Technical Skills Required
Benefits & Perks
Nice to Have
Job Description
Hiring via TechTree
This is a role that TechTree is recruiting for on behalf of one of its clients.
TechTree is an AI-driven recruitment platform working with high-growth companies.
When you apply, TechTree's AI Agent matches you not just to this role, but to other relevant opportunities across its network, so one application can unlock multiple roles.
Senior ML Engineer | Foundation Models & CUDA
TechTree's client is hiring a Senior ML Engineer to help build and scale a novel foundation model for automated software delivery in embedded systems.
This is a deeply technical role for an engineer who has built large-scale foundation models, developed custom CUDA kernels, and scaled distributed ML systems in production.
- Location: Brazil — remote
- Employment: Full-time
- Estimated compensation: £100,000–£150,000/year + equity
- Level: Mid-Senior / Senior
- Relocation: UK visa sponsorship may be available
- Lead development and production deployment of a large-scale foundation model
- Design and implement custom CUDA kernels for performance-critical workloads
- Optimise training and inference across GPUs and distributed infrastructure
- Architect scalable ML systems with demanding reliability and performance requirements
- Profile and improve data pipelines, training loops, inference, and serving infrastructure
- Evaluate modern architectures including Mixture-of-Experts and state-space models
- Build internal tooling, benchmarks, evaluation systems, and observability infrastructure
- Work closely with founders on technical strategy and optimisation priorities
- Improve model quality, latency, scalability, and compute efficiency
Looking to advance your Development & Programming career with relocation support? Explore Development & Programming Jobs with Relocation Packages that include comprehensive packages to help you move and settle in your new role.
- Deep professional experience with CUDA C/C++
- Strong Python skills
- Experience building and shipping large-scale foundation models
- Proven custom CUDA kernel development
- Experience with distributed training and inference
- Strong knowledge of modern deep learning architectures
- Production expertise with PyTorch or another major ML framework
- Experience operating ML workloads on AWS, Azure, or GCP
- Strong understanding of ML observability, evaluation, and production SLOs
- Experience delivering complex systems in fast-moving environments
Discover our full range of relocation jobs with comprehensive support packages to help you relocate and settle in your new location.
- Built a foundation model from 01
- Early-stage AI startup or top-tier research lab experience
- Experience with MoE, state-space models, and advanced inference optimisation
- Built tooling that significantly improved engineering or research productivity
- Experience optimising large-scale GPU workloads for cost and performance
- Work on a genuinely novel foundation-model problem
- Significant technical ownership from day one
- Direct influence over architecture and ML strategy
- Work closely with an experienced founding team
- Competitive compensation plus equity
- Remote opportunity for candidates based in Brazil
- Potential UK relocation support
Similar Jobs
Explore other opportunities that match your interests