Genesis Logo

Genesis

Training / AI Infrastructure

Posted 6 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in CA
Senior level
In-Office or Remote
Hiring Remotely in CA
Senior level
Design, build, and optimize distributed PyTorch training systems for multi-node GPU clusters. Profile and eliminate performance bottlenecks across data pipelines to GPU kernels, implement low-level CUDA/cuDNN/Triton kernels, tune CPU/GPU/memory/network utilization, and develop monitoring and debugging tools for large-scale training runs.
The summary above was generated by AI
What You’ll Do
  • Drive down wall-clock time to convergence by profiling and eliminating bottlenecks across the foundation model training stack stack, from data pipelines to GPU kernels

  • Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization

  • Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks

  • Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory management, data throughput, and networking

  • Develop monitoring and debugging tools for large-scale runs, enabling rapid diagnosis of performance regressions and failures

What You’ll Bring
  • Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years)

  • Production-grade expertise in Python

  • Low-level performance mastery: CUDA/cuDNN/Triton, CPU–GPU interactions, data movement, and kernel optimization

  • Scaling at the frontier: experience with PyTorch and training jobs using data, context, pipeline, and model parallelism

  • System-level mindset with a track record of tuning hardware–software interactions for maximum utilization

Similar Jobs

Senior level
Agency • Artificial Intelligence • Blockchain • Web3
Design, orchestrate, and optimize large-scale LLM pre-training across 1,000+ GPUs. Implement 3D parallelism, manage GPU clusters (SLURM/Kubernetes), optimize InfiniBand/RDMA networking and memory, and automate checkpointing and failure recovery for long training runs.
Top Skills: 3D ParallelismC++CudaDeepspeedGpuInfinibandKubernetesMegatron-LmPythonPyTorchRdmaSlurm
An Hour Ago
Easy Apply
Remote
Canada
Easy Apply
Senior level
Senior level
Cloud • Security • Software • Cybersecurity • Automation
Build and maintain backend services and APIs using Ruby and GraphQL to improve source code workflows. Ensure reliability, performance, testing, observability, and on-call support. Collaborate cross-functionally and use AI agents responsibly to accelerate delivery.
Top Skills: Ai Coding AgentsBackground ProcessingCachingGitGitalyGraphQLMonorepoObservabilityRubySQL
An Hour Ago
Easy Apply
Remote
Canada
Easy Apply
Mid level
Mid level
Cloud • Security • Software • Cybersecurity • Automation
Build and maintain backend services and APIs for source code workflows using Ruby and GraphQL. Improve performance, reliability, caching, and observability. Collaborate cross-functionally, write tests, participate in Tier 2 on-call, and use AI coding agents where appropriate.
Top Skills: GitGitalyGraphQLRubySQL

What you need to know about the Calgary Tech Scene

Employees can spend up to one-third of their life at work, so choosing the right company is crucial, not just for the job itself but for the company culture as well. While startups often offer dynamic culture and growth opportunities, large corporations provide benefits like career development and networking, especially appealing to recent graduates. Fortunately, Calgary stands out as a hub for both, recognized as one of Startup Genome's Top 100 Emerging Ecosystems, while also playing host to a number of multinational enterprises. In Calgary, job seekers can find a wide range of opportunities.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account