Job Overview : CUDA Engineering Expert (GPU Kernel Optimization, Remote)
| ๐ฐ Salary | $80โ$100 per hour |
| ๐ Location | Global (fully remote) |
| ๐ข Company | Mercor |
| ๐ผ Category | Software Development / GPU Engineering |
| ๐ Remote | Yes โ your own schedule |
| ๐ Contract Type | Hourly contract (independent contractor) |
| โฐ Commitment | Minimum 20 hours/week |
| ๐ธ Payment | Weekly via Stripe or Wise |
| ๐ฅ Hired this month | 84+ people |
| ๐ป Key Tech | CUDA, C++, Python, GPU Kernel Optimization |
About the role:
Mercor is seeking GPU kernel optimization experts to contribute to a project with a leading AI lab. This opportunity is designed for specialists with strong C++ skills, practical GPU programming experience, and the ability to improve kernel performance using profiler-guided analysis. You will help evaluate, optimize, and reason about GPU kernels across modern hardware environments.
What you will do:
- Analyze and optimize GPU kernels for performance, efficiency, and hardware utilization
- Use profiler metrics (L2 cache hit rate, occupancy, throughput) to guide kernel improvements
- Review GPU kernel implementations and identify bottlenecks
- Write, modify, and reason about C++17, Python, and GPU programming code
- Apply CUDA, HIP, shader programming, or related expertise to improve performance
- Document optimization decisions clearly, including when specific profiler metrics are or are not useful
What you need:
- Available to work at least 20 hours/week
- Fluent in core C++ features through C++17
- Working knowledge of Python and Git
- Fluent in at least one GPU programming model: CUDA, HIP, Slang, HLSL, GLSL, or related
- At least 1 year of professional or graduate-level research experience working with GPUs
- Strong understanding of GPU profiler performance metrics and how to use them to optimize kernels
- Ability to optimize GPU kernels without needing deep prior context on every algorithm
Nice to have (but not required):
- Experience with CUDA, HIP, CUDA C++ Core Libraries, inline PTX assembly, or tensor core-level optimization
- Experience optimizing kernels for NVIDIA Blackwell hardware
- Familiarity with NSight Compute
- Prior experience with GPU hardware organizations such as NVIDIA, AMD, or Qualcomm
- Open-source contributions related to GPU kernel optimization
Important notes:
- ๐ Global remote โ work from anywhere
- โฐ Minimum 20 hours/week โ flexible schedule
- ๐ฐ Premium rate โ $80-100/hour for GPU specialists
- ๐ Independent contractor โ not W-2
- ๐ง CUDA Expert Assessment required as part of application
Why this job is worth your time:
$80-100/hour to apply your GPU optimization expertise to cutting-edge AI research, with flexible hours and remote work.
