Responsibilities:
-
Kernel Optimization: Write and optimize custom CUDA kernels to accelerate our proprietary entropy simulation algorithms.
-
Architecture Design: Leverage the NVIDIA RAPIDS™ ecosystem and cuDF to build high-throughput data pipelines that bypass CPU bottlenecks.
-
Latency Engineering: Profile and debug system performance to achieve sub-millisecond latency for real-time industrial decision-making.
-
Hardware Integration: Collaborate with the AI team to deploy models effectively on H100/A100 clusters via our cloud partners.
-
Sovereign Logic: Ensure all compute processes adhere to our strict data sovereignty and security protocols.
Qualifications:
-
4+ years of experience in High-Performance Computing (HPC) or Systems Programming.
-
Deep expertise in C++ and CUDA programming.
-
Strong familiarity with Python internals and the NVIDIA software stack (TensorRT, Triton Inference Server).
-
Understanding of memory management, threading, and parallel algorithms.
-
B.Tech/M.Tech in Computer Science (bonus points for research background in Distributed Systems).
Compensation & Benefits:
-
Compensation: Competitive Market Salary + Phantom Equity Units.
-
Tech Stack: Full access to the NVIDIA Connect ecosystem and enterprise-grade compute resources.
-
Culture: Radical Autonomy. No micromanagement. You own the optimization layer.