
Fuse Energy
Dubai
AI Inference Engineer
DubaiRemote OK5–10 yrsFull-time
Responsibilities
Define and build the AI inference serving strategy and architecture from first principles to handle high-throughput, latency-sensitive workloads. Own the model-level optimization strategy and integrate low-level GPU/CUDA performance work into the serving layer.
Requirements
Requires 4+ years of experience building large-scale inference serving systems with deep expertise in optimization techniques like quantization and batching. Candidates must possess strong systems thinking and the ability to operate independently in a founding role.
Key skills
AI Inference ServingCUDAGPU Performance OptimizationvLLMTensorRT-LLMSGLangTriton Inference ServerQuantizationSpeculative DecodingKV-Cache ManagementRequest RoutingAutoscalingSystems ArchitectureKubernetesSlurmDistributed Systems
Benefits
Competitive Salary
Equity Sign-on Bonus
Biannual Bonus Scheme
Fully Expensed Tech
Breakfast And Dinner Allowance
Apply now
You'll apply on the employer's official page.