Mid-level/Full-time/EU Remote/United States / Singapore / Netherlands
LLM Inference Frameworks and Optimization Engineer
Together AIHiring SurgeEU RemoteLikely hiring
Role Opportunity Score
46Fair opportunity
Niche stack
Expected applicants: ~80–150
Together AI is hiring for a Mid LLM Inference Frameworks and Optimization Engineer with around 80-150 applicants expected, given their Greenhouse listing and sizable social following. Experience with vLLM, CUDA, or Triton represents a real filter, narrowing the field considerably.
Evidence Dossier
Remote Details
Scope: EU Remote
Role Summary
Design, develop, and optimize distributed inference engines that support multimodal and language models at scale, focusing on low-latency, high-throughput inference, GPU/accelerator optimizations, and software-hardware co-design for efficient large-scale deployment of LLMs and vision models.
Stack Required
vLLMCUDATritonPyTorchPythonC++
Compensation
$160,000 - $230,000 / year
startup equityhealth insurance
Requirements
Experience: 3+ years
Languages: English
Visa sponsorship not stated
Equity Offered
AI Tools Culture
AI mentionedLast seen on career page: 29d ago
Posted: 1mo ago