Senior AI & LLM Systems Engineer
Design custom model routing engines, high-throughput inference gateways, and autonomous multi-agent systems.
- Location
- Remote (Global / US / Pakistan)
- Level
- Senior
- Experience
- 3+ Years
- Type
- Full-Time
Stack
PyTorchvLLMPythonLangGraphQdrant / pgvectorFastAPITensorRT
What you will do
- Develop custom LLM orchestration backends with intelligent multi-model failover and speculative routing.
- Build high-precision RAG pipelines using hybrid semantic/dense embeddings and reranking algorithms.
- Deploy self-hosted open weights models on GPU clusters with vLLM, TensorRT-LLM, and quantization (AWQ/GPTQ).
- Instrument rigorous model evaluation benchmarks (evals) for hallucination detection and response latency.
What we are looking for
- 3+ years of AI/ML engineering experience with direct production exposure to LLMs and embedding architectures.
- Strong Python systems programming skills with FastAPI, asynchronous IO, and CUDA/PyTorch memory management.
- Deep understanding of vector databases, contextual retrieval mechanics, and agentic workflow frameworks.
- Experience optimizing inference throughput (tokens/sec) and GPU utilization.