Job Description
At Nexus Future Systems, we aren't just predicting the future; we are building it. Our mission is to define the 2026 AI Stack, delivering next-generation generative intelligence that powers enterprise-scale solutions.
We are seeking a visionary Senior Generative AI Engineer to join our elite engineering team. You will be at the forefront of deploying Large Language Models (LLMs) into production environments, optimizing latency, and ensuring robust, secure, and scalable AI infrastructure.
Responsibilities
- Architect & Deploy: Design and implement high-performance LLM inference pipelines and RAG (Retrieval-Augmented Generation) architectures for the 2026 standard.
- Model Optimization: Apply quantization, pruning, and distillation techniques to reduce latency and cost while maintaining model accuracy.
- MLOps Integration: Build and maintain CI/CD pipelines for model training, validation, and deployment using Kubernetes and modern cloud infrastructure.
- Performance Tuning: Conduct rigorous benchmarking and load testing to ensure systems handle millions of concurrent requests.
- Collaboration: Partner with product and data science teams to translate complex AI capabilities into user-friendly applications.
Qualifications
- Expertise: 5+ years of experience in Python, with deep proficiency in PyTorch or TensorFlow.
- LLM Mastery: Proven track record of working with open-source models (Llama 3, Mistral) and commercial APIs (GPT-4, Claude).
- System Design: Strong understanding of distributed systems, microservices, and vector database architectures (Pinecone, Milvus, Weaviate).
- Education: BS in Computer Science, Mathematics, or a related technical field (Master's preferred).
- Problem Solving: Ability to troubleshoot complex algorithmic issues and optimize for production environments.