Home Job Details
A
Information Technology 🏢 Full Time ⭐️ Verified

Senior Generative AI Architect (Next-Gen LLMs)

Apex Intelligence Labs
San Francisco
Estimated Salary
USD 180.000 – USD 260.000
Live Update
30 Juni 2026
Deadline
30 Jun 2027

Job Description

We are seeking a visionary Senior Generative AI Architect to join our elite engineering team in San Francisco. As we prepare for the technology landscape of 2026, we need a technical leader who can architect the next generation of Large Language Models (LLMs) and autonomous agents.

In this high-impact role, you will be at the forefront of AI innovation, bridging the gap between theoretical research and scalable production systems. You will define the architectural patterns that power our intelligent products, ensuring they are robust, efficient, and ready for the future.

Responsibilities

  • Lead Model Architecture: Design and implement state-of-the-art Generative AI models, focusing on scalability and efficiency for 2026-era requirements.
  • Optimize Inference Pipelines: Engineer high-performance systems for Large Language Models (LLMs) to reduce latency and cost while maximizing throughput.
  • Build Advanced RAG Systems: Develop Retrieval-Augmented Generation architectures to enhance accuracy and reduce hallucinations in complex data environments.
  • Fine-Tuning Strategy: Lead initiatives in Parameter-Efficient Fine-Tuning (PEFT) and custom model training using proprietary datasets.
  • Collaborative Innovation: Partner with product managers and engineers to integrate AI agents seamlessly into existing software ecosystems.
  • Research & Prototyping: Experiment with emerging architectures, including multimodal models and agentic workflows, to stay ahead of the industry curve.

Qualifications

  • Experience: 7+ years of software engineering with a deep specialization in Machine Learning and Deep Learning.
  • Technical Stack: Mastery of Python, PyTorch, TensorFlow, and CUDA programming.
  • Model Deployment: Proven track record of deploying LLMs (e.g., GPT, Llama) into production environments using Kubernetes and Docker.
  • Data Engineering: Strong proficiency in vector databases (Pinecone, Milvus) and embedding technologies.
  • Education: BS, MS, or PhD in Computer Science, Mathematics, or a related quantitative field.
  • Problem Solving: Demonstrated ability to solve complex optimization problems and scale systems to handle massive data loads.

Required Skills

Python PyTorch TensorFlow LLMs Generative AI RAG Fine-Tuning Machine Learning Deep Learning NLP CUDA Kubernetes Docker

Ready to Take This Challenge?

Make sure your resume is ready. Submit your application now before the deadline.

Apply Now

Related Jobs

Similar job recommendations for you

View All