Home Job Details
C
Artificial Intelligence 🏢 Full Time ⭐️ Verified

Senior AI Safety & Alignment Engineer 2026 | San Francisco, CA

Chronos Future Systems
San Francisco
Estimated Salary
USD 180.000 – USD 260.000
Live Update
29 Juni 2026
Deadline
29 Jun 2027

Job Description

We are seeking a visionary AI Safety & Alignment Engineer to join our elite research division. As we build the foundational models for the year 2026 and beyond, your work will be critical in ensuring artificial general intelligence remains safe, beneficial, and aligned with human values. You will operate at the intersection of cutting-edge machine learning and ethical philosophy.

Why Join Us?

We are backed by top-tier venture capital and are currently scaling our team to lead the industry in Agentic AI development. You will have the autonomy to shape research directions, access to state-of-the-art compute resources, and a culture that prioritizes long-term impact over short-term metrics.

Key Responsibilities:

  • Design and implement safety guardrails for Large Language Models (LLMs) to prevent misuse and ensure alignment with user intent.
  • Conduct adversarial red-teaming and stress-test models to identify and mitigate potential vulnerabilities in autonomous agents.
  • Develop and refine reinforcement learning from human feedback (RLHF) pipelines to optimize model behavior.
  • Collaborate with cognitive scientists and ethicists to define long-term value alignment frameworks.
  • Write and maintain high-performance Python code for model fine-tuning and evaluation benchmarks.
  • Publish research findings in top-tier AI conferences and contribute to open-source safety initiatives.

Qualifications:

  • Master’s or PhD in Computer Science, Artificial Intelligence, or a related technical field.
  • 5+ years of professional experience in Machine Learning, Deep Learning, or Natural Language Processing.
  • Strong proficiency in Python, PyTorch, and TensorFlow.
  • Experience with model interpretability, causal inference, or mechanistic interpretability.
  • Deep understanding of transformer architectures and current LLM capabilities.
  • Excellent problem-solving skills and the ability to thrive in a fast-paced, research-driven environment.

Responsibilities

  • Design and implement safety guardrails for Large Language Models (LLMs) to prevent misuse and ensure alignment with user intent.
  • Conduct adversarial red-teaming and stress-test models to identify and mitigate potential vulnerabilities in autonomous agents.
  • Develop and refine reinforcement learning from human feedback (RLHF) pipelines to optimize model behavior.
  • Collaborate with cognitive scientists and ethicists to define long-term value alignment frameworks.
  • Write and maintain high-performance Python code for model fine-tuning and evaluation benchmarks.
  • Publish research findings in top-tier AI conferences and contribute to open-source safety initiatives.

Qualifications

  • Master’s or PhD in Computer Science, Artificial Intelligence, or a related technical field.
  • 5+ years of professional experience in Machine Learning, Deep Learning, or Natural Language Processing.
  • Strong proficiency in Python, PyTorch, and TensorFlow.
  • Experience with model interpretability, causal inference, or mechanistic interpretability.
  • Deep understanding of transformer architectures and current LLM capabilities.
  • Excellent problem-solving skills and the ability to thrive in a fast-paced, research-driven environment.

Required Skills

Python Machine Learning Deep Learning NLP LLM PyTorch TensorFlow AI Safety RLHF Transformer Models Ethics in AI

Ready to Take This Challenge?

Make sure your resume is ready. Submit your application now before the deadline.

Apply Now

Related Jobs

Similar job recommendations for you

View All