Job Description
We are seeking a visionary AI Safety & Alignment Engineer to join our elite research division. As we build the foundational models for the year 2026 and beyond, your work will be critical in ensuring artificial general intelligence remains safe, beneficial, and aligned with human values. You will operate at the intersection of cutting-edge machine learning and ethical philosophy.
Why Join Us?
We are backed by top-tier venture capital and are currently scaling our team to lead the industry in Agentic AI development. You will have the autonomy to shape research directions, access to state-of-the-art compute resources, and a culture that prioritizes long-term impact over short-term metrics.
Key Responsibilities:
- Design and implement safety guardrails for Large Language Models (LLMs) to prevent misuse and ensure alignment with user intent.
- Conduct adversarial red-teaming and stress-test models to identify and mitigate potential vulnerabilities in autonomous agents.
- Develop and refine reinforcement learning from human feedback (RLHF) pipelines to optimize model behavior.
- Collaborate with cognitive scientists and ethicists to define long-term value alignment frameworks.
- Write and maintain high-performance Python code for model fine-tuning and evaluation benchmarks.
- Publish research findings in top-tier AI conferences and contribute to open-source safety initiatives.
Qualifications:
- Master’s or PhD in Computer Science, Artificial Intelligence, or a related technical field.
- 5+ years of professional experience in Machine Learning, Deep Learning, or Natural Language Processing.
- Strong proficiency in Python, PyTorch, and TensorFlow.
- Experience with model interpretability, causal inference, or mechanistic interpretability.
- Deep understanding of transformer architectures and current LLM capabilities.
- Excellent problem-solving skills and the ability to thrive in a fast-paced, research-driven environment.
Responsibilities
- Design and implement safety guardrails for Large Language Models (LLMs) to prevent misuse and ensure alignment with user intent.
- Conduct adversarial red-teaming and stress-test models to identify and mitigate potential vulnerabilities in autonomous agents.
- Develop and refine reinforcement learning from human feedback (RLHF) pipelines to optimize model behavior.
- Collaborate with cognitive scientists and ethicists to define long-term value alignment frameworks.
- Write and maintain high-performance Python code for model fine-tuning and evaluation benchmarks.
- Publish research findings in top-tier AI conferences and contribute to open-source safety initiatives.
Qualifications
- Master’s or PhD in Computer Science, Artificial Intelligence, or a related technical field.
- 5+ years of professional experience in Machine Learning, Deep Learning, or Natural Language Processing.
- Strong proficiency in Python, PyTorch, and TensorFlow.
- Experience with model interpretability, causal inference, or mechanistic interpretability.
- Deep understanding of transformer architectures and current LLM capabilities.
- Excellent problem-solving skills and the ability to thrive in a fast-paced, research-driven environment.