PurPassionDigital / Everywhere
The landscapeAI & Machine Learning EngineeringAI Safety & Alignment Researcher
AI & Machine Learning Engineering · Digital / Everywhere

AI Safety & Alignment Researcher

Unexpected
Protection · Vulnerable SafeThe pull to shield from harm
Pace
  • A steady rhythm with room to breathe
  • A hard push you keep up for a long stretch
  • Patient work over a long time, where showing up matters most
What your week looks likeQuiet stretches, then deadline storms
How much you move around at workScreen and chair, almost all day
Whether you can work from anywhereWork from anywhere with a signal or internet connection
How quickly you receive feedback on your workTakes a season or a project cycle
What you're actually working withNumbers, measurements, records — things you read on a screen / Concepts, theories, designs, stories — things you think up

Core
  • Following evidence toward hidden truth.
  • Systematic, methodical pursuit of understanding.
  • Working through a problem to its resolution.
  • Verifying whether something is true or works as claimed.
Also present
  • Fighting for someone who can't fight for themselves right now.
  • Systematic examination for hidden problems.
  • Conceptual architecture. Figuring out how something should work before it exists.
  • Constructing explanations for why things work the way they do.

The AI safety and alignment researcher works on ensuring that AI systems do what humans intend and do not cause harm — a problem that becomes more urgent and more difficult as AI systems become more capable. The primary pull is Protection: the researcher exists because powerful AI systems are, in a precise sense, dangerous — not because they are malicious, but because a system that is capable and misaligned, or capable and deployed without adequate safeguards, can cause harm at a scale that is difficult to reverse. The work ranges from theoretical (formalising what "alignment" means, proving safety properties of training methods) to empirical (red-teaming models, building evaluation frameworks, testing for dangerous capabilities) to applied (designing guardrails, content filters, and monitoring systems for deployed AI).

The daily texture depends on the specific sub-area. A mechanistic interpretability researcher might spend the day probing a neural network's internal representations to understand how it processes information. A red-teamer might spend the day trying to make a language model produce dangerous outputs. A governance researcher might spend the day drafting policy recommendations for regulators. The common thread is a focus on what can go wrong and how to prevent it, which requires a temperament that is comfortable with adversarial thinking — imagining failure modes, not just success cases.

The field is young, growing rapidly, and genuinely important. The question of whether humanity can build AI systems that are both powerful and safe is one of the defining technical and ethical challenges of this century, and the people working on it are acutely aware that they may not have as much time as they would like.

🦊
There's a guide here if you want one

Kitsune can talk through anything on this page — whether it might suit you, what to do next, questions this page doesn't answer. Everything here is yours to read either way.

This field barely existed a decade ago, and much of the foundational theory is still being developed. A researcher entering AI safety today is not joining a mature discipline with established methods and clear career paths — they are helping to build the discipline itself. This is intellectually exciting but practically uncertain: the methods that seem promising today may be superseded, and the career infrastructure (clear promotion paths, established journals, well-defined roles) is still forming.

The relationship between safety researchers and capability researchers is complicated. Safety work depends on understanding the systems it is trying to make safe, which means safety researchers need access to frontier models and close collaboration with the teams building them. But the incentive structures are different — capability teams are rewarded for making systems more powerful, safety teams for making them more constrained — and navigating this tension requires diplomatic skill alongside technical ability.

The emotional dimension is real. Working daily on questions about existential risk, potential misuse of powerful AI, and the possibility that the technology you are studying could cause large-scale harm is psychologically demanding in a way that most technical roles are not. The researchers who sustain themselves are those who can hold the gravity of the problem without being paralysed by it.

The field draws from multiple backgrounds: computer science and ML (for technical safety research), mathematics and philosophy (for theoretical alignment work), and policy and governance (for the regulatory dimension). A PhD is common but not universal — the field is new enough that demonstrated ability and research output matter more than credential pedigree. Organisations like Google DeepMind, Anthropic, OpenAI, and the UK AI Safety Institute are primary employers. Academic centres include the Centre for AI Safety, the Future of Humanity Institute (Oxford), and the Centre for the Study of Existential Risk (Cambridge). Internships and research fellowships (such as those offered by MATS, Redwood Research, and the Long-Term Future Fund) provide entry points. The UK AI Safety Institute, established in 2023, has created new government-adjacent roles that combine technical expertise with policy work.

AI Safety & Alignment Researcher · AI & Machine Learning Engineering · PurPassion