Kaiwen Zhou

AI Safety Researcher | Anthropic Safety Fellow (MATS) | PhD @ UC Santa Cruz | Published at ICLR · ICML · EMNLP · ACL | Post-training | AI agent.

Role
Ai Safety Research Fellow (Mats) at Anthropic
Location
Santa Cruz, CA, US
LinkedIn followers
500 followers

About Kaiwen Zhou

I make advanced AI systems safer. My research spans LLM agent red-teaming, safety alignment for reasoning models, and multimodal safety evaluation. My first-author papers have been published at ICLR, ICML, EMNLP, ACL, EACL, and ECCV.Currently Anthropic AI Safety Research Fellow via MATS on interpreting and monitoring misaligned reasoning in LLMs. Previously built an adversarial red-teaming framework at Microsoft that was deployed into Responsible AI product (EACL 2026).Core expertise: Post-training (SFT & RL) · Adversarial red-teaming · LLM/MLLM safety evaluation · AI agents · Embodied AIOpen to full-time Research Scientist / Research Engineer roles in AI safety and alignment. Let\'s connect: kevinz-01.github.io

Experience

  1. Ai Safety Research Fellow (Mats)

    Anthropic

    Jan 2026 — Present · 伯克利, CA, US

    Working with Anthropic\'s Alignment Science team to research interpretability and monitoring of misaligned reasoning processes in LLMs and agents.

Education

  • Zhejiang University

    Bachelor's degree, Statistics

    2021

  • University of California, Santa Cruz

    Doctor of Philosophy - PhD, Computer Science

    2026

Find verified contacts for anyone on LinkedIn

Unifers gives sales teams verified emails and direct dials, enriched profiles, and outreach that lands in the inbox.

Free plan included · No credit card required

This profile is compiled from publicly available professional sources. Unifers is not affiliated with or endorsed by LinkedIn. Request removal of this profile.

Kaiwen Zhou — Ai Safety Research Fellow (Mats) at Anthropic in Santa Cruz, CA, US | Unifers