Cover Image

We're a group of Princeton students working to reduce catastrophic risk from advanced AI.

Our Mission

We believe that AI presents a magnitude of risks and benefits unmatched by any previous technology. To realize the benefits, we must address the risks. Let's navigate the transition to advanced AI wisely.

Princeton AI Alignment is a community working to make the future better.

We support Princeton students and researchers through fellowships, speaker events, and mentorship. We help members learn the field, contribute in meaningful careers, and engage seriously with the high stakes of advanced AI.

Organizations our members have worked with

This is a list of some of the organizations our members have worked with. Not all organisations listed endorse or are affiliated with PAIA.

OpenAI
Anthropic
EleutherAI
Dedalus Labs
Schwarzman Scholars
Institute for Progress
Supervised Program for Alignment Research
Center for AI Safety
ML Alignment & Theory Scholars
Sentient Futures
Constellation
Kairos
ARENA
UChicago XLab

Recent papers our members have written

Are Large Language Models Sensitive to the Motives Behind Communication?

Wu et al.2025

Investigates whether LLMs can recognize and account for human communicative intentions when evaluating information. Finds that while LLMs can discount biased sources in controlled settings, they struggle with real-world sponsored content — and that prompting models to consider source incentives significantly improves alignment with rational decision-making. Published at NeurIPS 2025.

CCS-Lib: A Python package to elicit latent knowledge from LLMs

Laurito et al.2025

A Python package for implementing Contrast-Consistent Search (CCS) to extract truthful beliefs from language models, addressing the challenge of eliciting latent knowledge.

Prompt-Character Divergence: A Responsibility Compass for Human-AI Creative Collaboration

Maggie Wang, Wouter Haverals2025

A lightweight metric that quantifies semantic drift in AI-generated images, helping creators determine when outputs reflect their intent versus model-driven biases. Published at NeurIPS Creative AI Track 2025.

Dynamic Risk Assessment for Offensive Cybersecurity Agents

Wei et al.2025

A framework for dynamically assessing and managing risks in offensive cybersecurity agents, ensuring safe deployment of AI systems in security-critical contexts. Published at NeurIPS 2025 Datasets & Benchmarks Track.

Large Language Models Develop Novel Social Biases Through Adaptive Exploration

Wu et al.2025

Demonstrates that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist, resulting in highly stratified task allocations. These biases stem from exploration-exploitation trade-offs and are exacerbated by newer, larger models. Published at NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models (MTI-LLM).

Demo: Statistically Significant Results on Biases and Errors of LLMs Do Not Guarantee Generalizable Results

Liu et al.2025

Develops an infrastructure to probe medical chatbots using automatically generated queries across patient demographics, histories, and disorders. Finds that LLM annotators exhibit low agreement scores, and only specific LLM pairs yield statistically significant differences. Recommends using multiple LLM evaluators and publishing inter-LLM agreement metrics. Published at NeurIPS 2025 Workshop on GenAI for Health Potential, Trust, and Policy Compliance.

Applications open now until Sept. 11! Click here to apply.