
We're a group of Princeton students working to reduce catastrophic risk from advanced AI.
Our Mission
We believe that AI presents a magnitude of risks and benefits unmatched by any previous technology. To realize the benefits, we must address the risks. Let's navigate the transition to advanced AI wisely.
Princeton AI Alignment is a community working to make the future better.
We support Princeton students and researchers through fellowships, speaker events, and mentorship. We help members learn the field, contribute in meaningful careers, and engage seriously with the high stakes of advanced AI.
Recent papers our members have written
Are Large Language Models Sensitive to the Motives Behind Communication?
Wu et al. • 2025
Investigates whether LLMs can recognize and account for human communicative intentions when evaluating information. Finds that while LLMs can discount biased sources in controlled settings, they struggle with real-world sponsored content — and that prompting models to consider source incentives significantly improves alignment with rational decision-making. Published at NeurIPS 2025.
CCS-Lib: A Python package to elicit latent knowledge from LLMs
Laurito et al. • 2025
A Python package for implementing Contrast-Consistent Search (CCS) to extract truthful beliefs from language models, addressing the challenge of eliciting latent knowledge.
Prompt-Character Divergence: A Responsibility Compass for Human-AI Creative Collaboration
Maggie Wang, Wouter Haverals • 2025
A lightweight metric that quantifies semantic drift in AI-generated images, helping creators determine when outputs reflect their intent versus model-driven biases. Published at NeurIPS Creative AI Track 2025.
Dynamic Risk Assessment for Offensive Cybersecurity Agents
Wei et al. • 2025
A framework for dynamically assessing and managing risks in offensive cybersecurity agents, ensuring safe deployment of AI systems in security-critical contexts. Published at NeurIPS 2025 Datasets & Benchmarks Track.
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
Wu et al. • 2025
Demonstrates that LLMs can spontaneously develop novel social biases about artificial demographic groups even when no inherent differences exist, resulting in highly stratified task allocations. These biases stem from exploration-exploitation trade-offs and are exacerbated by newer, larger models. Published at NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models (MTI-LLM).
Demo: Statistically Significant Results on Biases and Errors of LLMs Do Not Guarantee Generalizable Results
Liu et al. • 2025
Develops an infrastructure to probe medical chatbots using automatically generated queries across patient demographics, histories, and disorders. Finds that LLM annotators exhibit low agreement scores, and only specific LLM pairs yield statistically significant differences. Recommends using multiple LLM evaluators and publishing inter-LLM agreement metrics. Published at NeurIPS 2025 Workshop on GenAI for Health Potential, Trust, and Policy Compliance.











