Simulating Human Moral Judgment in LLMs

2025-07-04 · technology · rough
abstractConstructs a benchmark from human moral responses to evaluate how closely large language models align with real-world ethical intuitions.

Idea

Create a dataset of moral dilemmas (like trolley problems, real-world ethical cases, etc.) and survey how different people respond to them. Then use that dataset to evaluate whether existing LLMs (GPT-4, Claude, etc.) mimic human responses, diverge in systematic ways, or exhibit biases. Explore implications for alignment.