~kris/9p

krisyotam.net

krisyotam.net/d/papers/simulating-moral-judgment-llms.html -rw-r--r-- 2.5 KiB
8bfa4b74 — Kris Yotam ci: declare SourceHut source for push builds a month ago
                                                                                
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
<!DOCTYPE html>
<html lang="en">
  <head>
    <meta charset="utf-8">
    <title>Simulating Human Moral Judgment in LLMs | Kris Yotam</title>
    <link rel="stylesheet" href="../../c/site.css">
    <meta name="description" content="Constructs a benchmark from human moral responses to evaluate how closely large language models align with real-world ethical intuitions.">
    <meta name="viewport" content="width=device-width, initial-scale=1">
  </head>
  <body>
    <div class="mode-container">
      <input type="radio" name="theme" id="theme-auto" checked>
      <label for="theme-auto" class="auto" title="Auto theme"></label>
      <input type="radio" name="theme" id="theme-light">
      <label for="theme-light" class="light" title="Light theme"></label>
      <input type="radio" name="theme" id="theme-dark">
      <label for="theme-dark" class="dark" title="Dark theme"></label>
    </div>
    <div class="top-bar" role="navigation">
      <a href="../../index.html" title="Main page"><img src="../../c/sigma.svg" alt="Main page"></a>
      <a href="../../about.html" title="About"><img src="../../c/news.svg" alt="About"></a>
      <a href="#" title="Code"><img src="../../c/code.svg" alt="Code"></a>
      <a href="../../pics.html" title="Photography"><img src="../../c/camera.svg" alt="Photography"></a>
      <a href="#" title="Pictures"><img src="../../c/picture.svg" alt="Pictures"></a>
      <a href="../../vids.html" title="Videos"><img src="../../c/video.svg" alt="Videos"></a>
      <a href="#" title="Donations"><img src="../../c/beer.svg" alt="Donations"></a>
      <a href="#" title="Social"><img src="../../c/elephant.svg" alt="Social"></a>
    </div>
    <h1 class="acad-title">Simulating Human Moral Judgment in LLMs</h1>
    <div class="acad-meta">2025-07-04 &middot; <a href="../../papers.html">technology</a> &middot; <span class="status-rough">rough</span></div>
    <fieldset class="wrap abstract"><legend>abstract</legend>Constructs a benchmark from human moral responses to evaluate how closely large language models align with real-world ethical intuitions.</fieldset>
    <div class="prose justify">
<h2>Idea</h2>
<p>Create a dataset of moral dilemmas (like trolley problems, real-world ethical cases, etc.) and survey how different people
respond to them. Then use that dataset to evaluate whether existing LLMs (GPT-4, Claude, etc.) mimic human responses, diverge
in systematic ways, or exhibit biases. Explore implications for alignment.</p>
    </div>
  </body>
</html>