Abstract
An emerging critical challenge for AI deployment is Value alignment: ensuring that AI systems make decisions that reflect human values. Many texts provide information about values, and policy documents in particular offer a rich source of value guidance. Large language models (LLMs) are powerful tools for analysing such text at scale. However, values are abstract, contextual, and subjective. Given the need to audit AI decision-making for safety and trust, we must understand how AI systems interpret and apply values. Neural systems, like LLMs, are opaque and often unpredictable, preventing reliable auditing and behaviour guarantees. There exists a fundamental tension: we want the scalability of neural systems to process diverse policy text, but we need to incorporate interpretability and predictability to enable verification of the interpreted values, and guarantee outcomes when the values are used in reasoning.This dissertation addresses this tension through a neuro-symbolic approach. We develop CAVA Bodega, the first complete pipeline for translating policy text into value-aligned decisions with explanations. CAVA Bodega integrates two novel frameworks: CAVA Reasoning, an argumentation framework for symbolic value-based reasoning across dynamic contexts and diverse stakeholders that provides the needed interpretability and predictability; and CAVA Press, a neural system for extracting value concepts from policy text at scale. Our research is grounded in a comprehensive survey of value alignment literature which provides thematic structure to the field and identifies key sub-problems.
Our contributions include multiple papers, novel frameworks with working implementations, open-source software, public datasets, and a critical insight: interpretability of symbolic components does not automatically translate to system-level explainability in neuro-symbolic architectures. This dissertation demonstrates the viability of neuro-symbolic approaches for value alignment, provides practical tools for building value-aware systems, and opens multiple directions for future research in value-aligned AI.
| Date of Award | 22 Apr 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Marina De Vos (Supervisor), Janina Hoffmann (Supervisor) & Andreas Theodorou (Supervisor) |
Keywords
- alternative format
- value alignment
- neuro-symbolic AI
- large language models
- logical framework
Cite this
- Standard