LKValues: Aligning Large Language Models with Sri Lankan Societal Values

This paper discusses how large language models (LLMs) often reflect Western values, which can lead to misunderstandings in diverse cultures like Sri Lanka.

Analyze with PDFdigest

This video presentation explains the key concepts from the paper in plain language.

Content & Liability Disclaimer

This article and its accompanying video are automated summaries derived from the original research paper by Unknown authors. The original research was conducted solely by the paper's authors; PDFdigest did not conduct any of the research and makes no claims of ownership over the underlying scientific work.

The video narration is generated by artificial intelligence and references the paper's authors for attribution. The video is not narrated by any of the paper's authors. This content may contain inaccuracies, omissions, or misinterpretations of the original research. First-person language (e.g., "we found", "our results") reflects the original authors' voice, not PDFdigest's. Always read the original paper for accurate, verified information before making any decisions based on this content.

This content is provided "as is" without any warranties, express or implied. Simulated systems OÜ, its officers, directors, employees, and agents shall not be liable for any direct, indirect, incidental, special, consequential, or punitive damages arising from your use of, reliance on, or access to this content, including but not limited to errors, omissions, or misinterpretations of the original research. This disclaimer applies to the fullest extent permitted by applicable law.

Key Takeaways
  1. 1 Models are fine-tuned on this dataset to develop two key capabilities: 1)Value-explanatory generation to produce locally contextualized value-based justifications that connect situations to Sri Lankan norms and social realities.
  2. 2 To improve generalizability and reduce privacy or memorization risks, scenarios were rewritten to avoid unnecessary names, dates, and overly specific event details while remaining faithful to the original news context.
  3. 3 The model was given the news content and its matched value tags, and was instructed to produce one scenario per value match when a clear situation was present.
  4. 4 For each match independently, identify whether there is a clear, distinct situation, event, action, scenario, or behavior in the content that relates to it.

Introduction

English-trained models exhibit mismatches in multilingual, non-Western societies. This motivates country-specific alignment resources that operationalize local societal values.

Sri Lanka illustrates the need for localized value alignment due to its multiethnic and multilingual nature.

Sri Lanka lacks an official national value framework or specific societal value benchmarks.

Methodology

We treat Sri Lankan-specific and Universal prompts as diagnostic framings of the same task. Appendix E provides Macro-F1, prompt sensitivity, and error analysis.

Study Design

The contrast suggests a task-transfer mismatch between generation and judgment.

Appendix E.1 provides further analysis.

Important Note

We note that lawful access to copyrighted content does not itself authorize further acts of exploitation beyond reading/viewing, so our processing is limited to research analysis and we do not redistribute the original articles Xalabarder .

Results & Findings

Large language models mediate everyday human decisions. We introduce a survey-driven pipeline for constructing Sri Lankan value alignment resources.

  • Large language models mediate everyday human decisions.
  • We introduce a survey-driven pipeline for constructing Sri Lankan value alignment resources.
  • We operationalize Sri Lankan societal values through a trilingual survey combining international frameworks and LLM-assisted elicitation.
  • We curate LKvaluesIT and LKvaluesBench datasets using the 40 survey-finalized values.
  • We construct a bilingual benchmark to test value-sensitive judgment.
Important Note

To improve generalizability and reduce privacy or memorization risks, scenarios were rewritten to avoid unnecessary names, dates, and overly specific event details while remaining faithful to the original news context.

Important Note

Model outputs are normalized with a caseinsensitive regular expression that extracts the first valid label from A, B, BOTH, or 0; outputs that cannot be normalized are counted as invalid.

How PDFdigest Helps You Understand Research

Instant Paper Analysis

Get structured summaries and key findings from dense PDFs in seconds.

Visual Explanations

Turn complex methods, figures, and results into clearer visual breakdowns.

AI-Powered Q&A

Ask focused questions and get answers grounded in the paper.

Try PDFdigest Free

Practical Applications

At the same time, the heatmap highlights where between-group contrasts may be larger in magnitude, but interpretation must be tempered by uncertainty: subgroup sample sizes are highly imbalanced (e.g., Sinhalese n = 162 vs. Burgher n = 3), producing substantially larger MOE for smaller groups and limiting the strength of inference for rare categories. Make it culturally contextual, but avoid overly specific names or dates when possible.

The two statements were created using the original correct answer and distractors, then edited so that only one statement, both statements, or neither statement could be correct.

This pattern suggests that the difficulty is not only conceptual, but also linguistic: even when models can recognize a value in English, they may fail to make the same value-sensitive judgment in Sinhala.

Related Work

This section reviews existing literature on aligning LLMs with diverse cultural values, noting the underrepresentation of Global South and low-resource contexts. It discusses various survey-based frameworks and methodologies for value elicitation and alignment.

Sri Lankan Value Identification Survey

Details the methodology of a trilingual survey designed to identify Sri Lankan societal values. It describes the process of selecting and validating values through a combination of established frameworks and LLM-assisted elicitation.

Dataset Curation

Describes the pipeline used to create the LKvaluesIT instruction dataset and the LKvaluesBench evaluation benchmark, focusing on the construction of culturally relevant scenario-based examples for LLM training.

LKvaluesIT Instruction Dataset

Explains the creation of a bilingual instruction dataset that operationalizes the identified Sri Lankan values through scenario-based examples, providing contextually relevant supervision for LLMs.

Figures Explained

The paper’s visual material highlights the workflow and the main system components.

  • Figure 2 :: Figure 2: End-to-end pipeline for curating the Sri Lankan value-aligned instruction and benchmark datasets, illustrating our human-guided, LLM-in-the-loop methodology from survey-driven value identification to bilingual scenario extraction and validation.
  • Figure 2: ); Pistilli et al. (2024) ( Daily Mirror, News First) covering 2009-2023, including major events such as the LTTE war, COVID-19 pandemic, and Easter Sunday attacks. After preprocessing, it contains 73,068 valid entries, yielding 15.19 million tokens when tokenized with the NLLB model.
  • Figure 5 :: Figure5: Heatmap of endorsement rates (%) for Sri Lankan societal values across ethnic subgroups, with each cell annotated as p ± MOE (percentage-point margin of error). Rows list values and columns correspond to Sinhalese, Tamil, Muslim, Burgher, and Other respondents; darker shading corresponds to lower endorsement, and lighter/brighter yellow corresponds to higher endorsement. Subgroup sample sizes are highly imbalanced (e.g., Sinhalese n = 162 vs. Burgher n = 3), yielding substantially wider MOE for smaller groups and thus greater uncertainty in between-group comparisons.
  • Figure 6 :: Figure 6: Examples from LKvaluesIT (English split).Each instance contains an instruction (scenario), a target value label, and a short value-grounded explanation.
  • Figure 8 :: Figure 8: System prompt used for scenario extraction in LKvaluesIT.The prompt instructs the model to convert value-tagged Sri Lankan news items into neutral, generalizable scenario-based instruction instances with a situation, value label, value-grounded explanation, and fixed “Supports” valence.
PDFDIGEST AI

Struggling to understand complex research papers?

Upload any PDF and get instant AI-powered explanations, summaries, and visual breakdowns. Turn dense academic writing into clear, actionable insights.

Upload a Paper

Frequently Asked Questions

The retained values are applicable to Sri Lanka’s cultural context based on citizen endorsement. We do not claim that all retained values are unique to Sri Lanka.

The task evaluates application of values to concrete scenarios. These gains are consistent with LoRA’s ability to preserve pretrained multilingual representations while learning task-specific updates .

Models are fine-tuned on this dataset to develop two key capabilities: 1)Value-explanatory generation to produce locally contextualized value-based justifications that connect situations to Sri Lankan norms and social realities. To improve generalizability and reduce privacy or memorization risks, scenarios were rewritten to.

Gender-disaggregated endorsements (Figure 4 ) largely preserve the same value ordering observed in the overall distribution, with the most widely endorsed values remaining high for both male and female subgroups. This pattern suggests that the difficulty is not only conceptual, but also.

To improve generalizability and reduce privacy or memorization risks, scenarios were rewritten to avoid unnecessary names, dates, and overly specific event details while remaining faithful to the original news context. Model outputs are normalized with a caseinsensitive regular expression that extracts the.

This paper discusses how large language models (LLMs) often reflect Western values, which can lead to misunderstandings in diverse cultures like Sri Lanka.

Related Research

Research

OLEDLM: A UNIFIED LANGUAGE MODEL FOR OLED MOLECULAR DESIGN

This paper introduces a new method for designing OLED materials using advanced AI techniques. It focuses on generating chemical structures that meet…

10 min read
Research

Enabling Rapid Calibration of BCI Systems that Detect Movement-Related Cortical Potentials in Children with Cerebral Palsy

This study explores a new way to help children with cerebral palsy improve their movement using brain-computer interfaces. By using advanced technology,…

10 min read
Research

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

This paper presents a new approach called OmniReasoner that helps models understand long audio and video content better by allowing them to…

10 min read