Analyzing Occupational Distribution Representation in Japanese Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in large language models have enabled users to generate fluent and seemingly convincing text, but they have uneven performance in different languages, which is associated with undesirable societal biases toward marginalized populations. |
| Approach: | They develop three Japanese language prompts to probe LLMs’ understanding of Japanese names and their association between gender and occupations. |
| Outcome: | The proposed models can associate Japanese names with correct gendered occupations when using constrained decoding, but with sampling or greedy decoding they prefer a small set of stereotypically genderes. |
Similar Papers
Hire Me or Not? Examining Language Model’s Behavior with Occupation Attributes (2025.coling-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have been widely integrated into production pipelines due to their impressive performance across multiple tasks. |
| Approach: | They construct a dataset using a standard occupation classification knowledge base and tested it on three families of LLMs. |
| Outcome: | The proposed framework analyzes LLMs’ behavior with respect to gender stereotypes in the context of occupation decision making. |
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans . |
| Approach: | They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs. |
| Outcome: | The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs. |
On the Mutual Influence of Gender and Occupation in LLM Representations (2025.acl-long)
Copied to clipboard
| Challenge: | We examine LLM representations of gender for first names in various occupational contexts to study how occupations and the gender perception of first names influence each other mutually. |
| Approach: | They examine LLM representations of gender for first names in various occupational contexts and examine how occupations and the gender perception of first names influence each other mutually. |
| Outcome: | The representations shift with the occupational context and are influenced by stereotypically feminine or masculine occupations. |
Job Unfair: An Investigation of Gender and Occupational Bias in Free-Form Text Completions by LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | a recent study has identified that LLMs are used in domains where they support or replace human decision-making . a systematic review of LLM outputs shows that many facets of social bias remain unaccounted for . |
| Approach: | They propose to disentangle gender and occupational biases in Italian and English as expressed by LLMs. |
| Outcome: | The proposed method captures gender and occupational biases in Italian and English . it also shows that models struggle with gender-neutral expressions, especially beyond English - the authors conclude . |
Measuring Normative and Descriptive Biases in Language Models Using Census Data (2023.eacl-main)
Copied to clipboard
| Challenge: | a new study examines how gender-based distributions of occupations are reflected in pre-trained language models. |
| Approach: | They propose a method to measure to what degree pre-trained language models are aligned to normative and descriptive occupational distributions. |
| Outcome: | The proposed method is language independent and can be extended to other dimensions of census data and demographic variables. |
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias (2025.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies have shown that large language models can cause harmful, human-like biases against various demographics. |
| Approach: | They propose a causal formulation for bias measurement in generative language models based on a list of desiderata for designing robust bias benchmarks and a bias-measuring procedure to investigate occupational gender bias. |
| Outcome: | The proposed framework is generalizable and can be extended to include other datasets. |
Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing research on stereotypes in large language models is limited and focuses on African Ameri- F. |
| Approach: | They propose to use global bias to probe a set of large language models via perplexity to determine how certain stereotypes are represented in the model's internal representations. |
| Outcome: | The proposed model amplifys harmful stereotypes and shows that the demographic groups associated with stereotypes remain consistent across model likelihoods and outputs. |
UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts (2024.eacl-srw)
Copied to clipboard
| Challenge: | Language models (LMs) often include societal biases encoded in the human-produced datasets used for their training. |
| Approach: | They evaluated six prominent language models: BERT, RoBERTa, DistilBERT, BERT- multilingual, XLM-RoBERT and DistilberT- multilinguistic. |
| Outcome: | The results show that the models generated by the models were stereotypically gendered and with a reduced bias in multilingual variants. |
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) focus on general domains, with fewer advancements in Japanese biomedical LLMs. |
| Approach: | They propose a benchmark for Japanese large language models with eight LLMs across four categories and 20 Japanese biomedical datasets for comparison. |
| Outcome: | The proposed benchmark includes eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks. |
The Linguistic Connectivities Within Large Language Models (2025.findings-acl)
Copied to clipboard
Dan Wang, Boxi Cao, Ning Bian, Xuanang Chen, Yaojie Lu, Hongyu Lin, Jia Zheng, Le Sun, Shanshan Jiang, Bin Dong, Xianpei Han
| Challenge: | Recent studies have discovered notable disparities in their performance across different languages. |
| Approach: | They conduct a systematic investigation into the behaviors of large language models across 27 different languages on 3 different scenarios and reveals a Linguistic Map correlates with the richness of available resources and linguistic family relations. |
| Outcome: | The proposed model demonstrates that there are significant disparities in performance across languages across 27 different languages on 3 different scenarios. |