Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)
Feriji: A French-Zarma Parallel Corpus, Glossary & Translator (2024.acl-srw)
Copied to clipboard
| Challenge: | MT has seen significant advances in recent years, but the representation of African languages in MT systems is underrepresented due to linguistic complexities and limited resources. |
| Approach: | They propose a first robust parallel French-Zarma corpus and a glossary for MT that contains 61,085 sentences in Zarma and 42,789 in French. |
| Outcome: | The proposed model improves the representation of the Zarma language, a dialect of Songhay, spoken by over 5 million people across Niger and neighboring countries. |
Pragmatic inference of scalar implicature by LLMs (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing Large Language Models (LLMs) engage in pragmatic inference of scalar implicature, such as some. |
| Approach: | They investigate how Large Language Models (LLMs) engage in pragmatic inference of scalar implicature, such as some. |
| Outcome: | The proposed models interpret some as pragmatic implicature not all in the absence of context, aligning with human language processing. |
Topic Modeling for Short Texts with Large Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can be used to solve topic modeling challenges for short texts by contextually learning the meanings of words. |
| Approach: | They propose two approaches to using Large Language Models (LLMs) for topic modeling: parallel prompting and sequential prompting. |
| Outcome: | The proposed methods identify more coherent topics than existing ones while maintaining the diversity of the induced topics. |
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods to translate spoken utterances from one language to another are unable to preserve speaker timbre of source speech. |
| Approach: | They propose a pipeline with style-transfer capability on the basis of self-supervised speech representations and codec units. |
| Outcome: | The proposed model achieves zero-shot cross-lingual style transfer on previously unseen source languages. |
BiasDPO: Mitigating Bias in Language Models through Direct Preference Optimization (2024.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been shown to be effective in complex language tasks, but their potential to perpetuate biases poses significant concerns. |
| Approach: | They propose a new framework employing Direct Preference Optimization to mitigate biases in LLMs. |
| Outcome: | The proposed model outperforms the baseline model on almost all bias benchmarks and achieves better performance than open-source models. |
Document Alignment based on Overlapping Fixed-Length Segments (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing studies show that web crawling can be used to obtain large-scale parallel corpora for NLP tasks. |
| Approach: | They propose a sentence-based segmentation method for document alignment . they compare it with a fixed-length segmentation technique to handle long-text encoding better. |
| Outcome: | The proposed method improves document alignment and recall by 1% to 10% on a document alignment task for Japanese-English and French-English datasets. |
ReMAG-KR: Retrieval and Medically Assisted Generation with Knowledge Reduction for Medical Question Answering (2024.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have significant potential for facilitating intelligent end-user applications in healthcare, but hallucinations remain an inherent problem with LLMs. |
| Approach: | They propose a pipeline to retrieve and medicalally-augmented-generation with knowledge reduction using cross-encoder re-ranking strategies to reduce the knowledge base. |
| Outcome: | The proposed pipeline reduces the knowledge base and improves inference time by 47%. |
Demystifying Instruction Mixing for Fine-tuning Large Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Instruction tuning is effective for aligning large language models with human instructions, but the procedure to optimizing the mixing of instruction datasets is still unclear. |
| Approach: | They categorize instructions into three primary types: NLP downstream tasks, coding, and general chat. |
| Outcome: | The proposed method improves performance of large language models (LLMs) but it is difficult to combine different instruction datasets to optimize overall performance. |
Fine-Tuning ASR models for Very Low-Resource Languages: A Study on Mvskoke (2024.acl-srw)
Copied to clipboard
| Challenge: | Recent advances in multilingual models for automatic speech recognition (ASR) have been able to achieve a high accuracy for languages with extremely limited resources. |
| Approach: | They examine the parameter efficiency of training an adapter for the Mvskoke language, an indigenous language of America. |
| Outcome: | The proposed model is parameter efficient and gives higher accuracy for a relatively small amount of data. |
Automating Qualitative Data Analysis with Large Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for qualitative data analysis are far from resembling a human's analysis outcome. |
| Approach: | They propose a method based on Large Language Models to tackle automated coding and make it as close as possible to the results of human researchers. |
| Outcome: | The proposed method is based on large language models and can be as close as possible to the results of human researchers. |
ANHALTEN: Cross-Lingual Transfer for German Token-Level Reference-Free Hallucination Detection (2024.acl-srw)
Copied to clipboard
| Challenge: | ANHALTEN is a new evaluation dataset that extends the English hallucination detection dataset to German. |
| Approach: | They propose a dataset that extends the English hallucination detection dataset to German . they show that larger context length leads to better halluciation detection in german . |
| Outcome: | ANHALTEN is the first evaluation dataset that extends the English hallucination detection dataset to German. |
Label-Aware Automatic Verbalizer for Few-Shot Text Classification in Mid-To-Low Resource Languages (2024.acl-srw)
Copied to clipboard
| Challenge: | Prompt-based learning has shown its effectiveness in few-shot text classification. |
| Approach: | They propose a prompt-based learning verbalizer that automatically selects a word to represent each class . they use the label name along with the conjunction "and" to induce the model to generate more effective words for the verbaliser. |
| Outcome: | The proposed method outperforms existing verbalizers on four Southeast Asian languages. |
Vector Spaces for Quantifying Disparity of Multiword Expressions in Annotated Text (2024.acl-srw)
Copied to clipboard
| Challenge: | We show that multiword expressions are a good study for linguistic diversity due to theiridiosyncratic nature. |
| Approach: | They train static MWE-aware word embeddings for verbal MWEs in 14 languages . they find that the disparity measure aggregatingthem at a global scale correlates with the number of types . |
| Outcome: | The proposed method is based on a set of vector spaces for VMWEs in 14 languages. |
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for multilingual framing differ from those used in English-speaking world . framers often use loaded vocabularies to create political images or favor a particular point of view . |
| Approach: | They use eight years of Russian-backed disinformation campaigns to examine framing . they find that disinformation campaign consistently favors specific framers . |
| Outcome: | The proposed method underperforms and shows high disagreements in Russian-language articles . the proposed method is based on eight years of Russian-backed disinformation campaigns . |
Assessing In-context Learning and Fine-tuning for Topic Classification of German Web Data (2024.acl-srw)
Copied to clipboard
| Challenge: | Using a few hundred annotated data points per topic, we detect content related to three German policies in a database of scraped webpages. |
| Approach: | They propose to use annotated data to train a binary classification task to detect topic-related content in a scraped database of webpages. |
| Outcome: | The proposed model detects content related to three German policies in a scraped database of scrapes of webpages using a few hundred annotated data points per topic. |
Knowledge Editing of Large Language Models Unconstrained by Word Order (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for identifying knowledge neurons for large language models have been challenging for black-boxed models . a new method is proposed to edit the knowledge held by the LLMs . |
| Approach: | They propose a method that identifies the knowledge neurons that encode the target knowledge and adjusts the parameters associated with these neurons to update the knowledge. |
| Outcome: | The proposed method outperforms existing methods on English and Japanese . it eliminates word order constraints and allows flexible locating regardless of the language . |
Exploring the Effectiveness and Consistency of Task Selection in Intermediate-Task Transfer Learning (2024.acl-srw)
Copied to clipboard
| Challenge: | Identifying beneficial tasks to transfer from is a critical step toward successful intermediate-task transfer learning. |
| Approach: | They propose a method that measures pairwise token similarity using maximum inner product search to improve task prediction. |
| Outcome: | The proposed method improves task prediction scores from 2.59% to 3.96% for tasks requiring reasoning abilities, but not for reasoning abilities. |
Does the structure of textual content have an impact on language models for automatic summarization? (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing models for automatic summarization of long sequences suffer from context limitation. |
| Approach: | They propose to take into account textual information coming from distinct passages from the long texts to be summarized. |
| Outcome: | The proposed model improves on the performance of LongFormer on English. |
Action Inference for Destination Prediction in Vision-and-Language Navigation (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing work on vision-and-language navigation focuses on spatial reasoning and semantic grounding of visual information, but there is still scope for improvement. |
| Approach: | They propose a VLN task of destination prediction for picking up a pedestrian that requires action inference from a crowd-sourced dataset. |
| Outcome: | The proposed model can reason about the effect of the next action and the next on the destination to a certain extent. |
A Computational Analysis and Exploration of Linguistic Borrowings in French Rap Lyrics (2024.acl-srw)
Copied to clipboard
| Challenge: | rap is a popular genre in the u.s. and has been used in countries far beyond the uk . linguistic borrowings are especially intriguing in countries such as the eu and europe . |
| Approach: | They manually annotate a lexicon of over 700 borrowings in the French language . they find that there are increases in the proportion of linguistic borrowings, interjections, and Niger-Congo borrowings . |
| Outcome: | The proposed method analyzes a corpus of over 8000 french rap song lyrics and shows that rap borrowings are increasing in prevalence and interjections are decreasing. |
On Improving Repository-Level Code QA for Large Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Commercial AI-assisted programming Chatbots may generate incorrect information when requests go beyond the model training data or require additional knowledge. |
| Approach: | They propose to implement different self-alignment processes and retrieval-augmented generation pipelines to improve the copilot performance. |
| Outcome: | The proposed model improves the copilot performance on repository-level semantics, dependency between files, and meta-information about the repository. |
Compromesso! Italian Many-Shot Jailbreaks undermine the safety of Large Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Despite efforts to align large language models with ethical guidelines, models can still be induced into unsafe behavior with jailbreaking. |
| Approach: | They investigate the impact of many-shot jailbreaking on LLMs in italian . they find models exhibit unsafe behaviors even with minimal exposure to harmful prompts . |
| Outcome: | The proposed model exhibits unsafe behaviors even with minimal exposure to harmful prompts, and this tendency rapidly escalates with more demonstrations. |
ViMedAQA: A Vietnamese Medical Abstractive Question-Answering Dataset and Findings of Large Language Model (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing abstractive question-answering datasets in Vietnamese are lacking . |
| Approach: | They propose to introduce a Vietnamese abstractive question-answering corpus to address this gap . they propose to use Vietnamese abstractives to generate answers to questions . |
| Outcome: | The proposed dataset examines the capability of large language models in the Vietnamese medical domain, including reasoning, memorizing and awareness of essential information. |
Rescue: Ranking LLM Responses with Partial Ordering to Improve Response Generation (2024.acl-srw)
Copied to clipboard
| Challenge: | Customizing LLMs for a specific task involves separating high-quality responses from lower-quality ones. Obtaining a large volume of expert-annotated data is costly for most tasks. |
| Approach: | They propose a method that trains the model to prioritize the best responses from a pool of candidates created for a task using ranking metrics. |
| Outcome: | The proposed method is more robust, less sensitive to noise, and can be achieved with limited human annotations or through heuristic methods. |
Basreh or Basra? Geoparsing Historical Locations in the Svoboda Diaries (2024.acl-srw)
Copied to clipboard
| Challenge: | In the historical domain, many geoparsing corpora are from large news collections. |
| Approach: | They propose a pipeline employing named entity recognition for geotagging and a map-based generate-and-rank approach incorporating candidate name augmentation and clustering of location context words for geocoding. |
| Outcome: | The proposed pipeline outperforms existing map-based geoparsers in terms of accuracy, lowest mean distance error, and number of locations correctly identified. |
Homophone2Vec: Embedding Space Analysis for Empirical Evaluation of Phonological and Semantic Similarity (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have shown that homophones with different semantic/syntactic contexts are easier for children to memorize. |
| Approach: | They propose a method for empirically evaluating the relationship between phonological and semantic similarity of linguistic units using embedding spaces. |
| Outcome: | The proposed method shows that Chinese character homophones have a positive semantic relationship at varying levels of sound-sharing. |
Trace-of-Thought Prompting: Investigating Prompt-Based Knowledge Distillation Through Question Decomposition (2024.acl-srw)
Copied to clipboard
| Challenge: | Trace-of-Thought Prompting allows small neural networks to emulate larger, teacher models with reduced computational demands. |
| Approach: | They propose a framework to distill critical reasoning capabilities from teacher models to student models . they use problem decomposition to enhance interpretability and facilitate human-in-the-loop interventions . |
| Outcome: | a new framework enables small neural networks to emulate the performance of larger, teacher models . it leverages problem decomposition to enhance interpretability and facilitate human-in-the-loop interventions . the proposed framework is available on github.com/trace-of-thought/trac-of_thought-prompting/main . |
Can LLMs Augment Low-Resource Reading Comprehension Datasets? Opportunities and Challenges (2024.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have demonstrated impressive zero-shot performance on a wide range of NLP tasks. |
| Approach: | They propose to use large language models to augment extractive reading comprehension datasets by fine-tuning their annotations and comparing their performance to human annotators. |
| Outcome: | The proposed model can be used to augment extractive reading comprehension datasets. |
Automatic Derivation of Semantic Representations for Thai Serial Verb Constructions: A Grammar-Based Approach (2024.acl-srw)
Copied to clipboard
| Challenge: | Using rich semantic representations for Thai Serial Verb Constructions (SVCs) is time-consuming and manual annotation is preferred. |
| Approach: | They propose to implement an HPSG analysis for Thai Serial Verb Constructions (SVCs) they use a DELPH-IN computational grammar to generate appropriate representations from syntactic features. |
| Outcome: | The proposed grammar increases verified coverage of Thai SVCs by 73% and decreases ambiguity by 46% on held-out data. |
Bridging Distribution Gap via Semantic Rewriting with LLMs to Enhance OOD Robustness (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for fine-tuning on indistribution data fail to provide robustness against distribution shifts limiting the practical deployment of LLMs in dynamic real-world scenarios. |
| Approach: | They propose a method that leverages the flexibility of LLMs to align both in-distribution (ID) and OOD data with the LLM's distributions. |
| Outcome: | The proposed method outperforms fine-tuning methods on OOD tasks and benchmark datasets. |
CoVoSwitch: Machine Translation of Synthetic Code-Switched Text Based on Intonation Units (2024.acl-srw)
Copied to clipboard
| Challenge: | Multilingual code-switching research is often hindered by the lack and linguistically biased status of available datasets. |
| Approach: | They synthesize code-switching data by replacing intonation units detected through PSST, a speech segmentation model fine-tuned from OpenAI’s Whisper, using a language-to-text translation dataset, CoVoST 2. |
| Outcome: | The proposed model outperforms two monolingual models and is better at code-switching translation into English than non-English. |
Beyond Abstracts: A New Dataset, Prompt Design Strategy and Method for Biomedical Synthesis Generation (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods to automate systematic reviews of papers are slow and incomplete . authors propose a new method to automating the systematic review process . |
| Approach: | They propose a method for automatic synthesis generation using a dataset and prompting-based method. |
| Outcome: | The proposed method improves the existing model and prompts the system to generate high-quality syntheses. |
Improving Sentence Embeddings with Automatic Generation of Training Data Using Few-shot Examples (2024.acl-srw)
Copied to clipboard
| Challenge: | Decoder-based large language models (LLMs) have shown high performance on many tasks in natural language processing. |
| Approach: | They propose to automatically generate an NLI dataset with an LLM and use it for fine-tuning of PromptEOL. |
| Outcome: | The proposed model outperforms existing models on STS tasks without large manually annotated datasets. |
Curriculum Learning for Small Code Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have shown that curriculum learning improves performance of code language models when it comes to complex tasks. |
| Approach: | They propose a novel code difficulty assessment metric and introduce a Novel Curriculum Learning schedule that enhances the performance of small decoder-only language models. |
| Outcome: | The proposed model improves on code execution tasks while its effect on code completion is less significant. |
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods to improve LLM performance have focused on sophisticating the model's step-by-step calculation. |
| Approach: | They propose a question analysis prompting strategy in which the model is prompted to explain the question in 'n' words before solving. |
| Outcome: | The proposed prompt outperforms state-of-the-art prompts on arithmetic and commonsense datasets and consistently ranks among the top-2 prompts. |
CheckersGPT: Learning World Models through Language Modeling (2024.acl-srw)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown impressive performance on various tasks, but the underlying process behind predicting the desired next token remains a black box. |
| Approach: | They train a GPT-style autoregressive language model using only the next character prediction objective and then train corresponding model with different layer sizes. |
| Outcome: | The proposed model shows a hint of learning a world model representation of the board positions on a simulated game of checkers and human gameplay dataset. |
In-Context Symbolic Regression: Leveraging Large Language Models for Function Discovery (2024.acl-srw)
Copied to clipboard
| Challenge: | State of the art Symbolic Regression (SR) methods build specialized models, while the application of Large Language Models (LLMs) remains largely unexplored. |
| Approach: | They propose a framework which iteratively refines a functional form with an LLM and determines its coefficients with an external optimizer. |
| Outcome: | The proposed method outperforms the best SR methods on four popular benchmarks while yielding simpler equations with better out of distribution generalization. |
STEP: Staged Parameter-Efficient Pre-training for Large Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for reducing computational costs during pre-training have been studied, but they often degrade performance under fair conditions. |
| Approach: | They propose a method that combines parameter-efficient tuning and staged training to reduce memory requirements while maintaining comparable performance. |
| Outcome: | The proposed method reduces memory requirements by 40.4% while maintaining comparable performance. |