Papers by Taro Watanabe
Copied to clipboard
| Challenge: | a higher-level law authorizes a lower-level to implement detailed provisions, which is called delegation. |
| Approach: | They propose a two-stage pipeline system for automatic delegation annotation in Japanese law . they extract keywords that indicate delegation using a named entity recognition approach . |
| Outcome: | The proposed system shows sufficient performance to assist manual annotation in practice. |
Copied to clipboard
| Challenge: | Existing studies show that MBR decoding improves model generation performance . however, the theoretical underpinnings of these results remain uncertain . |
| Approach: | They propose a theoretical interpretation of MBR decoding from the perspective of bias–diversity decomposition. |
| Outcome: | The proposed method improves the quality estimation of hypotheses by decomposing bias and diversity into two main factors. |
Copied to clipboard
| Challenge: | Existing knowledge probes for pre-trained language models exhibit quadratic time complexity, limiting the size of knowledge graphs used for probing. |
| Approach: | They propose an embedding-based relational probe that evaluates pre-trained language models' factual knowledge retrieval capabilities. |
| Outcome: | The proposed probe achieves effective time complexity of linear order O(n), supports rank-based evaluation metrics including Hit@k, handles multi-token entity names and enables probing whilst disambiguating homographic tail-entity names. |
Copied to clipboard
| Challenge: | Existing studies on multilingual fine-tuning with a fixed set of languages lack dynamic adaptability to new languages. |
| Approach: | They propose a modular fine-tuning pipeline that enables dynamic language adaptation for LLMs by first training English-centric adapters for each language separately and then merging them for arbitrary-direction translation. |
| Outcome: | The proposed pipeline achieves 86% performance over traditional fine-tuning on four languages, while training only 0.1% parameters and relying on English as a bridge language without catastrophic forgetting. |
Copied to clipboard
| Challenge: | Existing approaches to training or evaluating non-English dialogue datasets often introduce artifacts that reduce their naturalness and cultural appropriateness. |
| Approach: | They propose a structured framework for encoding, localizing, and generating multilingual dialogues from abstract intent representations. |
| Outcome: | The proposed framework outperforms translation models in Italian, German, and Chinese on cultural relevance, coherence, and situational appropriateness. |
Copied to clipboard
| Challenge: | Entity-based QA is a common framework for analyzing non-verbatim memorization, but typically query each entity using a single canonical surface form. |
| Approach: | They propose a dataset that pairs Wikidata factual triples with categorized entity surface forms . they examine surface-conditioned factual memorization and find that prediction outcomes change when only the entity surface form is changed. |
| Outcome: | The proposed dataset shows that large language models memorize factual knowledge when only the subject entity surface form is changed. |
Copied to clipboard
| Challenge: | Large language models inherit and amplify societal biases related to gender and race. |
| Approach: | They use a USChainMains dataset to evaluate group bias in Large Language Models . they found that LLMs recommend meals with higher levels of adverse nutrients for names associated with Black, Hispanic, or male individuals . |
| Outcome: | The proposed model scales improves overall recommendation healthfulness but is insufficient to eliminate the healthfulness gap between demographic groups. |
Copied to clipboard
| Challenge: | Minimum Bayes risk (MBR) decoding requires quadratic time since it computes the expected score between a translation hypothesis and all reference translations. |
| Approach: | They propose a centroid-based MBR decoding method that clusters the translations in the feature space and calculates the expected score using the centroids of each cluster. |
| Outcome: | The proposed method outperforms vanilla MBR decoding in translation quality by up to 0.5 COMET in the WMT’22 EnJa, EnDe, EnZh, and WMT'23 Enja translation tasks. |
Copied to clipboard
| Challenge: | a study shows that language models can explain vowel pronunciation based on tongue positions . a visual LM can explain the relationship between vowels and tongue positions, but it is unclear whether they align textual information with visual information. |
| Approach: | They created video and image datasets from MRI data to examine if LMs associate real tongue positions with vowel articulation. |
| Outcome: | The proposed model can explain vowel pronunciation and the correlation between vowels and tongue positions as textual knowledge. |
Copied to clipboard
| Challenge: | et al., 2006) considers geographic relatedness among geo-entity mentions in document-level geoparsing. |
| Approach: | They present a Japanese travelogue dataset that considers geographic relatedness among geo-entity mentions. |
| Outcome: | The proposed dataset includes 200 travelogue documents with rich geo-entity information . it shows that human activities, mobility, and events are often described with natural language expressions of locations or geographic entities (geo-entities) |
Copied to clipboard
| Challenge: | Decoding strategies affect the probability distribution underlying the output of a language model and can therefore affect both generation quality and uncertainty. |
| Approach: | They investigate the impact of decoding strategies on uncertainty estimation in large language models . |
| Outcome: | The proposed methods improve the uncertainty estimation of large language models by reducing repetition. |
Copied to clipboard
| Challenge: | Named Entity Recognition (NER) is a key task in NLP to find mentions of named entities and classify them into predefined categories. |
| Approach: | They investigated the impact of data augmentation on confidence calibration and uncertainty estimation in Named Entity Recognition (NER) tasks. |
| Outcome: | The data augmentation improves calibration and uncertainty in cross-genre and cross-lingual setting, especially in-domain setting. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation requires an enormous dataset, leaving the low-resource language (LRL) underdeveloped. |
| Approach: | They evaluated five languages using a parallel corpus of 1,000 instances each and found a zero-shot improvement of 7.4 from the baseline score of 7.1 to a score of 15.5 at best. |
| Outcome: | The proposed model improves performance in the linguistically diverse country of Indonesia by 7.4 from baseline score of 7.1 to 15.5 at best. |
Copied to clipboard
| Challenge: | Conventional approaches compare sentence probabilities directly, but large language models (LLMs) provide nuanced evaluation methods using prompts and templates. |
| Approach: | They propose to derive acceptability judgments from large language models using prompts and templates to comprehensively evaluate their grammatical knowledge. |
| Outcome: | The proposed methods excel in different linguistic phenomena, suggesting they access different aspects of the LLMs’ grammatical knowledge. |
Copied to clipboard
| Challenge: | Existing automatic evaluation metrics are based on procedures that diverge from human evaluation. |
| Approach: | They propose to aggregate automatic evaluation metrics to bridge this gap . they propose to use edit-based metrics, -gram based metrics and sentence-level metrics to find the best ranking system. |
| Outcome: | The proposed method outperforms existing metrics on the SEEDA benchmark and improves edit-based metrics, -gram based metrics and sentence-level metrics. |
Copied to clipboard
| Challenge: | Deep learning has demonstrated performance advantages in a wide range of natural language processing tasks. |
| Approach: | They propose to deepen the decoder layer in a Transformer model to reduce the difficulty of deep learning. |
| Outcome: | The proposed method can deepen the model on both the encoder and decoder at the same time, resulting in a deeper model and improved performance. |
Copied to clipboard
| Challenge: | citations that do not correspond to any existing work are a serious concern to scientific reliability and credibility. |
| Approach: | They analyze papers published at ACL, NAACL, and EMNLP in 2024 and 2025 . they identify 300 papers with at least one HalluCitation, most of which were published in 2025. |
| Outcome: | The authors analyze papers published at ACL, NAACL, and EMNLP in 2024 and 2025 . they find that nearly 300 papers contain at least one HalluCitation, most of which were published in 2025. |
Copied to clipboard
| Challenge: | Using large language models (LLMs) to generate human-like text has raised concerns about misuse, especially in low-resource languages like Urdu. |
| Approach: | They propose a dataset that contains documents, paragraphs, and sentences . they conducted human evaluations and automated evaluations . |
| Outcome: | The proposed dataset shows that distinguishing between human and machine-generated text is challenging for both humans and LLMs. |
Copied to clipboard
| Challenge: | Existing pre-trained language models outperform them in certain domains, indicating that there is significant potential for further improvement in this area. |
| Approach: | They propose to use pre-trained language models to evaluate ad texts from multiple perspectives within real-world advertising operations to define five tasks and construct a Japanese dataset. |
| Outcome: | The proposed benchmark outperforms existing pre-trained language models in several tasks, but humans outperformed them in certain domains. |
Copied to clipboard
| Challenge: | Existing methods for uncertainty estimation are inadequate for safety-critical applications. |
| Approach: | They propose a method that uses the distances from neighbors and the ratio of labels in neighbors to estimate uncertainty. |
| Outcome: | The proposed method outperforms baseline and density-based methods in calibration and uncertainty metrics. |
Copied to clipboard
| Challenge: | Pro-drop languages allow omissions of essential phrases or arguments . the presence of zero-pronouns affects downstream tasks of NLP . |
| Approach: | They propose a query-based method to identify zero-pronoun arguments . they use Japanese and Chinese datasets to evaluate the method . |
| Outcome: | The proposed method surpasses the sequence labeling baseline on Japanese and Chinese datasets. |
Copied to clipboard
| Challenge: | Existing studies attributed zero-shot translation to domination of central language, e.g. English, but we supplement this viewpoint with the strict dependence of non-centered languages. |
| Approach: | They propose a language-specific modeling method that adapts to non-centered languages to counteract the instability of zero-shot translation. |
| Outcome: | The proposed method performs better than baselines in centered data conditions and can easily fit non-centered data. |
Copied to clipboard
| Challenge: | Existing studies on lyrics translation have relied on fine-tuning open-source language models. |
| Approach: | They examine a multilingual lyrics translation dataset and apply prompting methods to large language models to evaluate singability. |
| Outcome: | The proposed methods improve singability and naturalness, compared to naive translation, the authors show . human evaluations using songs created from translated lyrics show that complex prompting strategies improve singable naturalness . |
Copied to clipboard
| Challenge: | Prior studies have shown that kNN-LM can retrieve long-tail contexts, leaving the model’s performance underexplored in estimating the probabilities of long-tailed target tokens. |
| Approach: | They investigate the behavior of kNN-LM on low-frequency tokens, examining prediction probability, retrieval accuracy, and token distribution in the datastore. |
| Outcome: | The proposed model improves the perplexity of given text by directly accessing a large datastore built from any text data during inference. |
Copied to clipboard
| Challenge: | averaging metric scores across languages is suspicious since translations of equal quality receive different scores across language. |
| Approach: | They propose a semi-automatically built dataset to benchmark translation metrics using MQM-defined errors and a normalization strategy to mitigate cross-lingual scoring bias. |
| Outcome: | The proposed model shows that translation metrics suffer from cross-lingual scoring bias . the proposed model is based on a semi-automatically built dataset covering nine translation directions . |
Copied to clipboard
| Challenge: | Large-scale Vision-Language Models (LVLMs) are being deployed in real-world settings that require visual inference. |
| Approach: | They evaluate LVLMs' ability to account for variation in color perception using the Ishihara Test. |
| Outcome: | The proposed models fail to reproduce the perceptual outcomes experienced by affected individuals and default to normative color perception. |
Copied to clipboard
| Challenge: | Existing studies evaluate sentences in isolation and do not consider how context influences LLM acceptability judgments. |
| Approach: | They examine how contextual cues affect model-generated acceptability ratings across multiple domains and several LLMs, using different forms of domain-specific contextual cueeds to situate sentences in intended usage settings. |
| Outcome: | The findings support the development of more context-aware evaluation frameworks. |
Copied to clipboard
| Challenge: | Existing approaches to cross-lingual vocabulary transfer face challenges when dealing with low-resource languages. |
| Approach: | They propose a dictionary-based crosslingual vocabulary transfer method that leverages bilingual dictionaries, which are available for many languages thanks to descriptive linguists. |
| Outcome: | The proposed method outperforms existing methods for low-resource languages. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation (MNMT) aims for arbitrary translations across multiple languages. |
| Approach: | They propose a method that inserts a set of tokens specifying the target language into the input sequence between the source and target tokens. |
| Outcome: | The proposed method outperforms existing models on a large-scale benchmark. |
Copied to clipboard
| Challenge: | Existing studies have focused on coarse-grained locations, but we focus on fine-grain POIs, which have many candidates with similar names. |
| Approach: | They develop a text embedding-based geocoding model and investigate (1) entry encoding representations and (2) hard negative mining approaches suitable for enhancing the model’s disambiguation ability. |
| Outcome: | The proposed model significantly improves its disambiguation ability and entry encoding representations. |
Copied to clipboard
| Challenge: | Existing benchmarks for linguistic knowledge of Indigenous languages of the Americas focus on high- and medium-resource languages with substantial digital presence. |
| Approach: | They propose a framework for probing large language models’ linguistic knowledge of Indigenous languages of the Americas using zero-shot prompting and few-shot probing. |
| Outcome: | The proposed framework evaluates models from five major families on 13 Indigenous languages including Bribri, Guarani, and Nahuatl. |
Copied to clipboard
| Challenge: | Simultaneous speech translation (SiST) begins translating before the entire source input is received. |
| Approach: | They propose a dataset that rearranges sentences into segmented monotonic data for simultaneous speech translation using the Large Language Model. |
| Outcome: | The proposed dataset improves quality and latency in siST translations by rearranging sentences into segmented monotonic data. |
Copied to clipboard
| Challenge: | Existing evaluation metrics for Grammatical error correction lack explainability . lack of explainability hinders researchers from analyzing strengths and weaknesses of models . |
| Approach: | They propose to assign sentence-level scores to individual edits to improve GEC performance . they use Shapley values, from cooperative game theory, to compute contribution of each edit . |
| Outcome: | The proposed method shows that the evaluation metrics are consistent across edits and human evaluations. |
Copied to clipboard
| Challenge: | k-nearest-neighbor machine translation (kNN-MT) is a new approach to improve NMT performance without additional training. |
| Approach: | They propose a method that integrates example-search into the decoding algorithm to improve neighbor token retrieval. |
| Outcome: | The proposed method achieves a speed-up of up to 132.2 times and an improvement in BLEU score of up 1.6 compared with kNN-MT in the WMT’19 translation task and the domain adaptation tasks in De-En and En-Ja. |
Copied to clipboard
| Challenge: | LVLMs are increasingly capable of responding in multiple languages . however, there is a lack of evaluation tools for LVLs that handle multiple languages. |
| Approach: | They used an extended dataset in multiple languages to evaluate LVLMs' ability to generate explanations in multiple language combinations. |
| Outcome: | The proposed dataset in multiple languages evaluates LVLMs' ability to generate explanations in other languages. |
Copied to clipboard
| Challenge: | Vision Language Models struggle with cultural-specific knowledge, especially in languages other than English and in underrepresented cultural contexts. |
| Approach: | They propose a visual question answering (VQA) dataset with text-image pairs across 30 languages and dialects and a training dataset. |
| Outcome: | The proposed model performs better with correct location context, but struggles with adversarial contexts and predicting specific regional cuisines and languages. |
Copied to clipboard
| Challenge: | Existing siMT corpora are limited due to high costs and limited annotator capabilities. |
| Approach: | They propose a method to convert ST corpora into interpretation-style corpors by fine-tuning models with Large Language Models. |
| Outcome: | The proposed method reduces latency while achieving better quality compared to other models. |
Copied to clipboard
| Challenge: | Existing research has analyzed regularization in language acquisition only by modeling word inflection directly, which is unnatural in light of human language acquisition. |
| Approach: | They hypothesize that language models that imitate errors children make during language acquisition have a learning process more similar to humans. |
| Outcome: | The proposed model shows child-like U-shaped learning curves clearly for certain verbs, but the preferences for types of overgeneralization did not fully match the observations in children. |
Copied to clipboard
| Challenge: | Large language models with instruction-following capabilities have revolutionized the field of artificial intelligence. |
| Approach: | They propose an annotation-free framework for empowering large language models with instruction-following capabilities. |
| Outcome: | The proposed framework generates multi-turn multimodal instruction-response conversations from a language model. |
Copied to clipboard
| Challenge: | Multilingual neural machine translation models support fine-tuning hundreds of languages simultaneously. |
| Approach: | They propose to fine-tune a language in its intrinsic subspace with a tiny fraction of entire parameters. |
| Outcome: | The proposed methods outperform full-parameter fine-tuning up to 2.25 spBLEU scores and reduce trainable parameters to 0.4% for high and medium-resource languages and 1.6% for low-resourced ones. |
Copied to clipboard
| Challenge: | Unsupervised image captioning is a challenging task that requires manual annotation. |
| Approach: | They propose a simple gating mechanism that is trained to align image features with the most reliable words in pseudo-captions. |
| Outcome: | The proposed method outperforms the previous methods without complex learning objectives. |
Copied to clipboard
| Challenge: | generative large language models (LLMs) compose sentences that include all given concepts but must generate sentences that adhere to the specified order. |
| Approach: | They propose a benchmark to evaluate compositional generalization and instruction-following abilities of generative large language models (LLMs) based on ordered coverage, which allows simultaneous evaluation of both abilities. |
| Outcome: | The proposed benchmark evaluates compositional generalization and instruction-following abilities of LLMs. |
Copied to clipboard
| Challenge: | Recent studies show that encoding more syntactic information does not lead to better performance. |
| Approach: | They propose a method to optimize pareto-optimal models by formalizing it as a multi-objective optimization problem. |
| Outcome: | The proposed method is better than a baseline method on two NLP tasks. |
Copied to clipboard
| Challenge: | Large-scale vision language models excel at generating factual content, but their ability to rank images from multiple perspectives has not been explored. |
| Approach: | They propose a framework to evaluate large-scale vision-language models by measuring their ability to rank image texts from multiple perspectives. |
| Outcome: | The proposed evaluation framework measures how closely LVLMs' judgments align with human interpretations. |
Copied to clipboard
| Challenge: | low-resource language research often hampered due to under-representation of how it is being used in reality. |
| Approach: | They propose to use a dataset comprising both Indonesian and English from personal travelogue articles . they used named and nominal expressions of four entity types related to travel . |
| Outcome: | The proposed dataset is more representative of how Indonesian language is being used in reality. |
Copied to clipboard
| Challenge: | Advancements in dialogue systems powered by large language models have outpaced the development of reliable evaluation metrics. |
| Approach: | They propose a benchmark to evaluate the robustness of reference-free dialogue metrics against four categories of adversarial attacks. |
| Outcome: | The proposed benchmarks show that the two axes of reliability are not always aligned . the findings motivate the development of nuanced evaluation frameworks to address real-world dialogue challenges. |
Copied to clipboard
| Challenge: | Knowledge Graph Completion (KGC) is a task that infers unseen relationships between entities . traditional embedding-based methods infer missing links using only training data . a pre-trained language model (PLM)-based KGC may be ineffective in practical applications . |
| Approach: | They propose to use knowledge Graph Completion (KGC) to infer unseen relationships . traditional embedding-based KGC methods infer missing links only from training data . they argue that pre-trained language models acquire inference abilities through pre-training . |
| Outcome: | The proposed method improves performance even though it does not use memorized knowledge. |
Copied to clipboard
| Challenge: | Existing reference-free automatic grammatical error correction methods do not correlate with human evaluation. |
| Approach: | They propose a reference-free automatic grammatical error correction evaluation method with enhanced gramma-ed capabilities. |
| Outcome: | The proposed method achieves highest correlation with human evaluations on a meta-evaluation dataset. |
Copied to clipboard
| Challenge: | a large part of human communication relies on nonverbal cues such as facial expressions, eye contact, and body language. |
| Approach: | They propose to validate whether video large language models can correctly interpret body language from short clips of body language. |
| Outcome: | The proposed model can correctly interpret emotions from short clips of body language. |
Copied to clipboard
| Challenge: | Experimental results show TableMBR outperforms the baseline, achieving relative improvements of up to 15% in F1 on Rotowire and 23% in accuracy on LiveSum. |
| Approach: | They propose a text-to-table task that generates structured data from unstructured text . they propose 'tableMBR' that maintains structural consistency through minimum Bayes risk decoding . |
| Outcome: | The proposed method outperforms the baseline and achieves relative improvements in F1 and LiveSum. |
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) is a sentence-level meaning representation based on predicate argument structure. |
| Approach: | They propose to use a dictionary to capture the structure of complex sentences . they train models on data derived from AMR and Wikipedia corpus . |
| Outcome: | The proposed model will be made public and the proposed patterns will be validated. |
Copied to clipboard
| Challenge: | Large-scale Vision-Language Models (LVLMs) output text from images and instructions, demonstrating advanced capabilities in text generation and comprehension. |
| Approach: | They propose to use artwork explanation generation task to quantitatively assess the understanding and utilization of artworks knowledge. |
| Outcome: | The proposed task evaluates the understanding and utilization of knowledge about artworks from images and titles and generates explanations using only images. |
Copied to clipboard
| Challenge: | a recent study examined the cross-lingual transferability of neural language models . previous studies focused on their first language acquisition . |
| Approach: | They propose to pretrain bilingual LMs with a scenario similar to human L2 acquisition . they find that pretraining accelerated their linguistic generalization in L2 . |
| Outcome: | The results show that pretraining bilingual LMs accelerates their linguistic generalizations . the results clarify their (non-)human-like L2 acquisition in particular aspects . |
Copied to clipboard
| Challenge: | Existing research on sentence-level paraphrase detection in Pashto has focused on English, but no work has been done on low-resource Pashtone. |
| Approach: | They propose to annotate sentences in Pashto to detect paraphrases . they will publicize a subset of 1,800 instances from their corpus, free from licensing issues. |
| Outcome: | The proposed corpus contains 6,727 sentences, encompassing 3,687 paraphrased and 3,040 non-paraphrased sentences. |
Copied to clipboard
| Challenge: | Existing studies treat travelogues as sequences of visited locations, but they lack a benchmark dataset. |
| Approach: | They propose to represent the trajectory as a graph that can capture the hierarchy as well as the visiting order and construct a benchmark dataset for the extraction. |
| Outcome: | The proposed dataset shows that even naive baseline systems can predict visited locations and the visiting order between them, while it is more challenging to predict the hierarchical relations. |
Copied to clipboard
| Challenge: | Phrase-level dense retrieval has shown many appealing characteristics in downstream NLP tasks. |
| Approach: | They propose a task formulation of dense retrieval, cross-lingual contextualized phrase retrieval . they extract pairs of cross-linguistic phrases using word alignment information . |
| Outcome: | The proposed task formulation surpasses baselines on the phrase retrieval task and a downstream task, i.e., machine translation, and achieves top-1 accuracy 13 points higher. |
Copied to clipboard
| Challenge: | Large language models (LLMs) remain unstable on long-context ranking. |
| Approach: | They propose a method that fuses explicit within-list positions with implicit cross-list preferences to score entities and return a top-k set. |
| Outcome: | Experimental results show that large language models remain unstable on long-context ranking . |
Copied to clipboard
| Challenge: | Existing knowledge about entities acquired from natural language models is not retained in pre-trained vision & language models. |
| Approach: | They propose a task to verify how knowledge about entities acquired from natural language is retained in Vision & Language (V&L) models. |
| Outcome: | The proposed model forgets part of its entity knowledge by pre-training to improve image related tasks. |
Copied to clipboard
| Challenge: | Word frequency is a key variable in psycholinguistics, useful for modeling human familiarity with words . a recent study shows that frequency from YouTube subtitles is comparable to and often better than the best available resources. |
| Approach: | They use YouTube subtitles to construct frequency norms for five languages . they find they are comparable to and often better than the best currently available resources . |
| Outcome: | The proposed method improves on the best currently available resources for Chinese, English, Indonesian, Japanese, and Spanish. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have been evaluated mostly on global or anglocentric subjects, often neglecting low-resource languages and culturally specific content. |
| Approach: | They evaluate 26 Large Language Models using a multiple-choice question answering benchmark for Sinhala. |
| Outcome: | The new benchmarks show that Claude 3.5 sonnet and GPT-4o achieve the highest average accuracies, but overall performance remains limited. |
Copied to clipboard
| Challenge: | Reference-free evaluation metrics for grammatical error correction have high correlation with human judgments, but they are not designed to evaluate adversarial systems that aim to obtain unjustifiably high scores. |
| Approach: | They propose adversarial attack strategies for four reference-free metrics . they propose SOME, Scribendi, IMPARA, and LLM-based metrics based on these metrics a . |
| Outcome: | The proposed attacks outperform the current state-of-the-art for four reference-free metrics . |
Copied to clipboard
| Challenge: | Document question answering is a task of question answering on given documents such as reports, slides, pamphlets, and websites. |
| Approach: | They propose a large-scale document-based QA dataset that requires both visual and textual information to answer questions. |
| Outcome: | The proposed dataset incorporates multiple categories of questions and unanswerable questions from the document for realistic question-answering applications. |
Copied to clipboard
| Challenge: | Morphological analysis (MA) and lexical normalization (LN) are important tasks for Japanese user-generated text. |
| Approach: | They construct a publicly available Japanese UGT corpus annotated with morphological and normalization information. |
| Outcome: | The proposed corpus shows low performance for non-general words and non-standard forms . morphological analysis is an important task in Japanese user-generated text . |
Copied to clipboard
| Challenge: | Simultaneous interpretation (SI) uses segmenting of source speech into chunks and translating them in order. |
| Approach: | They propose a variation of COMET that measures monotonicity for simultaneous interpretation . they train Simul-COMET on offline translation data and show stronger alignment with evaluation scores . |
| Outcome: | The proposed model shows stronger alignment with evaluation scores provided by interpreters than COMET. |
Copied to clipboard
| Challenge: | Existing methods for named entity recognition assume entities are not nested within other entities, so-called flat NER. |
| Approach: | They propose a layered method for nested named entity recognition . they use a set of hidden states to exclude the influence of the best path . |
| Outcome: | The proposed method performs better on ACE2004, ACE2005, and GENIA datasets. |
Copied to clipboard
| Challenge: | Existing methods to generate multiple translation candidates do not address the overcorrection problem, which discourages the model from generating synonymous expressions and leans toward gold standards, reducing the diversity in the candidates. |
| Approach: | They propose to introduce perturbed k-nearest neighbor machine translation (kNN-MT) to generate more diverse translations. |
| Outcome: | The proposed methods significantly improve candidate diversity and control diversity by tuning the perturbation’s magnitude. |
Copied to clipboard
| Challenge: | Large language models (LLMs) have advanced natural language processing by understanding, generating, and manipulating texts. |
| Approach: | They propose to use movie subtitle prompts to improve translation accuracy by incorporating movie meta-information into the models. |
| Outcome: | The proposed prompts improve translation accuracy and reduce computational effort. |
Copied to clipboard
| Challenge: | Recent studies highlight the use of Large Language Models (LLMs) for predicting response distributions as a cost-effective survey method. |
| Approach: | They examine whether LLMs can rationally estimate distributions when presented with explanations that are against commonsense. |
| Outcome: | The proposed models can rationally estimate distributions when presented with explanations that are against commonsense, but smaller or less human-optimized models follow explanations uncritically, compared to larger models that resist counterintuitive explanations by leveraging their pretraining-acquired knowledge. |
Copied to clipboard
| Challenge: | grammatical information annotation requires high human resources and is not trivial due to language mismatches and out-of-vocabulary problem. |
| Approach: | They propose to incorporate grammatical information without supervising annotation by induced latent phrase structure and synchronized phrase structures in encoder and decoder to enhance explainability. |
| Outcome: | The proposed method produces better performance and explainability in translation and alignment tasks without extra resources. |
Copied to clipboard
| Challenge: | Japanese input method editors (IMEs) allow users to input Japanese text using a limited set of characters such as the kana syllabary. |
| Approach: | They propose a simple decoding policy to enable simultaneous kana-kanji conversion in Japanese IMEs inspired by simultaneous machine translation. |
| Outcome: | The proposed approach achieves a better quality-latency trade-off than baselines while being more practical due to its ability to directly handle streaming input. |
Copied to clipboard
| Challenge: | Currently, multilingual datasets are created through translation, which cannot evaluate such language-specific aspects. |
| Approach: | They propose to curate a dataset for language-specific knowledge and commonsense . they propose to use multilingual commonsensiaq to leverage language models for a more efficient construction . |
| Outcome: | The proposed method reduces the creation cost by using multilingual LMs to create QAs . the proposed approach is based on the construction process of CSQA but with language models . |
Copied to clipboard
| Challenge: | Existing methods to improve summarization quality are limited to using source as guidance . reranking can be effective, but there are limitations, such as relying on reference-free metrics and rely on a single metric. |
| Approach: | They propose a model that reranks model-generated summaries by considering consistency to the source document and consensus among the other candidates. |
| Outcome: | The proposed system is competitive with existing methods, with human evaluations further confirming that it is superior. |
Copied to clipboard
| Challenge: | Experimental results show that n-gram models can achieve satisfactory performance on a large proportion of testing cases. |
| Approach: | They propose to learn a neural LM that fits the residual between an n-gram LM and the real-data distribution. |
| Outcome: | The proposed model achieves additional performance gains over popular standalone models on three typical language tasks. |
Copied to clipboard
| Challenge: | Web banner advertisements are often selected manually because of human preferences . a new benchmark evaluates the degree of alignment with human preferences in two tasks . |
| Approach: | a benchmark was developed to evaluate the human preference-driven banner selection process using vision-language models. |
| Outcome: | The proposed benchmark assesses the degree of alignment with human preferences in two tasks using vision-language models. |
Copied to clipboard
| Challenge: | Entity linking (EL) is the task of mapping named entities in text to canonical entries in a knowledge base. |
| Approach: | They propose a unified library for using and developing entity linking systems . a strong emphasis is placed on usability, making it highly extensible . |
| Outcome: | a new library aims to disambiguate named entities in text by mapping them to canonical entries in a knowledge base. |
Copied to clipboard
| Challenge: | Minimum Bayes risk (MBRS) decoding is a decision rule of text generation tasks that outperforms conventional maximum a posteriori (MAP) decoders by selecting high-quality outputs based on quality or preference rather than probability. |
| Approach: | They propose to use minimum bayes risk (MBRS) decoding to determine outputs based on quality rather than probability. |
| Outcome: | MBRS is an MIT-licensed open-source project with a focus on speed, reproducibility, and extensibility. |
Copied to clipboard
| Challenge: | Recent advances in language models have shown remarkable progress in various tasks. |
| Approach: | They introduce a dataset that incorporates wug words and inject them into pretraining data and evaluate them on evaluation data. |
| Outcome: | The proposed model does not induce grammatical knowledge even after repeated exposure to instances with the same structure but differing only in lexical items from evaluation instances in certain language phenomena. |
Copied to clipboard
| Challenge: | Existing work on identifying target-irrelevant information relies on locally normalized attention without considering possible labels at other time steps. |
| Approach: | They propose to extend local normalized attention to leverage structural information for refinement . they propose to use two implementation tricks to accelerate CRF computation and an initialization trick for Chinese character embeddings . |
| Outcome: | The proposed method can be extended to include Chinese character embeddings and two implementation tricks to accelerate CRF computation. |
Copied to clipboard
| Challenge: | Existing datasets for multiword expressions are inconsistently annotated, limited to a single type of MWE, or limited in size. |
| Approach: | They propose to use a new interface to generate MWE annotations for the first time in a dataset of MWE identification. |
| Outcome: | The proposed model outperforms existing models on the DiMSUM dataset. |
Copied to clipboard
| Challenge: | Existing instruction following datasets lack logical coherence across turns, narrow topical breadth and heavy manual effort. |
| Approach: | They propose a pipeline that leverages LLMs’ reasoning capabilities to assemble rich, topic-related single-instruction data into multi-turn dialogues and produce chains that are logically coherent, progressively deepen in content, and span diverse domains without fixed templates or extensive human annotation. |
| Outcome: | The proposed pipeline improves the performance of existing LLMs by integrating multiple topic-related data into multi-turn dialogues without fixed templates or extensive human annotation. |
Copied to clipboard
| Challenge: | Existing information extraction systems are not able to accurately capture organizational changes. |
| Approach: | They propose a task to extract corporate history events related to organizational changes by identifying company names before and after each event, as well as the corresponding date. |
| Outcome: | The proposed task is designed to identify company names before and after an event, as well as the corresponding date. |
Copied to clipboard
| Challenge: | Recent studies have focused on scaling the context size of large language models (LLMs) however, the enormous inference costs of LLMs limit their applications. |
| Approach: | They propose a method which uses attention scores and the l 1 norm to evaluate token importance. |
| Outcome: | Extensive experiments on LLaMA2-7B-chat and Vicuna-v1.5-7B show that the proposed method outperforms attention-score-only baselines in over 12 tasks. |
Copied to clipboard
| Challenge: | a library for using and developing grammatical error correction (GEC) evaluation metrics is released under the MIT license . |
| Approach: | They propose a library for using and developing grammatical error correction (GEC) evaluation metrics through a unified interface. |
| Outcome: | The proposed method is based on a unified evaluation framework with a strong focus on API usage and extensible. |
Copied to clipboard
| Challenge: | a new method for nominal coordination boundary identification is proposed . it uses pre-trained word embeddings to measure similarities of words and detects the span of coordination . |
| Approach: | They propose a method for nominal coordination boundary identification that uses pre-trained word embeddings to measure similarities of words and detects the span of coordination. |
| Outcome: | The proposed method can identify coordination boundaries without training on labeled data . it is comparable to a recent supervised method for the case when the coordinator conjoins simple noun phrases. |