| Challenge: | Existing computer-based writing assistance systems do not help non-native English speakers find better expressions than those they already know . existing FE dictionaries are limited in example sentences and are difficult to predict the category labels from user input, forcing users to manually designate the category when searching. |
| Approach: | They propose a framework for semantic searches of formulaic expressions and a method to leverage existing dictionaries and domain sentence corpora. |
| Outcome: | The proposed method leverages existing dictionaries and domain sentence corpora to find expressions that are not accurate. |
Similar Papers
Communicative-Function-Based Sentence Classification for Construction of an Academic Formulaic Expression Database (2021.eacl-main)
Copied to clipboard
| Challenge: | Formulaic expressions (FEs) are used in scientific papers. |
| Approach: | They propose to use top–down approach to assign CF labels to sentences and then extract FEs from them using a CF-labelled sentence dataset. |
| Outcome: | The proposed method can be used to build FE databases of disciplines that are different from the training data. |
MathAlign: Linking Formula Identifiers to their Contextual Natural Language Descriptions (2020.lrec-1)
Copied to clipboard
Maria Alexeeva, Rebecca Sharp, Marco A. Valenzuela-Escárcega, Jennifer Kadowaki, Adarsh Pyarelal, Clayton Morrison
| Challenge: | Existing approaches to extract mathematical concepts and their descriptions are useful for a variety of tasks, including math information retrieval and accessibility efforts to make scientific documents available to the visually impaired. |
| Approach: | They propose a rule-based approach which extracts LaTeX representations of formula identifiers and links them to their in-text descriptions, given only the original PDF and the location of the formula of interest. |
| Outcome: | The proposed approach extracts LaTeX representations of formula identifiers and links them to their in-text descriptions, given only the original PDF and the location of the formula of interest. |
FormulaReasoning: A Dataset for Formula-Based Numerical Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for numerical reasoning often lack explicit knowledge of formulas . current datasets do not provide process supervision information, resulting in incomplete reasoning . |
| Approach: | They propose a benchmark for formula-based numerical reasoning with 5,324 questions . they provide annotations in English and Chinese and a formula database as an external knowledge source . |
| Outcome: | The proposed model includes 5,324 questions requiring calculations grounded in external physics principles. |
Dictionary-Aided Translation for Handling Multi-Word Expressions in Low-Resource Languages (2024.findings-acl)
Copied to clipboard
| Challenge: | Multi-word expressions (MWEs) are a challenging task in natural language processing . they are defined as combinations of at least two words with distinct lexical, morphological, syntactic, semantic or statistical characteristics. |
| Approach: | They propose a method leveraging an available out-of-context lexicon to improve translations . they propose to use a dictionary-aided translation to better translate multi-word expressions based on human annotations. |
| Outcome: | The proposed method improves translations comparable to those of a human speaker. |
An Evaluation Dataset for Identifying Communicative Functions of Sentences in English Scholarly Papers (2020.lrec-1)
Copied to clipboard
| Challenge: | Formulaic expressions are used by authors of scientific papers because they convey specific communicative functions in the rhetorical structure of papers. |
| Approach: | They created a manually annotated dataset to detect formulaic expressions in sentences using a seed list of labelled formulaic words. |
| Outcome: | The proposed dataset can detect communicative functions in sentences using a seed list of labelled expressions from scholarly papers in the ACL Anthology. |
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)
Copied to clipboard
| Challenge: | WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them. |
| Approach: | This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings. |
| Outcome: | The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets. |
A Large Automatically-Acquired All-Words List of Multiword Expressions Scored for Compositionality (L18-1)
Copied to clipboard
| Challenge: | Existing literature on semantically idiosyncratic multiword expressions is limited to English . idiomatic expressions are phraseological units consisting of more than one lexeme and exhibit some kind of idiom. |
| Approach: | They propose to make available a large automatically-acquired all-words list of English multiword expressions scored for compositionality. |
| Outcome: | The proposed list improves the BLEU scores of the English multiword expressions. |
Dedicated Language Resources for Interdisciplinary Research on Multiword Expressions: Best Thing since Sliced Bread (2020.lrec-1)
Copied to clipboard
| Challenge: | Multiword expressions are challenging for disciplines like NLP, psycholinguistics and second language acquisition due to their more or less fixed character. |
| Approach: | They propose to develop tools and language resources that are crucial for multifaceted research. |
| Outcome: | The proposed tools and language resources are crucial for this kind of multifaceted research. |
Textual Coverage of Eventive Entries in Lexical Semantic Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment. |
| Approach: | They propose to quantify coverage gaps in lexical semantic resources when applied to running texts taken from the internet. |
| Outcome: | The proposed resources cover eventive entries (verbs, predicates, etc.) of well-known lexical semantic resources when applied to running texts taken from the internet. |
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)
Copied to clipboard
| Challenge: | Terms are notoriously difficult to identify, both automatically and manually. |
| Approach: | They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information . |
| Outcome: | The proposed method provides a tool for evaluation and rich source of information about terms. |