Using Formulaic Expressions in Writing Assistance Systems (C18-1)

Copied to clipboard

Challenge: Existing computer-based writing assistance systems do not help non-native English speakers find better expressions than those they already know . existing FE dictionaries are limited in example sentences and are difficult to predict the category labels from user input, forcing users to manually designate the category when searching.
Approach: They propose a framework for semantic searches of formulaic expressions and a method to leverage existing dictionaries and domain sentence corpora.
Outcome: The proposed method leverages existing dictionaries and domain sentence corpora to find expressions that are not accurate.

Similar Papers

Communicative-Function-Based Sentence Classification for Construction of an Academic Formulaic Expression Database (2021.eacl-main)

Copied to clipboard

Challenge: Formulaic expressions (FEs) are used in scientific papers.
Approach: They propose to use top–down approach to assign CF labels to sentences and then extract FEs from them using a CF-labelled sentence dataset.
Outcome: The proposed method can be used to build FE databases of disciplines that are different from the training data.
MathAlign: Linking Formula Identifiers to their Contextual Natural Language Descriptions (2020.lrec-1)

Copied to clipboard

Challenge: Existing approaches to extract mathematical concepts and their descriptions are useful for a variety of tasks, including math information retrieval and accessibility efforts to make scientific documents available to the visually impaired.
Approach: They propose a rule-based approach which extracts LaTeX representations of formula identifiers and links them to their in-text descriptions, given only the original PDF and the location of the formula of interest.
Outcome: The proposed approach extracts LaTeX representations of formula identifiers and links them to their in-text descriptions, given only the original PDF and the location of the formula of interest.
FormulaReasoning: A Dataset for Formula-Based Numerical Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Existing datasets for numerical reasoning often lack explicit knowledge of formulas . current datasets do not provide process supervision information, resulting in incomplete reasoning .
Approach: They propose a benchmark for formula-based numerical reasoning with 5,324 questions . they provide annotations in English and Chinese and a formula database as an external knowledge source .
Outcome: The proposed model includes 5,324 questions requiring calculations grounded in external physics principles.
Dictionary-Aided Translation for Handling Multi-Word Expressions in Low-Resource Languages (2024.findings-acl)

Copied to clipboard

Challenge: Multi-word expressions (MWEs) are a challenging task in natural language processing . they are defined as combinations of at least two words with distinct lexical, morphological, syntactic, semantic or statistical characteristics.
Approach: They propose a method leveraging an available out-of-context lexicon to improve translations . they propose to use a dictionary-aided translation to better translate multi-word expressions based on human annotations.
Outcome: The proposed method improves translations comparable to those of a human speaker.
An Evaluation Dataset for Identifying Communicative Functions of Sentences in English Scholarly Papers (2020.lrec-1)

Copied to clipboard

Challenge: Formulaic expressions are used by authors of scientific papers because they convey specific communicative functions in the rhetorical structure of papers.
Approach: They created a manually annotated dataset to detect formulaic expressions in sentences using a seed list of labelled formulaic words.
Outcome: The proposed dataset can detect communicative functions in sentences using a seed list of labelled expressions from scholarly papers in the ACL Anthology.
A Survey on Automatically-Constructed WordNets and their Evaluation: Lexical and Word Embedding-based Approaches (L18-1)

Copied to clipboard

Challenge: WordNets are lexical databases in which groups of synonyms are stored according to the semantic relationships between them.
Approach: This paper describes various approaches to constructing WordNets automatically by leveraging traditional lexical resources and newer trends such as word embeddings.
Outcome: The proposed methods leverage traditional lexical resources and newer trends such as word embeddings to build and evaluate WordNets.
A Large Automatically-Acquired All-Words List of Multiword Expressions Scored for Compositionality (L18-1)

Copied to clipboard

Challenge: Existing literature on semantically idiosyncratic multiword expressions is limited to English . idiomatic expressions are phraseological units consisting of more than one lexeme and exhibit some kind of idiom.
Approach: They propose to make available a large automatically-acquired all-words list of English multiword expressions scored for compositionality.
Outcome: The proposed list improves the BLEU scores of the English multiword expressions.
Dedicated Language Resources for Interdisciplinary Research on Multiword Expressions: Best Thing since Sliced Bread (2020.lrec-1)

Copied to clipboard

Challenge: Multiword expressions are challenging for disciplines like NLP, psycholinguistics and second language acquisition due to their more or less fixed character.
Approach: They propose to develop tools and language resources that are crucial for multifaceted research.
Outcome: The proposed tools and language resources are crucial for this kind of multifaceted research.
Textual Coverage of Eventive Entries in Lexical Semantic Resources (2024.lrec-main)

Copied to clipboard

Challenge: Several English, German, Spanish and Czech lexical semantic resources (which, for the most part, focus on verbs and predicates) have been selected for this experiment.
Approach: They propose to quantify coverage gaps in lexical semantic resources when applied to running texts taken from the internet.
Outcome: The proposed resources cover eventive entries (verbs, predicates, etc.) of well-known lexical semantic resources when applied to running texts taken from the internet.
A Gold Standard for Multilingual Automatic Term Extraction from Comparable Corpora: Term Structure and Translation Equivalents (L18-1)

Copied to clipboard

Challenge: Terms are notoriously difficult to identify, both automatically and manually.
Approach: They propose a method to annotate terms manually from a comparable corpus . they show that the gold standard provides a tool for evaluation and a rich source of information .
Outcome: The proposed method provides a tool for evaluation and rich source of information about terms.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations