Papers by Gerald Penn

14 papers
Rationally Reappraising ATIS-based Dialogue Systems (P19-1)

Copied to clipboard

Challenge: Recent state-of-the-art neural models have obtained F1-scores near 98% on the task of slot filling.
Approach: They propose to fix annotation errors in ATIS and propose a rule-based grammar for slot filling that achieves a 95.82% F1 score.
Outcome: The proposed grammar achieves a 95.82% F1-score on the ATIS domain.
Does BERT Rediscover a Classical NLP Pipeline? (2022.coling-1)

Copied to clipboard

Challenge: Existing theories of BERT's structure lack conclusive empirical support . however, there is scepticism about the premises of probing itself .
Approach: They propose a new probe called GridLoc that can take into account token positions, training rounds, and random seeds.
Outcome: The proposed probe detects other, stronger regularities suggesting appeals to layer depth may not be the preferable mode of explanation for BERT’s inner workings.
Tiny Budgets, Big Gains: Parameter Placement Strategy in Parameter Super-Efficient Fine-Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods such as LoRA and VeRA use memory-efficient methods to fine-tune large language models.
Approach: They propose a method that uses only 1–5% of the standard LoRA parameters and achieves state-of-the-art performance across a wide range of tasks.
Outcome: The proposed method achieves state-of-the-art performance across a wide range of tasks using only 1–5% of the standard LoRA parameters.
Inside-Outside Algorithm for Probabilistic Product-Free Lambek Categorial Grammar (2025.coling-main)

Copied to clipboard

Challenge: Many studies have discovered hidden syntactic structures within language models without the guidance of explicit rules.
Approach: They propose an inside-outside algorithm for Probabilistic Lambek Categorical Grammar.
Outcome: The proposed algorithm is used in the estimation of probabilistic context-free grammars.
LCGbank: A Corpus of Syntactic Analyses Based on Proof Nets (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies have focused on statistical syntactic parsing with proof nets . however, there has been a paucity of corpora in formalisms for which proof net is applicable .
Approach: They propose a corpus of syntactic analyses based on Lambek categorial grammar . they leverage the relationship between LCG and CCG to address this problem .
Outcome: The proposed method exploits the relationship between LCG and CCG to build an English-language corpus of syntactic analyses based on proof nets . the results suggest that the proposed method is weakly context-free equivalent and NP-complete .
Temporal Histories of Epidemic Events (THEE): A Case Study in Temporal Annotation for Public Health (2020.lrec-1)

Copied to clipboard

Challenge: Current EBS estimates the occurrence time of events based on coarse metadata such as document publication time.
Approach: They propose a temporal annotation standard THEE-TimeML and a corpus TheeBank . they document the corpus annotation process and demonstrate the immediate benefit .
Outcome: The proposed standards are based on the existing timeML and the corpus TheeBank . the proposed standards demonstrate the immediate benefit to public health applications .
The Chinese Remainder Theorem for Compact, Task-Precise, Efficient and Secure Word Embeddings (2021.eacl-main)

Copied to clipboard

Challenge: a new method for compressing word vector embeddings into integers is being developed . a high precision approach to compressing words into integer results in negligible performance gains .
Approach: They propose a method for compressing word vector embeddings into integers using the Chinese Reminder Theorem.
Outcome: The proposed method speeds up addition by 48.27% and compresses GloVe word embedding libraries by 25.86%.
Reanalyzing the Most Probable Sentence Problem: A Case Study in Explicating the Role of Entropy in Algorithmic Complexity (2021.eacl-main)

Copied to clipboard

Challenge: Existing descriptive complexity measures are ineffective at describing algorithms' behaviour, and can make an apparently tractable problem seem NP-complete.
Approach: They propose to use statistical measures to give an updated analysis of the complexity of the NP-complete most probable sentence problem for pCFGs.
Outcome: The proposed method can be applied to word sense disambiguation and inference tasks.
LLM-supertagger: Categorial Grammar Supertagging via Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that LLMs are underperforming in classification tasks due to their decoder-based nature.
Approach: They propose a method that significantly boosts LLMs' performance in supertagging for both Combinatory Categorial Grammar (CCG) and Lambek Categorian Grammar (LCG).
Outcome: The proposed method outperforms LSTM and encoder-based models and achieves state-of-the-art performance.
A Generative Model for Lambek Categorial Sequents (2024.lrec-main)

Copied to clipboard

Challenge: generative models such as PLC+ generate grammatical sentences with a high probability of being grammatized.
Approach: They propose a generative model, PLC+, for generating Lambek Categorial Grammar(LCG) sequents.
Outcome: The proposed model generates Lambek Categorial Grammar(LCG) sequents and is more robust to probabilistic context-free grammars.
ConTempo: A Unified Temporally Contrastive Framework for Temporal Relation Extraction (2024.findings-acl)

Copied to clipboard

Challenge: Temporal relation extraction (TRE) is a task of classifying temporal relations between events conveyed in narratives.
Approach: They propose a Temporally Contrastive learning model that increases the model’s awareness of the meaning of temporal relations by leveraging their symmetric or antisymmetric properties.
Outcome: The proposed model improves the model's representation of meaning of temporal relations and its ability to integrate with the underlying temporal calculus.
Sheaf Discovery with Joint Computation Graph Pruning and Flexible Granularity (2025.emnlp-main)

Copied to clipboard

Challenge: Experimental results show that DiscoGP extracts sheaves that preserve 93-100% of a model’s performance while comprising only 1-7% of the original weights and connections.
Approach: They propose a framework for extracting self-contained modular units within neural language models (LMs) they use a gradient-based pruning algorithm to prune the original LM to a sparse skeleton .
Outcome: The proposed framework preserves 93-100% of the original model's performance while preserving only 1-7% of the model''s original weights and connections.
Decomposed scoring of CCG dependencies (2023.acl-short)

Copied to clipboard

Challenge: a standard evaluation of supertagging errors can result in disproportionate penalization of supertaggers . comparative categorial grammar (ccg) supertaggers can adjust for their own errors to keep sentences parsable .
Approach: They propose a decomposed scoring method based on subcategorial labels to address this problem.
Outcome: The proposed method penalizes supertagging errors and obfuscates erroneous dependencies . the proposed method is based on subcategorial labels .
FAB: The French Absolute Beginner Corpus for Pronunciation Training (2020.lrec-1)

Copied to clipboard

Challenge: French Absolute Beginner corpus is intended for the development and study of Computer-Assisted Pronunciation Training (CAPT) tools for absolute beginner learners.
Approach: They introduce the French Absolute Beginner (FAB) speech corpus which is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners.
Outcome: The proposed corpus is intended for the development and study of Computer-Assisted Pronunciation Training tools for absolute beginner learners.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations