Papers by Sandeep Soni
Predicting Long-Term Citations from Short-Term Linguistic Influence (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to quantify linguistic influence in timestamped documents are not informative about extent to which a paper affected subsequent publications. |
| Approach: | They propose to quantify linguistic influence in timestamped document collections by estimating a Hawkes process with a low-rank parameter matrix and identify lexical and semantic changes using contextual embeddings and word frequencies. |
| Outcome: | The proposed method is based on an online evaluation with incremental temporal training/test splits, in comparison with a strong baseline that includes predictors for initial citation counts, topics, and lexical features. |
Grounding Characters and Places in Narrative Text (2023.acl-long)
Copied to clipboard
| Challenge: | Prior work has analyzed characters and locations from text independently without grounding characters to their locations in narrative time. |
| Approach: | They propose a task to assign a spatial relationship category for every character and location co-mention within a window of text, taking into account linguistic context, narrative tense, and temporal scope. |
| Outcome: | The proposed model allows to test hypotheses on mobility and domestic space . women as characters tend to occupy more interior space than men, the model shows . |
Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4 (2023.emnlp-main)
Copied to clipboard
| Challenge: | a recent study has shown that open AI models memorize a wide collection of copyrighted materials . however, these models also present a challenge for establishing the validity of results . |
| Approach: | They propose to use a name cloze membership inference query to infer books that are known to ChatGPT and GPT-4. |
| Outcome: | The proposed model performs better on memorized books than on non-memorized books for downstream tasks. |