Papers by Yossi Matias

5 papers
Learning and Evaluating a Differentially Private Pre-trained Language Model (2021.findings-emnlp)

Copied to clipboard

Challenge: Contextual language models have improved performance but can lead to information leakage .
Approach: They propose a differentially-private word-piece algorithm that allows training a tailored domain-specific vocabulary while maintaining privacy.
Outcome: The proposed model can guarantee privacy while maintaining good model performance.
Audio De-identification - a New Entity Recognition Task (N19-2)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is an important step in de-identification (de-ID) of medical records, many of which are recorded conversations between a patient and a doctor.
Approach: They propose to use Named Entity Recognition (NER) to detect audio spans with entity mentions in medical records and then use it to evaluate the results.
Outcome: The proposed pipeline is based on a large labeled segment of the Switchboard and Fisher audio datasets and compares it with a benchmark.
Breaking the Language Barrier: Can Direct Inference Outperform Pre-Translation in Multilingual LLM Applications? (2024.naacl-short)

Copied to clipboard

Challenge: Existing studies have focused on pre-translation, but there is still need for it . authors say that it is not universally necessary to translate large language models .
Approach: They re-evaluate the need for pre-translation in the context of PaLM2 models . authors found that PaLM2-L consistently outperforms pre-translated in 94 out of 108 languages .
Outcome: The proposed model outperforms pre-translation in 94 out of 108 languages and 6 benchmarks . authors argue that pre-translated inputs can be used to improve performance .
TRUE: Re-evaluating Factual Consistency Evaluation (2022.naacl-main)

Copied to clipboard

Challenge: Grounded text generation systems often generate factual inconsistencies, hindering their real-world applicability.
Approach: They propose a method to assess factual consistency metrics on standardized texts . they recommend NLI and question generation-and-answering-based methods as starting points .
Outcome: The proposed method is more actionable and interpretable than previous methods.
Location Not Found: Exposing Implicit Local and Global Biases in Multilingual LLMs (2026.acl-long)

Copied to clipboard

Challenge: Multilingual large language models have minimized the fluency gap between languages, but they are exposed to the risk of biases as knowledge and norms may propagate across languages.
Approach: They propose a test set with 2,156 questions in 12 languages to quantify models' biases . they show a global bias towards answers relevant to the US-locale .
Outcome: The proposed model can answer locale-ambiguous questions in 12 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations