Papers by Norman Sadeh

6 papers
Creation and Analysis of an International Corpus of Privacy Laws (2024.lrec-main)

Copied to clipboard

Challenge: a corpus of 1,043 privacy laws, regulations, and guidelines covers 183 jurisdictions . prior efforts to study privacy law in the form of privacy policies have lacked a large-scale collection .
Approach: They propose a corpus of 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions.
Outcome: The Privacy Law Corpus covers 1,043 privacy laws, regulations, and guidelines covering 183 jurisdictions.
Question Answering for Privacy Policies: Combining Computational and Legal Perspectives (D19-1)

Copied to clipboard

Challenge: Privacy policies are long and complex documents that are difficult for users to read and understand.
Approach: They present a corpus of 1750 questions about privacy policies of mobile applications and over 3500 expert annotations of relevant answers.
Outcome: The proposed corpus of 1750 questions on privacy policies shows that a strong neural baseline underperforms human performance by almost 0.3 F1 on PrivacyQA.
A Tale of Two Regulatory Regimes: Creation and Analysis of a Bilingual Privacy Policy Corpus (2022.lrec-1)

Copied to clipboard

Challenge: With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages.
Approach: They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA.
Outcome: The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices.
Breaking Down Walls of Text: How Can NLP Benefit Consumer Privacy? (2021.acl-long)

Copied to clipboard

Challenge: Privacy policies are long and complex documents that are difficult for users to read and comprehend.
Approach: They propose language technologies to help users reclaim control over their privacy . they highlight many remaining opportunities to develop more precise or nuanced language technologies .
Outcome: The proposed language technologies can address the privacy information gap . they can be more precise or nuanced in the way they use the text of privacy policies.
Stress Test Evaluation for Natural Language Inference (C18-1)

Copied to clipboard

Challenge: Existing models perform well at standard datasets for NLI, achieving impressive results across different genres of text.
Approach: They propose to use automatic stress tests to evaluate models' ability to make inferential decisions.
Outcome: The proposed model performs well across genres of text, but lacks the ability to make inferential decisions.
Supervised and Unsupervised Methods for Robust Separation of Section Titles and Prose Text in Web Documents (D18-1)

Copied to clipboard

Challenge: a web text structure is underutilized, but its visual organization is useful for NLP tasks . a flexible system for extracting hierarchical section titles and prose organization is developed .
Approach: a new system extracts hierarchical section titles and prose organization from web documents . the system uses features from syntax, semantics, discourse and markup to build two models .
Outcome: a new system extracts the hierarchical section titles and prose organization of web documents . the system achieves an overall precision of 0.82 and a recall of 0.98 on three domains of web text .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations