Challenge: Documentlevel NLI is an important problem for many tasks including verification of factual correctness of documents.
Approach: They propose a document-level natural language inference model that builds a hierarchical document graph enriched through inter-sentence relations and performs paragraph pruning using the novel SubGraph Pooling layer.
Outcome: The proposed model performs on a legal judicial reasoning task with a dataset enriched with document graphs and a proposed evidence selection algorithm.

Similar Papers

DocNLI: A Large-scale Dataset for Document-level Natural Language Inference (2021.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on sentence-level inference, which limits its application in downstream NLP problems.
Approach: They propose to construct a large-scale dataset for document-level NLI that can be used to study NLP problems.
Outcome: The proposed model performs well on popular sentence-level benchmarks and generalizes well to out-of-domain NLP tasks that rely on inference at document granularity.
R2F: A General Retrieval, Reading and Fusion Framework for Document-level Natural Language Inference (2022.emnlp-main)

Copied to clipboard

Challenge: Document-level natural language inference (DOCNLI) is a new task in natural language processing.
Approach: They propose a document-level natural language inference framework that fuses sentence-level tasks into a set of sentence-based tasks.
Outcome: The proposed framework improves interpretability and performance with evidence.
ContractNLI: A Dataset for Document-level Natural Language Inference for Contracts (2021.findings-emnlp)

Copied to clipboard

Challenge: Contract review is a time-consuming procedure that costs companies millions of dollars each year . linguistic characteristics of contracts, such as negations by exceptions, contribute to the difficulty of this task .
Approach: They propose a document-level natural language inference (NLI) task for contracts . they annotate and release the largest corpus to date consisting of 607 annotated contracts a linguistically rich system is proposed .
Outcome: The proposed system is based on a contract review task that includes 607 annotated contracts.
Cross-Domain Modeling of Sentence-Level Evidence for Document Retrieval (D19-1)

Copied to clipboard

Challenge: Existing test collections provide only document-level relevance judgments, and documents exceed the length that BERT was designed to handle.
Approach: They propose to aggregate sentence-level evidence to rank news articles using BERT . they also leverage passage-level relevance judgments available in other domains to fine-tune BERT models that capture cross-domain notions of relevance.
Outcome: The proposed model aggregates sentence-level evidence to rank documents on three standard test collections.
Eider: Empowering Document-level Relation Extraction with Efficient Evidence Extraction and Inference-stage Fusion (2022.findings-acl)

Copied to clipboard

Challenge: Document-level relation extraction (DocRE) aims to extract semantic relations among entity pairs in a document.
Approach: They propose an evidence-enhanced framework that empowers document-level relation extraction (DocRE) Eider efficiently extracts evidence and effectively fuses extracted evidence in inference.
Outcome: The proposed framework outperforms state-of-the-art methods on three benchmark datasets.
Evaluating BERT for natural language inference: A case study on the CommitmentBank (D19-1)

Copied to clipboard

Challenge: Natural language inference datasets can identify premise-hypothesis relationship without observing premise . recasting of the CommitmentBank for NLI creates hypotheses that stand in entailment/contradiction/neutral relationship with premise.
Approach: They propose to recast the CommitmentBank for NLI to stand in certain relationships with the premise . hypotheses are complements of clause-embedding verbs in each premise, rethinking the CommittedBank .
Outcome: The proposed model performs well on the CommitmentBank with 85% F1 . however, the model does not capture the full complexity of pragmatic reasoning, authors say .
Stretching Sentence-pair NLI Models to Reason over Long Documents and Clusters (2022.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in modeling and datasets demonstrate promising performance for NLI.
Approach: They explore the direct zero-shot applicability of NLI models to real applications . they analyze the robustness of models to longer and out-of-domain inputs .
Outcome: The proposed models are robust to longer and out-of-domain inputs and can perform on full documents.
Modeling Document-Level Context for Event Detection via Important Context Selection (2021.emnlp-main)

Copied to clipboard

Challenge: Existing methods for Event Detection (ED) do not encode long-range document-level context . e.g., BERT cannot encode long text-level contextual information .
Approach: They propose a method to model document-level context for Event Detection using transformer-based language models.
Outcome: The proposed model can predict event prediction of target sentence in document-level context . the proposed model is effective on multiple benchmark datasets .
Enhancing Knowledge Selection via Multi-level Document Semantic Graph (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods view knowledge selection as a sentence matching or classification. Existing techniques can’t capture the semantic relationships within complex documents.
Approach: They propose a method that can construct multi-level document semantic graph from the grounding document and store semantic relationships within the documents effectively.
Outcome: The proposed method can store semantic relationships within documents effectively and efficiently and achieve state-of-the-art results on public datasets.
DocCGen: Document-based Controlled Code Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) produce state-of-the-art performance on natural language to code generation for resource-rich general-purpose languages like C++, Java, and Python.
Approach: They propose a framework that breaks the NL-to-Code generation task into two steps . they use library documentation to detect the correct libraries and schema rules extracted from the documentation to constrain the decoding .
Outcome: The proposed framework improves different sized language models across all six evaluation metrics, reducing syntactic and semantic errors in structured code.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations