Papers by Weijia Jia
GAN Driven Semi-distant Supervision for Relation Extraction (N19-1)
Copied to clipboard
| Challenge: | Existing methods for relation extraction are limited to costly hand-labeled training sets and hard to be extended to large-scale relations. |
| Approach: | They propose a semi-distant supervision approach for relation extraction by constructing a small accurate dataset and properly leveraging numerous instances without relation labels. |
| Outcome: | The proposed approach achieves significant improvements over baselines on real-world datasets. |
Active Testing: An Unbiased Evaluation Method for Distantly Supervised Relation Extraction (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for distantly supervised relation extraction suffer from low quality of test set, which leads to considerable biased performance evaluation. |
| Approach: | They propose a method to evaluate distantly supervised relation extraction using noisy test sets and manual annotations. |
| Outcome: | Experiments on a widely used benchmark show that the proposed method can yield approximately unbiased evaluations for distantly supervised relation extractors. |
Self-distilled Transitive Instance Weighting for Denoised Distantly Supervised Relation Extraction (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to reducing wrongly labeled instances are based on a bag-level setting . however, sentence-level training is vulnerable to the noise brought by DS, which limits its application. |
| Approach: | They propose a transitive instance weighting mechanism integrated with the self-distilled BERT backbone to generate dynamic instance weights for denoised sentence-level training. |
| Outcome: | The proposed method can tackle wrongly labeled instances and prevent overfitting. |
Regularized Attentive Capsule Network for Overlapped Relation Extraction (2020.coling-main)
Copied to clipboard
| Challenge: | Existing methods to extract relations from distant supervision contain low-quality instances with noisy words and overlapped relations. |
| Approach: | They propose a Regularized Attentive Capsule Network to better identify overlapped relations in informal sentences . they embed multi-head attention into the capsule network as the low-level capsules . |
| Outcome: | Extensive experiments show that the proposed model improves relation extraction. |
Improving Abstractive Document Summarization with Salient Information Modeling (P19-1)
Copied to clipboard
| Challenge: | Abstractive document summarization is a task of natural language generation which generates fluent summaries with salient information automatically. |
| Approach: | They propose to incorporate a Gaussian focal bias on attention scores into an encoder to enhance the perception of local context and to distinguish salient information precisely. |
| Outcome: | The proposed framework outperforms state-of-the-art models on the CNN/Daily Mail benchmark and is based on a focus-attention mechanism and two new extensions. |
Automatic Slide Updating with User-Defined Dynamic Templates and Natural Language Instructions (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing automation methods follow fixed template filling and cannot support dynamic updates for diverse, user-authored decks. |
| Approach: | They propose a framework that combines multimodal slide parsing, natural language instruction grounding, and tool-augmented reasoning for tables, charts, and textual conclusions. |
| Outcome: | The proposed framework updates content while preserving layout and style while maintaining a strong reference baseline on DynaSlide. |
Distantly Supervised Relation Extraction using Multi-Layer Revision Network and Confidence-based Multi-Instance Learning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Distantly supervised relation extraction is used in knowledge bases but its low quality and noisy sentences are present in sentence bags. |
| Approach: | They propose a multi-layer revision network which emphasizes inner-sentence correlations before extracting relevant information within sentences. |
| Outcome: | The proposed method improves on two New York Times datasets. |
Neural Relation Extraction via Inner-Sentence Noise Reduction and Transfer Learning (D18-1)
Copied to clipboard
| Challenge: | Existing methods for extracting relations are slow and lack precision . a novel approach to extract relations is proposed to reduce noise between sentences . |
| Approach: | They propose a word-level distant supervised approach for relation extraction using New York Times and Freebase. |
| Outcome: | The proposed method improves the area of precision/call(PR) from 0.35 to 0.39 over the state-of-the-art methods. |
Exploring Sentence Community for Document-Level Event Extraction (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing approaches to document-level event extraction neglect the complex logic structures in long texts. |
| Approach: | They propose a framework that exploits the relationship between sentences to extract multiple events by sentence community detection using graph attention networks. |
| Outcome: | The proposed framework achieves competitive results over state-of-the-art methods on the large-scale document-level event extraction dataset. |
ODTQA-FoRe: An Open-Domain Tabular Question Answering Dataset for Future Data Forecasting and Reasoning (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing tabular question answering systems cannot perform future-oriented numerical prediction . open-domain tabular questions are a popular approach for QA tasks . |
| Approach: | They propose a task that covers time-series forecasting and forecast-based reasoning scenarios using real estate data. |
| Outcome: | The proposed framework decomposes the problem into three collaborative roles that synthesize the results to construct a precise and consistent final answer. |
ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to tabular QA are limited to closed-domain scenarios . existing approaches do not solve the core challenge of generating correct answers without user clarification . |
| Approach: | They propose a benchmark to tackle underspecified or uncertain queries in tabular question answering . they propose ODUTQA-MDC task and a multi-agent framework to detect ambiguities . |
| Outcome: | The proposed framework excels at detecting ambiguities, clarifying them through dialogue, and refining answers. |
ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering (2026.acl-long)
Copied to clipboard
| Challenge: | Real estate agents are labor-intensive, difficult to scale, and prone to interest-driven bias. |
| Approach: | They propose a large-scale benchmark of 29,270 real-estate instances with machine-verifiable supervision for intermediate steps . they propose 'hIRE-Agent' framework that integrates heterogeneous evidence into an understand–plan–execute architecture as a strong baseline . |
| Outcome: | Experiments show that HIRE-Agent integrates heterogeneous evidence . the framework is able to integrate a front-end parser, planning Supervisor, and execution Specialists . |