Papers by Daniel Lo

7 papers
S2ORC: The Semantic Scholar Open Research Corpus (2020.acl-main)

Copied to clipboard

Challenge: Academic papers are an increasingly important textual domain for natural language processing (NLP) research.
Approach: They propose to aggregate 81.1M English-language academic papers into a unified source . they hope this resource will facilitate research and development of tools for text mining over academic text.
Outcome: The proposed corpus includes metadata, abstracts, bibliographic references, and structured full text for 8.1M open access papers.
PaperMage: A Unified Toolkit for Processing, Representing, and Manipulating Visually-Rich Scientific Documents (2023.emnlp-demo)

Copied to clipboard

Challenge: Existing tools for working with scientific documents are limited and documents are often in difficult-to-use PDF formats.
Approach: They propose an open-source Python toolkit for analyzing and processing visually-rich scientific documents.
Outcome: PaperMage provides turn-key recipes for common scientific document processing use-cases.
Dynamic Stashing Quantization for Efficient Transformer Training (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive performance on a range of Natural Language Processing (NLP) tasks.
Approach: They propose a dynamic quantization strategy that reduces the amount of memory operations and reduces arithmetic cost by 20.95 on two translation tasks and three classification tasks.
Outcome: The proposed model reduces the amount of arithmetic operations by 20.95 and the number of DRAM operations by 2.55 on two translation tasks and three classification tasks.
ACCoRD: A Multi-Document Approach to Generating Diverse Descriptions of Scientific Concepts (2022.emnlp-demos)

Copied to clipboard

Challenge: Current systems that automatically define unfamiliar terms only surface a single "best" description for all users, which may not be accessible for all readers, given varying background knowledge.
Approach: They propose an end-to-end system that generates sets of descriptions of scientific concepts . ACCoRD corpus includes 1,275 labeled contexts and 1,787 expert-authored concept descriptions .
Outcome: The proposed system produces diverse descriptions of concepts in terms of reference concepts.
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on text-based agents, neglecting many natural tasks that require visual information to effectively solve.
Approach: They propose a benchmark to assess the performance of multimodal web agents . they use visual and textual inputs to process and interpret natural language instructions .
Outcome: a new benchmark assesses the performance of multimodal agents on visually grounded tasks . the benchmark identifies limitations of text-only agents and offers insights towards building stronger agents for the web .
TLDR: Extreme Summarization of Scientific Documents (2020.findings-emnlp)

Copied to clipboard

Challenge: TLDR generation requires expert background knowledge and understanding of complex domain-specific language.
Approach: They propose a learning strategy that exploits titles as an auxiliary training signal.
Outcome: The proposed method improves upon strong baselines under both automated metrics and human evaluations.
ArxivDIGESTables: Synthesizing Scientific Literature into Tables using Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Using language models (LMs) can generate literature review tables by decomposing it into separate schema and value generation steps.
Approach: They propose a framework that leverages language models to perform literature review table generation by decomposing it into separate schema and value generation steps.
Outcome: The proposed framework decomposes the task into two sub-tasks: schema generation and value generation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations