Challenge: a system that organizes scientific knowledge into a hierarchical concept structure is needed to enable efficient exploration of Web-scale knowledge.
Approach: They propose a system that organizes scientific knowledge into a hierarchical concept structure . system allows researchers to identify hundreds of thousands of scientific concepts . it also allows researchers tagging scientific publications into millions of concepts based on text and graph structure based model .
Outcome: The proposed system builds the most comprehensive cross-domain scientific concept ontology published to date, with more than 200 thousand concepts and over one million relationships.

Similar Papers

SciConceptMiner: A system for large-scale scientific concept discovery (2021.acl-demo)

Copied to clipboard

Challenge: SciConceptMiner is a self-supervised system for the capture of scientific concepts . the system is scalable to the size of documents and the number of topics it can model .
Approach: They propose a self-supervised system for the automatic capture of scientific concepts from academic publications and semi-structured data.
Outcome: The proposed system achieves high accuracy (94.7%) with more than 740K scientific concepts.
A Summarization System for Scientific Documents (D19-3)

Copied to clipboard

Challenge: a qualitative user study identified the most valuable scenarios for scientific content consumption.
Approach: They propose a system that retrieves and summarizes scientific documents for a given information need.
Outcome: The proposed system ingested 270,000 scientific papers and validated with human experts.
Knowledge Navigator: LLM-guided Browsing Framework for Exploratory Search in Scientific Literature (2024.findings-emnlp)

Copied to clipboard

Challenge: Knowledge Navigator organizes retrieved documents into a navigable, two-level hierarchy of named and descriptive topics and subtopics.
Approach: They propose to organize retrieved scientific documents into a navigable, two-level hierarchy of named and descriptive topics and subtopics.
Outcome: The proposed system provides an overall view of the research themes in a domain while also enabling iterative search and deeper knowledge discovery within specific subtopics.
A Web Scale Entity Extraction System (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing systems for large-scale entity extraction are limited by the scale and variety of data available on internet platforms.
Approach: They propose to build an entity extraction system for multiple document types at large scale using multi-modal Transformers.
Outcome: The proposed system extracts multiple types of entities from multiple document types at large scale using multi-modal Transformers.
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations of large language models fail to reflect fine-grained capabilities . existing benchmarks are manually curated or domain-generic, limiting scalability and alignment with real use cases.
Approach: They propose a framework that allows custom construction of benchmarks from large-scale scientific data to evaluate application-specific scientific capabilities in LLMs.
Outcome: The proposed framework reveals fine-grained differences in scientific capabilities that standard benchmarks overlook . it allows custom construction of benchmarks from large-scale scientific data to evaluate application-specific capabilities in LLMs.
Faceted Hierarchy: A New Graph Type to Organize Scientific Concepts and a Construction Method (D19-53)

Copied to clipboard

Challenge: faceted concept hierarchy is a structure of parent-child relationships . concepts are expected to be organized in a hierarchical structure for student learning .
Approach: They propose a faceted concept hierarchy that aims to build facets from scientific literature.
Outcome: The proposed hierarchy is more complete than "type-of" relations, and resolves conflicts by maintaining the acyclic structure of a hierarchy.
Will This Idea Spread Beyond Academia? Understanding Knowledge Transfer of Scientific Concepts across Text Corpora (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing research on knowledge transfer focuses on documents as unit of analysis and follow their transfer into practice for a specific scientific domain.
Approach: They analyze scientific concepts from corpora and use them to predict knowledge transfer . they find that only a small proportion of these ideas will be used in inventions .
Outcome: The proposed model predicts the use of scientific concepts in clinical trials and inventions.
Automatic Detection of Cross-Disciplinary Knowledge Associations (P18-3)

Copied to clipboard

Challenge: Currently, scientists tend to deal with fragments of the literature according to their specialisation, resulting in important and hidden associations among fragmented knowledge.
Approach: a doctoral thesis examines cross-disciplinary knowledge associations hidden in scientific literature . the aim is to identify most promising research pathways by analysing existing scientific literature.
Outcome: The proposed approach suggests most promising research pathways by analysing the existing scientific literature.
Tree-KG: An Expandable Knowledge Graph Construction Framework for Knowledge-intensive Domains (2025.acl-long)

Copied to clipboard

Challenge: Knowledge graphs are a useful tool for organizing complex data in knowledge-intensive domains.
Approach: They propose an expandable framework that combines structured domain texts with advanced semantic techniques to create a tree-like graph from textbooks.
Outcome: The proposed framework surpasses competing methods in the text-Annotated dataset with high scores on the Text-Annalytated data.
Datasets for Scientific Literature Understanding: A Survey (2026.findings-acl)

Copied to clipboard

Challenge: Empowering machines to understand scientific literature is crucial for accelerating scientific discovery and advancing the AI for Science paradigm.
Approach: They propose a systematic taxonomy that organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.
Outcome: The proposed taxonomy organizes resources spanning structural understanding, text understanding, multimodal understanding and pre-training/instruction fine-tuning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations