Chemical Compounds Knowledge Visualization with Natural Language Processing and Linked Data (L18-1)
Copied to clipboard
| Challenge: | Existing systems for chemical compounds extraction and registration depend on human labor . CAS databases are being created, but information written in other languages is not exploited well . |
| Approach: | They propose a visualization system for chemical compounds extracted from Japanese texts and chemical compound databases represented as Linked Data (LD) system integrates extracted results with existing chemical compound knowledge to provide different views of chemical compounds. |
| Outcome: | The proposed system integrates extraction results with existing chemical compound knowledge to provide different views of chemical compounds. |
Similar Papers
NLP for Chemistry – Introduction and Recent Advances (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | This tutorial will provide an introductory overview to a relatively underrepresented application domain: chemistry. |
| Approach: | This tutorial will provide an introductory overview to a number of recent applications of natural language processing to chemistry. |
| Outcome: | This tutorial will provide an overview of the latest applications of natural language processing to chemistry. |
From Generalist to Specialist: A Survey of Large Language Models for Chemistry (2025.coling-main)
Copied to clipboard
| Challenge: | Existing studies on pretraining of LLMs on extensive web-based texts are insufficient for advanced scientific discovery, especially in chemistry. |
| Approach: | They outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs and conceptualize chemistry LLM agents using chemistry tools. |
| Outcome: | The proposed models are based on domain-specific chemistry knowledge and multi-modal information and are capable of accelerating scientific research. |
Representing Compounding with OntoLex. An Evaluation of Vocabularies for Word Formation Resources (2024.lrec-main)
Copied to clipboard
| Challenge: | OntoLex is a de facto standard for the modelling of lexical resources in the framework of Linguistic Linked Open Data. |
| Approach: | They propose to use OntoLex to convert Linked Open Data into compounds by using the RDF model. |
| Outcome: | The proposed model can be applied to all resources harmonized in that format, potentially allowing for the conversion into Linked Open Data of a large amount of structured data. |
Language + Molecules (2024.eacl-tutorials)
Copied to clipboard
| Challenge: | In the last year, instruction-following language models have surged in popularity. |
| Approach: | This tutorial will provide an introduction to applying natural language-driven solutions to chemistry problems. |
| Outcome: | This tutorial will provide an introduction to this area of research. it requires no knowledge outside mainstream NLP, and it will enable participants to begin exploring relevant research. |
MolTRES: Improving Chemical Language Representation Learning for Molecular Property Prediction (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for chemical representation learning often lead to overfitting and limited scalability due to early convergence. |
| Approach: | They propose a framework to train Transformers on SMILES sequences to learn from structural examples and integrate external materials embedding to enrich molecular representations. |
| Outcome: | The proposed model outperforms state-of-the-art models on molecular property prediction tasks. |
Transformer-based Approach for Predicting Chemical Compound Structures (2020.aacl-main)
Copied to clipboard
| Challenge: | Existing methods to predict chemical compound structures from their names are limited and use handcrafted rules. |
| Approach: | They propose a Transformer-based model that predicts SMILES strings from chemical compound names instead of handcrafted rules. |
| Outcome: | The proposed model achieves higher F-measures than the existing model and the existing one. |
Lost in Translation: Chemical Language Models and the Misunderstanding of Molecule Structures (2024.findings-emnlp)
Copied to clipboard
Veronika Ganeeva, Andrey Sakhovskiy, Kuzma Khrabrov, Andrey Savchenko, Artur Kadurin, Elena Tutubalina
| Challenge: | chemistry and natural language processing (NLP) have advanced drug discovery. |
| Approach: | They propose a framework for assessment of Chemistry LMs of different natures that relies on augmentations that preserve an underlying chemical. |
| Outcome: | The proposed framework relies on augmentations that preserve an underlying chemical, such as kekulization and cycle replacements. |
Massively Translingual Compound Analysis and Translation Discovery (L18-1)
Copied to clipboard
| Challenge: | Morphological compounding is one of the most common and productive methods of word formation across the world's languages. |
| Approach: | They propose a model for compounding using bilingual dictionaries and no annotated training data . they also release a massively multilingual dataset of compound words and their decompositions . |
| Outcome: | The proposed model generates novel translations of English concepts on a multilingual dataset . the model can be applied to a wide range of languages and is highly reproducible. |
Reaction Miner: An Integrated System for Chemical Reaction Extraction from Textual Data (2023.emnlp-demo)
Copied to clipboard
Ming Zhong, Siru Ouyang, Yizhu Jiao, Priyanka Kargupta, Leo Luo, Yanzhen Shen, Bobby Zhou, Xianrui Zhong, Xuan Liu, Hongxiang Li, Jinfeng Xiao, Minhao Jiang, Vivian Hu, Xuan Wang, Heng Ji, Martin Burke, Huimin Zhao, Jiawei Han
| Challenge: | Reaction Miner is a system designed to extract chemical reactions from raw scientific PDFs. |
| Approach: | They propose a system that extracts chemical reactions directly from raw scientific PDFs. |
| Outcome: | The proposed system can extract chemical reactions from raw scientific PDFs. |
A Survey of Large Language Models for Text-Guided Molecular Discovery: From Molecule Generation to Optimization (2026.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are introducing a paradigm shift in molecular discovery by enabling text-guided interaction with chemical spaces through natural language and symbolic notations. |
| Approach: | They analyze the current LLM learning paradigms to tackle four critical evaluation dimensions that have emerged as critical dimensions in recent studies. |
| Outcome: | The proposed models are able to interact with chemical spaces through natural language and symbolic notations, and have emerging extensions to incorporate multi-modal inputs. |