GLProtein: Global-and-Local Structure Aware Protein Representation Learning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Despite advances in protein sequence analysis, there remains potential for further exploration in integrating protein structural information. |
| Approach: | They propose a framework that integrates global structural similarity and local amino acid details to enhance protein pre-training. |
| Outcome: | The proposed framework outperforms existing methods in several bioinformatics tasks. |
Similar Papers
Protein Large Language Models: A Comprehensive Survey (2025.findings-emnlp)
Copied to clipboard
Yijia Xiao, Wanjia Zhao, Junkai Zhang, Yiqiao Jin, Han Zhang, Zhicheng Ren, Renliang Sun, Haixin Wang, Guancheng Wan, Pan Lu, Xiao Luo, Yu Zhang, James Zou, Yizhou Sun, Wei Wang
| Challenge: | Existing studies focus on specific aspects or applications, but this study provides a comprehensive overview of Protein-specific large language models. |
| Approach: | This paper proposes a structured taxonomy of state-of-the-art ProteinLLMs . they analyze how they leverage large-scale protein sequence data for improved accuracy . |
| Outcome: | The proposed model covers their architectures, training datasets, evaluation metrics, and diverse applications. |
Rethinking Text-based Protein Understanding: Retrieval or LLM? (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies have focused on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment. |
| Approach: | They propose a retrieval-enhanced method which significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios. |
| Outcome: | The proposed method significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios. |
LM2Protein: A Structure-to-Token Protein Large Language Model (2025.findings-emnlp)
Copied to clipboard
| Challenge: | RNA-binding proteins are critical for various molecular functions, relying on their precise tertiary structures. |
| Approach: | They propose a method to integrate protein 3D structural data within a sequence processing framework. |
| Outcome: | The proposed method achieves high sequence recovery in inverse folding and protein-conditioned RNA design. |
Protein2Text: Resampling Mechanism to Translate Protein Sequences into Human-Interpretable Text (2025.naacl-industry)
Copied to clipboard
Ala Jararweh, Oladimeji Macaulay, David Arredondo, Yue Hu, Luis E Tafoya, Kushal Virupakshappa, Avinash Sahu
| Challenge: | Proteins play critical roles in biological systems, yet 99.7% of 227 million known protein sequences remain uncharacterized due to the limitations of experimental methods. |
| Approach: | They propose a multimodal large language model that interprets protein sequences and generates informative text to address open-ended questions about protein functions and attributes. |
| Outcome: | The proposed model outperforms existing models in open-ended question-answering tasks. |
A Primer in BERTology: What We Know About How BERT Works (2020.tacl-1)
Copied to clipboard
| Challenge: | a new study examines the current state of knowledge about the BERT model . the model is a stack of transformer encoder layers that are based on multiple self-attention ''heads'' |
| Approach: | They present a survey of over 150 studies of the popular Transformer-based model BERT . they discuss the current state of knowledge about how BERT works and how it is represented . |
| Outcome: | The proposed model is based on the Transformer-based model with state-of-the-art results . the proposed model has little cognitive motivation and is too small to perform ablation studies . |
InstructProtein: Aligning Human and Protein Language via Knowledge Instruction (2024.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are a promising new approach to understanding biological sequences such as proteins. |
| Approach: | They propose an LLM that can generate protein sequences in human and protein languages by pre-training an Lm on protein and natural language corpora and supervised instruction tuning to facilitate alignment. |
| Outcome: | The proposed model outperforms state-of-the-art LLMs on protein-text generation tasks by a large margin. |
Document Structure in Long Document Transformers (2024.eacl-long)
Copied to clipboard
| Challenge: | Existing long-document Transformers do not learn representations of document structure during pretraining. |
| Approach: | They propose to use long-document Transformers to acquire an internal representation of document structure during pre-training and evaluate the effects of structure infusion on QASPER and Evidence Inference. |
| Outcome: | The proposed models acquire implicit understanding of document structure during pre-training, which can be enhanced by structure infusion, leading to improved end-task performance. |
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training (2024.acl-long)
Copied to clipboard
Le Zhuo, Zewen Chi, Minghao Xu, Heyan Huang, Jianan Zhao, Heqi Zheng, Conghui He, Xian-Ling Mao, Wentao Zhang
| Challenge: | Experimental results demonstrate that ProtLLM achieves superior performance against protein-specialized baselines on protein-centric tasks and induces zero-shot and in-context learning capabilities on protein language tasks. |
| Approach: | They propose a cross-modal large language model (LLM) that can handle protein-centric and protein-language tasks by using a dynamic protein mounting mechanism. |
| Outcome: | The proposed model can predict proteins from a vast pool of candidates and can also predict natural language and biological papers. |
ProtT3: Protein-to-Text Generation for Text-based Protein Understanding (2024.acl-long)
Copied to clipboard
| Challenge: | Language Models excel in understanding textual descriptions of proteins, but struggle to process texts. |
| Approach: | They propose a framework for Protein-to-Text Generation for Text-based Protein Understanding that integrates a PLM as its protein understanding module. |
| Outcome: | The proposed framework surpasses existing baselines and is highly efficient in protein-to-text generation. |
DocPolarBERT: A Pre-trained Model for Document Understanding with Relative Polar Coordinate Encoding of Layout Structures (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing models that take text block positions into account are not efficient for document understanding. |
| Approach: | They propose a layout-aware BERT model that takes into account text block positions in relative polar coordinate system rather than the Cartesian one. |
| Outcome: | The proposed model eliminates the need for absolute positional embeddings on a dataset more than six times smaller than the widely used IIT-CDIP corpus. |