Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)
PAI-Diffusion: Constructing and Serving a Family of Open Chinese Diffusion Models for Text-to-image Synthesis on the Cloud (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing diffusion models fail to address the challenges of generating high-quality images from textual descriptions due to its large vocabulary size and complex character relationships. |
| Approach: | They propose a framework that integrates Chinese diffusion models with Alibaba Cloud's Platform for AI and enables the generation of contextually relevant images. |
| Outcome: | The proposed framework integrates with Alibaba Cloud’s Platform for AI, providing accessible and scalable solutions. |
OpenVNA: A Framework for Analyzing the Behavior of Multimodal Language Understanding System under Noisy Scenarios (2024.acl-demos)
Copied to clipboard
| Challenge: | OpenVNA is an open-source framework for analyzing the behavior of multimodal language understanding systems under noisy conditions. |
| Approach: | They propose to use OpenVNA to analyze behavior of multimodal language understanding systems under noisy conditions. |
| Outcome: | The proposed framework provides high flexibility and extensibility, enabling customization with user-defined noise types and models. |
XNLP: An Interactive Demonstration System for Universal Structured NLP (2024.acl-demos)
Copied to clipboard
| Challenge: | Structured Natural Language Processing (XNLP) is an important subset of NLP that entails understanding the underlying semantic or syntactic structure of texts. |
| Approach: | They propose a XNLP demonstration system that leverages LLM to achieve universal XnLP with one model for all with high generalizability. |
| Outcome: | The proposed system advances in multiple aspects, including universal XNLP modeling, high performance, interpretability, scalability, and interactivity. |
Towards the TopMost: A Topic Modeling System Toolkit (2024.acl-demos)
Copied to clipboard
| Challenge: | Current topic models adopt totally different datasets, implementations, and evaluations, hindering their research progress and applications. |
| Approach: | They propose a Topic Modeling System Toolkit that covers a broader spectrum of topic modeling scenarios with their complete lifecycles. |
| Outcome: | The proposed toolkit covers a broader spectrum of topic modeling scenarios with their complete lifecycles, including datasets, preprocessing, models, training, and evaluations. |
Wordflow: Social Prompt Engineering for Large Language Models (2024.acl-demos)
Copied to clipboard
| Challenge: | Large language models (LLMs) require well-crafted prompts for effective use. |
| Approach: | They propose a social prompt engineering paradigm that leverages social computing techniques to facilitate collaborative prompt design. |
| Outcome: | The proposed paradigm leverages social computing techniques to facilitate prompt design. |
LM Transparency Tool: Interactive Tool for Analyzing Transformer Language Models (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing tools focus on isolated parts of the decision-making process, but LM-TT makes the entire prediction process transparent. |
| Approach: | They present an open-source toolkit for analyzing the internal workings of Transformer-based language models. |
| Outcome: | The LM Transparency Tool makes the entire prediction process transparent . it shows the importance of specific component at each step . |
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot (2024.acl-demos)
Copied to clipboard
| Challenge: | EmpathyEar is an open-source, avatar-based multimodal empathetic chatbot . currently, ERG systems rely on text, sound, and vision . |
| Approach: | They propose an open-source, avatar-based multimodal empathetic chatbot to fill the gap in traditional text-only ERG systems. |
| Outcome: | The proposed system enables users to generate emotional responses to user queries . it can also generate avatars with talking faces and synchronized speeches . |
OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models (2024.acl-demos)
Copied to clipboard
Iat Long Iong, Xiao Liu, Yuxuan Chen, Hanyu Lai, Shuntian Yao, Pengbo Shen, Hao Yu, Yuxiao Dong, Jie Tang
| Challenge: | OpenWebAgent integrates large language models and large multimodal models to improve web automation. |
| Approach: | They propose to integrate large language models and large multimodal models into an open toolkit to optimize web automation. |
| Outcome: | The open toolkit integrates both large language models (LLMs) and large multimodal models (LMMs) it enables the development of powerful, task-oriented web agents, significantly enhancing user experience and operational efficiency on the web. |
EasyEdit: An Easy-to-use Knowledge Editing Framework for Large Language Models (2024.acl-demos)
Copied to clipboard
Peng Wang, Ningyu Zhang, Bozhong Tian, Zekun Xi, Yunzhi Yao, Ziwen Xu, Mengru Wang, Shengyu Mao, Xiaohan Wang, Siyuan Cheng, Kangwei Liu, Yuansheng Ni, Guozhou Zheng, Huajun Chen
| Challenge: | Large Language Models (LLMs) suffer from knowledge cutoff or fallacy issues, which means they are unaware of unseen events or generate text with incorrect facts owing to outdated/noisy data. |
| Approach: | They propose an easy-to-use knowledge editing framework for Large Language Models that allows users to easily edit updated knowledge and adjust undesired behavior while minimizing the impact on unrelated inputs. |
| Outcome: | The proposed framework surpasses traditional fine-tuning in terms of reliability and generalization. |
EasyInstruct: An Easy-to-use Instruction Processing Framework for Large Language Models (2024.acl-demos)
Copied to clipboard
Yixin Ou, Ningyu Zhang, Honghao Gui, Ziwen Xu, Shuofei Qiao, Runnan Fang, Lei Li, Zhen Bi, Guozhou Zheng, Huajun Chen
| Challenge: | Large Language Models (LLMs) have improved performance across tasks and domains . instruction tuning is a crucial technique to enhance the capabilities of LLMs - but there is no standard open-source instruction processing framework available for the community . |
| Approach: | They propose an open-source instruction tuning framework for Large Language Models that modularizes instruction generation, selection, prompting and their combination and interaction. |
| Outcome: | The proposed framework is open-source and available on Github. |
BotEval: Facilitating Interactive Human Evaluation (2024.acl-demos)
Copied to clipboard
| Challenge: | Using language models to perform complex interactive tasks is becoming more common with the rapid progress in natural language processing (NLP) models. |
| Approach: | They develop an evaluation toolkit that enables human-bot interactions as part of the evaluation process. |
| Outcome: | The evaluation toolkit enables human-bot interactions as part of the evaluation process, rather than making judgements for a static input. |
GenGO: ACL Paper Explorer with Semantic Features (2024.acl-demos)
Copied to clipboard
| Challenge: | Scholarly document processing (SDP) is a powerful tool for researchers to process knowledge stored in research papers. |
| Approach: | They propose a system that allows researchers to search papers published in ACL conferences with metadata and text embeddings. |
| Outcome: | The proposed system is simple and efficient to reduce maintenance and financial costs and is extensible to support open development and transparency. |
NLP-KG: A System for Exploratory Search of Scientific Literature in Natural Language Processing (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing systems for scientific literature search are typically tailored to keyword-based lookup searches, limiting possibilities for exploration. |
| Approach: | They propose a feature-rich system that supports the exploration of research literature in unfamiliar natural language processing fields. |
| Outcome: | The proposed system supports the exploration of research literature in unfamiliar natural language processing fields. |
LocalRQA: From Generating Data to Locally Training, Testing, and Deploying Retrieval-Augmented QA Systems (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing tools for augmented question-answering do not support researchers and developers to customize the training, testing, and deployment process. |
| Approach: | They propose an open-source toolkit that features a wide selection of model training algorithms, evaluation methods, and deployment tools curated from the latest research. |
| Outcome: | The proposed framework trains and deploys 7B-models with the same performance as OpenAI’s text-ada-002 and GPT-4-turbo. |
JORA: JAX Tensor-Parallel LoRA Library for Retrieval Augmented Fine-Tuning (2024.acl-demos)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) face significant memory constraints when fine-tuning large prompt sequences. |
| Approach: | They propose a framework for PEFT-compatible fine-tuning of large language models, leveraging distributed training. |
| Outcome: | The proposed framework improves performance 12x compared to Hugging Face/DeepSpeed implementation with four GPUs while consuming less than half the VRAM per GPU. |
LinguaLinked: Distributed Large Language Model Inference on Mobile Devices (2024.acl-demos)
Copied to clipboard
| Challenge: | Recent research shows that large language models demonstrate enhanced capabilities in various language tasks. |
| Approach: | They introduce a system for decentralized, distributed LLM inference on mobile devices . they use optimized model assignment technique to segment LLMs and linear optimization to align segments with each device . |
| Outcome: | The proposed system performs well on high-end to low-end Android devices. |
IMGTB: A Framework for Machine-Generated Text Detection Benchmarking (2024.acl-demos)
Copied to clipboard
| Challenge: | MGTD methods are needed in many areas, such as prevention of disinformation spreading, plagiarism, impersonation and identity theft. |
| Approach: | They propose a framework for machine-generated text detection that integrates custom methods and evaluation datasets into existing frameworks. |
| Outcome: | The proposed framework simplifies the benchmarking of machine-generated text detection methods by easy integration of custom (new) methods and evaluation datasets. |
DrugWatch: A Comprehensive Multi-Source Data Visualisation Platform for Drug Safety Information (2024.acl-demos)
Copied to clipboard
| Challenge: | Drug safety research is crucial for maintaining public health, but resources available to the public are limited. |
| Approach: | They propose an easy-to-use and interactive multi-source information visualisation platform for drug safety study. |
| Outcome: | The proposed platform provides a one-stop information analysis, retrieval, and annotation service. |
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety (2024.acl-demos)
Copied to clipboard
Chuang Liu, Linhao Yu, Jiaxuan Li, Renren Jin, Yufei Huang, Ling Shi, Junhui Zhang, Xinmeng Ji, Tingting Cui, Liutao Liutao, Jinwang Song, Hongying Zan, Sun Li, Deyi Xiong
| Challenge: | a rapid development of Chinese large language models poses big challenges for efficient LLM evaluation. |
| Approach: | They propose an evaluation testbed that benchmarks Chinese LLMs across capability, alignment and safety. |
| Outcome: | The evaluation platform OpenEval benchmarks Chinese LLMs across capability, alignment and safety. |
AutoRE: Document-Level Relation Extraction with Large Language Models (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing methods for relation extraction are limited to Sentence-level Relation Extraction (SentRE) tasks. |
| Approach: | They propose an end-to-end DocRE model that adopts a novel RE extraction paradigm named RHF (Relation-Head-Facts) Unlike existing approaches, AutoRE does not rely on the assumption of known relation options, making it more reflective of real-world scenarios. |
| Outcome: | The proposed model surpasses TAG by 10.03% and 9.03% on the dev and test set. |
LinkTransformer: A Unified Package for Record Linkage with Transformer Language Models (2024.acl-demos)
Copied to clipboard
| Challenge: | Large language models (LLMs) are used for many computational analyses, but approximate string matching packages are not widely used in social science applications. |
| Approach: | The open-source package LinkTransformer provides an end-to-end software for performing record linkage and other data cleaning tasks with transformer LLMs. |
| Outcome: | The open-source package LinkTransformer outperforms standard methods in a variety of languages and settings. |
DocPilot: Copilot for Automating PDF Edit Workflows in Documents (2024.acl-demos)
Copied to clipboard
| Challenge: | Document workflow copilot system that can understand user intent and execute tasks accordingly to help users streamline their workflows. |
| Approach: | They propose an AI-assisted document workflow copilot system capable of understanding user intent and executing tasks accordingly. |
| Outcome: | The proposed system can understand user intent and execute tasks accordingly to help users streamline their workflows. |
UltraEval: A Lightweight Platform for Flexible and Comprehensive Evaluation for LLMs (2024.acl-demos)
Copied to clipboard
Chaoqun He, Renjie Luo, Shengding Hu, Ranchi Zhao, Jie Zhou, Hanghao Wu, Jiajie Zhang, Xu Han, Zhiyuan Liu, Maosong Sun
| Challenge: | Existing evaluation platforms are complex and poorly modularized, hindering seamless incorporation into researcher’s workflows. |
| Approach: | They propose a lightweight evaluation framework characterized by lightweight, comprehensiveness, modularity, and efficiency that integrates models, data, and metrics into a unified evaluation workflow. |
| Outcome: | The proposed evaluation framework is lightweight, comprehensive, modular, and efficient. |
PyFoma: a Python finite-state compiler module (2024.acl-demos)
Copied to clipboard
| Challenge: | Finite-state models can be used to constrain output of neural networks to prevent text generation that fails to adhere to a specific format. |
| Approach: | They propose to build finite-state automata from regular expressions, string rewriting rules, right-linear grammars, or low-level state/transition manipulation. |
| Outcome: | The module is designed for teaching finite-state models and finite models. |
VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning (2024.acl-demos)
Copied to clipboard
Cheng Niu, Yang Guan, Yuanhao Wu, Juno Zhu, Juntong Song, Randy Zhong, Kaihua Zhu, Siliang Xu, Shizhe Diao, Tong Zhang
| Challenge: | generative artificial intelligence has exacerbated the challenge of distinguishing genuine news from fabricated stories. |
| Approach: | They propose a retrieval-augmented system that extracts the core facts from a given piece of news and conducts an internet-wide search to identify corroborating or conflicting reports. |
| Outcome: | The proposed system has demonstrated state-of-the-art accuracy in the realm of fake news detection. |
string2string: A Modern Python Library for String-to-String Algorithms (2024.acl-demos)
Copied to clipboard
| Challenge: | Notable algorithms include the Smith-Waterman algorithm for pairwise local alignment, the Hirschberg algorithm for global alignment, and the Wagner-Fischer algorithm for edit distance. |
| Approach: | **string2string** is an open-source library that offers efficient algorithms for string-to-string problems. |
| Outcome: | **string2string** is an open-source library that offers efficient algorithms for string-to-string problems. |
Proofread: Fixes All Errors with One Tap (2024.acl-demos)
Copied to clipboard
Renjie Liu, Yanxiang Zhang, Yun Zhu, Haicheng Sun, Yuanbo Zhang, Michael Huang, Shanqing Cai, Lei Meng, Shumin Zhai
| Challenge: | Extensive experiments on a human-labeled golden set showed our tuned PaLM2-XS model achieved 85.56% good ratio. |
| Approach: | They propose a two-stage tuning approach to acquire the dedicated Large Language Model for the feature, followed by a reinforcement learning approach for targeted refinement. |
| Outcome: | The proposed model achieves 85.56% good quality on Rewrite and proofread tasks on human-labeled golden sets. |
SeaLLMs - Large Language Models for Southeast Asia (2024.acl-demos)
Copied to clipboard
Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Zhiqiang Hu, Chenhui Shen, Yew Ken Chia, Xingxuan Li, Jianyu Wang, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, Hang Zhang, Lidong Bing
| Challenge: | Existing large language models favor high-resource languages, such as English, at the expense of low-resourced and regional languages. |
| Approach: | They propose a series of language models that specifically focuses on Southeast Asian languages. |
| Outcome: | SeaLLM models outperform ChatGPT-3.5 in non-Latin languages by large margins . linguistic disparity impedes access to state-of-the-art AI technologies for non-English-speaking populations . |
Fundus: A Simple-to-Use News Scraper Optimized for High Quality Extractions (2024.acl-demos)
Copied to clipboard
| Challenge: | Fundus is a news scraper that extracts news articles from the web with just a few lines of code. |
| Approach: | They introduce Fundus, a news scraper that enables users to obtain news articles with just a few lines of code. |
| Outcome: | The proposed news scraper optimizes for quality and provides a unified interface for newspapers. |
CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM (2024.acl-demos)
Copied to clipboard
| Challenge: | Traditional systems in this field usually accept keywords as user inputs, resulting in limited control over content. |
| Approach: | They propose a Chinese classical poetry generation system based on token-free LLMs that allow unrestricted user instructions to be used. |
| Outcome: | The proposed system outperforms traditional systems including Jiuge and GPT-4 in format accuracy and content quality. |
ITAKE: Interactive Unstructured Text Annotation and Knowledge Extraction System with LLMs and ModelOps (2024.acl-demos)
Copied to clipboard
| Challenge: | Unstructured text data contains a large amount of valuable knowledge, but there are many tools that do not meet the needs of actual business. |
| Approach: | They propose an unstructured text annotation and knowledge extraction system that integrates Large Language Models and ModelOps to improve model supervision and performance. |
| Outcome: | The proposed system integrates large language models and ModelOps to improve performance in low-resource contexts. |
LEGENT: Open Platform for Embodied Agents (2024.acl-demos)
Copied to clipboard
Zhili Cheng, Zhitong Wang, Jinyi Hu, Shengding Hu, An Liu, Yuge Tu, Pengkai Li, Lei Shi, Zhiyuan Liu, Maosong Sun
| Challenge: | Existing integrations of large language models and large multimodal models are limited . Existing platforms for developing embodied agents are limited and limited based on open-source software. |
| Approach: | They propose an open platform for developing embodied agents using LLMs and LMMs. |
| Outcome: | The proposed platform surpasses GPT-4V in embodied tasks with its model training on LEGENT data. |
Variationist: Exploring Multifaceted Variation and Bias in Written Language Data (2024.acl-demos)
Copied to clipboard
| Challenge: | Existing tools that inspect and visualize language data are limited in their capabilities. |
| Approach: | They propose a highly-modular, extensible, and task-agnostic tool that inspects language variation and bias across multiple variables, language units, and diverse metrics. |
| Outcome: | The proposed tool can inspect and visualize language variation and bias across variables, language units, and diverse metrics that go beyond descriptive statistics. |
An LLM-based Knowledge Synthesis and Scientific Reasoning Framework for Biomedical Discovery (2024.acl-demos)
Copied to clipboard
Oskar Wysocki, Magdalena.wysocka@cruk.manchester.ac.uk Magdalena.wysocka@cruk.manchester.ac.uk, Danilo Carvalho, Alex Bogatu, Danilo.miranda@idiap.ch Danilo.miranda@idiap.ch, Maxime.delmas@idiap.ch Maxime.delmas@idiap.ch, Harriet.unsworth@cruk.manchester.ac.uk Harriet.unsworth@cruk.manchester.ac.uk, Andre Freitas
| Challenge: | BioLunar integrates Large Language Models (LLMs) to facilitate scientific reasoning across distributed evidence spaces. |
| Approach: | They present a tool for supporting biological analyses using Large Language Models (LLMs) it integrates large language models to facilitate scientific reasoning across distributed evidence spaces . |
| Outcome: | The proposed tool exemplifies the potential of the integration between LLMs, specialised databases and biomedical tools to support expert-level knowledge synthesis and discovery. |
CogMG: Collaborative Augmentation Between Large Language Model and Knowledge Graph (2024.acl-demos)
Copied to clipboard
| Challenge: | Large language models (LLMs) are susceptible to generating hallucinated content and often encompass factually inaccurate information. |
| Approach: | They propose a framework that leverages knowledge graphs to address the limitations of Large Language Models (LLMs) they identify and decompose required knowledge triples that are not present in the KG, enriching them and aligning updates with real-world demands. |
| Outcome: | The proposed framework reduces hallucinations and increases factual accuracy in QA scenarios while retaining the same quality of knowledge. |
ELLA: Empowering LLMs for Interpretable, Accurate and Informative Legal Advice (2024.acl-demos)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have shown impressive performance in various tasks, showing great potential for specific domains, such as law (Lai et al., 2023), finance (Zeng e e al. 2023) and law (Lam elms, 2024). |
| Approach: | They propose to use large language models to provide interpretable, accurate, and informative legal advice by visually presenting the correlation between legal articles and LLM's response by calculating their similarities. |
| Outcome: | The proposed model provides users with an intuitive legal basis for the responses and retrieves relevant legal cases for user reference. |
LLMBox: A Comprehensive Library for Large Language Models (2024.acl-demos)
Copied to clipboard
Tianyi Tang, Hu Yiwen, Bingqian Li, Wenyang Luo, ZiJing Qin, Haoxiang Sun, Jiapeng Wang, Shiyi Xu, Xiaoxue Cheng, Geyang Guo, Han Peng, Bowen Zheng, Yiru Tang, Yingqian Min, Yushuo Chen, Jie Chen, Ranchi Zhao, Luran Ding, Yuhao Wang, Zican Dong, Xia Chunxuan, Junyi Li, Kun Zhou, Xin Zhao, Ji-Rong Wen
| Challenge: | a library to facilitate the development, use, and evaluation of large language models (LLMs) is presented. |
| Approach: | They propose a unified library to facilitate the development, use and evaluation of large language models (LLMs). |
| Outcome: | The proposed library is based on extensive experiments in a variety of evaluation settings. |
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models (2024.acl-demos)
Copied to clipboard
| Challenge: | Efficient fine-tuning of large language models requires non-trivial efforts to implement these methods on different models. |
| Approach: | They propose a framework that democratizes the fine-tuning of large language models by integrating a suite of efficient training methods into one framework. |
| Outcome: | The proposed framework is able to scale to 100+ LLMs without coding and receives over 25,000 stars and 3,000 forks. |