ILCM - A Virtual Research Infrastructure for Large-Scale Qualitative Data (L18-1)
Copied to clipboard
Andreas Niekler, Arnim Bleier, Christian Kahmann, Lisa Posch, Gregor Wiedemann, Kenan Erdogan, Gerhard Heyer, Markus Strohmaier
| Challenge: | iLCM project develops integrated research environment for qualitative data analysis . text mining and text mining tools are extended by "Open Research Computing" |
| Approach: | iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture. |
| Outcome: | iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture. |
Similar Papers
Automating Qualitative Data Analysis with Large Language Models (2024.acl-srw)
Copied to clipboard
| Challenge: | Existing methods for qualitative data analysis are far from resembling a human's analysis outcome. |
| Approach: | They propose a method based on Large Language Models to tackle automated coding and make it as close as possible to the results of human researchers. |
| Outcome: | The proposed method is based on large language models and can be as close as possible to the results of human researchers. |
S2ORC: The Semantic Scholar Open Research Corpus (2020.acl-main)
Copied to clipboard
| Challenge: | Academic papers are an increasingly important textual domain for natural language processing (NLP) research. |
| Approach: | They propose to aggregate 81.1M English-language academic papers into a unified source . they hope this resource will facilitate research and development of tools for text mining over academic text. |
| Outcome: | The proposed corpus includes metadata, abstracts, bibliographic references, and structured full text for 8.1M open access papers. |
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus (2025.acl-long)
Copied to clipboard
| Challenge: | a dataset of over 1.1M podcast transcripts is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. |
| Approach: | They propose to build a large-scale open dataset of podcast transcripts that includes metadata, speaker roles, audio features and speaker turns for a subset of 370K episodes. |
| Outcome: | The proposed dataset is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. |
The D-WISE Tool Suite: Multi-Modal Machine-Learning-Powered Tools Supporting and Enhancing Digital Discourse Analysis (2023.acl-demo)
Copied to clipboard
| Challenge: | The D-WISE Tool Suite addresses limitations of current DH tools due to the ever-increasing amount of heterogeneous, unstructured, and multi-modal data in which discourses of contemporary societies are encoded. |
| Approach: | They propose to use D-WISE Tool Suite to analyze heterogeneous, unstructured, and multi-modal data in the Digital Humanities (DH) |
| Outcome: | The proposed tool leverages state-of-the-art machine learning technologies from Natural Language Processing and Com-puter Vision to ensure its usability for modernDH research. |
Argument Mining with Fine-Tuned Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Argument Mining (AM) pipelines use fine-tuned large language models (LLMs) . initial approaches employ supervised machine learning algorithms, such as Maximum Entropy classifiers and Logistic Regressions. |
| Approach: | They propose to model the three main AM sub-tasks as text generation tasks and fine-tune eight popular quantized and non-quantized large language models (LLMs) on the benchmark PE, AbstRCT, and CDCP datasets. |
| Outcome: | The proposed pipeline achieves state-of-the-art across all AM sub-tasks and datasets, showing significant improvements over previous benchmarks. |
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation. |
| Approach: | They propose a generic workflow for LLM-driven synthetic data generation. |
| Outcome: | The proposed workflows highlight gaps in existing research and outline avenues for future studies. |
INDUS: Effective and Efficient Language Models for Scientific Applications (2024.emnlp-industry)
Copied to clipboard
Bishwaranjan Bhattacharjee, Aashka Trivedi, Masayasu Muraoka, Muthukumaran Ramasubramanian, Takuma Udagawa, Iksha Gurung, Nishan Pantha, Rong Zhang, Bharath Dandala, Rahul Ramachandran, Manil Maskey, Kaylin Bugbee, Michael Little, Elizabeth Fancher, Irina Gerasimov, Armin Mehrabian, Lauren Sanders, Sylvain Costes, Sergi Blanco-Cuaresma, Kelly Lockhart, Thomas Allen, Felix Grezes, Megan Ansdell, Alberto Accomazzi, Yousef El-Kurdi, Davis Wertheimer, Birgit Pfitzmann, Cesar Berrospi Ramis, Michele Dolfi, Rafael Lima, Panagiotis Vagenas, S. Mukkavilli, Peter Staar, Sanaz Vahidinia, Ryan McGranaghan, Tsengdar Lee
| Challenge: | Large language models trained on general domain corpora showed remarkable results on natural language processing tasks. |
| Approach: | They develop a suite of large language models trained on general domain corpora that address NLP tasks and smaller versions of them created using knowledge distillation. |
| Outcome: | The proposed models outperform general-purpose and domain-specific encoders on new and existing tasks and in industrial settings. |
ET: A Workstation for Querying, Editing and Evaluating Annotated Corpora (2021.emnlp-demo)
Copied to clipboard
| Challenge: | Using annotated corpora is costly for humans alone and requires a large amount of time and effort to manipulate. |
| Approach: | They propose to use annotated corpora as a tool for linguistic research and natural language processing. |
| Outcome: | The proposed work is based on two integrated environments – Interrogatório and Julgamento . the open-source environment is used in several linguistic and NLP-related studies . |
Large Language Models for Scientific Information Extraction: An Empirical Study for Virology (2024.findings-eacl)
Copied to clipboard
| Challenge: | Scholarly communication in the digital age is facing significant challenges due to the overwhelming volume of publications. |
| Approach: | They propose to use Wikipedia infoboxes and structured Amazon product descriptions to create structured scholarly contribution summaries using text generation capabilities of LLMs. |
| Outcome: | The proposed model can be applied to complex IE tasks within terse domains like Science with 1000x fewer parameters than the state-of-the-art GPT-davinci. |
From Parameters to Performance: A Data-Driven Study on LLM Structure and Development (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large language models have revolutionized a wide range of domains, driving significant advancements in both technology and real-world applications. |
| Approach: | They present a large-scale dataset encompassing diverse open-source LLM structures and their performance across multiple benchmarks. |
| Outcome: | The proposed model validates the relationship between structural configurations and performance across multiple benchmarks and further corroborates the findings using mechanistic interpretability techniques. |