Challenge: iLCM project develops integrated research environment for qualitative data analysis . text mining and text mining tools are extended by "Open Research Computing"
Approach: iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture.
Outcome: iLCM project develops integrated research environment for analysis of structured and unstructured data in a "Software as a Service" architecture.

Similar Papers

Automating Qualitative Data Analysis with Large Language Models (2024.acl-srw)

Copied to clipboard

Challenge: Existing methods for qualitative data analysis are far from resembling a human's analysis outcome.
Approach: They propose a method based on Large Language Models to tackle automated coding and make it as close as possible to the results of human researchers.
Outcome: The proposed method is based on large language models and can be as close as possible to the results of human researchers.
S2ORC: The Semantic Scholar Open Research Corpus (2020.acl-main)

Copied to clipboard

Challenge: Academic papers are an increasingly important textual domain for natural language processing (NLP) research.
Approach: They propose to aggregate 81.1M English-language academic papers into a unified source . they hope this resource will facilitate research and development of tools for text mining over academic text.
Outcome: The proposed corpus includes metadata, abstracts, bibliographic references, and structured full text for 8.1M open access papers.
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus (2025.acl-long)

Copied to clipboard

Challenge: a dataset of over 1.1M podcast transcripts is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020.
Approach: They propose to build a large-scale open dataset of podcast transcripts that includes metadata, speaker roles, audio features and speaker turns for a subset of 370K episodes.
Outcome: The proposed dataset is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020.
The D-WISE Tool Suite: Multi-Modal Machine-Learning-Powered Tools Supporting and Enhancing Digital Discourse Analysis (2023.acl-demo)

Copied to clipboard

Challenge: The D-WISE Tool Suite addresses limitations of current DH tools due to the ever-increasing amount of heterogeneous, unstructured, and multi-modal data in which discourses of contemporary societies are encoded.
Approach: They propose to use D-WISE Tool Suite to analyze heterogeneous, unstructured, and multi-modal data in the Digital Humanities (DH)
Outcome: The proposed tool leverages state-of-the-art machine learning technologies from Natural Language Processing and Com-puter Vision to ensure its usability for modernDH research.
Argument Mining with Fine-Tuned Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Argument Mining (AM) pipelines use fine-tuned large language models (LLMs) . initial approaches employ supervised machine learning algorithms, such as Maximum Entropy classifiers and Logistic Regressions.
Approach: They propose to model the three main AM sub-tasks as text generation tasks and fine-tune eight popular quantized and non-quantized large language models (LLMs) on the benchmark PE, AbstRCT, and CDCP datasets.
Outcome: The proposed pipeline achieves state-of-the-art across all AM sub-tasks and datasets, showing significant improvements over previous benchmarks.
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) provide a data-centric solution to alleviate limitations of real-world data with synthetic data generation.
Approach: They propose a generic workflow for LLM-driven synthetic data generation.
Outcome: The proposed workflows highlight gaps in existing research and outline avenues for future studies.
ET: A Workstation for Querying, Editing and Evaluating Annotated Corpora (2021.emnlp-demo)

Copied to clipboard

Challenge: Using annotated corpora is costly for humans alone and requires a large amount of time and effort to manipulate.
Approach: They propose to use annotated corpora as a tool for linguistic research and natural language processing.
Outcome: The proposed work is based on two integrated environments – Interrogatório and Julgamento . the open-source environment is used in several linguistic and NLP-related studies .
Large Language Models for Scientific Information Extraction: An Empirical Study for Virology (2024.findings-eacl)

Copied to clipboard

Challenge: Scholarly communication in the digital age is facing significant challenges due to the overwhelming volume of publications.
Approach: They propose to use Wikipedia infoboxes and structured Amazon product descriptions to create structured scholarly contribution summaries using text generation capabilities of LLMs.
Outcome: The proposed model can be applied to complex IE tasks within terse domains like Science with 1000x fewer parameters than the state-of-the-art GPT-davinci.
From Parameters to Performance: A Data-Driven Study on LLM Structure and Development (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models have revolutionized a wide range of domains, driving significant advancements in both technology and real-world applications.
Approach: They present a large-scale dataset encompassing diverse open-source LLM structures and their performance across multiple benchmarks.
Outcome: The proposed model validates the relationship between structural configurations and performance across multiple benchmarks and further corroborates the findings using mechanistic interpretability techniques.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations