Challenge: Automated minuting is a rather unstructured writing activity and can be difficult due to a variety of factors including the quality of automatic speech recorders, availability of public meeting data, subjective knowledge of the minuter, etc.
Approach: They propose a dataset on automatic minuting which includes transcripts from ASRs and minuted by annotators.
Outcome: The proposed dataset covers more than 160 hours of meeting content.

Similar Papers

A Sliding-Window Approach to Automatic Creation of Meeting Minutes (2021.naacl-srw)

Copied to clipboard

Challenge: Existing methods to extract utterances and keyphrases from transcripts are lacking in meeting minutes.
Approach: They propose a sliding-window approach to automatic generation of meeting minutes . they use a neural abstractive abstractive to navigate through the raw transcript .
Outcome: The proposed approach is evaluated on natural transcripts and two versions of automatic transcripts.
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)

Copied to clipboard

Challenge: a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results .
Approach: They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach .
Outcome: The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts.
MeetingBank: A Benchmark Dataset for Meeting Summarization (2023.acl-long)

Copied to clipboard

Challenge: a lack of annotated meeting corpora hinders the development of meeting summarization technology.
Approach: They present a new benchmark dataset of city council meetings over the past decade . they use a divide-and-conquer approach to divide professionally written minutes into shorter passages .
Outcome: The proposed dataset provides a testbed for various meeting summarization systems and allows the public to gain insight into how council decisions are made.
Abstractive Meeting Summarization: A Survey (2023.tacl-1)

Copied to clipboard

Challenge: Recent advances in deep learning have improved language generation systems, opening the door to improved forms of abstractive summarization.
Approach: They propose to use neural encoder-decoder architectures to generate abstractive meeting summarizations that are particularly well-suited for multi-party conversation.
Outcome: The proposed system could be used in a wide variety of real-world contexts, from business meetings to medical consultations to customer service calls.
SlovakSum: A Large Scale Slovak Summarization Dataset (2024.lrec-main)

Copied to clipboard

Challenge: Existing datasets with hundreds and thousands of documents are mainly in the English language, but the available data is small or non-existent.
Approach: They propose to use a large Slovak news summarization dataset to evaluate its performance . the dataset contains headlines, short abstracts, and full source text .
Outcome: The proposed dataset is compared with a standard ROUGE metric and a mT5 model to evaluate its performance.
LR-Sum: Summarization for Less-Resourced Languages (2023.findings-acl)

Copied to clipboard

Challenge: LR-Sum contains human-written summaries for 40 languages, many of which are less-resourced.
Approach: They propose to use a permissively-licensed dataset to analyze human-written summaries for 40 languages.
Outcome: The proposed dataset contains human-written summaries for 40 languages . authors describe abstractive and extractive summarization experiments .
The Discussion Tracker Corpus of Collaborative Argumentation (2020.lrec-1)

Copied to clipboard

Challenge: The Discussion Tracker corpus is an annotated dataset of transcripts of spoken, multi-party argumentation transcribed from 985 minutes of audio .
Approach: They analyze 29 multi-party arguments transcribed from 985 minutes of audio . they provide descriptive statistics and code for predicting each dimension separately.
Outcome: The Discussion Tracker corpus was collected in high school English classes and annotated for argument moves, specificity, specificities and collaboration dimensions.
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)

Copied to clipboard

Challenge: Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible.
Approach: They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries.
Outcome: The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country.
DocRED: A Large-Scale Document-Level Relation Extraction Dataset (P19-1)

Copied to clipboard

Challenge: Existing relation extraction methods focus on extracting intra-sentence relations for single entities.
Approach: They propose a relation extraction dataset from Wikipedia and Wikidata with three features . document-level relation extraction is a task to identify relational facts between entities .
Outcome: The proposed dataset is the largest human-annotated dataset for document-level RE from plain text.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization (D19-54)

Copied to clipboard

Challenge: Existing work on abstractive dialogue summarizations has focused on news summarizing but there is no such comprehensive dataset.
Approach: They propose to use a chat-dialogues corpus with abstractive dialogue summaries to generate a short version of text that covers the main points succinctly.
Outcome: The proposed dataset achieves higher ROUGE scores than the model-generated summaries of news, compared with human evaluators' judgement.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations