ELITR Minuting Corpus: A Novel Dataset for Automatic Minuting from Multi-Party Meetings in English and Czech (2022.lrec-1)
Copied to clipboard
| Challenge: | Automated minuting is a rather unstructured writing activity and can be difficult due to a variety of factors including the quality of automatic speech recorders, availability of public meeting data, subjective knowledge of the minuter, etc. |
| Approach: | They propose a dataset on automatic minuting which includes transcripts from ASRs and minuted by annotators. |
| Outcome: | The proposed dataset covers more than 160 hours of meeting content. |
Similar Papers
A Sliding-Window Approach to Automatic Creation of Meeting Minutes (2021.naacl-srw)
Copied to clipboard
| Challenge: | Existing methods to extract utterances and keyphrases from transcripts are lacking in meeting minutes. |
| Approach: | They propose a sliding-window approach to automatic generation of meeting minutes . they use a neural abstractive abstractive to navigate through the raw transcript . |
| Outcome: | The proposed approach is evaluated on natural transcripts and two versions of automatic transcripts. |
Large Corpus of Czech Parliament Plenary Hearings (2020.lrec-1)
Copied to clipboard
| Challenge: | a corpus of Czech parliament plenary sessions is a valuable resource for future research . only a few public datasets are available in the Czech language . end-to-end approaches require extensive training data to produce competitive results . |
| Approach: | They present a corpus of Czech parliament plenary sessions which is a large corpus . they combine a traditional approach with a more traditional approach . |
| Outcome: | The proposed model architectures can be used to train and evaluate speech recognition systems on a large corpus of speech data and transcripts. |
MeetingBank: A Benchmark Dataset for Meeting Summarization (2023.acl-long)
Copied to clipboard
| Challenge: | a lack of annotated meeting corpora hinders the development of meeting summarization technology. |
| Approach: | They present a new benchmark dataset of city council meetings over the past decade . they use a divide-and-conquer approach to divide professionally written minutes into shorter passages . |
| Outcome: | The proposed dataset provides a testbed for various meeting summarization systems and allows the public to gain insight into how council decisions are made. |
Abstractive Meeting Summarization: A Survey (2023.tacl-1)
Copied to clipboard
| Challenge: | Recent advances in deep learning have improved language generation systems, opening the door to improved forms of abstractive summarization. |
| Approach: | They propose to use neural encoder-decoder architectures to generate abstractive meeting summarizations that are particularly well-suited for multi-party conversation. |
| Outcome: | The proposed system could be used in a wide variety of real-world contexts, from business meetings to medical consultations to customer service calls. |
SlovakSum: A Large Scale Slovak Summarization Dataset (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing datasets with hundreds and thousands of documents are mainly in the English language, but the available data is small or non-existent. |
| Approach: | They propose to use a large Slovak news summarization dataset to evaluate its performance . the dataset contains headlines, short abstracts, and full source text . |
| Outcome: | The proposed dataset is compared with a standard ROUGE metric and a mT5 model to evaluate its performance. |
LR-Sum: Summarization for Less-Resourced Languages (2023.findings-acl)
Copied to clipboard
| Challenge: | LR-Sum contains human-written summaries for 40 languages, many of which are less-resourced. |
| Approach: | They propose to use a permissively-licensed dataset to analyze human-written summaries for 40 languages. |
| Outcome: | The proposed dataset contains human-written summaries for 40 languages . authors describe abstractive and extractive summarization experiments . |
The Discussion Tracker Corpus of Collaborative Argumentation (2020.lrec-1)
Copied to clipboard
| Challenge: | The Discussion Tracker corpus is an annotated dataset of transcripts of spoken, multi-party argumentation transcribed from 985 minutes of audio . |
| Approach: | They analyze 29 multi-party arguments transcribed from 985 minutes of audio . they provide descriptive statistics and code for predicting each dimension separately. |
| Outcome: | The Discussion Tracker corpus was collected in high school English classes and annotated for argument moves, specificity, specificities and collaboration dimensions. |
Using Automatic Speech Recognition in Spoken Corpus Curation (2020.lrec-1)
Copied to clipboard
| Challenge: | Automatic Speech Recognition (ASR) is a new way to make audio-visual data accessible. |
| Approach: | They propose to use automatic speech recognition (ASR) to make audio-visual data accessible by systematic queries. |
| Outcome: | The proposed system has higher recognition scores for the north of Germany vs. lower scores for south of the country. |
DocRED: A Large-Scale Document-Level Relation Extraction Dataset (P19-1)
Copied to clipboard
Yuan Yao, Deming Ye, Peng Li, Xu Han, Yankai Lin, Zhenghao Liu, Zhiyuan Liu, Lixin Huang, Jie Zhou, Maosong Sun
| Challenge: | Existing relation extraction methods focus on extracting intra-sentence relations for single entities. |
| Approach: | They propose a relation extraction dataset from Wikipedia and Wikidata with three features . document-level relation extraction is a task to identify relational facts between entities . |
| Outcome: | The proposed dataset is the largest human-annotated dataset for document-level RE from plain text. |
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization (D19-54)
Copied to clipboard
| Challenge: | Existing work on abstractive dialogue summarizations has focused on news summarizing but there is no such comprehensive dataset. |
| Approach: | They propose to use a chat-dialogues corpus with abstractive dialogue summaries to generate a short version of text that covers the main points succinctly. |
| Outcome: | The proposed dataset achieves higher ROUGE scores than the model-generated summaries of news, compared with human evaluators' judgement. |