Papers with in-house

9 papers
Computer Assisted Annotation of Tension Development in TED Talks through Crowdsourcing (D19-59)

Copied to clipboard

Challenge: Using a neural network, we annotate whether tension is increasing, decreasing, or staying unchanged.
Approach: They propose a machine-assisted method for the identification of tension development using a neural network based prediction model.
Outcome: The proposed method is compared with other methods in in-house and crowdsourced environments.
TaskMix: Data Augmentation for Meta-Learning of Spoken Intent Understanding (2022.findings-aacl)

Copied to clipboard

Challenge: Meta-Learning requires a large number of training tasks to learn representations that transfer well to unseen tasks.
Approach: They propose a method which synthesizes new tasks by linearly interpolating existing tasks.
Outcome: The proposed method outperforms baselines and does not degrade performance even when it is high.
Improve Speech Translation Through Text Rewrite (2025.coling-industry)

Copied to clipboard

Challenge: Recent advances in speech translation (ST) research have focused on the unique characteristics of spontaneous speech, including accents and presentation quality.
Approach: They propose to transform transcribed speech into a cleaner style more in line with the expectations of translation models built from written text.
Outcome: Experiments on public and in-house translation models show that the proposed model can be effectively distilled into a standalone translation model.
Neural Cross-Lingual Relation Extraction Based on Bilingual Word Embedding Mapping (D19-1)

Copied to clipboard

Challenge: Relation extraction (RE) is an important information extraction task that seeks to detect and classify semantic relationships between entities.
Approach: They propose a bilingual word embedding mapping approach for cross-lingual RE model transfer . they use a small bilingual dictionary with only 1K word pairs to embed word pairs .
Outcome: The proposed approach achieves very good performance on target and target languages . it uses bilingual word embedding mapping to transfer a source-language model .
Multimodal Context Carryover (2022.emnlp-industry)

Copied to clipboard

Challenge: Existing voice-only dialogue systems lack multimodality support, which can lead to costly system redesigns.
Approach: They propose to augment existing voice-only dialogue systems with additional multimodal components to facilitate quick delivery of visual modality support with minimal changes.
Outcome: The proposed framework improves visual modality support with minimal changes on an in-house multi-modal visual navigation data set.
Probing the Depths of Language Models’ Contact-Center Knowledge for Quality Assurance (2024.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in large Language Models (LMs) have significantly enhanced their capabilities across various domains, including natural language understanding and domain knowledge.
Approach: They propose methods to transfer domain-specific knowledge to smaller models by leveraging evaluation plans generated by more knowledgeable models with optional human-in-the-loop refinement to enhance the capabilities of smaller models.
Outcome: The proposed models improve 18.95% on an in-house QA dataset on a contact-center quality assurance task.
Task-Driven and Experience-Based Question Answering Corpus for In-Home Robot Application in the House3D Virtual Environment (2022.lrec-1)

Copied to clipboard

Challenge: Question answering is an important part of natural language processing (NLP)
Approach: They propose to use TEQA to investigate the ability of agent task experience understanding for the long-term household task.
Outcome: The proposed corpus aims to investigate the ability of task experience understanding of agents for the daily question answering scenario on the ALFRED dataset.
RSC: A Romanian Read Speech Corpus for Automatic Speech Recognition (2020.lrec-1)

Copied to clipboard

Challenge: Romanian language is under-resourced due to the lack of acoustic and linguistic resources.
Approach: They propose to use a Romanian speech corpus to train automatic speech recognition algorithms based on the spoken hotword detection mechanism.
Outcome: The read speech corpus is a speech recognition system that can perform automatic speech recognition and speech synthesis using state-of-the-art speech recognition toolkit.
KazQAD: Kazakh Open-Domain Question Answering Dataset (2024.lrec-main)

Copied to clipboard

Challenge: KazQAD contains just under 6,000 unique questions with extracted short answers and nearly 12,000 passage-level relevance judgements.
Approach: They introduce a Kazakh open-domain question answering dataset that can be used in reading comprehension and full ODQA settings.
Outcome: The proposed dataset can be used in reading comprehension and full ODQA settings, as well as for information retrieval experiments.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations