Papers with TTS
Copied to clipboard
| Challenge: | Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
| Approach: | Silver es provides high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
| Outcome: | Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
Copied to clipboard
| Challenge: | SpiRit-LM is a foundation multimodal language model that freely mixes text and speech. |
| Approach: | They propose a multimodal language model that freely mixes text and speech . they extend the model to the speech modality by continuously training it on text and language units. |
| Outcome: | The proposed model can learn new tasks in a few-shot fashion across modalities. |
Copied to clipboard
| Challenge: | Autoregressive (AR) models can only generate target sequence word-by-word due to the AR mechanism and suffer from slow inference. |
| Approach: | This tutorial provides an introduction to non-autoregressive sequence generation. |
| Outcome: | This tutorial explains how to generate non-autoregressive sequence generation models. |
Copied to clipboard
| Challenge: | Neural codec language models (or codec LMs) are emerging as a powerful framework for text-to-speech (TTS) despite the close interdependence of codecs and LM, research on codec and lms has largely remained siloed. |
| Approach: | They propose a frame-wise codec encoder that improves both LM log-likelihood and TTS metrics . they also propose LM codebook level dropout to efficiently navigate a portion of codec-LM design space . |
| Outcome: | The proposed codec-LM co-design improves intelligibility, audio quality and speaker control compared to a siloed baseline. |
Copied to clipboard
| Challenge: | Developing Text Normalization systems for Text-to-Speech (TTS) on new languages is hard. |
| Approach: | They propose a novel architecture to facilitate Text Normalization systems for TTS on new languages . they use a granular tokenization mechanism that enables the system to learn majority of classes . |
| Outcome: | The proposed architecture performs comparable with the state-of-the-art systems on English . the proposed system learns most classes from training data and precodes them for other classes . |
Copied to clipboard
| Challenge: | Non-autoregressive (NAR) models generate all tokens in parallel, resulting in faster generation speed compared to autoregressive models. |
| Approach: | They propose to use knowledge distillation and source-target alignment to bridge the gap between NAR and autoregressive models in various tasks. |
| Outcome: | The proposed techniques can speed up NAR models in some tasks but not all . the proposed techniques reduce target token dependency while allowing for faster inference . |
Copied to clipboard
| Challenge: | Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over speech attributes. |
| Approach: | They provide a review of controllable TTS methods from traditional control techniques to emerging approaches using natural language prompts. |
| Outcome: | The proposed methods are based on models, strategies, and features, and summarize challenges, datasets, and evaluations. |
Copied to clipboard
| Challenge: | Existing approaches to speech-to-speech translation rely on cascaded pipelines . current approaches rely only on text representations, but they suffer from errors and latency . a new direct speech translation framework is proposed to bridge linguistic gaps . |
| Approach: | They propose a sequence-to-sequence direct speech translation framework that can translate speech from one Indian language to another without relying on intermediate text representations. |
| Outcome: | The proposed framework can translate speech from one Indian language to another without relying on intermediate text representations. |
Copied to clipboard
| Challenge: | Using self-training or text-to-speech (TTS) to improve low-resource ASR performance is costly and can lead to catastrophic forgetting. |
| Approach: | They examine whether data augmentation techniques could help improve low-resource ASR performance . they use self-training to generate transcriptions, which are combined with original data to train new system . |
| Outcome: | The proposed approach yields a 20.5% reduction in WER compared to a system trained on 24 minutes of manually transcribed speech. |
Copied to clipboard
| Challenge: | a new framework automates deployment and debugging of AI projects . complexity of environment configurations, dependency conflicts, and debuggering issues hinder scalability and adoption. |
| Approach: | They propose an end-to-end framework that automates AI project deployment . they conducted experiments on 30 AI deployment cases to evaluate its effectiveness . |
| Outcome: | The proposed framework reduces deployment time and improves success rates by reducing human intervention. |
Copied to clipboard
| Challenge: | Currently, the amount of data needed for TTS is limited to the vast majority of the spoken languages. |
| Approach: | They propose to use language agnostic meta learning procedure to learn speaking a new language with just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers. |
| Outcome: | The proposed approach is able to learn speaking a new language using just 5 minutes of training data while retaining the ability to infer the voice of even unseen speakers in the newly learned language. |
Copied to clipboard
| Challenge: | MELLE is a novel language modeling approach for text-to-speech synthesis that generates continuous tokens from text . authors demonstrate that it reduces the need for vector quantization and improves model robustness . |
| Approach: | They propose to autoregressively generate continuous mel-spectrogram frames directly from text condition, bypassing vector quantization. |
| Outcome: | The proposed model achieves superior performance across multiple metrics and is more streamlined. |
Copied to clipboard
| Challenge: | Neural text-to-speech (TTS) models typically rely on extensive transcribed speech datasets and intricate training pipelines. |
| Approach: | They propose a framework for zero-shot multi-speaker text-to-speech using retrieval methods which leverage the linear relationships between SSL features. |
| Outcome: | The proposed framework achieves comparable performance to state-of-the-art models trained on large training datasets. |
Copied to clipboard
| Challenge: | Text-to-speech (TTS) models require data availability and quality of training data. |
| Approach: | They propose an end-to-end tool to generate high-quality datasets for text-to speech models . language-specific phoneme distribution is integrated into sample selection, they argue . |
| Outcome: | The proposed tool aims to streamline the dataset creation process for voice-based technologies by integrating language-specific phonemes into sample selection and quality assurance of recordings. |
Copied to clipboard
| Challenge: | a new spoken dialogue system with single-stage training is demonstrating its low latency and high quality . SLAM-Omni achieves zero-shot timbre control by modeling spoken language with semantic tokens . |
| Approach: | They propose a timbre-controllable, end-to-end voice interaction system with single-stage training. |
| Outcome: | The proposed system outperforms previous models on 4 GPUs with limited data. |
Copied to clipboard
| Challenge: | Using TTS, Reasoning Models (RMs) are able to perform tasks such as math and coding with limited results. |
| Approach: | They evaluate 12 Reasoning Models across a diverse suite of MT benchmarks, examining three scenarios: direct translation, forced-reasoning extrapolation, and post-editing. |
| Outcome: | The proposed approach improves translation quality on three domains, with inconsistent results for general-purpose RMs and performance plateauing. |
Copied to clipboard
| Challenge: | a recent study has shown that text-to-speech systems can capture human-like emotion, but they lack the ability to predict emotion in speech. |
| Approach: | They propose to use 8 large language models for identifying emotion in text and 2 audio models for emotion in speech to investigate the correlation between emotion and speech. |
| Outcome: | The proposed models perform well on emotion recognition from situational text and audiobooks, but show weak correlation for Valence only. |
Copied to clipboard
| Challenge: | End-to-end speech-totext translation (ST) models require large amounts of data to train, but their size is considerably smaller than text-based MT data. |
| Approach: | They propose a method to convert MT data to ST data via text-to-speech systems. |
| Outcome: | The proposed method improves translation quality by an average of 1.83 BLEU score while performing equally well as TTS-generated speech in improving translation quality. |
Copied to clipboard
| Challenge: | Existing methods for generating speech from facial images rely on pre-trained visual encoders and fine-tune them to align with speech embeddings. |
| Approach: | They propose to derive corresponding voices from facial images using face-to-voice synthesis, which derives corresponding voice from facial image. |
| Outcome: | The proposed approach significantly improves face-voice congruence and synthesis stability. |
Copied to clipboard
| Challenge: | Recent advances in text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre. |
| Approach: | They propose a zero-shot text-to-speech framework that enables disentangled control of style and timbre by conditioning on two distinct reference utterances. |
| Outcome: | The proposed framework achieves high-fidelity synthesis and competitive zero-shot naturalness while supporting consistent and independent manipulation of style and timbre. |
Copied to clipboard
| Challenge: | Injecting fillers into spoken dialogue systems has a rich history of study . ambiguity of filler occurrence and inter-speaker difference make modeling and evaluation difficult. |
| Approach: | They propose an objective score for filler insertion using sampling-based sampling . they build three models trained on two single-speaker spontaneous corpora and evaluate them with FPP and perceptual tests. |
| Outcome: | The proposed model is useful in analysis but does not correlate well with perceptual MOS. |
Copied to clipboard
| Challenge: | a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented. |
| Approach: | They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods. |
| Outcome: | The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents. |
Copied to clipboard
| Challenge: | Grapheme-to-phoneme conversion datasets suffer from the long-tail problem . context learning for polyphonic characters often stems from a single dimension . |
| Approach: | They propose a model for long-tailed polyphone disambiguation in Mandarin that decouples representation and classification learnings. |
| Outcome: | The proposed model can decouple representation and classification learnings . it achieves transition learning of context from local to global . |
Copied to clipboard
| Challenge: | Recent advances in text-to-speech systems have been driven by large, multi-domain speech corpora. |
| Approach: | They propose a large-scale 5.3K hours of expressive speech drawn from character quotations . they fine-tune a flow-matching model and train from scratch . |
| Outcome: | The proposed model improves expressivity and intelligibility while training from scratch improves expressiveness of an autoregressive model. |
Copied to clipboard
| Challenge: | Variational Auto-Encoders (VAEs) for learning disentangled latent representations in speech fail to learn latent clusters of speaker attributes when trained on limited or noisy datasets. |
| Approach: | They propose a Variational Auto-Encoder (VAE) that minimizes mutual information between latent variables and learns controllable latent representations in speech data. |
| Outcome: | The proposed model reduces the cluster overlap of speaker attributes by 30% over LSTM-VAE. |
Copied to clipboard
| Challenge: | A popular idea in Computer Assisted Language Learning (CALL) is to use multimodal annotated texts to support reading. |
| Approach: | They propose to use an open source platform to create good quality audio for L2 learning . they use four passages from LARA versions of Saint-Exupèry’s “Le petit prince” to instantiate the 2x2 cross product of dialogue, not-dialogue and humour, not humor. |
| Outcome: | The proposed method is based on a web form and ten languages. |
Copied to clipboard
| Challenge: | Text-to-speech synthesis (TTS) has seen rapid progress in recent years, but still suffers from latencies. |
| Approach: | They propose a neural incremental TTS approach that synthesizes speech in an online fashion, playing a segment of audio while generating the next. |
| Outcome: | Experiments on English and Chinese TTS show that the proposed approach achieves similar speech naturalness compared to full sentence TTS, but with a constant (1-2 words) latency. |
Copied to clipboard
| Challenge: | a text-to-speech voice building tool is available for low-resourced languages . the tool allows researchers to run voice training experiments and listen to the resulting voice . |
| Approach: | They propose an opensource text-to-speech (TTS) voice building tool that focuses on simplicity, flexibility, and collaboration. |
| Outcome: | The proposed tool can help improve TTS research especially for low-resourced languages . it can be used to run voice training experiments and listen to the resulting synthesized voice . |
Copied to clipboard
| Challenge: | Recent advances have adapted this paradigm to Multimodal Foundation Models (MFMs), unlocking their potential in multimodal reasoning and generation. |
| Approach: | They propose a taxonomy framework that categorizes existing methodologies into three distinct strategies: sampling-based, feedback-based and search-based approaches. |
| Outcome: | The proposed framework categorizes existing methodologies into three distinct strategies: sampling-based, feedback-based and search-based approaches. |
Copied to clipboard
| Challenge: | a recent article outlines five projects that address contemporary Romanian language . the authors argue that a constant accumulation of human expertise is needed to develop complex projects. |
| Approach: | a new article gives a general overview of five AI language-related projects at the Romanian Academy . they focus on the creation of a contemporary Romanian language text and speech corpus and language related applications . |
| Outcome: | a new article gives an overview of five AI language-related projects at the Romanian Academy . the projects address contemporary Romanian language, as well as language related applications . |
Copied to clipboard
| Challenge: | Existing two-pass direct speech-to-speech translation models require parallel speech data to train, which is challenging to collect. |
| Approach: | They propose a two-pass direct speech-to-speech translation (S2ST) model that decomposes the task into speech- to-text translation (s2TT) and text-tospech (TTS) they propose 'composer' S2ST model that integrates pretrained S2TT and TTS models into a direct S2 ST model. |
| Outcome: | The proposed model integrates pretrained S2TT and TTS models into a direct S2ST model without parallel speech data. |
Copied to clipboard
| Challenge: | The current state-of-the-art training procedures involve retired pilots that train as virtual plane pilots and process the spoken prompts to form that can be entered into software that simulates the plane movement on the radar screen. |
| Approach: | They describe the process of creating domain-specific speech corpora containing air traffic control (ATC) communication prompts. |
| Outcome: | The proposed system could be used for training air traffic controllers in the Czech Republic. |
Copied to clipboard
| Challenge: | Existing methods for expressive text-to-speech only implicitly learn prosody with masked token reconstruction tasks. |
| Approach: | They propose a cross-modal contrastive pre-training framework that learns from prosody variance of the same text token under different contexts. |
| Outcome: | The proposed framework can learn from prosody variance of a text token under different contexts. |
Copied to clipboard
| Challenge: | Text-to-speech systems that scale up the amount of training data have certain limitations: they require a large amount of data, which increases costs, and overlook prosody similarity. |
| Approach: | They propose a zero-shot multi-task TTS system that can perform TTS or speech style transfer in zero- shot and cross-lingual conditions. |
| Outcome: | The proposed system outperforms other TTS systems trained with the same small amount of data and achieves zero-shot performance comparable to data-driven systems. |
Copied to clipboard
| Challenge: | Existing verification approaches, such as Process Reward Models, are computationally expensive and limited to specific domains. |
| Approach: | They propose a transformer-based probe that uses internal states of frozen LLMs to estimate credibility of reasoning steps during generation. |
| Outcome: | The proposed probes match or exceed PRMs that are up to 810 larger. |
Copied to clipboard
| Challenge: | Recent advances in generative language modeling applied to discrete speech tokens presented a new avenue for text-to-speech (TTS) synthesis. |
| Approach: | They propose to use generative language modeling to generate text-to-speech (TTS) outputs by a discrete token-based model. |
| Outcome: | The proposed model is rated higher in naturalness and context appropriateness in listening tests compared to a conventional TTS. |
Copied to clipboard
| Challenge: | Autoregressive (AR) Transformer-based sequence models have difficulty generalizing to sequences longer than those seen during training. |
| Approach: | They propose a system that provides cross-attention operations with relative location information. |
| Outcome: | The proposed system matches the naturalness and expressiveness of a baseline T5-based system while eliminating problems with repeated or dropped words. |
Copied to clipboard
| Challenge: | VoiceCraft is a token-infilling neural codec language model for speech editing and zero-shot text-to-speech evaluation. |
| Approach: | They introduce a token infilling neural codec language model that performs on speech editing and zero-shot text-to-speech tasks. |
| Outcome: | The proposed model outperforms previous models on speech editing and zero-shot text-to-speech tasks. |
Copied to clipboard
| Challenge: | Existing multilingual TTS datasets are limited in speech generation fields due to lack of quality data. |
| Approach: | They propose to use 30,000 hours of high-quality speech data across 3 languages . they filter out low-quality text-text pairs and concatenate short transcripts . |
| Outcome: | The proposed dataset comprises 30,000 hours of high-quality speech data, across 3 languages with multiple speakers and styles, suitable for various speech tasks such as TTS and ASR. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have rapidly advanced on coding tasks, enabling increasingly capable software engineering agents for real-time code editing and bug fixing. |
| Approach: | They propose to use a rubric checklist to create a context-grounded rubric for SWE agents. |
| Outcome: | The proposed rubrics achieve a score of 54.2% on Qwen3-Coder-30B-A3B and 40.6% on Qween3-332B . |
Copied to clipboard
| Challenge: | Existing work on speech-to-speech translation (S2ST) systems rely on text representation, but they are text-centric. |
| Approach: | They introduce a massively multilingual-to-English speech-tospeech translation corpus . they synthesize the translation text from the Common Voice speech corpus and CoVoST 2 into English . |
| Outcome: | The proposed corpus outperforms existing models on CoVoST 2 by 5.8 BLEU . the proposed model outperformed the previous state-of-the-art model without extra data . |
Copied to clipboard
| Challenge: | Existing zero-shot text-to-speech systems require a few seconds of unseen speaker voice prompts to generate high-quality voices. |
| Approach: | They propose a zero-shot text-to-speech system based on mobile devices . they use a discrete speech codec to integrate hierarchical information from the codec . |
| Outcome: | The proposed system achieves RTF of 0.09 on a single A100 GPU and has been successfully deployed on mobile devices. |
Copied to clipboard
| Challenge: | Guided by Gut (GG) is an efficient self-guided TTS framework for Large Language Models (LLMs) that performs step-by-step reasoning at a low cost without any reward models or verifiers. |
| Approach: | They propose a self-guided TTS framework that enables LLMs to perform step-by-step reasoning at a low cost without any reward models or verifiers. |
| Outcome: | Empirical evaluations show that GG performs better than TTS with PRMs while reducing GPU memory usage by up to 10. |
Copied to clipboard
| Challenge: | Existing systems lack a self-emotion determination mechanism to drive the streaming text-to-speech (TTS) synthesis. |
| Approach: | They propose an emotion-planning framework that determines the emotion prior to the textual generation, grounding the downstream emotional TTS in a streaming manner. |
| Outcome: | The proposed framework outperforms baselines on DailyDialog, EmoryNLP, IMEOCAP, and MELD on emotional alignment, contextual coherence, and expressive fluency. |
Copied to clipboard
| Challenge: | Neural text-to-speech (TTS) systems limited to predefined speaker styles or specific sets of speaker IDs. |
| Approach: | They propose a network that can adapt adapter parameters to new speakers . they compare two domain adaptation settings and find it to be very efficient . |
| Outcome: | The proposed Adapters improve speech synthesis performance on two domains and compare them with baselines. |
Copied to clipboard
| Challenge: | Spontaneous speech is unscripted and created on the fly by the speaker, whereas read speech is pre-planned. |
| Approach: | They propose a tool that allows developers to select a varied, representative set of utterances from a spoken genre to be used for evaluation of TTS for a given domain. |
| Outcome: | The proposed tool can be used to evaluate TTS for a given domain using visualisation and tree-based algorithm. |
Copied to clipboard
| Challenge: | rynSpeech corpus is a dataset that can be used for both Text-to-Speecher (TTS) and Automatic Speech Recognition (ASR) speakers of many African languages have no access to voice-enabled applications in their native languages. |
| Approach: | They propose a dataset to collect Yorùbá speech data that can be used for both TTS and ASR tasks. |
| Outcome: | The proposed dataset can generate a good quality model with as little as 5 hours of speech . the results are consistent with previous studies on the Yorùbá language . |
Copied to clipboard
| Challenge: | DisCo-Speech is a zero-shot controllable text-to-speech framework . standard codecs entangle timbre and prosody, which hinders independent control in continuation-based LMs. |
| Approach: | They propose a disentangled speech codec and an LM-based generator to solve this problem . they propose fusion and reconstruction that merges content and prosody into unified tokens . |
| Outcome: | DisCo-Speech achieves competitive voice cloning and superior zero-shot prosody control. |
Copied to clipboard
| Challenge: | Existing research has overlooked the efficiency of TTS from a latency-sensitive perspective. |
| Approach: | They propose two approaches to achieve latency-optimal TTS by branch-wise parallelism and sequence-wise parallelism. |
| Outcome: | The proposed approach achieves latency-optimal TTS for large models . branch-wise parallelism and sequence-wise parallelism are key approaches . |
Copied to clipboard
| Challenge: | Recent training-based TTS methods, such as continued reinforcement learning, have surged in popularity, while training-free TTS approaches are gradually fading from prominence. |
| Approach: | They propose a fine-grained sequential scaling method guided by process verification that integrates training-free TTS methods with other classical parallel scaling methods at the step level. |
| Outcome: | Experiments on five instruction-tuned large language models (LLMs) show that training-free TTS methods can extend reasoning performance boundaries. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have limited fault localization capabilities due to limited context length. |
| Approach: | They propose a hierarchical localization reward model to evaluate and select the most accurate fault localization candidates from the outputs of LLMs. |
| Outcome: | The proposed model improves the final line-level localization recall by 12% on the SWE-Bench-Lite dataset. |
Copied to clipboard
| Challenge: | Existing systems suffer from linguistic collapse when pursuing high intensity or fail to meet target emotional levels. |
| Approach: | They propose an inference framework that introduces a neutral prosody bias and a uniform Classifier-Free Guidance that distorts the acoustic manifold, leading to artifacts. |
| Outcome: | The proposed framework achieves superior linguistic accuracy and expressiveness without model retraining. |
Copied to clipboard
| Challenge: | Existing methods for optimizing reasoning quality are limited by overthinking. |
| Approach: | They propose a method that allocates thinking budgets to critical reasoning steps by tracking and aggregating step-wise uncertainty over time. |
| Outcome: | The proposed method reduces computation by over 45% on average while improving accuracy by 0.33–3.46%. |
Copied to clipboard
| Challenge: | Existing controllable Text-to-Speech methods limited to inter-utterance-level control . utterance expressiveness remains a challenge in building human-like TTS synthesis systems . |
| Approach: | They propose a training-free controllable framework for pretrained zero-shot TTS to enable intra-utterance emotion and duration expression. |
| Outcome: | The proposed framework achieves state-of-the-art intra-utterance consistency while maintaining baseline-level speech quality. |
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) have reshaped the landscape of reasoning tasks. |
| Approach: | They propose a method that enhances LLM reasoning without finetuning by using test-time scaling. |
| Outcome: | The proposed method outperforms baseline models in both budget and model size. |
Copied to clipboard
| Challenge: | Existing approaches to improve LLM reasoning are limited in complex domains and lack external grounding makes verifiers unreliable on computation-intensive tasks. |
| Approach: | They propose a framework that transforms reward modeling into a multi-turn, tool-augmented deliberative process. |
| Outcome: | The proposed framework surpasses state-of-the-art ORMs by 25.2% under parallel and sequential TTS. |
Copied to clipboard
| Challenge: | Audience Response System (ARS) evaluations are not well understood for text-to-speech synthesis (TTS) evaluation is a key weakness in the field and needs to adapt to be better-suited for this new generation of voices. |
| Approach: | They revisit three published TTS studies and perform an ARS-based evaluation on the stimuli used in each study. |
| Outcome: | The results show that Audience Response System (ARS) is highly useful for evaluating long and continuous stimuli. |
Copied to clipboard
| Challenge: | Existing approaches to test-time scaling are limited due to the quality of candidate responses. |
| Approach: | They propose a new metric to quantify the relative improvement of self-refinement beyond majority voting. |
| Outcome: | The proposed method achieves state-of-the-art performance across five benchmarks over other methods. |
Copied to clipboard
| Challenge: | Generative audio modeling has been fragmented into specialized tasks such as text-to-speech (TTS), text- to-music (TTM), and text-ta (TTA) specialized models require reference audio for timbre cloning and strict phoneme alignment, whereas TTA models generate unstructured textures from open-ended captions. |
| Approach: | They propose a unified flow-matching framework capable of synthesizing speech, music, sound effects . they propose 'token injection mechanism' that projects unstructured environmental sounds into structured temporal latent space . |
| Outcome: | The proposed framework achieves state-of-the-art performance in instruction-based TTS and TTM while maintaining competitive fidelity in TTA. |
Copied to clipboard
| Challenge: | Fact-checking real-world claims requires multistep reasoning and numerical reasoning . large language models are unable to understand nuance of numerical aspects . |
| Approach: | They propose scaling test-time compute (TTS) for large language models to solve this problem . they train a verifier model to navigate the space of possible reasoning paths . |
| Outcome: | The proposed approach achieves 1.8x higher efficiency than standard TTS while delivering 18.8% performance improvement over single-shot verification methods. |
Copied to clipboard
| Challenge: | Entropy-Guided Stepwise Scaling (EGSS) is a novel TTS framework for software engineering tasks. |
| Approach: | They propose an entropy-guided stepwise scaling framework that balances efficiency and effectiveness through entropic-guide encoding and robust test-suite augmentation. |
| Outcome: | EGSS boosts performance by 5–10% across all evaluated models, and reduces inference-time token usage by over 28% . compared to existing methods, EGS reduces token usage and reduce inference time by over 20% . |
Copied to clipboard
| Challenge: | Existing methods for testing time scales treat reasoning traces or tokens equally, ignoring substantial variations in trajectory quality and localized logical failures. |
| Approach: | They propose a chronological reasoning scorer that models each trajectory as a time series. |
| Outcome: | The proposed method achieves relative improvements of 34.21% over Pass@128 and 22.70% over Maj@135 on HMMT25, highlighting its effectiveness. |
Copied to clipboard
| Challenge: | despite its linguistic significance, the Wu dialect of Chinese has long been hindered by the lack of large-scale speech data, standardized evaluation benchmarks, and publicly available models. |
| Approach: | They propose to use WenetSpeech-Wu as a large-scale, multi-dimensionally annotated open-source speech corpus for the Wu dialect of Chinese. |
| Outcome: | The proposed dataset includes 8,000 hours of speech data and strong open-source models . the proposed dataset is competitive and empirically validated . |
Copied to clipboard
| Challenge: | Recent advances in spontaneous text-to-speech (TTS) have enabled the realistic generation of creaky voice, a voice quality known for its diverse pragmatic and paralinguistic functions. |
| Approach: | They used a creaky voice detection tool and a neural TTS engine to control creaky phonation in a spontaneous speech corpus to investigate the effect of creaky voices on perceived certainty, valence, sarcasm, and turn finality. |
| Outcome: | The proposed model enables the realistic synthesis of creaky voice in perceptual tests without formal training. |
Copied to clipboard
| Challenge: | Existing approaches to test-time scaling rely on external verifiers and one-shot independent sampling. |
| Approach: | They propose a test-time scaling framework that reallocates a fixed inference budget into iterative sample–filter–diversify–select cycles. |
| Outcome: | ConMA outperforms baselines on multiple benchmarks while converging early with only 18 samples on average, substantially reducing inference cost. |
Copied to clipboard
| Challenge: | Existing methods to scale complex, open-ended tasks with unverifiable rewards are not scalable to multi-stage pipelines. |
| Approach: | They propose a process-based refinement framework that scales inference across stages of a multi-agent pipeline, instead of refining a single output over time. |
| Outcome: | The proposed framework scales inference across stages of a multi-agent pipeline, instead of refining a single output over time as in prior work. |
Copied to clipboard
| Challenge: | Recent studies show that Test-Time Scaling (TTS) can improve reasoning performance without retraining the model. |
| Approach: | They propose a step-wise pruning strategy that identifies and removes redundant chains using inter-chain similarity at the thought level. |
| Outcome: | The proposed method reduces inference latency and KVC usage by up to 45% and 26% with R1-Distill while maintaining or improving accuracy. |
Copied to clipboard
| Challenge: | Existing autoregressive models for dialogue generation suffer from high latency and stability issues. |
| Approach: | They propose a non-autoregressive (NAR) zero-shot spoken dialogue generation model based on flow-matching. |
| Outcome: | The proposed model outperforms existing models in speech generation due to poor speech intelligibility and turn-taking precision. |
Copied to clipboard
| Challenge: | Parallel test-time scaling is a pivotal approach for enhancing large language models. |
| Approach: | They propose two uncertainty-inspired stochastic strategies for parallel test-time scaling for latent reasoning models and a Latent Reward Model for aggregation. |
| Outcome: | The proposed model scales well with compute and enables effective trajectory selection. |
Copied to clipboard
| Challenge: | Existing approaches often fail to leverage the linguistic intelligence of Large Language Models (LLMs) Existing models lack the ability to follow text instructions for controllable Text-to-Speech (TTS). |
| Approach: | They propose a framework where an LLM acts as a conductor, understanding user instructions and generating a textual plan - explicit vocal features. |
| Outcome: | The proposed model outperforms open- and closed-source models in speech synthesis and achieves zero-shot cross-lingual generalization. |