


default search action
Computer Speech & Language, Volume 99
Volume 99, 2026
- Ander González-Docasal

, Juan Camilo Vásquez-Correa
, Haritz Arzelus
, Aitor Álvarez
, Santiago Andres Moreno-Acevedo
:
Do modern speech LLMs and re-scoring techniques improve bilingual ASR performance for Basque and Spanish in domain-specific contexts? 101905 - Defu Lan, Hai Cheng:

MS-Swinformer and DMTL: Multi-scale spatial fusion and dynamic multi-task learning for speech emotion recognition. 101908 - Hanyu Ding, Wenlong Dong, Qirong Mao:

Keyword Mamba: Spoken keyword spotting with state space models. 101909 - Ihab Asaad, Maxime Jacquelin, Olivier Perrotin, Laurent Girin, Thomas Hueber

:
Is self-supervised learning enough to fill in the gap? A study on speech inpainting. 101922 - Zohre Foroushi

, Richard M. Dansereau
:
Enhanced audio-visual speech enhancement with posterior sampling methods in recurrent variational autoencoders. 101923 - Cheng Yu

, Vahid Ahmadi Kalkhorani, Buye Xu, DeLiang Wang:
Audiovisual speech enhancement and voice activity detection using generative and regressive visual features. 101924 - Xinlu He

, Jacob Whitehill
:
Survey of end-to-end multi-speaker automatic speech recognition for monaural audio. 101925 - Vikramjit Mitra

, Anirban Chatterjee, Ke Zhai, Helen Weng, Ayuko Hill, Nicole Hay, Christopher Webb, Jamie Cheng, Erdrin Azemi:
Leveraging saliency-based pre-trained foundation model representations to uncover breathing patterns in speech. 101926 - Jaime Bellver-Soler

, Anmol Guragain
, Samuel Ramos-Varela, Ricardo de Córdoba, Luis Fernando D'Haro:
Speech emotion recognition using multimodal LLMs and quality-controlled TTS-based data augmentation for Iberian languages. 101927 - Sanjeevkumar Angadi, Saili Hemant Sable, Tejaswini Zope, Rajani Amol Hemade, Vaibhavi Umesh Avachat:

Artificial protozoa lotus effect algorithm enabled cognitive brain optimal model for sentiment analysis utilizing multimodal data. 101929 - Xuan Shi, Tiantian Feng, Jay Park

, Christina Hagedorn, Louis Goldstein, Shrikanth Narayanan:
Speech acoustics to rt-MRI articulatory dynamics inversion with video diffusion model. 101928 - Jay Kejriwal

, Stefan Benus
, Lina Maria Rojas-Barahona:
Entrainment detection using DNN. 101930 - Jianqiang Zhang

, Yushui Geng, Peng Zhang, Fuqiang Wang, Xiaoming Wu:
One-class neural network with hybrid pooling on dual-band frequency for spoofing speech detection. 101937 - Myeong-Ha Hwang

, Jikang Shin, Junseong Bang:
V-APA: A Voice-driven Agentic Process Automation System. 101938 - Xabier de Zuazo

, Eva Navas, Ibon Saratxaga, Mathieu Bourguignon, Nicola Molinaro:
Decoding phone pairs from MEG signals across speech modalities. 101939 - Da-Hee Yang, Joon-Hyuk Chang:

An experimental study of diffusion-based general speech restoration with predictive-guided conditioning. 101940 - Ayub Othman Abdulrahman:

Pitch-Aware multi-feature fusion for classifying statements, questions, and exclamations in low-resource languages. 101941 - Yanlu Xie, Huihang Zhong, Xuhui Lan, Wenwei Dong:

Mispronunciation detection and diagnosis based on large language models. 101942 - Ladislav Mosner, Oldrich Plchot, Lukás Burget, Chunlei Zhang, Jan Cernocký, Meng Yu:

Trainable multi-channel front-ends for joint beamforming and speaker embedding extraction. 101944 - Shiqing Zhang, Chen Chen, Dandan Wang, Xin Tao

, Xiaoming Zhao:
Two-stage multiple instance learning networks with attention-based hybrid aggregation for speech emotion recognition. 101946 - Juan Ignacio Álvarez-Trejos

, Sara Barahona, Laura Herrera-Alarcón, Jérémie Touati, Alicia Lozano-Diez
:
On the use of DiaPer models and matching algorithm for RTVE speaker diarization 2024 dataset. 101948 - Yujiang Liu, Lijun Fu, Xiaojun Xia:

Medical related word enhancement framework: A new method for large language model in medical dialogue generation. 101949 - Nikolaos Malamas

, Andreas L. Symeonidis, John B. Theocharis:
QuAVA: A privacy-aware architecture for conversational desktop Content Retrieval systems. 101950 - Wenzhe Jia, Yuhang Wang, Yahui Kang:

Emotion-guided cross-modal alignment for multimodal depression detection. 101951

manage site settings
To protect your privacy, all features that rely on external API calls from your browser are turned off by default. You need to opt-in for them to become active. All settings here will be stored as cookies with your web browser. For more information see our F.A.Q.


Google
Google Scholar
Semantic Scholar
Internet Archive Scholar
CiteSeerX
ORCID













