arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2202.10290v3 [eess.AS] 17 Mar 2022

Speaker Adaptation Using Spectro-Temporal Deep Features for Dysarthric and Elderly Speech RecognitionThanks: Mengzhe Geng, Zi Ye, Tianzi Wang, Guinan Li and Shujie Hu are with the Chinese University of Hong Kong, China (email: {mzgeng,zye,twang,gnli,sjhu}@se.cuhk.edu.hk).
Xurong Xie is with Institute of Software, Chinese Academy of Sciences, Beijing, China (email: xurong@iscas.ac.cn).
Xunying Liu is with the Chinese University of Hong Kong, China and the corresponding author (email: xyliu@se.cuhk.edu.hk).
Helen Meng is with the Chinese University of Hong Kong, China (email: hmmeng@se.cuhk.edu.hk).

Mengzhe Geng    Xurong Xie    Zi Ye    Tianzi Wang    Guinan Li    Shujie Hu Affiliation: Xunying Liu, , Helen Meng, 
Abstract

Despite the rapid progress of automatic speech recognition (ASR) technologies targeting normal speech in recent decades, accurate recognition of dysarthric and elderly speech remains highly challenging tasks to date. Sources of heterogeneity commonly found in normal speech including accent or gender, when further compounded with the variability over age and speech pathology severity level, create large diversity among speakers. To this end, speaker adaptation techniques play a key role in personalization of ASR systems for such users. Motivated by the spectro-temporal level differences between dysarthric, elderly and normal speech that systematically manifest in articulatory imprecision, decreased volume and clarity, slower speaking rates and increased dysfluencies, novel spectro-temporal subspace basis deep embedding features derived using SVD speech spectrum decomposition are proposed in this paper to facilitate auxiliary feature based speaker adaptation of state-of-the-art hybrid DNN/TDNN and end-to-end Conformer speech recognition systems. Experiments were conducted on four tasks: the English UASpeech and TORGO dysarthric speech corpora; the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets. The proposed spectro-temporal deep feature adapted systems outperformed baseline i-Vector and x-Vector adaptation by up to 2.63% absolute (8.63% relative) reduction in word error rate (WER). Consistent performance improvements were retained after model based speaker adaptation using learning hidden unit contributions (LHUC) was further applied. The best speaker adapted system using the proposed spectral basis embedding features produced the lowest published WER of 25.05% on the UASpeech test set of 16 dysarthric speakers.

Index Terms: 
Speaker Adaptation, Disordered speech recognition, Elderly Speech Recognition

I Introduction

Despite the rapid progress of automatic speech recognition (ASR) techonologies targeting normal speech in recent decades [1, 2, 3, 4, 5, 6, 7, 8], accurate recognition of dysarthric and elderly speech remains highly challenging tasks to date [9, 10, 11, 12, 13, 14, 15, 16]. Dysarthria is caused by a range of speech motor control conditions including cerebral palsy, amyotrophic lateral sclerosis, stroke and traumatic brain injuries [17, 18, 19, 20, 21]. In a wider context, speech and language impairments are also commonly found among older adults experiencing natural ageing and neurocognitive disorders, for example, Alzheimer’s disease [22, 23]. People with speech disorders often experience co-occurring physical disabilities and mobility limitations. Their difficulty in using keyboard, mouse and touch screen based user interfaces makes voice based assistive technologies more natural alternatives [24, 25] even though speech quality is degraded. To this end, in recent years there has been increasing interest in developing ASR technologies that are suitable for dysarthric and elderly speech [26, 10, 27, 28, 29, 30, 31, 32, 33, 12, 34, 35, 36, 37, 38, 39, 40, 41, 16, 8, 42, 43, 44, 45, 46, 47, 48, 15, 49, 50].

Dysarthric and elderly speech bring challenges on all fronts to current deep learning based automatic speech recognition technologies predominantly targeting normal speech recorded from healthy, non-aged users. In addition to the scarcity of such data, their large mismatch against healthy speech and the difficulty in collecting them on a large scale from impaired and elderly speakers due to mobility issues, the need of modelling the prominent heterogeneity among speakers is particularly salient. Sources of variability commonly found in normal speech including accent or gender, when further compounded with those over age and speech pathology severity, create large diversity among dysarthric and elderly speakers  [39, 51]. The deficient abilities in controlling the articulators and muscles responsible for speech production lead to abnormalities in dysarthric and elderly speech manifested across many fronts including articulatory imprecision, decreased volume and clarity, increased dysfluencies, changes in pitch and slower speaking rate [52]. In addition, the temporal or spectral perturbation based data augmentation techniques [53, 54, 37] that are widely used in current systems to circumvent data scarcity further contribute to speaker-level variability. To this end, speaker adaptation techniques play a key role in personalization of ASR systems for such users. Separate reviews over conventional speaker adaptation techniques developed for normal speech and those for dysarthric or elderly speech are presented in the following Section I-A and I-B.

I-A Speaker Adaptation for Normal Speech

Speaker adaptation techniques adopted by current deep neural networks (DNNs) based ASR systems targeting normal speech can be divided into three major categories: 1) auxiliary speaker embedding feature based methods that represent speaker dependent (SD) features via compact vectors [55, 56, 57, 58], 2) feature transformation based approaches that produce speaker independent (SI) canonical features at the acoustic front-ends [59, 60, 61, 62, 63] and 3) model based adaptation techniques that compensate the speaker-level variability by often incorporating additional SD transformations that are applied to DNN parameters or hidden layer outputs [64, 65, 66, 67, 68, 69].

In the auxiliary speaker embedding feature based approaches, speaker dependent (SD) features such as speaker codes [55] and i-Vectors [56, 57] are concatenated with acoustic features to facilitate speaker adaptation during both ASR system training and evaluation. The estimation of SD auxiliary features can be performed independently of the remaining recognition system components. For example, i-Vectors [56, 57] are learned from Gaussian mixture model (GMM) based universal background models (UBMs). The SD auxiliary features can also be jointly estimated with the back-end acoustic models, for example, via an alternating update between them and the remaining SI DNN parameters in speaker codes [55]. Auxiliary feature based speaker adaptation methods benefit from both their low complexity in terms of the small number of SD feature parameters to be estimated, and their flexibility allowing them to be incorporated into a wide range of ASR systems including both hybrid DNN-HMM systems and recent end-to-end approaches [70].

In feature transformation based speaker adaptation, feature transforms are applied to acoustic front-ends to produce canonical, speaker invariant inputs. These are then fed into the back-end DNN based ASR systems to model the remaining sources of variability, for example, phonetic and phonological context dependency in speech. Feature-space maximum likelihood linear regression (f-MLLR) transforms [63] estimated at speaker-level from GMM-HMM based ASR systems [59, 61] are commonly used. In order to account for the vocal tract length differences between speakers, physiologically motivated vocal tract length normalization (VTLN) can also be used as feature transformation [60, 62]. Speaker-level VTLN normalized features can be obtained using either piecewise linear frequency warping factors directly applied to the spectrum, or affine linear transformations akin to f-MLLR.

In model based adaptation approaches, separately designed speaker-dependent DNN model parameters are used to account for speaker-level variability. In order to ensure good generalization and reduce the risk of overfitting to limited speaker-level data, a particular focus of prior researches has been on deriving compact forms of SD parameter representations. These are largely based on linear transforms that are incorporated into various parts of DNN acoustic models. These include the use of SD linear input networks (LIN) [59, 67], linear output networks (LON) [21], linear hidden networks (LHN) [66], learning hidden unit contributions (LHUC) [71, 68, 69], parameterized activation functions (PAct) [72, 69], speaker-independent (SI) and SD factored affine transformations [73], and adaptive interpolation of outputs of basis sub-networks [74, 75]. In addition to only modelling speaker-level variability in the test data during recognition, the estimation of SD parameters in both the system training and evaluation stages leads to more powerful speaker adaptive training (SAT) [65] approaches, allowing a joint optimization of both the SD and SI parameters during system training.

I-B Speaker Adaptation for Dysarthric and Elderly Speech

In contrast, only limited research on speaker adaptation techniques targeting dysarthric and elderly speech recognition has been conducted so far. Earlier works in this direction were mainly conducted in the context of traditional GMM-HMM acoustic models. The application of maximum likelihood linear regression (MLLR) and maximum a posterior (MAP) adaptation to such systems were investigated in [76, 77, 9, 78]. MLLR was further combined with MAP adaptation in speaker adaptive training (SAT) of SI GMM-HMM in [11]. F-MLLR based SAT training of GMM-HMM systems was investigated in [79]. Regularized speaker adaptation using Kullback-Leibler (KL) divergence was studied for GMM-HMM systems in [80].

More recent researches applied model adaptation techniques to a range of state-of-the-art DNN based dysarthric and elderly speech recognition systems. Normal to dysarthric speech domain adaptation approaches using direct parameter fine-tuning were investigated in both lattice-free maximum mutual information (LF-MMI) trained time delay neural network (TDNN) [40, 43] based hybrid ASR systems and end-to-end recurrent neural network transducer (RNN-T) [36, 48] systems. In order to mitigate the risk of overfitting to limited speaker-level data during model based adaptation, more compact learning hidden unit contributions (LHUC) based dysarthric speaker adaptation was studied in [12, 37, 41] while Bayesian learning inspired domain speaker adaptation approaches have also been proposed in [81].

One main issue associated with previous researches on dysarthric and elderly speaker adaptation is that the systematic, fine-grained speaker-level diversity attributed to speech impairment severity and aging is not considered. Such diversity systematically manifests itself in a range of spectro-temporal characteristics including articulatory imprecision, decreased volume and clarity, breathy and hoarse voice, increased dysfluencies as well as slower speaking rate.

In order to address this issue, novel deep spectro-temporal embedding features are proposed in this paper to facilitate auxiliary speaker embedding feature based adaptation for dysarthric and elderly speech recognition. Spectral and temporal basis vectors derived by singular value decomposition (SVD) of dysarthric or elderly speech spectra were used to structurally and intuitively represent the spectro-temporal key attributes found in such data, for example, an overall decrease in speaking rate and speech volume as well as changes in the spectral envelope. These two sets of basis vectors were then used to construct DNN based speech pathology severity or age classifiers. More compact, lower dimensional speaker specific spectral and temporal embedding features were then extracted from the resulting DNN classifiers’ bottleneck layers, before being further utilized as auxiliary speaker embedding features to adapt start-of-the-art hybrid DNN [41], hybrid TDNN [3] and end-to-end (E2E) Conformer [6] ASR systems.

Experiments were conducted on four tasks: the English UASpeech [82] and TORGO [83] dysarthric speech corpora; the English DementiaBank Pitt [84] and Cantonese JCCOCC MoCA [85] elderly speech datasets. Among these, UASpeech is by far the largest available and widely used dysarthric speech database, while DementiaBank Pitt is the largest publicly available elderly speech corpus. The proposed spectro-temporal deep feature adapted systems outperformed baseline i-Vector [56] and x-Vector [86] adapted systems by up to 2.63%2.63\% absolute (8.63%8.63\% relative) reduction in word error rate (WER). Consistent performance improvements were retained after model based speaker adaptation using learning hidden unit contributions (LHUC) was further applied. The best speaker adapted system using the proposed spectral basis embedding features produced the lowest published WER of 25.05%25.05\% on the UASpeech test set of 1616 dysarthric speakers. Speech pathology severity and age prediction performance as well as further visualization using t-distributed stochastic neighbor embedding (t-SNE) [87] indicate that our proposed spectro-temporal deep features can more effectively learn the speaker-level variability attributed to speech impairment severity and age than conventional i-Vector [56] and x-Vector [86]. The main contributions of this paper are summarized below:

1) To the best of our knowledge, this paper presents the first use of spectro-temporal deep embedding features to facilitate speaker adaptation for dysarthric and elderly speech recognition. In contrast, there were no prior researches published to date on auxiliary features based speaker adaptation targeting such data. Existing speaker adaptation methods for dysarthric and elderly speech use mainly model based approaches [76, 77, 9, 78, 79, 11, 12, 36, 40, 37]. Speaker embedding features were previously only studied for speech impairment assessment [88, 89, 90].

2) The proposed spectro-temporal deep features are inspired and intuitively related to the latent variability of dysarthric and elderly speech. The spectral basis embedding features are designed to learn characteristics such as volume reduction, changes of spectral envelope, imprecise articulation as well as breathy and hoarse voice, while the temporal basis embedding features to capture patterns such as increased dysfluencies and pauses. The resulting fine-grained, factorized representation of diverse impaired speech characteristics serves to facilitate more powerful personalized user adaptation for dysarthric and elderly speech recognition.

3) The proposed spectro-temporal deep feature adapted systems achieve statistically significant performance improvements over baseline i-Vector or x-Vector adapted hybrid DNN/TDNN and end-to-end (E2E) Conformer systems by up to 2.63%2.63\% absolute (8.63%8.63\% relative) word error rate (WER) reduction on four dysarthric or elderly speech recognition tasks across two languages. These findings serve to demonstrate the efficacy and genericity of our proposed spectro-temporal deep features for dysarthric and elderly speaker adaptation.

The rest of this paper is organized as follows. The derivation of spectro-temporal basis vectors using SVD speech spectrum decomposition is presented in Section II. The extraction of spectro-temporal deep embedding features and their incorporation into hybrid DNN/TDNN and end-to-end Conformer based ASR systems for speaker adaptation are proposed in Section III. A set of implementation issues affecting the learning of spectro-temporal deep embedding features are discussed in Section IV. Section V presents the experimental results and analysis. Section VI draws the conclusion and discusses possible future works.

II Speech Spectrum Subspace Decomposition

Spectro-temporal subspace decomposition techniques provide a simple and intuitive solution to recover speech signals from noisy observations by modelling the combination between these two using a linear system [91]. This linear system can then be solved using signal subspace decomposition schemes, for example, singular value decomposition (SVD) [91, 90] or non-negative matrix factorization (NMF) methods [92, 93, 94], both of which are performed on the time-frequency speech spectrum.

An example SVD decomposition of a mel-scale filter-bank based log amplitude spectrum is shown in Fig. 1a and 1b. Let 𝐒𝐫\mathbf{S_{r}} represent a C×TC\times T dimensional mel-scale spectrogram of utterance rr with CC filter-bank channels and TT frames. The SVD decomposition [91] of 𝐒𝐫\mathbf{S_{r}} is given by:

𝐒𝐫=𝐔𝐫𝚺𝐫𝐕𝐫T\mathbf{S_{r}=U_{r}{\Sigma_{r}}V_{r}^{\mathrm{T}}} (1)

where the set of column vectors of the C×CC\times C dimensional left singular 𝐔𝐫\mathbf{U_{r}} matrix and the row vectors of the T×TT\times T dimensional right singular 𝐕𝐫T\mathbf{V_{r}^{\mathrm{T}}} matrix are the bases of the spectral and temporal subspaces respectively. Here 𝚺𝐫\mathbf{\Sigma_{r}} is a C×TC\times T rectangular diagonal matrix containing the singular values sorted in a descending order, which can be further absorbed into a multiplication with 𝐕𝐫T\mathbf{V_{r}^{\mathrm{T}}} for simplicity. In order to obtain more compact representation of the two subspaces, a low-rank approximation [93] obtained by selecting the top-dd principal spectral and temporal basis vectors can be used. In this work, the number of principal components dd is empirically set to vary from 22 to 1010.

Refer to caption
(a) DYS v.s. CTL from UASpeech corpus
Refer to caption
(b) PAR v.s. INV from DBANK corpus
Fig. 1: Example subspace decomposition of mel-spectrogram of: (a) a pair of normal (CTL, left upper) and dysarthric (DYS, left lower) utterances of word “python” to obtain top d=4d=4 spectral and temporal basis vectors (circled in red in 𝐔\mathbf{U} and 𝐕T\mathbf{V^{\mathrm{T}}}) of the UASpeech [82] corpus; and (b) a pair of non-aged clinical investigator (INV, right upper) and elderly participant (PAR, right lower) utterances of word “okay” to obtain top d=3d=3 spectral and temporal basis vectors (circled in red in 𝐔\mathbf{U} and 𝐕T\mathbf{V^{\mathrm{T}}}) of the DementiaBank Pitt (DBANK) [84] dataset.

The SVD decomposition shown in Fig. 1 intuitively separates the speech spectrum into two sources of information that can be related to the underlying sources of variability in dysarthric and elderly speech: a) time-invariant spectral subspaces that can be associated with an average utterance-level description of dysarthric or elderly speakers’ characteristics such as an overall reduction of speech volume, changes in the spectral envelope shape, weakened formats due to articulation imprecision as well as hoarseness and energy distribution anomaly across frequencies due to difficulty in breath control. For example, the comparison between the spectral basis vectors extracted from a pair of dysarthric and normal speech utterances of the identical content “python” in Fig. 1a shows that the dysarthric spectral basis vectors exhibit a pattern of energy distribution over mel-scale frequencies that differs from that obtained from the normal speech spectral bases. Similar trends can be found between the spectral basis vectors of non-aged and elderly speech utterances of the same word content “okay” shown in Fig. 1b. b) time-variant temporal subspaces that are considered more related to sequence context dependent features such as decreased speaking rate as well as increased dysfluencies and pauses, for example, shown in the contrast between the temporal basis vectors separately extracted from normal and dysarthric speech in Fig. 1a and those from non-aged and elderly speech in Fig. 1b, where the dimensionality of the temporal subspace captures the speaking rate and duration.

SVD spectrum decomposition is performed in an unsupervised fashion. In common with other unsupervised feature decomposition methods such as NMF, it is theoretically non-trivial to produce a perfect disentanglement [95] between the time-invariant and variant speech characteristics encoded by the spectral and temporal basis vectors respectively, as both intuitively represent certain aspects of the underlying speaker variability associated with speech pathology severity and age.

For the speaker adaptation task considered in this paper, the ultimate objective is to obtain more discriminative feature representations to capture dysarthric and elderly speaker-level diversity attributed to speech impairment severity and age. To this end, further supervised learning of deep spectro-temporal embedding features is performed by constructing deep neural network based speech pathology severity or age classifiers taking the principal spectral or temporal basis vectors as their inputs. These are presented in the following Section III.

III Spectro-Temporal Deep Features

This section presents the extraction of spectro-temporal deep embedding features and their incorporation into hybrid DNN/TDNN and end-to-end Conformer based ASR systems for auxiliary feature based speaker adaptation.

In order to obtain sufficiently discriminative feature representations to capture dysarthric and elderly speaker-level diversity associated with the underlying speech impairment severity level and age information, further supervised learning of deep spectro-temporal embedding features is performed by constructing deep neural network based speech pathology severity or age classifiers. The principal SVD decomposed utterance-level spectral or temporal basis vectors are used as their inputs. More compact, lower dimensional speaker specific spectral and temporal embedding features are then extracted from the resulting impairment severity or age DNN classifiers’ bottleneck layers, before being further used as auxiliary embedding features for speaker adaptation of ASR systems. An overall system architecture flow chart covering all the three stages including SVD spectrum decomposition, deep spectral and temporal embedding features extraction and ASR system adaptation using such features is illustrated in Fig. 2.

Fig. 2: Overall system architecture including from left to right: a) front-end mel-filter bank feature extraction (in grey, top left); b) SVD spectrum decomposition (circled in green, top middle); c) DNN based speech impairment severity or age classification and deep spectro-temporal embedding feature extraction (in light blue, top right); d) auxiliary feature based ASR system adaptation (in orange, bottom).

III-A Extraction of Spectro-Temporal Deep Features

When training the speech impairment severity or age classification DNNs to extract deep spectro-temporoal embedding features, the top-dd principal spectral or temporal basis are used as input features to train the respective DNNs sharing the same model architecture shown in Fig. 3, where either speech pathology severity based on, for example, the speech intelligibility metrics provided by the UASpeech [82] corpus, or the binary aged v.s. non-aged speaker annotation of the DementiaBank Pitt [84] dataset, are used as the output targets.

The DNN classifier architecture is a fully-connected neural network containing four hidden layers, the first three of which are of 20002000 dimensions, while the last layer contains 2525 dimensions. Each of these hidden layers contains a set of neural operations performed in sequence. These include affine transformation (in green), rectified linear unit (ReLU) activation (in yellow) and batch normalization (in orange), while the outputs of the first layer are connected to those of the third layer via a skip connection. Linear bottleneck projection (in light green) is also applied to the inputs of the middle two hidden layers while dropout operation (in grey) is used on the outputs of the first three hidden layers. Softmax activation (in dark green) is used in the last layer. Further fine-grained speaker-level information can be incorporated into the training cost via a multitask learning (MTL) [74] interpolation between the cross-entropy over speech intelligibility level or age, and that computed over speaker IDs. The outputs of the 2525-dimensional bottleneck (BTN) layer are extracted as compact neural embedding representations of the spectral or temporal basis vectors (bottom right in Fig. 3).

Fig. 3: An example DNN based speech intelligibility or age classifier containing a bottleneck layer to extract spectral and temporal embedding features for speaker adaptation.

When training the DNN speech impairment severity or age classifier using the SVD temporal basis vectors as the input, a frame-level sliding window of 2525 dimensions was applied to the top-dd selected temporal basis vectors. Their corresponding 2525-dimensional mean and standard deviation vectors were then computed to serve as the “average” temporal basis representations of fixed dimensionality. This within utterance windowed averaging of temporal basis vectors allows dysarthric or elderly speakers who speak of different word contents but exhibit similar patterns of temporal context characteristics such as slower speaking rate and increased pauses to be mapped consistently to the same speech impairment severity or age label. This flexible design is in contrast to conventional speech intelligibility assessment approaches that often require the contents spoken by different speakers to be the same [96, 97, 90]. It not only facilitates a more practical speech pathology assessment scheme to be applied to unrestricted speech contents of unknown duration, but also the extraction of fixed size temporal embedding features for ASR system adaptation.

The speaker-level speech impairment severity or age information can be then captured by the resulting DNN embedding features. For example, visualization using t-distributed stochastic neighbour embedding (t-SNE) [87] reveals the speaker-level spectral basis neural embedding features averaged over those obtained over all utterances of the same non-aged clinical investigator (in red) or elderly participant (in green) of the DementiaBank Pitt [84] corpus shown in Fig. 4c demonstrate much clearer age discrimination than the comparable speaker-level i-Vectors and x-Vectors shown in Fig 4a and Fig. 4b respectively. Similar trends can also be found on the Cantonese JCCOCC MoCA [85] corpus designed by a similar data collection protocol based on neuro-physiological interviews comparable to the English DementiaBank Pitt corpus11 1 Due to the relatively smaller number of speakers included in the UASpeech [82] and TORGO [83] dysarthric speech corpora (29 and 15 respectively), t-SNE visualization of spectral embedding features are performed on the English DementiaBank Pitt and Cantonese JCCOCC MoCA datasets containing 688688 and 369369 speakers each..

Refer to caption
(a) DBANK-iVector
Refer to caption
(b) DBANK-xVector
Refer to caption
(c) DBANK-BTN Feats.
Refer to caption
(d) JCMoCA-iVector
Refer to caption
(e) JCMoCA-xVector
Refer to caption
(f) JCMoCA–BTN Feats.
Fig. 4: T-SNE plot of i-Vectors, x-Vectors and spectral DNN embedding features obtained on the English DementiaBank Pitt (DBANK) corpus with 688688 speakers (444444 non-aged clinical investigators in red and 244244 aged participants in green) and the Cantonese JCCOCC MoCA (JCMoCA) corpus with 369369 speakers (211211 non-aged clinical investigators in red and 158158 aged participants in green).

III-B Use of Spectro-Temporal Deep Features

The compact 2525-dimensional spectral and temporal basis embedding features extracted from the DNN speech impairment severity or age classifiers’ bottleneck layers presented above in Section III-A are concatenated to the acoustic features at the front-end to facilitate auxiliary feature based speaker adaptation of state-of-the-art ASR systems based on hybrid DNN [41], hybrid lattice-free maximum mutual information (LF-MMI) trained time delay neural network (TDNN) [3] or end-to-end (E2E) Conformer models [6], as shown in Fig 5. For hybrid DNN and TDNN systems, model based adaptation using learning unit contributions (LHUC)  [68] can optionally be further applied on top of auxiliary feature based speaker adaptation, as shown in Fig. 5a and Fig. 5b respectively.

(a)
(b)
(c)
Fig. 5: Incorporation of spectral and temporal deep features at the front end of (a) hybrid DNN [41], (b) hybrid TDNN [3] and (c) Conformer [6] ASR systems for auxiliary feature based speaker adaptation. For hybrid DNN and TDNN systems, adaption configuration (i) leads to systems using auxiliary feature based adaptation only, while selecting configuration (ii) leads to systems with additional speaker adaptive training using LHUC-SAT [68].

IV Implementation Details

In this section, several key implementation issues associated with the learning and usage of spectro-temporal deep embedding features are discussed. These include the choices of spectro-temporal basis embedding neural network output targets when incorporating speech intelligibility measures or age, the smoothing of the resulting embedding features extracted from such embedding DNNs to ensure the homogeneity over speaker-level characteristics, and the number of principal spectral and temporal basis vectors required for the embedding networks. Ablation studies were conducted on the UASpeech dysarthric speech corpus [82] and the DementiaBank Pitt elderly speech corpus [84]. After speaker independent and speaker dependent speed perturbation based data augmentation [37, 15], their respective training data contain approximately 130.1130.1 hours and 58.958.9 hours of speech. After audio segmentation and removal of excessive silence, the UASpeech evaluation data contains 99 hours of speech while DementiaBank development and evaluation sets of 2.52.5 hours and 0.60.6 hours of speech respectively were used. Mel-scale filter-bank (FBK) based log amplitude spectra of 4040 channels are used as the inputs of singular value decomposition (SVD) in all experiments of this paper.

IV-A Choices of Embedding Network Targets

In the two dysarthric speech corpora, speech pathology assessment measures are provided for each speaker. In the UASpeech data, the speakers are divided into several speech intelligibility subgroups: “very low”, “low”, “mid” and “high” [82]. In the TORGO corpus, speech impairment severity measures based on “severe”, “moderate” and “mild” are provided [83]. In the two elderly speech corpora, the role of each speaker during neuro-physiological interview for cognitive impairment assessment is annotated. Each interview is based on a two-speaker conversation involving a non-aged investigator and another aged, elderly participant [84, 85].

By default, the speech intelligibility metrics provided by the UASpeech corpus, or the binary aged v.s. non-aged speaker annotation of the DementiaBank Pitt dataset, are used as the output targets in the following ablation study over embedding target choices. In order to incorporate further speaker-level information, a multitask learning (MTL) [74] style cost function featuring interpolation between the cross-entropy error computed over the speech intelligibility level or age labels, and that computed over speaker IDs can be used.

As is shown in the results obtained on the UASpeech [82] data in Table I, using both the speech intelligibility and speaker ID labels as the embedding targets in multi-task training produced lower word error rates (WERs) across all severity subgroups than using speech intelligibility output targets only (Sys.7 v.s. Sys.6 in Table I). The results obtained on the DementiaBank Pitt [84] data in Table II suggest that there is no additional benefit in adding the speaker information during the embedding process (Sys.7 v.s. Sys.6 in Table II). Based on these trends, in the main experiments of the following Section V, the embedding network output targets exclusively use both speech severity measures and speaker IDs on the UASpeech and TORGO [83] dysarthric speech datasets, while only binary aged v.s. non-aged labels are used on the DementiaBank Pitt and Cantonese JCCOCC MoCA [85] elderly speech datasets.

IV-B Smoothing of Embedding Features

For auxiliary feature based adaptation techniques including the spectral and temporal basis deep embedding representations considered in this paper, it is vital to ensure the speaker-level homogeneity to be consistently encoded in these features. As both forms of embedding features are computed on individual utterances, additional smoothing is required to ensure such homogeneity, for example, an overall reduction of speech volume of a dysarthric or elderly speaker’s data, to be consistently retained in the resulting speaker embedding representations. To this end, two forms of speaker embedding smoothing are considered in this paper. The first is based on a simple averaging of all utterance-level spectral or temporal embedding features for each speaker. The second smoothing method is based on Latent Dirichlet allocation (LDA) [98] based clustering of utterance-level spectral or temporal embedding features. Following earlier researches [99], a 100100-component Gaussian Mixture Model (GMM) is trained first to quantize the utterance-level spectral or temporal embedding features of the same speaker, before LDA clustering is applied to produce speaker-level features of varying dimensions ranging from 1010, 2525 to 5050.

A general trend observed in the results of both Table I and II is that using spectral embedding feature smoothing, whether via a simple speaker-level averaging (Sys.6 in Table I and II) or LDA clustering (Sys.3-5 in Table I and II), produced better performance than directly using non-smoothed spectral embedding features (Sys.2 in Table I and II). Across both the UASpeech and DementiaBank Pitt tasks, the simpler speaker-level averaging based smoothing (Sys.6 in Table I and II) consistently outperform LDA clustering (Sys.3-5 in Table I and II), and is subsequently used in all experiments of the following Section V.

TABLE I: Ablation study on the augmented UASpeech corpus [82] with 130.1130.1h training data. “SB”, “TB” and “STB” are in short for spectral or temporal basis vectors and spectral plus temporal basis vectors. “Seve.” and “SpkId” stand for speech impairment severity group and speaker ID. “Dim./(d)” denote the dimensionality and the number of principal spectral or temporal vectors. “LDA-1010”, “LDA-2525” and “LDA-5050” denote Latent Dirichlet allocation based clustering features of 1010, 2525 and 5050 dimensions obtained on the embedding features. “Avg.” stands for speaker-level averaging of the embedding features. “O.V.” stands for “overall”.
Sys. Embed. Network Subspace Avg. WER
Input Target VL L M H O.V.
Basis Dim./(d) Seve. SpkId
1 / 66.45 28.95 20.37 9.62 28.73
2 SB 80/(2) embed. 66.95 32.38 21.86 11.11 30.52
3 +LDA-10 64.22 28.10 19.13 9.07 27.62
4 +LDA-25 62.92 29.25 20.00 8.51 27.62
5 +LDA-50 62.62 29.22 20.03 8.42 27.52
6 embed. 62.70 28.65 18.60 8.60 27.18
7 61.55 27.52 17.31 8.22 26.26
8 TB 250/(5) embed. 68.52 32.24 20.98 9.46 30.09
9 STB 330/(2,5) 61.24 27.77 17.45 8.31 26.32
10 SB 40/(1) embed. 70.49 49.19 27.05 11.93 36.90
11 120/(3) 64.50 33.01 19.47 9.82 29.27
12 160/(4) 72.22 47.52 20.27 9.63 34.76
13 200/(5) 67.71 34.05 19.96 9.28 30.14
14 400/(10) 69.82 45.98 33.23 13.83 37.76
15 800/(20) 74.68 45.82 28.72 11.98 37.25
16 1600/(40) 71.39 44.38 29.07 11.23 36.00
TABLE II: Ablation study on the augmented DementiaBank Pitt corpus [84] with 58.958.9h training data. “Age” and “SpkId” stand for speaker age group and speaker ID. “Dev” and “Eval” stand for the development and evaluation sets. “INV” and “PAR” denote non-aged clinical investigator and aged participant. Other naming conventions follow Table II.
Sys. Embed. Network Subspace Avg. WER
Input Target Dev Eval O.V.
Basis Dim./(d) Age SpkId INV PAR INV PAR
1 / 19.91 47.93 19.76 36.66 33.80
2 SB 160/(4) embed. 19.88 45.91 17.54 33.72 32.43
3 +LDA-10 19.31 45.50 19.31 45.50 32.25
4 +LDA-25 19.86 45.78 19.86 45.78 32.85
5 +LDA-50 20.40 46.30 20.40 46.30 33.59
6 embed. 18.61 43.84 17.98 33.82 31.12
7 18.49 44.24 18.53 34.01 31.28
8 TB 160/(4) embed. 19.28 45.35 20.75 34.18 32.14
9 STB 660/(4,10) 20.10 46.00 20.53 35.31 32.91
10 SB 40/(1) embed. 18.98 44.07 19.87 33.38 31.35
11 80/(2) 18.68 44.72 17.54 34.03 31.52
12 120/(3) 18.36 44.39 19.64 34.71 31.44
13 200/(4) 18.93 44.54 18.42 33.80 31.54
14 400/(10) 19.38 44.18 18.87 34.41 31.69
15 800/(20) 19.90 45.47 19.31 35.04 32.54
16 1600/(40) 19.82 45.55 20.53 34.12 32.42

IV-C Number of Spectral and Temporal Basis Vectors

In this part of the ablation study on implementation details, the effect of the number of principal spectral or temporal basis vectors on system complexity and performance is analyzed. Consider selecting the top-dd principal SVD spectral and temporal basis components, the input feature dimensionality of the spectral basis embedding (SBE) DNN network is then expressed as 40×d40\times d, for example, 8080 dimensions when d=2d=2. The temporal basis embedding (TBE) network is 50×d50\times d including both the 2525 dimensional mean and the 2525 dimensional standard deviation vectors both computed over a frame-level sliding window of 2525 dimensions for each of the selected top-dd principal temporal basis vector, for example, 250250 dimensions when d=5d=5. The input dimensionality of the comparable spectro-temporal basis embedding (STBE) network modelling both forms of bases is then 40×ds+50×dt40\times d_{s}+50\times d_{t}, if further allowing the number of principal spectral components dsd_{s} and that of the temporal components dtd_{t} to be separately adjusted.

In the experiments of this section, dsd_{s} and dtd_{t} are empirically adjusted to be 22 and 55 for dysarthric speech (Sys.2-9 in Table I) while 44 and 1010 for elderly speech (Sys.2-9 in Table II). These settings were found to produce the best adaptation performance when the corresponding set of top principal spectral or temporal basis vectors were used to produce the speaker embedding features. For example, as the results shown in both Table I and II for the UASpeech and DemmentiaBank Pitt datasets, varying the number of principal spectral components from 11 to 4040 (the corresponding input feature dimensionality ranging from 4040 to 16001600, Sys.10-16 in Table I and II) suggests the optimal number of spectral basis vectors is generally set to be 22 for the dysarthric speech data (Sys.7 in Table I) and 44 for the elderly speech data (Sys.6 in Table II) when considering both word error rate (WER) and model complexity.

V Experiments

In this experiment section, the performance of our proposed deep spectro-temporal embedding feature based adaptation is investigated on four tasks: the English UASpeech [82] and TORGO [83] dysarthric speech corpora as well as the English DementiaBank Pitt [84] and Cantonese JCCOCC MoCA [85] elderly speech datasets. The implementation details discussed in Section IV are adopted. Data augmentation featuring both speaker independent perturbation of dysarthric or elderly speech and speaker dependent speed perturbation of control healthy or non-aged speech following our previous works [37, 15] is applied on all of these four tasks. A range of acoustic models that give state-of-the-art performance on these tasks are chosen as the baseline speech recognition systems, including hybrid DNN [41], hybrid lattice-free maximum mutual information (LF-MMI) trained time delay neural network (TDNN) [3] and end-to-end (E2E) Conformer [6] models. Performance comparison against conventional auxiliary embedding feature based speaker adaptation including i-Vector [56] and x-Vector [86] is conducted. Model based speaker adaptation using learning hidden unit contributions (LHUC) [68] is further applied on top of auxiliary feature based speaker adaptation. Section V-A presents the experiments on the two dysarthric speech corpora while Section V-B introduces experiments on the two elderly speech datasets. For all the speech recognition results measured in word error rate (WER) presented in this paper, matched pairs sentence-segment word error (MAPSSWE) based statistical significance test [100] was performed at a significance level α=0.05\alpha=0.05.

V-A Experiments on Dysarthric Speech

V-A1 the UASpeech Corpus

The UASpeech [82] corpus is the largest publicly available and widely used dysarthric speech dataset [82]. It is an isolated word recognition tasks containing approximately 103103 hours of speech recorded from 2929 speakers, among whom 1616 are dysarthric speakers and 1313 are control healthy speakers. It is further divided into 3 blocks Block 1 (B1), Block 2 (B2) and Block 3 (B3) per speaker, each containing the same set of 155155 common words and a different set of 100100 uncommon words. The data from B1 and B3 of all the 2929 speakers are treated as the training set which contains 69.169.1 hours of audio and 9919599195 utterances in total, and the data from B2 collected of all the 1616 dysarthric speakers (excluding speech from control healthy speakers) are used as the test set containing 22.622.6 hours of audio and 2652026520 utterances in total.

After removing excessive silence at both ends of the speech audio segments using a HTK [101] trained GMM-HMM system [12], a combined total of 30.630.6 hours of audio data from B1 and B3 (9919599195 utterances) were used as the training set, while 99 hours of speech from B2 (2652026520 utterances) were used for performance evaluation. Data augmentation featuring speed perturbing the dysarthric speech in a speaker independent fashion and the control healthy speech in a dysarthric speaker dependent fashion was further conducted [37] to produce a 130.1130.1 hours augmented training set (399110399110 utterances, perturbing both healthy and dysarthric speech). If perturbing dysarthric data only, the resulting augmented training set contains 65.965.9 hours of speech (204765204765 utterances).

V-A2 the TORGO Corpus

The TORGO [83] corpus is a dysarthric speech dataset containing 88 dysarthric and 77 control healthy speakers with a totally of approximately 13.513.5 hours of audio data (1639416394 utterances). It consists of two parts: 5.85.8 hours of short sentence based utterances and 7.77.7 hours of single word based utterances. Similar to the setting of the UASpeech corpus, a speaker-level data partitioning was conducted combining all 77 control healthy speakers’ data and two-thirds of the 88 dysarthric speakers’ data into the training set (11.711.7 hours). The remaining one-third of the dysarthric speech was used for evaluation (1.81.8 hours). After removal of excessive silence, the training and test sets contains 6.56.5 hours (1454114541 utterances) and 11 hour (18921892 utterances) of speech respectively. After data augmentation with both speaker dependent and speaker independent speed perturbation [37, 102], the augmented training set contains 34.134.1 hours of data (6181361813 utterances).

V-A3 Experiment Setup for the UASpeech Corpus

Following our previous work [37, 41], the hybrid DNN acoustic models containing six 20002000-dimensional and one 100100-dimensional hidden layers were implemented using an extension to the Kaldi toolkit [103]. As is shown in Fig. 5a, each of its hidden layer contains a set of neural operations performed in sequence. These include affine transformation (in green), rectified linear unit (ReLU) activation (in yellow) and batch normalization (in orange). Linear bottleneck projection (in light green) is applied to the inputs of the five intermediate hidden layers while dropout operation (in grey) is applied on the outputs of the first six hidden layers. Softmax activation (in dark green) is applied in the output layer. Two skip connections feed the outputs of the first hidden layer to those of the third and those of the fourth to the sixth respectively. Multi-task learning (MTL) [74] was used to train the hybrid DNN system with frame-level tied triphone states and monophone alignments obtained from a HTK [101] trained GMM-HMM system. The end-to-end (E2E) Conformer systems were implemented using the ESPnet toolkit [104]22 2 88 encoder layers + 44 decoder layers, feed-forward layer dim = 10241024, attention heads = 44, dim of attention heads = 256256, interpolated CTC+AED cost. to directly model grapheme (letter) sequence outputs. 8080-dimensional mel-scale filter-bank (FBK) + Δ\Delta features were used as input for both hybrid DNN and E2E Conformer systems while a 99-frame context window was used in the hybrid DNN system. The extraction of i-Vector33 3 Kaldi: egs/wsj/s5/local/nnet3/run_ivector_common.sh and x-Vector44 4 Kaldi: egs/sre16/v1/local/nnet3/xvector/tuning/run_xvector_1a.sh for UASpeech as well as the three other tasks follow the Kaldi recipe. Following the configurations given in [9, 12], a uniform language model with a word grammar network was used in decoding. Using the spectral basis embedding (SBE) features (d=2d=2) and temporal basis embedding (TBE) features (d=5d=5) trained on the UASpeech B1 plus B3 data considered here for speaker adaptation, their corresponding dysarthric v.s. control binary utterance-level classification accuracies measured on the B2 data of all 2929 speakers are 99.499.4% and 90.290.2% respectively.

V-A4 Experiment Setup for the TORGO Corpus

The hybrid factored time delay neural network (TDNN) systems containing 77 context slicing layers were trained following the Kaldi [103] chain system setup, as illustrated in Fig. 5b. The setup of the E2E graphemic Conformer system was the same as that for UASpeech. 4040-dimensional mel-scale FBK features were used as input for both hybrid TDNN and E2E Conformer systems while a 33-frame context window was used in the hybrid TDNN system. A 33-gram language model (LM) trained by all the TORGO transcripts with a vocabulary size of 1.61.6k was used during recognition with both the hybrid TDNN and E2E Conformer systems.

V-A5 Performance Analysis

TABLE III: Performance comparison between the proposed spectral and temporal basis embedding feature based adaptation against i-Vector, x-Vector and LHUC adaptation on the UASpeech test set of 1616 dysarthric speakers. “66M” and “2626M” refer to the number of model parameters. “DYS” and “CTL” in “Data Aug.” column standard for perturbing the dysarthric and the normal speech respectively for data augmentation. “SBE” and “TBE” denote spectral basis and temporal basis embedding features. “VL/L/M/H” refer to intelligibility subgroups. {\dagger} denotes a statistically significant improvement (α=0.05\alpha=0.05) is obtained over the comparable baseline i-Vector adapted systems (Sys. 2, 7, 12, 17, 22, 27 and 32).
Sys. Model (# Para.a) Data Aug. # Hrs Adapt. Feat. LHUC SAT WER%
VL L M H O.V.
1 Hybrid DNN (6M) 30.6 69.82 32.61 24.53 10.40 31.45
2 i-Vector 67.25 32.70 22.56 10.11 30.46
3 x-Vector 66.29 30.00 22.23 9.29 29.40
4 SBE 64.43 29.71 19.84 8.57 28.05
5 SBE+TBE 64.54 29.13 18.90 8.69 27.83
6 64.39 29.88 20.27 8.95 28.29
7 i-Vector 64.31 29.46 18.39 8.70 27.72
8 x-Vector 64.95 29.28 19.56 8.58 27.99
9 SBE 63.40 28.90 18.64 8.13 27.24
10 SBE+TBE 63.01 28.10 18.25 8.22 26.90
11 Hybrid DNN (6M) DYS 65.9 68.43 29.60 21.37 10.44 29.79
12 i-Vector 66.06 31.16 20.27 8.86 28.95
13 x-Vector 64.22 29.00 21.27 9.23 28.53
14 SBE 62.56 28.33 18.21 9.01 27.13
15 SBE+TBE 61.40 28.93 18.74 9.00 27.14
16 60.99 28.20 18.86 8.41 26.69
17 i-Vector 62.15 28.78 18.52 8.30 26.98
18 x-Vector 63.58 28.91 20.33 8.29 27.66
19 SBE 60.98 27.29 17.96 8.54 26.32
20 SBE+TBE 60.23 27.87 18.27 8.67 26.41
21 Hybrid DNN (6M) DYS + CTL 130.1 66.45 28.95 20.37 9.62 28.73
22 i-Vector 65.52 30.63 19.27 8.60 28.42
23 x-Vector 64.50 28.00 19.94 8.48 27.82
24 SBE 61.55 27.52 17.31 8.22 26.26
25 SBE+TBE 61.24 27.77 17.45 8.31 26.32
26 62.50 27.26 18.41 8.04 26.55
27 i-Vector 61.60 28.61 17.94 8.06 26.63
28 x-Vector 62.83 28.84 18.09 7.93 26.93
29 SBE 59.30 26.25 16.25 7.60 25.05
30 SBE+TBE 60.92 27.11 17.00 7.52 25.73
31 Conformer (19M) DYS + CTL 130.1 66.77 49.39 46.47 42.02 50.03
32 i-Vector 69.32 51.05 47.23 41.81 51.07
33 x-Vector 70.57 52.78 48.72 42.42 52.27
34 SBE 65.95 47.79 44.94 41.32 48.91
35 SBE+TBE 65.57 48.13 45.49 41.45 49.06

The performance of the proposed spectro and temporal deep feature based adaptation is compared with that obtained using conventional i-Vector [56] and x-Vector [86] based adaptation, as shown in Table III. Sys.1-30 were trained using the hybrid DNN system [41], where Sys.1-10 were trained on the 30.630.6h non-augmented training set, Sys.11-20 on the 65.965.9h training set augmented by speed perturbing the dysarthric speech only, and Sys.21-30 on the 130.1130.1h training set augmented by speed perturbing both the dysarthric and control healthy speech [37]. Sys.31-35 were trained using the E2E Conformer system on the 130.1130.1h augmented training set. The following trends can be observed:

i) The proposed spectral and temporal deep feature adapted systems consistently outperformed the comparable baseline speaker independent (SI) systems across all speech intelligibility subgroups with different amount of training data and baseline ASR system settings (Sys.4-5 v.s. Sys.1, Sys.14-15 v.s. Sys.11, Sys.24-25 v.s. Sys.21 and Sys.34-35 v.s. Sys.31) by up to 3.62%3.62\% absolute (11.51%11.51\% relative) statistically significant reduction in overall WER (Sys.5 v.s. Sys.1).

ii) When compared with conventional i-Vector and x-Vector based adaptation, our proposed spectro-temporal deep feature adapted systems consistently produced lower WERs across very low, low and mid speech intelligibility subgroups (Sys.4-5 v.s. Sys.2-3, Sys.14-15 v.s. Sys.12-13, Sys.24-25 v.s. Sys.22-23 and Sys.34-35 v.s. Sys.32-33). A statistically significant overall WER reduction of up to 2.63%2.63\% absolute (8.63%8.63\% relative) was obtained (Sys.5 v.s. Sys.2).

iii) When further combined with model based speaker adaptation via LHUC, the spectro-temporal deep feature adapted systems consistently achieved better performance than the systems with no auxiliary feature based adaptation or i-Vector and x-Vector based adaptation (Sys.9-10 v.s. Sys.6-8, Sys.19-20 v.s. Sys.16-18 and Sys.29- 30 v.s. Sys.26-28). A statistically significant reduction in overall WER by up to 1.58%1.58\% absolute (5.93%5.93\% relative) was obtained (Sys.29 v.s. Sys.27).

iv) Compared with using spectral basis embedding features (SBE) only in adaptation, using both spectral and temporal embedding features (SBE+TBE) leads to comparable performance but no consistent benefit. For example, marginal performance improvements were obtained on the non-augmented 30.630.6h training set (Sys.5 v.s. Sys.4 and Sys.10 v.s. Sys.9) while small performance degradation found on the other augmented training sets for both hybrid DNN (Sys.15 v.s. Sys.14, Sys.20 v.s. Sys.19, Sys.25 v.s. Sys.24 and Sys.30 v.s. Sys.29) and E2E Conformer systems (Sys.35 v.s. Sys.34). Based on these observations, only spectral basis embedding feature (SBE) based adaptation are considered in the remaining experiments of this paper.

A comparison between previously published systems on the UASpeech corpus and our system is shown in Table IV. To the best of our knowledge, this is the lowest WER obtained by ASR systems published so far on the UASpeech test set of 16 dysarthric speakers in the literature.

TABLE IV: A comparison between published systems on UASpeech and our system. Here “DA” refers to data augmentation and “GAN” stands for generative adversarial network.
Systems WER%
Sheffield-2013 Cross domain augmentation [10] 37.50
Sheffield-2015 Speaker adaptive training [11] 34.80
CUHK-2018 DNN System Combination [12] 30.60
Sheffield-2020 Fine-tuning CNN-TDNN speaker adaptation [40] 30.76
CUHK-2020 DNN + DA + LHUC SAT [37] 26.37
CUHK-2021 LAS + CTC + Meta-learning + SAT [45] 35.00
CUHK-2021 QuartzNet + CTC + Meta-learning + SAT [45] 30.50
CUHK-2021 DNN + GAN DA [16] 25.89
CUHK-2021 NAS DNN + DA + LHUC SAT + AV fusion [41] 25.21
DA + SBE Adapt + LHUC SAT (Table III, Sys.29) 25.05

On a comparable set of experiments conducted on the on the TORGO [83] corpus using the 34.134.1h augmented training set shown in Table V, trends similar to those found the on UASpeech task in Table III can be observed. Compared with the i-Vector adapted systems, a statistically significant overall WER reduction by up to 1.27%1.27\% absolute (14%14\% relative) (Sys.4 v.s. Sys.2) can be obtained using the spectral embedding feature (SBE) adapted TDNN systems. The SBE adapted Conformer system outperformed its i-Vector baseline statistically significantly by 1.98%1.98\% absolute (14.01%14.01\% relative) (Sys.12 v.s. Sys.10).

TABLE V: Performance comparison between the proposed spectral basis embedding feature based adaptation against i-Vector, x-Vector and LHUC adaptation on the TORGO test set of 88 dysarthric speakers. “1010M” and “1818M” refer to the number of model parameters. “DYS + CTL” in “Data Aug.” column denotes perturbing both dysarthric and normal speech in data augmentation. “SBE” denote spectral basis embedding features. “Seve./Mod./Mild” refer to the speech impairment severity levels: severe, moderate and mild. {\dagger} denotes a statistically significant improvement (α=0.05\alpha=0.05) is obtained over the comparable baseline i-Vector adapted systems (Sys. 2, 6 and 10).
Sys. Model (# Para.) Data Aug. # Hrs Adapt. Feat. LHUC WER%
Seve. Mod. Mild O.V.
1 Hybrid TDNN (10M) DYS + CTL 34.1 12.80 8.78 3.64 9.47
2 i-Vector 13.82 5.92 2.40 9.07
3 x-Vector 12.76 5.31 3.17 8.60
4 SBE 11.67 4.59 2.86 7.80
5 12.60 8.78 3.64 9.36
6 i-Vector 13.74 5.82 2.55 9.04
7 x-Vector 12.68 5.20 3.02 8.50
8 SBE 11.71 4.29 2.86 7.76
9 Conformer (18M) DYS + CTL 34.1 21.66 6.22 4.10 13.67
10 i-Vector 22.15 7.44 3.94 14.13
11 x-Vector 20.52 7.55 3.71 13.25
12 SBE 18.69 6.83 3.71 12.15

V-B Experiments on Elderly Speech

V-B1 the DementiaBank Pitt Corpus

The DementiaBank Pitt [84] corpus contains approximately 3333 hours of audio data recorded over interviews between the 292292 elderly participants and the clinical investigators. It is further split into a 27.227.2h training set, a 4.84.8h development and a 1.11.1h evaluation set for ASR system development. The evaluation set is based on exactly the same 4848 speakers’ Cookie (picture description) task recordings as those in the ADReSS55 5 http://www.homepages.ed.ac.uk/sluzfil/ADReSS/ [105] test set, while the development set contains the remaining recordings of these speakers in other tasks if available. The training set contains 688688 speakers (244244 elderly participants and 444444 investigators), while the development set includes 119119 speakers (4343 elderly participants and 7676 investigators) and the evaluation set contains 9595 speakers (4848 elderly participants and 4747 investigators). After removal of excessive silence [15], the training set contains 15.715.7 hours of audio data (2968229682 utterances) while the development and evaluation sets contain 2.52.5 hours (51035103 utterances) and 0.60.6 hours (928928 utterances) of audio data respectively. Data augmentation featuring speaker independent speed perturbation of elderly speech and elderly speaker dependent speed perturbation of non-aged investigators’ speech [15] produced an 58.958.9h augmented training set (112830112830 utterances).

V-B2 the JCCOCC MoCA Corpus

The Cantonese JCCOCC MoCA corpus contains conversations recorded from cognitive impairment assessment interviews between 256 elderly participants and the clinical investigators [85]. The training set contains 369369 speakers (158158 elderly participants and 211211 investigators) with a duration of 32.432.4 hours. The development and evaluation sets each contains speech recorded from 4949 elderly speakers. After removal of excessive silence, the training set contains 32.132.1 hours of speech (9544895448 utterances) while the development and evaluation sets contain 3.53.5 hours (1367513675 utterances) and 3.43.4 hours (1341413414 utterances) of speech respectively. After data augmentation following approaches similar to those adopted on the DementiaBank Pitt corpus [15], the augmented training set consists of 156.9156.9 hours of speech (389409389409 utterances).

V-B3 Experiment Setup for the DementiaBank Pitt Corpus

Following the Kaldi [103] chain system setup, the hybrid TDNN system shown in Fig. 5b contain 1414 context slicing layers with a 33-frame context. 4040-dimensional mel-scale FBK features were used as input for all systems. For both the hybrid TDNN and E2E graphemic Conformer systems66 6 1212 encoder layers + 1212 decoder layers, feed-forward layer dim = 20482048, attention heads = 44, dim of attention heads = 256256, interpolated CTC+AED cost., a word level 44-gram LM was trained following the settings of our previous work [15] and a 3.83.8k word recognition vocabulary covering all the words in the DementiaBank Pitt corpus was used in recognition. Using the spectral basis embedding (SBE) features (d=4d=4) considered here for speaker adaptation, the corresponding aged v.s. non-aged (participant v.s. investigator) utterance-level classification accuracy on the combined development plus evaluation set is 84.9%84.9\%.

V-B4 Experiment Setup for the JCCOCC MoCA Corpus

The architecture of the hybrid TDNN and E2E graphemic (character) Conformer systems were the same as those for the DementiaBank Pitt corpus above. 4040-dimensional mel-scale FBK features were used as input for all systems. A word level 44-gram language model with Kneser-Ney smoothing was trained on the transcription of the JCCOCC MoCA corpus (610610k words) using the SRILM toolkit [106] and a 5.25.2k recognition vocabulary covering all the words in the JCCOCC MoCA corpus was used.

V-B5 Performance Analysis

The performance comparison between the proposed spectral deep feature based adaptation against traditional i-Vector [56] and x-Vector [86] based adaptation using either hybrid TDNN [3] or E2E Conformer [6] systems on the DimentiaBank Pitt corpus with the 58.958.9h augmented training set is shown in Table VI, where trends similar to those found on the the dysarthric speech experiments in Table III and V can be observed:

i) The proposed spectral basis embedding feature (SBE) adapted systems consistently outperform the comparable baseline speaker independent (SI) systems with or without model based speaker adaptation using LHUC (Sys.4 v.s. Sys.1, Sys.8 v.s. Sys.5 and Sys.12 v.s. Sys.9) by up to 3.17%3.17\% absolute (9.81%9.81\% relative) overall WER reduction (Sys.8 v.s. Sys.5).

ii) When compared with conventional i-Vector and x-Vector based adaptation, our proposed SBE feature adapted systems consistently produced lower WERs with or without model based speaker adaptation using LHUC (Sys.4 v.s. Sys.2-3, Sys.8 v.s. Sys.6-7 and Sys.12 v.s. Sys.10-11). A statistically significant overall WER reduction of 2.57%2.57\% absolute (8.1%8.1\%relative) was obtained (Sys.8 v.s. Sys.6).

TABLE VI: Performance comparison between the proposed spectral basis embedding feature based adaptation against i-Vector, x-Vector and LHUC adaptation on the DementiaBank Pitt corpus. “1818M” and “5252M” refer to the number of model parameters. “SBE” denote spectral basis embedding features. “Dev” and “Eval” stand for the development and evaluation sets. “INV” and “PAR” refer to clinical investigator and elderly participant. {\dagger} denotes a statistically significant improvement (α=0.05\alpha=0.05) is obtained over the comparable baseline i-Vector adapted systems (Sys. 2, 6 and 10).
Sys. Model (# Para.) Data Aug. # Hrs Adapt. Feat. LHUC SAT WER%
Dev Eval O.V.
INV PAR INV PAR
1 Hybrid TDNN (18M) 58.9 19.91 47.93 19.76 36.66 33.80
2 i-Vector 19.97 46.76 18.20 37.01 33.37
3 x-Vector 18.01 46.42 18.76 37.62 32.56
4 SBE 18.61 43.84 17.98 33.82 31.12
5 19.26 45.49 18.42 35.44 32.33
6 i-Vector 18.62 44.70 17.98 35.38 31.73
7 x-Vector 17.93 45.76 16.76 36.11 31.95
8 SBE 17.41 40.94 17.98 31.89 29.16
9 Conformer (52M) 58.9 20.97 48.71 19.42 36.93 34.57
10 i-Vector 21.48 48.32 17.42 37.79 34.71
11 x-Vector 20.83 48.53 32.29 43.10 35.88
12 SBE 20.44 47.70 17.31 36.11 33.76
TABLE VII: Performance comparison between the proposed spectral basis embedding feature based adaptation against i-Vector, x-Vector and LHUC adaptation on the JCCOCC MoCA corpus development (Dev) and evaluation (Eval) sets containing elderly speakers only. “1818M” and “5353M” refer to the number of model parameters. “SBE” denote spectral basis embedding features. {\dagger} denotes a statistically significant improvement (α=0.05\alpha=0.05) is obtained over the comparable baseline i-Vector adapted systems (Sys. 2, 6 and 10).
Sys. Model (# Para.) Data Aug. # Hrs Adapt. Feat. LHUC CER%
Dev Eval O.V.
1 Hybrid TDNN (18M) 156.9 26.87 23.71 25.28
2 i-Vector 25.46 22.80 24.12
3 x-Vector 25.06 21.93 23.49
4 SBE 24.43 21.68 23.05
5 25.77 22.94 24.35
6 i-Vector 24.73 22.12 23.42
7 x-Vector 24.18 21.48 22.82
8 SBE 23.59 21.42 22.50
9 Conformer (53M) 156.9 33.08 31.24 32.15
10 i-Vector 33.76 31.83 32.79
11 x-Vector 33.79 32.25 33.02
12 SBE 32.08 30.75 31.41

Table VII shows the performance comparison of our proposed spectral basis embedding feature based adaptation against i-Vector, x-Vector and LHUC adaptation on the JCCOCC MoCA [85] data. Trends similar to those previously found on the DementiaBank Pitt corpus in Table VI can be observed in Table VII. Compared with the i-Vector adapted systems, a statistically significant overall WER reduction by up to 1.07%1.07\% absolute (4.44%4.44\% relative) (Sys.4 v.s. Sys.2) can be obtained using the spectral embedding feature adapted hybrid TDNN systems. The SBE adapted E2E Conformer system outperformed its i-Vector baseline statistically significantly by 1.38%1.38\% absolute (4.21%4.21\% relative) (Sys.12 v.s. Sys.10).

VI Discussion and Conclusions

This paper proposes novel spectro-temporal deep feature based speaker adaptation approaches for dysarthric and elderly speech recognition. Experiments were conducted on two dysarthric and two elderly speech datasets including the English UASpeech and TORGO dysarthric speech corpora as well as the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets. The best performing spectral basis embedding feature adapted hybrid DNN/TDNN and end-to-end Conformer based ASR systems consistently outperformed their comparable baselines using i-Vector and x-Vector adaptation across all four tasks covering both English and Cantonese. Experimental results suggest the proposed spectro-temporal deep feature based adaptation approaches can effectively capture speaker-level variability attributed to speech pathology severity and age, and facilitate more powerful personalized adaptation of ASR systems to cater for the needs of dysarthric and elderly users. Future researches will focus on fast, on-the-fly speaker adaptation using spectro-temporal deep features.

Acknowledgment

This research is supported by Hong Kong Research Grants Council GRF grant No. 14200218, 14200220, Theme based Research Scheme T45-407/19N, Innovation & Technology Fund grant No. ITS/254/19, PiH/350/20 and Shun Hing Institute of Advanced Engineering grant No. MMT-p1-19.

References

  • [1] L. Bahl et al., “Maximum mutual information estimation of hidden Markov model parameters for speech recognition,” in ICASSP, 1986.
  • [2] A. Graves et al., “Speech recognition with deep recurrent neural networks,” in ICASSP, 2013.
  • [3] V. Peddinti et al., “A time delay neural network architecture for efficient modeling of long temporal contexts,” in INTERSPEECH, 2015.
  • [4] W. Chan et al., “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in ICASSP, 2016.
  • [5] Y. Wang et al., “Transformer-based acoustic modeling for hybrid speech recognition,” in ICASSP, 2020.
  • [6] A. Gulati et al., “Conformer: Convolution-augmented Transformer for Speech Recognition,” in INTERSPEECH, 2020.
  • [7] S. Hu et al., “Neural Architecture Search For LF-MMI Trained Time Delay Neural Networks,” in ICASSP, 2021.
  • [8] S. Hu et al., “Bayesian Learning of LF-MMI Trained Time Delay Neural Networks for Speech Recognition,” IEEE T AUDIO SPEECH, vol. 29, 2021.
  • [9] H. Christensen et al., “A comparative study of adaptive, automatic recognition of disordered speech,” in INTERSPEECH, 2012.
  • [10] H. Christensen et al., “Combining in-domain and out-of-domain speech data for automatic recognition of disordered speech,” in INTERSPEECH, 2013.
  • [11] S. Sehgal et al., “Model adaptation and adaptive training for the recognition of dysarthric speech,” in SLPAT, 2015.
  • [12] J. Yu et al., “Development of the CUHK Dysarthric Speech Recognition System for the UA Speech Corpus,” in INTERSPEECH, 2018.
  • [13] S. Hu et al., “The CUHK Dysarthric Speech Recognition Systems for English and Cantonese,” in INTERSPEECH, 2019.
  • [14] S. Liu et al., “Exploiting cross-domain visual feature generation for disordered speech recognition,” in INTERSPEECH, 2020.
  • [15] Z. Ye et al., “Development of the CUHK Elderly Speech Recognition System for Neurocognitive Disorder Detection Using the Dementiabank Corpus,” in ICASSP, 2021.
  • [16] Z. Jin et al., “Adversarial Data Augmentation for Disordered Speech Recognition,” in INTERSPEECH, 2021.
  • [17] T. L. Whitehill et al., “Speech errors in Cantonese speaking adults with cerebral palsy,” CLIN LINGUIST PHONET, vol. 14, no. 2, 2000.
  • [18] T. Makkonen et al., “Speech deterioration in amyotrophic lateral sclerosis after manifestation of bulbar symptoms,” INT J LANG COMM DIS, vol. 53, no. 2, 2018.
  • [19] S. Scott et al., “Speech therapy for Parkinson’s disease,” J MACH LEARN RES, vol. 46, no. 2, 1983.
  • [20] P. Jerntorp et al., “Stroke registry in Malmö, Sweden,” STROKE, vol. 23, no. 3, 1992.
  • [21] W. Lanier, Speech disorders. Greenhaven Publishing LLC, 2010.
  • [22] K. C. Fraser et al., “Linguistic features identify Alzheimer’s disease in narrative speech,” J ALZHEIMERS DIS, vol. 49, no. 2, 2016.
  • [23] J. Wiley, “Alzheimer’s disease facts and figures,” Alzheimers Dement, vol. 17, no. 3, 2021.
  • [24] K. Hux et al., “Accuracy of three speech recognition systems: Case study of dysarthric speech,” AUGMENT ALTERN COMM, vol. 16, no. 3, 2000.
  • [25] V. Young et al., “Difficulties in automatic speech recognition of dysarthric speakers and implications for speech-based applications used by the elderly: A literature review,” ASSIST TECHNOL, vol. 22, no. 2, 2010.
  • [26] R. Vipperla et al., “Ageing voices: The effect of changes in voice parameters on ASR performance,” EURASIP J AUDIO SPEE, vol. 2010, 2010.
  • [27] F. Rudzicz et al., “Speech recognition in Alzheimer’s disease with personal assistive robots,” in SLPAT, 2014.
  • [28] L. Zhou et al., “Speech Recognition in Alzheimer’s Disease and in its Assessment,” in INTERSPEECH, 2016.
  • [29] B. Vachhani et al., “Deep Autoencoder Based Speech Features for Improved Dysarthric Speech Recognition,” in INTERSPEECH, 2017.
  • [30] M. J. Kim et al., “Dysarthric Speech Recognition Using Convolutional LSTM Neural Network,” in INTERSPEECH, 2018.
  • [31] A. König et al., “Fully automatic speech-based analysis of the semantic verbal fluency task,” DEMENT GERIATR COGN, vol. 45, no. 3-4, 2018.
  • [32] N. M. Joy et al., “Improving acoustic models in TORGO dysarthric speech database,” IEEE T NEUR SYS REH, vol. 26, no. 3, 2018.
  • [33] L. Tóth et al., “A speech recognition-based solution for the automatic detection of mild cognitive impairment from spontaneous speech,” CURR ALZHEIMER RES, vol. 15, no. 2, 2018.
  • [34] S. Liu et al., “Exploiting Visual Features Using Bayesian Gated Neural Networks for Disordered Speech Recognition,” in INTERSPEECH, 2019.
  • [35] B. Mirheidari et al., “Dementia detection using automatic analysis of conversations,” COMPUT SPEECH LANG, vol. 53, 2019.
  • [36] J. Shor et al., “Personalizing ASR for Dysarthric and Accented Speech with Limited Data,” in INTERSPEECH, 2019.
  • [37] M. Geng et al., “Investigation of Data Augmentation Techniques for Disordered Speech Recognition,” in INTERSPEECH, 2020.
  • [38] Y. Lin et al., “Staged Knowledge Distillation for End-to-End Dysarthric Speech Recognition and Speech Attribute Transcription,” in INTERSPEECH, 2020.
  • [39] I. Kodrasi et al., “Spectro-Temporal Sparsity Characterization for Dysarthric Speech Detection,” IEEE T AUDIO SPEECH, vol. 28, 2020.
  • [40] F. Xiong et al., “Source Domain Data Selection for Improved Transfer Learning Targeting Dysarthric Speech Recognition,” in ICASSP, 2020.
  • [41] S. Liu et al., “Recent Progress in the CUHK Dysarthric Speech Recognition System,” IEEE T AUDIO SPEECH, vol. 29, 2021.
  • [42] X. Xie et al., “Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition,” in INTERSPEECH, 2021.
  • [43] R. Takashima et al., “Two-step acoustic model adaptation for dysarthric speech recognition,” in ICASSP, 2020.
  • [44] E. Hermann et al., “Dysarthric speech recognition with lattice-free MMI,” in ICASSP, 2020.
  • [45] D. Wang et al., “Improved end-to-end dysarthric speech recognition via meta-learning based model re-initialization,” in ISCSLP, 2021.
  • [46] P. Wang et al., “A Study into Pre-training Strategies for Spoken Language Understanding on Dysarthric Speech,” in INTERSPEECH, 2021.
  • [47] B. MacDonald et al., “Disordered speech data collection: Lessons learned at 1 million utterances from project euphonia,” in INTERSPEECH, 2021.
  • [48] J. R. Green et al., “Automatic Speech Recognition of Disordered Speech: Personalized models outperforming human listeners on short phrases,” in INTERSPEECH, 2021.
  • [49] Y. Pan et al., “Using the Outputs of Different Automatic Speech Recognition Paradigms for Acoustic-and BERT-Based Alzheimer’s Dementia Detection Through Spontaneous Speech,” in INTERSPEECH, 2021.
  • [50] M. Geng et al., “Spectro-Temporal Deep Features for Disordered Speech Assessment and Recognition,” in INTERSPEECH, 2021.
  • [51] B. L. Smith et al., “Temporal characteristics of the speech of normal elderly adults,” J SPEECH LANG HEAR R, vol. 30, no. 4, 1987.
  • [52] R. D. Kent et al., “What dysarthrias can tell us about the neural control of speech,” J PHONETICS, vol. 28, no. 3, 2000.
  • [53] B. Vachhani et al., “Data Augmentation Using Healthy Speech for Dysarthric Speech Recognition,” in INTERSPEECH, 2018.
  • [54] F. Xiong et al., “Phonetic analysis of dysarthric speech tempo and applications to robust personalised dysarthric speech recognition,” in ICASSP, 2019.
  • [55] O. Abdel-Hamid et al., “Fast speaker adaptation of hybrid NN/HMM model for speech recognition based on discriminative learning of speaker code,” in ICASSP, 2013.
  • [56] G. Saon et al., “Speaker adaptation of neural network acoustic models using i-vectors,” in ASRU, 2013.
  • [57] A. Senior et al., “Improving DNN speaker independence with i-vector inputs,” in ICASSP, 2014.
  • [58] H. Huang et al., “An investigation of augmenting speaker representations to improve speaker normalisation for dnn-based speech recognition,” in ICASSP, 2015.
  • [59] V. V. Digalakis et al., “Speaker adaptation using constrained estimation of Gaussian mixtures,” IEEE T SPEECH AUDI P, vol. 3, no. 5, 1995.
  • [60] L. Lee et al., “Speaker normalization using efficient frequency warping procedures,” in ICASSP, 1996.
  • [61] M. J. Gales, “Maximum likelihood linear transformations for HMM-based speech recognition,” COMPUT SPEECH LANG, vol. 12, no. 2, 1998.
  • [62] L. F. Uebel et al., “An investigation into vocal tract length normalisation,” in EUROSPEECH, 1999.
  • [63] F. Seide et al., “Feature engineering in context-dependent deep neural networks for conversational speech transcription,” in ASRU, 2011.
  • [64] J. Neto et al., “Speaker-adaptation for hybrid HMM-ANN continuous speech recognition system,” in EUROSPEECH, 1995.
  • [65] T. Anastasakos et al., “A compact model for speaker-adaptive training,” in ICSLP, 1996.
  • [66] R. Gemello et al., “Linear hidden transformations for adaptation of hybrid ANN/HMM models,” SPEECH COMMUN, vol. 49, no. 10-11, 2007.
  • [67] B. Li et al., “Comparison of discriminative input and output transformations for speaker adaptation in the hybrid NN/HMM systems,” in INTERSPEECH, 2010.
  • [68] P. Swietojanski et al., “Learning hidden unit contributions for unsupervised acoustic model adaptation,” IEEE T AUDIO SPEECH, vol. 24, no. 8, 2016.
  • [69] C. Zhang et al., “DNN speaker adaptation using parameterised sigmoid and ReLU hidden activation functions,” in ICASSP, 2016.
  • [70] Z. Tüske et al., “On the limit of English conversational speech recognition,” in INTERSPEECH, 2021.
  • [71] P. Swietojanski et al., “Learning hidden unit contributions for unsupervised speaker adaptation of neural network acoustic models,” in SLT, 2014.
  • [72] C. Zhang et al., “Parameterised sigmoid and ReLU hidden activation functions for DNN acoustic modelling,” in INTERSPEECH, 2015.
  • [73] Y. Zhao et al., “Low-rank plus diagonal adaptation for deep neural networks,” in ICASSP, 2016.
  • [74] C. Wu et al., “Multi-basis adaptive neural network for rapid adaptation in speech recognition,” in ICASSP, 2015.
  • [75] T. Tan et al., “Cluster adaptive training for deep neural network based acoustic model,” IEEE T AUDIO SPEECH, vol. 24, no. 3, 2015.
  • [76] A. BABA et al., “Elderly Acoustic Models for Large Vocabulary Continuous Speech Recognition,” IEICE T INF SYST, vol. 85, no. 3, 2002.
  • [77] K. T. Mengistu et al., “Adapting acoustic and lexical models to dysarthric speech,” in ICASSP, 2011.
  • [78] M. J. Kim et al., “Dysarthric speech recognition using dysarthria-severity-dependent and speaker-adaptive models,” in INTERSPEECH, 2013.
  • [79] C. Bhat et al., “Recognition of Dysarthric Speech Using Voice Parameters for Speaker Adaptation and Multi-Taper Spectral Estimation,” in INTERSPEECH, 2016.
  • [80] M. Kim et al., “Regularized speaker adaptation of KL-HMM for dysarthric speech recognition,” IEEE T NEUR SYS REH, vol. 25, no. 9, 2017.
  • [81] J. Deng et al., “Bayesian Parametric and Architectural Domain Adaptation of LF-MMI Trained TDNNs for Elderly and Dysarthric Speech Recognition,” in INTERSPEECH, 2021.
  • [82] H. Kim et al., “Dysarthric speech database for universal access research,” in INTERSPEECH, 2008.
  • [83] F. Rudzicz et al., “The TORGO database of acoustic and articulatory speech from speakers with dysarthria,” LANG RESOUR EVAL, vol. 46, no. 4, 2012.
  • [84] J. T. Becker et al., “The natural history of Alzheimer’s disease: description of study cohort and accuracy of diagnosis,” ARCH NEUROL-CHICAGO, vol. 51, no. 6, 1994.
  • [85] S. S. Xu et al., “Speaker Turn Aware Similarity Scoring for Diarization of Speech-Based Cognitive Assessments.”
  • [86] D. Snyder et al., “X-vectors: Robust dnn embeddings for speaker recognition,” in ICASSP, 2018.
  • [87] L. Van der Maaten et al., “Visualizing data using t-SNE,” J MACH LEARN RES, vol. 9, no. 11, 2008.
  • [88] G. An et al., “Automatic recognition of unified Parkinson’s disease rating from speech with acoustic, i-vector and phonotactic features,” in INTERSPEECH, 2015.
  • [89] N. Garcia et al., “Multimodal I-vectors to Detect and Evaluate Parkinson’s Disease,” in INTERSPEECH, 2018.
  • [90] P. Janbakhshi et al., “Subspace-based Learning for Automatic Dysarthric Speech Detection,” IEEE SIGNAL PROC LET, 2020.
  • [91] A.-J. Van Der Veen et al., “Subspace-based signal analysis using singular value decomposition,” P IEEE, vol. 81, no. 9, 1993.
  • [92] D. D. Lee et al., “Learning the parts of objects by non-negative matrix factorization,” Nature, vol. 401, no. 6755, 1999.
  • [93] C. Févotte et al., “Nonnegative matrix factorization with the Itakura-Saito divergence: With application to music analysis,” NEURAL COMPUT, vol. 21, no. 3, 2009.
  • [94] D. Wang et al., “Online non-negative convolutive pattern learning for speech signals,” IEEE T SIGNAL PROCES, vol. 61, no. 1, 2012.
  • [95] F. Locatello et al., “A Sober Look at the Unsupervised Learning of Disentangled Representations and their Evaluation,” J MACH LEARN RES, vol. 21, 2020.
  • [96] D. J. Berndt et al., “Using dynamic time warping to find patterns in time series,” in KDD, 1994.
  • [97] T. Bocklet et al., “Automatic evaluation of parkinson’s speech-acoustic, prosodic and voice related cues,” in INTERSPEECH, 2013.
  • [98] D. M. Blei et al., “Latent dirichlet allocation,” J MACH LEARN RES, vol. 3, no. Jan, 2003.
  • [99] M. Doulaty et al., “Latent dirichlet allocation based organisation of broadcast media archives for deep neural network adaptation,” in ASRU, 2015.
  • [100] L. Gillick et al., “Some statistical issues in the comparison of speech recognition algorithms,” in ICASSP, 1989.
  • [101] S. Young et al., “The HTK book,” Cambridge university engineering department, vol. 3, no. 175, 2002.
  • [102] S. Hu et al., “Exploiting Cross Domain Acoustic-to-articulatory Inverted Features for Disordered Speech Recognition,” in ICASSP, 2022.
  • [103] D. Povey et al., “The Kaldi speech recognition toolkit,” in ASRU, 2011.
  • [104] S. Watanabe et al., “ESPnet: End-to-End Speech Processing Toolkit,” in INTERPSEECH, 2018.
  • [105] S. Luz et al., “Alzheimer’s Dementia Recognition through Spontaneous Speech: The ADReSS Challenge,” in INTERSPEECH, 2020.
  • [106] A. Stolcke, “SRILM-an extensible language modeling toolkit,” in ICSLP, 2002.