A Discriminative Hierarchical PLDA-based Model for Spoken Language Recognition

Authors: L. Ferrer, D. Castan, M. McLaren and A. Lawson

Abstract: Spoken language recognition (SLR) refers to the automatic process used to determine the language present in a speech sample. SLR is an important task in its own right, for example, as a tool to analyze or categorize large amounts of multi-lingual data. Further, it is also an essential tool for selecting downstream applications in a work flow, for example, to chose appropriate speech recognition or machine translation models. SLR systems are usually composed of two stages, one where an embedding representing the audio sample is extracted and a second one which computes the final scores for each language. In this work, we approach the SLR task as a detection problem and implement the second stage as a probabilistic linear discriminant analysis (PLDA) model. We show that discriminative training of the PLDA parameters gives large gains with respect to the usual generative training. Further, we propose a novel hierarchical approach where two PLDA models are trained, one to generate scores for clusters of highly-related languages and a second one to generate scores conditional to each cluster. The final language detection scores are computed as a combination of these two sets of scores. The complete model is trained discriminatively to optimize a cross-entropy objective. We show that this hierarchical approach consistently outperforms the non-hierarchical one for detection of highly related languages, in many cases by large margins. We train our systems on a collection of datasets including over 100 languages, and test them both on matched and mismatched conditions, showing that the gains are robust to condition mismatch.

More information: https://arxiv.org/abs/2201.01364

Andres Juarez2022-12-16T13:23:41-03:00 16/diciembre/2022|Papers|

Activity homogeneity: a measure for comparing time discretization and state quantization in ODE simulation

A note on busy beaver bounds

EnCodecMAE: Leveraging Neural Codecs for Universal Audio Representation Learning

Are Optimal Algorithms Still Optimal? Rethinking Sorting in LLM-Based Pairwise Ranking with Batching and Caching

Modal Abstractions for Smart Contract Validation

Integrating Bayesian and neural networks models for eye movement prediction in hybrid search

Algorithms to prove the maximum number of MUBs in arbitrary dimensión

Rauzy complexity and block entropy

Hybrid resource allocation control in cyber-physical systems: a novel simulation-driven methodology with applications to UAVs

Mapping Semantic Segmentation to Point Clouds Using Structure from Motion for Forest Analysis

A multi-scale agent-based model of aerosol-mediated indoor infections in heterogeneous scenarios

Non-crossing H-graphs: a generalization of proper interval graphs admitting FPT algorithms

The discrepancy estimate of the Champernowne constant

No Need for Ad-hoc Substitutes: The Expected Cost is a Principled All-purpose Classification Metric

Low-cost algorithms for clinical notes phenotype classification to enhance epidemiological surveillance: A case study

A Discriminative Hierarchical PLDA-based Model for Spoken Language Recognition

Compartir en las redes

Related Posts