skip to main content
Invitado
Mi portal
Mi Cuenta
Cerrar sesión
Identificarse
This feature requires javascript
Tags
Periódicos Eletrónicos
Libros Eletrónicos
Bases de Datos
Bibliotecas de USP
Ayuda
Ayuda
Idioma:
Inglés
Castellano
Portugués (Brasil)
This feature required javascript
This feature requires javascript
Primo Advanced Search
Búsqueda General
Búsqueda General
Colección Física
Colecciones Físicas
Producción Intelectual USP
Producción USP
Primo Advanced Search Query Term
Input search text:
Show Results with:
criteria input
Cualquiera
Show Results with:
Cualquiera
Primo Advanced Search prefilters
Tipo de material:
criteria input
Todos los registros
Búsqueda General
Búsqueda Sencilla
This feature requires javascript
Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification
Li, Jingyu ; Tian, Yusheng ; Tan, Lee
arXiv.org, 2022-10
Ithaca: Cornell University Library, arXiv.org
Texto completo disponible
Citas
Citado por
Recurso en línea
Detalles
Comentarios y Etiquetas
Servicios adicionales
Veces citado
This feature requires javascript
Acciones
Agregar a Mi Portal
Eliminar de Mi Portal
Correo Electrónico
Imprimir
Enlae permanente
Cita bibliográfica
EasyBib
EndNote
RefWorks
Delicious
Exportación RIS
Exportar BibTeX
This feature requires javascript
Título:
Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification
Autor:
Li, Jingyu
;
Tian, Yusheng
;
Tan, Lee
Materias:
Artificial neural networks
;
Computer Science - Sound
;
Verification
Es parte de:
arXiv.org, 2022-10
Descripción:
Deep convolutional neural networks (CNNs) have been applied to extracting speaker embeddings with significant success in speaker verification. Incorporating the attention mechanism has shown to be effective in improving the model performance. This paper presents an efficient two-dimensional convolution-based attention module, namely C2D-Att. The interaction between the convolution channel and frequency is involved in the attention calculation by lightweight convolution layers. This requires only a small number of parameters. Fine-grained attention weights are produced to represent channel and frequency-specific information. The weights are imposed on the input features to improve the representation ability for speaker modeling. The C2D-Att is integrated into a modified version of ResNet for speaker embedding extraction. Experiments are conducted on VoxCeleb datasets. The results show that C2DAtt is effective in generating discriminative attention maps and outperforms other attention methods. The proposed model shows robust performance with different scales of model size and achieves state-of-the-art results.
Editor:
Ithaca: Cornell University Library, arXiv.org
Idioma:
Inglés
Enlaces
View paper in arXiv
This feature requires javascript
This feature requires javascript
Volver a la lista de resultados
Anterior
Resultado
9
Siguiente
This feature requires javascript
This feature requires javascript
Buscando en bases de datos remotas, por favor espere
Buscando por
en
scope:(USP_VIDEOS),scope:("PRIMO"),scope:(USP_FISICO),scope:(USP_EREVISTAS),scope:(USP),scope:(USP_EBOOKS),scope:(USP_PRODUCAO),primo_central_multiple_fe
Mostrar lo que tiene hasta ahora
This feature requires javascript
This feature requires javascript