Audiology - Communication Research
https://www.audiolcommres.org.br/article/doi/10.1590/2317-6431-2025-3117pt
Audiology - Communication Research
Original Article

Processamento de arquivos de áudio para composição de testes de reconhecimento de fala em português

Processing of audio files for the construction of speech recognition tests in portuguese

Kauana Knappmann; Stephan Paul

Downloads: 0
Views: 42

Resumo

Objetivo: desenvolver códigos computacionais para modificação automatizada de grandes quantidades de sentenças gravadas, que realizem modificações de formato, filtragens, simulem o processamento dos sinais sonoros em implantes cocleares e ajustem as médias quadráticas de amplitude, a fim de equalizar o volume sonoro percebido entre sentenças.

Métodos: para os diferentes processamentos pretendidos, foram desenvolvidos códigos em Python, usando as interfaces Spyder e pacotes tais como o pydub, soundfile, os e numpy. Os códigos foram testados em dois conjuntos de arquivos de áudio gravados previamente em português brasileiro, nos formatos .MP3 e .WAV.

Resultados: foram implementados códigos para modificação do formato dos arquivos, ajuste de fade-in e fade-out, filtragem de passa-alta, vocoderização opcional e ajuste das médias quadráticas das amplitudes. Os testes dos códigos desenvolvidos em dois conjuntos de sentenças disponíveis em .WAV e .MP3 na língua portuguesa demonstraram resultados consistentes com o esperado.

Conclusão: desenvolveram-se códigos na linguagem Python para modificação de maneira automatizada de arquivos de áudio, disponíveis no site GitHub para adaptações e aprimoramentos por terceiros.

Palavras-chave

Audiologia; Processamento de sinais assistido por computador; Testes auditivos; Implante coclear; Simulação por computador

Abstract

Purpose: To develop computational codes for the automated modification of large quantities of sentence recordings, capable of performing format modifications, filtering, simulating sound signal processing in cochlear implants, and adjusting root mean square amplitude to equalize perceived volume between sentences.

Methods: Python codes were developed for the intended processes, using the Spyder interface and packages such as pydub, soundfile, os, and numpy. The codes were tested on two sets of previously recorded audio files in Brazilian Portuguese, in .MP3 and .WAV formats.

Results: Codes were implemented for 1) file format modification, 2) fade-in and fade-out adjustment, 3) high-pass filtering, 4) optional vocoderization, and 5) adjustment of root mean square amplitude. Testing the developed codes on two sets of sentence recordings available in .WAV and .MP3 formats in Portuguese showed consistent results as expected.

Conclusion: Python codes were developed for the automated modification of audio files, available on the GitHub website for further adaptations and improvements by third parties.

Keywords

Audiology; Signal processing computer-assisted; Hearing tests; Cochlear implantation; Computer simulation

References

1 Força MT, Cal RVR, Santos SR, Pereira LC, Boas GPV, Zell RGA, et al. Análise comparativa da percepção da perda auditiva com o resultado da audiometria em pacientes adultos e idosos do Hospital Bettina Ferro de Souza/PA. BJHR. 2020;3(6):17457-73. https://doi.org/10.34119/bjhrv3n6-162.

2 WHO: World Health Organization. World report on hearing [Internet]. Geneva: WHO; 2021 [citado em 2025 junho 25]. Disponível em: https://www.who.int/publications/i/item/9789240020481

3 Costa LD, Vaucher AVA, Pagliarin KC, Costa MJ. Teste de palavras no ruído: desenvolvimento, validação e valores de referência. CoDAS. 2024;36(3):e20230091. https://doi.org/10.1590/2317-1782/20242023091en. PMid:38836822.

4 Campos Salvato C, Araújo SRS, Muller R, Soares AD, Chiari BM. Correlação entre reconhecimento de fala, tempo de privação auditiva e tempo de uso de Implante Coclear em usuários com surdez pós-lingual. Distúrb Comun. 2020;32(3):396-405. https://doi.org/10.23925/2176-2724.2020v32i3p396-405.

5 Ferreira MC, Zamberlan-Amorim NE, Wolf AE, Reis ACMB. Influence of different types of noise on sentence recognition in normally hearing adults. Rev CEFAC. 2021;23(5):e2121. https://doi.org/10.1590/1982-0216/20212352121.

6 Wilson BS, Dorman MF. Cochlear implants: current designs and future possibilities. JRRD. 2008;45(5):695-730. https://doi.org/10.1682/JRRD.2007.10.0173. PMid:18816422.

7 Gifford RH, Shallop JK, Peterson AM. Speech recognition materials and ceiling effects: considerations for cochlear implant programs. Audiol Neurootol. 2008;13(3):193-205. https://doi.org/10.1159/000113510. PMid:18212519.

8 Johnson PA, McNamara DM, Ziarani AK. A novel VOCODER for cochlear implants. In: 30th Annual International Conference of the IEEE Engineering in Medicine and Biology Society; 2008; Vancouver, BC. New York: IEEE; 2008. p. 4732-5.

9 Pinheiro MMC, Vieira MG, Vieira LM, Koerich I, Rosseto I, Lazzarotto-Volcão C, et al. Adaptação de listas de sentenças para avaliação da percepção da fala. CoDAS. 2022;34(1):e20200301. https://doi.org/10.1590/2317-1782/20202020301. PMid:35019063.

10 Sbompato AF, Corteletti LCBJ, Moret ADLM, Jacob RTDS. Hearing in Noise Test Brazil: standardization for young adults with normal hearing. Braz J Otorhinolaryngol. 2015;81(4):384-8. https://doi.org/10.1016/j.bjorl.2014.07.018. PMid:26130593.

11 Costa MJ. Listas de sentenças do português. Santa Maria: Pallotti; 1998. 48 p.

12 Valente SLO. Elaboração de listas de sentenças construídas na língua portuguesa [dissertação]. São Paulo: Pontifícia Universidade Católica; 1998.

13 Wilson RH, McArdle RA, Smith SL. An evaluation of the BKB-SIN, HINT, QuickSIN, and WIN materials on listeners with normal hearing and listeners with hearing loss. J Speech Lang Hear Res. 2007;50(4):844-56. https://doi.org/10.1044/1092-4388(2007/059). PMid:17675590.

14 Spahr AJ, Dorman MF, Litvak LM, Van Wie S, Gifford RH, Loizou PC, et al. Development and validation of the AzBio sentence lists. Ear Hear. 2012;33(1):112-7. https://doi.org/10.1097/AUD.0b013e31822c2549. PMid:21829134.

15 Massa ST, Ruckenstein MJ. Comparing the performance plateau in adult cochlear implant patients using HINT and AzBio. Otol Neurotol. 2014;35(4):598-604. https://doi.org/10.1097/MAO.0000000000000264. PMid:24557031.

16 Jung C, Shah KV, Ferraro T, Bigelow DC, Ruckenstein MJ, Peng-Hwa T. Non-english validation of the azbio sentence test: a systematic review of current applications and early adoption patterns. Otol Neurotol. 2026;47(1):e8391. https://doi.org/10.1097/MAO.0000000000004663. PMid:41145396.

17 Nilsson M, Soli SD, Sullivan JA. Development of the Hearing In Noise Test for the measurement of speech reception thresholds in quiet and in noise. J Acoust Soc Am. 1994;95(2):1085-99. https://doi.org/10.1121/1.408469. PMid:8132902.

18 Upadhyay N, Karmakar A. A multi-band speech enhancement algorithm exploiting iterative processing for enhancement of single channel speech. J Signal Inf Process. 2013;4(2):197-211. https://doi.org/10.4236/jsip.2013.42027.

19 Karoui C, James C, Barone P, Bakhos D, Marx M, Macherey O. Searching for the sound of a cochlear implant: evaluation of different vocoder parameters by cochlear implant users with single-sided deafness. Trends Hear. 2019;23:2331216519866029. https://doi.org/10.1177/2331216519866029. PMid:31533581.

20 Välimäki V, Reiss J. All about audio equalization: solutions and frontiers. Appl Sci. 2016;6(5):129. https://doi.org/10.3390/app6050129. ]

21 Raj VA, Dhas MDK. Analysis of audio signal using various transforms for enhanced audio processing. Int J Health Sci. 2022;6(2):8890-7.

22 Berglund B, Hassmén P, Job RFS. Sources and effects of low-frequency noise. J Acoust Soc Am. 1996;99(5):2985-3002. https://doi.org/10.1121/1.414863. PMid:8642114.

23 Lupşa-Tătaru L. Customizing audio fades with a view to real-time processing. Applied Computer Science. 2019;15(4):16-26. https://doi.org/10.35784/acs-2019-27.

24 Ricketts TA, Bentler RA. The effect of test signal type and bandwidth on the categorical scaling of loudness. J Acoust Soc Am. 1996;99(4):2281-7. https://doi.org/10.1121/1.415415. PMid:8730074.

25 Emily Shannon Fu Foundation. AngelSim (TigerCIS): Cochlear Implant and Hearing Loss Simulator [Internet]. 2013 [citado em 2025 junho 25]. Disponível em: http://angelsim.emilyfufoundation.org/

26 Cychosz M, Winn MB, Goupell MJ. How to vocode: uing channel vocoders for cochlear-implant research. J Acoust Soc Am. 2024;155(4):2407-37. https://doi.org/10.1121/10.0025274. PMid:38568143.

27 The Python Standard Library [Internet]. 2001 [citado em 2025 junho 25]. Disponível em: https://docs.python.org/3/library/index.html

28 Kent RD, Read C. Acoustic analysis of speech. 2nd ed. San Diego: Singular Publishing Group; 2002.

29 Greenwood DD. Auditory masking and the critical band. J Acoust Soc Am. 1961;33(4):484-502. https://doi.org/10.1121/1.1908699.

30 Forinash K, Christian W. Sound: an interactive ebook [Internet]. 2012 [citado em 2025 junho 25]. Disponível em: https://www.compadre.org/books/SoundBook
 


Submitted date:
11/12/2025

Accepted date:
04/01/2026

6a72225da953957f467c9e23 acr Articles
Links & Downloads

Audiol. Commun. Res.

Share this page
Page Sections