Please use this identifier to cite or link to this item:
http://hdl.handle.net/10603/459881
Title: | Automatic Dialect Identification in Ao, a Low Resource Language |
Researcher: | Tzudir, Moakala |
Guide(s): | Prasanna, S R M and Sarmah, Priyankoo |
Keywords: | Engineering Engineering and Technology Engineering Electrical and Electronic |
University: | Indian Institute of Technology Guwahati |
Completed Date: | 2023 |
Abstract: | Dialect Identification (DID) is a significant research problem widely explored in major languages like Arabic, Chinese, and Spanish. DID can serve as a frontend for many applications like Automatic Speech Recognition (ASR) that may require special dialect-specific enhancements for improved performance. This thesis proposes an automatic DID system for Ao, an under-resourced language of India. Ao is a Tibeto-Burman language spoken in Nagaland. It is a tonal language with three lexical tones: high, mid, and low. Chungli, Mongsen, and Changki are the three dialects of Ao that differ in their respective tone assignment on lexical words. Four principal contributions are made in this thesis. The first contribution of this thesis is creating a manually collected and annotated novel speech dataset to foster research on the Ao language. The second contribution of the thesis is a detailed acoustic study of the unexplored tone dynamics of the dialects of Ao. Based on the analysis, a tonal feature ($F_0$) to capture the dialect-specific tone information is proposed. The DID performance improves when the proposed tonal feature is combined with other spectral features. As the third contribution, this thesis explores three excitation source features in the DID task. The source features studied are Residual Mel Frequency Cepstral Coefficient (RMFCC), Integrated Linear Prediction Residual Log Mel Spectrogram (ILPR-LMS), and Linear Prediction (LP)-gammatonegram. A notable performance improvement is observed when the source information is combined with the vocal tract information. The fourth contribution of this thesis is the exploration of prosody-related characteristics of speech signals. The prosodic features are observed to provide significant performance improvements in classifying the dialects of Ao. The thesis work is concluded by combining all the proposed approaches to build an efficient DID system for Ao. Among many hurdles in studying under-resourced languages like Ao, the need for more data is the most prominent. Nevertheless, the contributions of this thesis may bridge some of those gaps and spur future research in this direction. |
URI: | http://hdl.handle.net/10603/459881 |
Appears in Departments: | DEPARTMENT OF ELECTRONICS AND ELECTRICAL ENGINEERING |
Files in This Item:
File | Description | Size | Format | |
---|---|---|---|---|
01_fulltext.pdf | Attached File | 9.78 MB | Adobe PDF | View/Open |
04_abstract.pdf | 98.08 kB | Adobe PDF | View/Open | |
80_recommendation.pdf | 292.12 kB | Adobe PDF | View/Open |
Items in Shodhganga are licensed under Creative Commons Licence Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0).
Altmetric Badge: