Multilingual prediction of Alzheimer’s disease through domain adaptation and concept-based language modeling

Fraser, K. C., Linz, N., Li, B., Fors, K. L., Rudzicz, F., König, A., … & Kokkinakis, D. (2019, June). Multilingual prediction of Alzheimer’s disease through domain adaptation and concept-based language modelling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 3659-3670).

Abstract

There is growing evidence that changes in speech and language may be early markers of dementia, but much of the previous NLP work in this area has been limited by the size of the available datasets. Here, we compare sev- eral methods of domain adaptation to augment a small French dataset of picture descriptions (= 57) with a much larger English dataset (= 550), for the task of automatically distin- guishing participants with dementia from con- trols. The first challenge is to identify a set of features that transfer across languages; in addition to previously used features based on information units, we introduce a new set of features to model the order in which information units are produced by dementia patients and controls. These concept-based language model features improve classification performance in both English and French separately, and the best result (AUC = 0.89) is achieved using the multilingual training set with a combination of information and language model features.