Skip to main content
QUICK REVIEW

[Paper Review] Classification of Autistic and Non-Autistic Children's Speech: A Cross-Linguistic Study in Finnish, French, and Slovak

Sofoklis Kakouros, Ida-Lotta Myllylä|arXiv (Cornell University)|Mar 6, 2026
Autism Spectrum Disorder Research0 citations
TL;DR

The study trains simple acoustic-prosodic classifiers (88 openSMILE features) to distinguish ASD vs TD speech in Finnish, French, and Slovak, evaluating within-language performance and cross-language generalization with pooled and LOCO setups.

ABSTRACT

We present a cross-linguistic study of speech in autistic and non-autistic children speaking Finnish, French, and Slovak. We combine supervised classification with within-language and cross-corpus transfer experiments to evaluate classification performance within and across languages and to probe which acoustic cues are language-specific versus language-general. Using a large set of acoustic-prosodic features, we implement speaker-level classification benchmarks as an analytical tool rather than to seek state-of-the-art performance. Within-language models, evaluated with speaker-level cross-validation, yielded heterogeneous results. The Finnish model performed best (Accuracy 0.84, F1 0.88), followed by Slovak (Accuracy 0.63, F1 0.68) and French (Accuracy 0.68, F1 0.56). We then tested cross-language generalization. A model trained on all pooled corpora reached an overall Accuracy of 0.61 and F1 0.68. Leave-one-corpus-out experiments, which test transfer to an unseen language, showed moderate success when testing on Slovak (F1 0.70) and Finnish (F1 0.78), but poor transfer to French (F1 0.42). Feature-importance analyses across languages highlighted partially shared, but not fully language-invariant, acoustic markers of autism. These findings suggest that some autism-related speech cues generalize across typologically distinct languages, but robust cross-linguistic classifiers will likely require language-aware modeling and more homogeneous recording conditions.

Motivation & Objective

  • Assess within-language ASD vs TD discrimination using prosodic features across three languages.
  • Evaluate cross-language generalization of classifiers trained on one or more languages.
  • Identify language-general versus language-specific acoustic markers of autism via feature importance analyses.
  • Use speaker-level cross-validation to obtain realistic generalization estimates.
  • Explore how recording conditions and language shape cross-linguistic robustness of classifiers.

Proposed method

  • Extract 88-dimensional acoustic-prosodic features with openSMILE eGeMAPS utterance-level functionals for each IPU.
  • Train XGBoost and Random Forest classifiers to distinguish ASD vs TD at the speaker level.
  • Perform within-language cross-validation to assess language-specific cues.
  • Conduct cross-linguistic analyses with pooled multilingual training and leave-one-corpus-out (LOCO) evaluation.
  • Analyze feature importance with tree-based importance measures, TreeSHAP, and permutation importance.
  • Contrast language-specific vs pooled results to identify language-general cues.

Experimental results

Research questions

  • RQ1Can ASD vs TD speech be discriminated within each language using prosodic features?
  • RQ2How well do models trained on one language generalize to other languages in pooled and LOCO setups?
  • RQ3Which acoustic-prosodic features serve as language-general markers of autism, and which are language-specific?

Key findings

  • Finnish within-language accuracy 0.84 and F1 0.88; Slovak accuracy 0.63 and F1 0.68; French accuracy 0.68 and F1 0.56.
  • Cross-language pooled model achieves accuracy 0.61 and F1 0.68 on average.
  • LOCO transfer shows Finnish F1 0.78 and Slovak F1 0.70 as relatively successful, but French F1 0.42 remains poor.
  • Across languages, F0 (pitch) distribution consistently helps distinguish ASD from TD; language-specific cues involve spectral tilt, global spectral shape, dynamics, formant structure, and intensity.
  • Pooled multilingual analysis reveals a shared, language-general set of cues, including pitch and spectral-shape features, though language-specific emphases remain evident.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.