Skip to main content
QUICK REVIEW

[Paper Review] Learning from Child-Directed Speech in Two-Language Scenarios: A French-English Case Study

Liel Binyamin, Elior Sulem|arXiv (Cornell University)|Mar 13, 2026
Language Development and Disorders0 citations
TL;DR

The paper systematically studies compact multilingual models (English-French) trained on child-directed speech and multi-domain data, comparing monolingual, bilingual, and cross-lingual pretraining and evaluating on semantic and grammatical tasks across languages.

ABSTRACT

Research on developmentally plausible language models has largely focused on English, leaving open questions about multilingual settings. We present a systematic study of compact language models by extending BabyBERTa to English-French scenarios under strictly size-matched data conditions, covering monolingual, bilingual, and cross-lingual settings. Our design contrasts two types of training corpora: (i) child-directed speech (about 2.5M tokens), following BabyBERTa and related work, and (ii) multi-domain corpora (about 10M tokens), extending the BabyLM framework to French. To enable fair evaluation, we also introduce new resources, including French versions of QAMR and QASRL, as well as English and French multi-domain corpora. We evaluate the models on both syntactic and semantic tasks and compare them with models trained on Wikipedia-only data. The results reveal context-dependent effects: training on Wikipedia consistently benefits semantic tasks, whereas child-directed speech improves grammatical judgments in monolingual settings. Bilingual pretraining yields notable gains for textual entailment, with particularly strong improvements for French. Importantly, similar patterns emerge across BabyBERTa, RoBERTa, and LTG-BERT, suggesting consistent trends across architectures.

Motivation & Objective

  • Motivate developmentally plausible multilingual language modeling beyond English in constrained-resource settings.
  • Systematically compare monolingual, bilingual, and cross-lingual pretraining with carefully size-matched corpora (CDS and multi-domain).
  • Introduce French variants of evaluation datasets (QAMR, QASRL) and bilingual resources to enable fair cross-linguistic testing.
  • Assess both grammatical (syntactic) and semantic understanding (QA, entailment) in English and French.
  • Provide architectural generalization checks across multiple small models to ensure robustness of observed patterns.

Proposed method

  • Use BabyBERTa as the core compact model and retrain on two data scales: ~2.5M tokens (CDS) and ~10M tokens (multi-domain).
  • Construct parallel monolingual, bilingual, and cross-lingual pretraining settings with strictly size-matched English and French corpora.
  • Evaluate on syntactic (CLAMS) and semantic tasks (SQuAD/FQuAD, QAMR, QASRL, XNLI) with language-specific fine-tuning; baseline comparisons include RoBERTa-base and CamemBERT-base.
  • Create French versions of QAMR and QASRL and develop English/French multi-domain corpora for balanced cross-linguistic testing.
  • Examine cross-architecture generalization by replicating analyses with RoBERTa, T5-tiny, and LTG-BERT.

Experimental results

Research questions

  • RQ1Do competencies transfer across languages under monolingual, bilingual, and cross-lingual pretraining?
  • RQ2How do CDS and multi-domain corpora influence grammatical vs. semantic performance in English and French?
  • RQ3Does bilingual pretraining yield consistent gains across tasks, particularly for the weaker language (French)?
  • RQ4Are the observed effects robust across different model architectures (compact vs. larger baselines)?
  • RQ5What is the effect of combining CDS with Wikipedia data on semantic and transfer-sensitive tasks?

Key findings

  • Bilingual pretraining yields notable gains for textual entailment (XNLI), especially benefiting French.
  • Wikipedia training favors semantic tasks (QA, entailment), while child-directed speech training favors grammatical competence in monolingual settings.
  • Exposure to CDS interacts positively with Wikipedia training, improving semantic and transfer-sensitive tasks, particularly for French.
  • At smaller data scales (≈2.5M CDS), bilingual exposure provides semantic benefits; at larger scales (≈10M multi-domain), monolingual dominance grows but bilingual gains persist for some tasks like XNLI.
  • Patterns are consistent across multiple architectures (BabyBERTa, RoBERTa, LTG-BERT, T5-tiny), indicating robustness of the observed effects.
  • Small models can reach meaningful semantic capabilities with developmentally plausible data, approaching performance trends of larger models under constrained resources.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.