ResearcharXivNEW

LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure

Li 2026-08-13
Fanfei LiJana ZellerManuel Prada-Corral

Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabular

Read on arXiv
Data aggregated and editorially reviewed by TrendMing.

Key Contributions

  • Modern language models are trained on heterogeneous web-scale text corpora.
  • Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize.
  • To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S.

Research Themes

AIResearch