ResearcharXivNEW
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
Li 2026-08-13
Fanfei LiJana ZellerManuel Prada-Corral
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabular
Read on arXivData aggregated and editorially reviewed by TrendMing.
Key Contributions
- Modern language models are trained on heterogeneous web-scale text corpora.
- Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize.
- To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S.
Research Themes
AIResearch