Human children continue to surpass large language models (LLMs) like ChatGPT and Claude in acquiring language fluency, despite AI systems processing vastly more data. Four years after ChatGPT’s release, LLMs can converse naturally but require an enormous amount of text—up to 100,000 times more than a child hears by their first birthday—to reach comparable fluency, according to technologyreview.com.
Michael C. Frank, a cognitive scientist at Stanford University, explained that while LLMs have made impressive progress, they still need to consume the entire sum of human knowledge to replicate the milestone of language mastery that children achieve within about a year in everyday environments. This discrepancy is known as the data efficiency gap, highlighting how children learn language with remarkable efficiency compared to machines.
The data efficiency gap underscores a fundamental difference between human and machine learning. While LLMs rely on massive datasets and computational power, children acquire language through natural interaction and limited exposure. This gap presents a challenge for AI developers aiming to create models that learn more like humans, potentially reducing the enormous data requirements and energy consumption associated with current AI training methods.
The article from technologyreview.com emphasizes that despite advances in AI language models, human children remain the only entities capable of mastering language fluently with minimal data input, a milestone that has persisted for at least 100,000 years of human communication history.