Less Data, Smarter Models
Apple Research, in collaboration with the National University of Singapore, just published work that sounds like it shouldn't work: throwing away training data makes language models memorize more facts, not fewer. The paper, "Cram Less to Fit More," accepted at ICLR 2026, formalizes fact memorization from an information-theoretic perspective and demonstrates that the standard approach to LLM training — feed the model everything you have — is provably suboptimal when it comes to factual knowledge retention. The…