Commentary As A Substitute For Pretraining: Mixing Secondary Text Into A Small Language Model For A Philosophical Corpus

Mixing Secondary Text Into A Small Language Model For A Philosophical Corpus

In this whitepaper, we explore whether secondary text, such as commentary and explanatory material, can improve the performance of small language models trained on specialized domains. It examines different commentary-to-primary-text mixtures, evaluates their impact on model performance, and investigates whether the observed benefits carry over to pretrained models.

We train small language models on the collected works of Sri Aurobindo, a corpus of 6.1M tokens of English philosophical prose, and study whether adding secondary material recorded commentary on those works—improves the model on the primary text. Mixing commentary into the training data at a 20% sampling rate reduces held-out perplexity on primary text by 3.4% relative to training on primary text alone. A sweep over five mixture ratios shows an interior optimum at 20%; both lower and higher proportions are worse, and training on commentary alone gives perplexity 1425.6. The effect is consistent across three random seeds, with non-overlapping groups.

The improvement does not transfer to a pretrained model. Fine-tuning GPT-2-small on the same data with the same 20% mixture makes it 0.69% worse in bits per byte than fine-tuning on primary text alone, again with non-overlapping seed groups. Replacing the commentary with secondary text by a different author, in a different register, changes this penalty by less than a tenth of its size, so the penalty is associated with the training budget being spent on non-target text rather than with any property of the commentary. We conclude that the secondary text supplies volume rather than content: it helps when data quantity is the binding constraint and hurts when it is not. We also find that marking each training sequence with a provenance token indicating primary or secondary source has no measurable effect, and that a 466-item benchmark of ontological cloze items does not distinguish any of the models we train, though all score above chance.

Download PDF

    Chatbot Aria

    Hello, I am Aria!

    Would you like to know anything in particular? I am happy to assist you.