Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

About

Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome during training. To address this, we present DatedGPT, a family of twelve 1.3B-parameter language models trained from scratch on approximately 100 billion tokens each with strict annual data cutoffs spanning 2013 to 2024, together with DatedInstruct, an instruction dataset grounded in each year's documents to prevent leakage during post-training. The models are competitive with open models of similar scale, and perplexity-based probing confirms that each model's knowledge is bounded by its cutoff year. On stock return prediction over 61,000 firm-day news headlines, DatedGPT-instruct achieves an annualised Sharpe ratio of $3.20$ under the lookahead-bias-free setup. Lookahead-biased models, whose training data covers the outcome period, add a lookahead premium of $26.4$ b.p. per standard deviation, significant at the 1% level. The series thus enables direct analysis of lookahead bias in financial forecasting. We provide an interactive web demo that allows users to query and compare responses from models across different cutoff years, available at www.datedgpt.com.

Yutong Yan, Raphael Tang, Zhenyu Gao, Wenxi Jiang, Yao Lu• 2026

Related benchmarks

TaskDatasetResultRank
Commonsense ReasoningHellaSwag
HellaSwag Accuracy54.6
897
Instruction FollowingIFEval--
854
Physical Commonsense ReasoningPIQA
Accuracy71.8
724
Question AnsweringARC Easy
Accuracy71.6
597
Multitask Language UnderstandingMMLU
Accuracy26.3
568
Question AnsweringARC-E
Accuracy52
544
Multi-task Language UnderstandingMMLU
Accuracy25.3
353
Question AnsweringARC-C
Accuracy35.2
283
Science Question AnsweringARC-C
Accuracy34.8
268
Science Question AnsweringARC-E
Accuracy52
240
Showing 10 of 15 rows

Other info

Follow for update