Source: https://mayphus.org/gpt124m-restartable-training/ Title: GPT-124M from scratch: benchmarking and restart recovery Metadata: {"category":"artifact","date":"2026-10-05","id":"gpt124m-restartable-training","kind":"article","language":"en","locale":"en","route":"/gpt124m-restartable-training/","slug":"gpt124m-restartable-training","tags":["language-models","training","checkpoints","kaggle"],"type":"article"} # GPT-124M from scratch: benchmarking and restart recovery I am training a small GPT model on free Kaggle GPUs. The useful progress so far is in making the work measurable and restartable: a timed single-GPU pilot, a checkpoint that preserves the training state, and a recovery test in a completely fresh saved session. This is a training-engineering record as of October 5, 2026. The longer run and final evaluation are still unfinished. ## What starts from scratch The model has 124,439,808 randomly initialized parameters. Its GPT-2 architecture uses 12 layers, 12 attention heads, a width of 768, and a 1,024-token context. No pretrained model weights are loaded. I use the existing GPT-2 tokenizer and public FineWeb-Edu data. Starting the model weights from scratch does not mean inventing a new tokenizer or collecting a new corpus. ## A pilot with an explicit timing boundary The single-T4 pilot processed 6,881,280 tokens in total. Its timed portion covered 6,750,208 tokens in 904.14 seconds, including checkpoint work. That gives approximately 7,465.9 tokens per second. Reported peak GPU memory was 4.46 GiB. | Single-T4 measurement | Result | | --- | ---: | | Total pilot tokens | 6,881,280 | | Tokens inside the timing window | 6,750,208 | | Timed duration, including checkpoint work | 904.14 seconds | | Throughput inside that window | 7,465.9 tokens/second | | Reported peak GPU memory | 4.46 GiB | The timing window matters. Dividing the total pilot token count by a duration that measured only part of the run would produce a different and misleading rate. On a held-out set of 65,536 tokens, loss fell from approximately 10.99 to 8.41. Generated samples were still repetitive. These are early diagnostics, not a capability evaluation or evidence that the model is ready for practical use. ## A checkpoint must preserve the next step Saving model weights alone is not enough to resume the same training trajectory. The optimizer, learning-rate schedule, mixed-precision scaler, random-number generators, and data position all influence what happens next. The full-state replay test restored the model, optimizer, scheduler, scaler, RNG state, and data offset. The tested continuation reproduced the expected result exactly. I then tested a separate failure boundary: a completely fresh Kaggle session loaded the saved checkpoint and resumed successfully. That matters because a checkpoint that works only while the original process is alive has not demonstrated useful recovery. These results establish the tested replay and saved-session recovery cases. They do not establish identical behavior across every hardware configuration or every possible interruption. ## Preparing the longer run A separate training token store contains 1.1 billion token positions. A document-hash holdout separates the selected evaluation documents from training. No additional near-duplicate decontamination has been established, so this split should not be described as a fully decontaminated benchmark. An initial two-T4 distributed-data-parallel test reached approximately 14,787 tokens per second across the two GPUs. That is an initial measurement, not an established sustained rate for the longer run. The planned one-billion-token training run and final evaluation remain unfinished. I am leaving out moving training counters here: an old progress counter is easy to mistake for current state, and it does not establish completion. For now, the concrete result is a measured pilot and a training state that survived both exact replay and a fresh-session restart. The next useful evidence is completion of the longer run, evaluation with clear data boundaries, and samples assessed beyond a small loss diagnostic.