For the first decade of deep learning, the pursuit of Artificial General Intelligence (AGI) operated under an empirical dogma: the pre-training scaling laws. Frontier research laboratories operated under the assumption that maximizing compute, dataset token volume, and neural network parameter counts would yield broad human-level cognition. If a transformer-based model swallowed enough petabytes of human […]