https://www.lesswrong.com/posts/PagGF8roBJmjLunsX/competitiv...
I think the idea is interesting though, although I wonder if training time for LoRA is such a bottleneck to deserve its own, extremely narrowly scoped, leaderboard. Maybe if it was more tasks or more models we could hope that it transfers? With a single task, and a single model, I’d be afraid of this overfitting pretty heavily.
For NanoGPT, I think the idea always was that the ideas can be transferred to much larger models, or serve as stepping stones for investigations on larger models.
What is LoRA in this context? the communication protocol? Or another term appropriated by LLMs? Why the speedrun?