Thoughts About Scaling Law

(twitter.com)

6 points | by tosh 10 hours ago ago

3 comments

  • kzrdude 8 hours ago ago
  • jarbus 4 hours ago ago

    Maybe just me, but I have a feeling that we are currently stuck at 100-200k effective context window due to a parameter limitation, specifically the hidden dimension of the model, which would be an incredibly expensive dimension to increase vs adding optimizations elsewhere. I don’t think we are done scaling parameters yet.

    • kzrdude an hour ago ago

      Why is it connected to number of parameters?