The Next 3x in Inference Won't Come from Faster Kernels

(twitter.com)

2 points | by floathub 5 hours ago ago

1 comments

  • floathub 5 hours ago ago

    One unique idea for solving the chip shortage: pack more models into each existing GPU by having them share the resources. :-)