Show HN: Shoehorn – Quantize any model down to run on your machine

(notactuallytreyanastasio.github.io)

46 points | by rhgraysonii 2 days ago ago

9 comments

  • akshay_akula a day ago ago

    This is interesting. I wonder how it could work with something like https://github.com/JustVugg/colibri.

  • hmokiguess 2 days ago ago
    • rhgraysonii 2 days ago ago

      LLMFit tells you what can run on something. I built something quite similar to their search into Shoehorn now.

  • mbuchel-hn 2 days ago ago

    does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?

    • rhgraysonii 2 days ago ago

      Yes that is exactly what this does.

      • kennywinker 2 days ago ago

        Could you explain what happens when you try to shoehorn a 2.4T parameter model into a 24gb m4 mac?

        • akshay_akula a day ago ago

          Wondering the same thing but for 48gb M5 Max.

  • jaylane 2 days ago ago

    tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running

    • rhgraysonii 2 days ago ago

      If you could post an issue if you still have the error around that would be awesome.