5 comments

  • sawfwair a day ago ago

    some real numbers on my m4 max -

    image gen (zimage-nano), 1024x1024: 58s

    image -> textured mesh (trellis.2): 2m 49s

    sfx generate (5s clip): 3.6s

    music generate (8s, ace-step): 15s

    speech synth: 13s | transcribed back: 2.2s

    video gen w/ audio(ltx unified-av) 4s,768x512: 2m 48s

    text chat (laguna xs2.1): 102 tok/s - https://mlx.fast leaderboard

  • ks2048 15 hours ago ago

    FYI, the title makes no sense in isolation.

    • sawfwair 15 hours ago ago

      totally fair! i spent a while trying to get clever and compress what i wanted to say and finally just hit submit but prob lost too much - local inference runtime, one cli that runs image/video/music/speech/3d/etc on your own machine without package hell. would def edit if I still could after your feedback, but window is closed. Thank you!

      • ks2048 14 hours ago ago

        At least add “models”. Doing “music” on your local machine can mean many things.

        • sawfwair 12 hours ago ago

          Yes, missed that completely... Great point!