What if AI worked at 1.000.000 tokens per seconds?

(echohive.ai)

7 points | by echohive42 12 hours ago ago

10 comments

  • gchamonlive 11 hours ago ago

    I don't want to bash on the article, but a much more interesting take with llm that worked at 1kk TPS would the the countless amount of real time applications you could build with it.

    I like electronic music, I dance to it a lot. What if I had a LLM that had enough throughput and low latency to do the reverse, take my dancing and generate music.

    We are still stuck stargazing raw artificial cognitive intelligence, but real progress will be when we stop perceiving these as external intelligence and just an extension of ourselves.

    • echohive42 10 hours ago ago

      I think you are right and I think we are close to that speed, with codex ultrafast we are near 500 TPS but only with Astra and that is not financially viable. I would guess we will have wide spread near 1000 TPS with 6.1 sol level model with enough usage to, without worry, do the kind of things you describe in about 6 months max

  • sixtyj 10 hours ago ago

    “If” is temporary. More TPS is the question of time. In such usecase human in the loop will be eliminated as nobody, even anybody with fastest recognition, will be the slowest part.

    How many industries are? How many use cases are possible?

    If you solve everything, what will people do? Sitting on the beach, sipping drinks?

    Massive degradation of human brains and rise of neurodegenerative illnesses as unused device starts to be erroneous?

    1,000,000 TPS is good for some use cases but not publicly available…

    • danmaz74 10 hours ago ago

      Higher TPS by itself wouldn't remove the humans in the loop. Humans in the loop aren't there to make things faster, but to make them right - current LLMs can't do that on their own in most real world projects.

      On the contrary, I would say that higher TPS would make humans in the loop much more productive, because the current long wait times mean that you lose focus and flow.

    • echohive42 10 hours ago ago

      maybe we will all hang once again at the clubhouse like app, not sure if you experienced it in its heyday. it was amazing! :)

  • ed 10 hours ago ago

    At 1m tps, LLM’s could generate UI’s in realtime (a 5k webpage would take 5ms). Applications will have every little in common with today’s stacks when that happens. (For starters, rendering, event handling and data storage could move to the context.)

    • echohive42 10 hours ago ago

      yeah very true. I asked Opus 5.5 max if even such a thing is possible and it thought so, albeit after a lot of hurdles, especially memory related ones being overcome.

      it also thought it could happen by 2033

      • hsellentin 10 hours ago ago

        well 60fps is 16ms frame time so maybe we could get that by 2029 :)

  • paul_irolla 28 minutes ago ago

    [flagged]