500B Tokens Later: Letting AI Agents Decompile a First-Person Shooter

(momo5502.com)

117 points | by davikr 12 hours ago ago

83 comments

  • brandonpelfrey 12 hours ago ago

    Unless you explicitly need byte-matching decompilation, there are significantly faster ways to produce a decompilation/C which is functionally equivalent. I need to post about this. What's been working for me is that for every function, Agent A is tasked with writing some code which is semantically equivalent to the original assembly, but not necessarily exactly the same. Agent A also writes tests. Agent A submits the implementation of the function and tests to the harness for it to judge. The harness runs both the original function and the submitted function in a virtual machine/simulator/emulator (the tests define function inputs and starting state). The harness will only accept the implementation if 1) the read/write sequence to RAM is identical to the original function's, and 2) there must be complete line and branch coverage of the original function being decompiled.

    I've found this to be robust for decompiling games, while giving the agents enough freedom to write code that is readable and not waste a ton of time making sure e.g. instruction ordering, register assignments, etc. are all exactly the same. For me, having byte-matching decompilation is only one way to produce a decompilation I know is faithful to the original. This "high-level decompilation" process I just described is something agents can do much more quickly.

    • j2kun 12 hours ago ago

      Functional equivalence here, of course, depends on the completeness of the test suite, where byte-identical compiled artifacts does not.

      (For example, your approach would not necessarily catch all the same overflow behaviors; the OP expressly claimed that "replicating all bugs" was also important, and many bugs are caused by certain overflow behaviors)

      • nine_k 10 hours ago ago

        > catch all the same overflow behaviors

        So you're looking not just for functional but also dysfunctional equivalence %)

        • 8note 9 hours ago ago

          no, thats still functional here. the bugs have to be the same, such that speed runs could still run correctly

      • SubiculumCode 12 hours ago ago

        Byte exact seems only of interest to preserve known bugs etc for cheats/shortcuts/etc.

        • j2kun 12 hours ago ago

          That may be true, but I hate it when people repeat the false idea that functional equivalence requires only a test suite that has full branch/line coverage. Call me triggered :)

          That said, I would probably follow this same approach if I were to do this, but with extensive randomized testing as well.

          • hedgehog 11 hours ago ago

            You can do the process in stages. Do the first decompilation mechanically (no LLM), use a SMT solver to show it builds to an equivalent binary to the original, and then use LLM to clean up the code into something idiomatic with the benefit of a correct binary built with the new toolchain. This helps when you want to port across languages or toolchains, and helps protect against toolchain bugs.

        • kg 9 hours ago ago

          For anything with recorded replays or multiplayer you need to preserve known and unknown bugs for compatibility reasons, not just for cheating.

          • cowlevel 2 hours ago ago

            Or a speedrunning community

    • cowlevel 2 hours ago ago

      Only an LLM can give reasonably meaningful names to most functions and variables automatically, though.

    • cedws 4 hours ago ago

      Could you talk more about how you’ve used this technique? I want to start on a similar kind of AI-driven decomp project and I’m looking for any kind of edge I can get.

    • SubiculumCode 12 hours ago ago

      So you restricted it to implementing the same function (same inputs,outputs, dependencies as original?) and prevented the agents from making design decisions by keeping it's scope restricted?

      • brandonpelfrey 11 hours ago ago

        Yes. It can gain more context, but this has been enough. Note, there is also a notion of adversarial review layered on top in which it tries to poke holes in the test plan "you didn't handle this case of XYZ". It isn't actually perfect as a parallel thread said it may miss things like wrapping behaviors. In practice, it's very effective.

  • aetherspawn 11 hours ago ago

    The reason this cost so much is because the AI has the ridiculous goal of getting identical assembly output.

    The agents would have had to mess around with compiler versions, optimisation options, and the phase of the moon as well.

    If you just went for functional equivalence, it would probably cost 10x or 100x less tokens.

    Another false economy was using Sonnet instead of a more intelligent model like Sol 6.1 (1), which would have cost more per token, but is 100x or so better at reverse engineering and coding and therefore can chew through the source code much quicker and make fewer mistakes, meaning less work needing to be scrapped.

    In my testing doing a similar task, I ran multiple sonnet for weeks and burnt through ~$1000 in tokens to get 20% completion and output that was pretty bad. After switching to Sol 6.1, it finished the whole task in around 2 days, cost around $50, and it did it with zero supervision and a single /goal.

    (1): struggle to use Opus for reverse engineering, too many safeguards. OAI has virtually none, and uses way less tokens so is more economical.

    • Gigachad 11 hours ago ago

      Identical output isn’t ridiculous. It’s pretty much a requirement to ensure the game actually is the same in every way. These decomp projects are pitched as a high performance alternative to emulation.

      No one is going to use it if it’s a kind of close but not really reimplementation.

      Testing functional equivalence is also pretty much impossible. How would you for example test the new one has exactly the same bugs which haven’t been discovered yet. Or doesn’t introduce new ones? This stuff matters for speed runners.

      • aetherspawn 11 hours ago ago

        It’s ridiculous because something as simple as the compiler picking different registers is going to make zero functional difference but give a false negative on assembly compare.

        Yet the C code can’t pick what registers to use, so the poor agent is probably shuffling the code around randomly for hours or days until it matches.

        That’s probably why the agent dropped down into inline assembly in the first place (the author complained about this), because I bet it’s thinking trace was that this is futile.

        Compilers themselves are not even deterministic and running them multiple times creates different assembly.

        • hypercube33 an hour ago ago

          Isn't MW2 rooted in id tech engines? Last I checked there was raw assembly inside the code base. I would have started with quake 2/3 whatever source code and tried to match that up with Claude first - some of it has to have survived (though I doubt much)

        • Neywiny 2 hours ago ago

          I think I disagree when it comes to bug reproduction. Doesn't seem I'm alone in that.

        • dezgeg 2 hours ago ago

          How do you know the original code didn't use some undefined variable on stack, but happened to consistently work because something else spilled a register containing "suitable" value?

        • boricj 6 hours ago ago

          > Yet the C code can’t pick what registers to use, so the poor agent is probably shuffling the code around randomly for hours or days until it matches.

          Humans too do that for matching decompilation. Or at least I imagine so, given that I personally refuse to do that.

          People have different goals and will use different techniques to achieve them. The video game reverse-engineering/decompilation community isn't a hive-mind, I went ahead and created ghidra-delinker-extension because I had my own ideas on how to do that.

    • stavros 4 hours ago ago

      Identical assembly output is the only way to get functional equivalence, otherwise you're getting equivalence in some percentage of situations, but not 100%. Your bug is my feature etc.

  • nvme0n1p1 11 hours ago ago

    > The avid reader of my blog might have noticed that I had previously written two posts that have since been removed. Everyone else might now be wondering which game I am talking about. To both of you I can only say that corporate America was here to ruin our fun.

    Call of Duty: Modern Warfare 2 (2009)

    https://web.archive.org/web/20260925153118/https://momo5502....

    https://web.archive.org/web/20260925153131/https://momo5502....

    Come at me, corporate America.

    • heresie-dabord 3 hours ago ago

      Let us note without irony that Corporate America is indeed here to ruin your fun. It is their call of duty to wage this type of modern warfare.

    • TeMPOraL 3 hours ago ago

      A C&D-130 in the air, heading for you.

    • ares623 9 hours ago ago

      Friendly reminder that corporate America came at Aaron Swartz.

  • stevefan1999 an hour ago ago

    Reverse engineering dynamic dispatches like vtable and fat pointers/slices and traits would be hell. Different padding and compiler settings also contribute a lot.

    One of the few problem is that the decompiler is not always reliable. That forces you to go read the assembly, and it is not apparent to decipher the right kind of feng shui, not even from human before the LLM era. I used to play CTFs and my conclusion is exactly that.

    This is extremely apparent when there are self-modifying code (e.g. JIT) is involved. You need a stepping debugger to read the right control flow, because the code will diverge based on the instruction pointer and regions you jumped into. At this point static analysis like IDA and Ghidra stopped working.

  • edg5000 11 hours ago ago

    I sense the approach overcomplicates things. I wonder how long this would have taken in a single session. Maybe this is actually a textbook example of something where subagents make sense, but when I first started LLMs I was often overcomplicating the workflow with all kinds of orchestration. Now I just use one agent, it better allows controlling the output even if the agent works slightly longer. Most time is spent by me writing prompts and reviewing work anyway (for me at least).

    • hansvm 10 hours ago ago

      The best models have O(1k tokens per second). A billion tokens, sequentially, takes 1-2 weeks. 500B takes 500x that ... not fast. The problem might not have required 500B tokens, but if it needed anything within a couple orders of magnitude then something like the given approach (or anything else yielding equivalent results in exchange for parallelism) was mandatory.

      • speedstyle 10 hours ago ago

        Yeah, I don't think the problem needed billions of tokens

        • hansvm 9 hours ago ago

          That's plausibly true. I've definitely seen absurd token costs abused and wasted. I've also seen simple problems require absurd token counts regardless of prompt quality. Which factors made this problem require 1000x fewer tokens than they used?

          • WASDx 6 hours ago ago

            When you see these absurd numbers, they are re-counting cached input for every turn. So a simple tool call when your context size is at 500k counts as another 500k to the sum.

            Total Opus 5.5 token usage on OpenRouter last week is 5000B, presumably not counting cached input.

  • WheelsAtLarge 11 hours ago ago

    Interesting, if all software can be decompiled and copied what is the future of software. Will all software be SaaS? A time where the majority of PCs will be terminals? Game consoles are almost there. It's only a small jump for all software to go that way.

    Edit: Here's a possibility.

    The future of software is agent only software. We ask for a result, agent asks questions from us, agent uses the specialized software, user gets result. We subscribe to an AI assistant and specialized agents. We are almost there,at least the start. The future of PC's as we know them are numbered. OSs,CLI,compilers and whatever will melt into AI assistants. Say goodbye to writing software for people.

    • edg5000 10 hours ago ago

      I wonder about this too. Although, cracking was always possible. So what changed is that people can now modify, improve proprietary software with zero technical knowledge. I can now tell Astra or Fable (maybe not Fable because of safeguards?) to add a feature to Adobe Photoshop and it might actually succeed even when prompting purely on a functional level (e.g. the functionality I want out of the program). It would take days, but I think it would succeed. Call it PhotoBench?

      • patrickk 6 hours ago ago

        Eventually Fable-level models will be abliterated and hosted on Openrouter too, like so many others have been. Remember Chinese models are only a few months behind American ones at this point and are trained on their outputs, and can be abliterated after release.

        Someone has vibe coded the entire Adobe suite btw: https://getartcraft.com/apps

        https://abliteration.ai/

    • vagab0nd 10 hours ago ago

      Non-SaaS, close sourced software is already dead, no? I think pretty soon we'll all be using bespoke software written by the AIs. I already use throwaway software designed for each job because it's so cheap to do so.

    • unsnap_biceps 11 hours ago ago

      A future of all SaaS only works if you can't just tell a LLM what you want and get a custom implementation. I've always done a lot of personal projects, but my velocity has increased dramatically and I've replaced a number of projects that I used to pay for with custom ones that work well enough. I can't imagine that SaaS is going to be a viable model unless it involves a community that wants a unified experience, like a multiplayer game.

      • bitwize 11 hours ago ago

        Even then, interact with a SaaS enough and you may be able to "distill" a spec out of it that's sufficient enough for an LLM to replicate it closely enough. If you feed the LLM screenshots or video of people using it, it will be able to recreate it even more accurately.

        • georgemcbay 9 hours ago ago

          > If you feed the LLM screenshots or video of people using it, it will be able to recreate it even more accurately.

          You could just have the LLM agent use the software itself and then duplicate the functionality it finds, don't even need the screenshots or video.

          Unless the software is doing something extremely novel, the LLM can just do this all itself after you point it toward a url for the software with instructions to clone what it finds there.

          It'll be interesting to see if SaaS companies and/or providers of internet services like Cloudflare try to stop this sort of agentic cloning of SaaS products.

          Like, obviously they would kinda want to if/when this sort of thing becomes commonplace, but how do you do it while allowing your software to be agent-friendly for non-cloning uses, and who is going to want to tell their customers they can't use agents with their SaaS for fear of their software being cloned? Kind of a bad situation with either option.

    • Daishiman 11 hours ago ago

      Most software organizations pay for is effectively a SaaS or something where software is a minor part of the artifact and support is where the real money is.

  • petetnt 6 hours ago ago

    Great article that really demonstrates why the copious amounts of AI generated PC ports do not fit under ”preservation” guise, which to me has been the point of doing these ports in the first place and what many of the enthusiasts before the LLM wave were aiming for.

    > There are no noticeable bugs and all features of the original game are present.

    >

    > The remaining functions have been reworked repeatedly. While they still don’t match byte for byte, we believe their semantics are correct.

    As there are no way to confirm this outside of ”believinh”, what you are left is with a end product that might contain thousands of micro changes that essentially make the game something else than it was intended to be.

    • fxtentacle 5 hours ago ago

      Fully agree.

      Since even Fable 5.1 cannot reliably remember what it wrote into "memory" files before the last compaction, why would we trust it to not forget about some side-effect when handling assembly that is so verbose it'll surely exceed the context and, hence, necessitate compaction.

  • Neywiny 2 hours ago ago

    I do wonder if LTO would obfuscate this a bit more. Presumably if you can't chunk it into small compilation units, you can't find a way to make C from the machine code.

  • bob1029 6 hours ago ago

    Reverse engineering existing game binaries seems kind of silly when a lot of the most important AAA knowledge is already available to the general public.

    https://github.com/ValveSoftware/halflife/blob/master/pm_sha...

    The LLMs have presumably already consumed information like this as part of their training sets. Converting between Hammer and Unity scale is a fairly trivial linear operation. You can dump the BSPs to recover geometry and rapidly accelerate map development using timings that are already known to work.

    Attacking the raw binary directly is certainly impressive, but it's totally unnecessary.

  • thway15269037 11 hours ago ago

    I struggle to understand what legal leverage they used to threaten him to remove every detail about the game. Can someone post the game name and company name?

    So, if you reverse-engineer game X and post reverse-engineered code, what exactly do you infringe, how and in which jurisdiction? What changes if it is done via LLM?

    (I understand that LLM decompilation is absolutely out of hand right now and something surely will come to trample the fun. But what and when? I suppose american LLMs will have their system prompt updated to forbid any reversing help and report suspicious activity straight to legal hotline)

    • xnx 11 hours ago ago

      Call of Duty: Modern Warfare 2 (2009) (from a previous version of the page)

    • rasz 10 hours ago ago

      >what legal leverage

      PIF owns EA, when they invite you to a meeting you might start to worry about foil lined room and carpentry tools lying around

    • bitwize 11 hours ago ago

      > So, if you reverse-engineer game X and post reverse-engineered code, what exactly do you infringe, how and in which jurisdiction? What changes if it is done via LLM?

      In 1986 there was a federal court case, Whelan Associates Inc. v. Jaslow Dental Laboratory, in which it was ruled that the "structure, sequence, and organization" of a computer program was protected by copyright, and thus independently produced software could be found infringing if it copied these elements, even if the code were not copied (or mechanically translated) verbatim. This led to a six-year period in which computer software enjoyed generous copyright protection, such that "clones" of copyrighted software were effectively infringing. It wouldn't be until the early nineties that other court rulings would tighten the rules again, notably Computer Associates International, Inc. v. Altai Inc.. The 3-step "abstraction-filtration-comparison" test has been used by most courts since 1992 to determine whether nonliteral parts of program code are eligible for copyright protection, and whether another, independently written program is infringing.

      HOWEVER, the Whelan standard was never actually overturned or stricken from U.S. law due to legislation or litigation! And companies have sued and won under the Whelan standard! Most notably, Oracle in their copyright and patent case against Google regarding Java APIs in Android. Google ultimately prevailed, but only because the Supreme Court ruled that Google's use of the APIs was fair use; they did not overturn the Federal Circuit's finding that the "structure, sequence, and organization" of the Java declaration code was ineligible for copyright! And I doubt that the Supreme Court would similarly smile on a reverse-engineered complete video game!

      I believe that these reverse-engineered projects infringe copyright, under the Whelan standard and perhaps under the stricter Altai standard as well. You do not get a free pass because the original code was in C and yours is in Rust; or because you created a slightly different version of each function in the original code.

      • nullpoint420 10 hours ago ago

        Yet AI companies distilling copyrighted books and repositories is somehow okay?

        And let's say there's a future where an AI model could zero-shot the game itself. What then?

      • thway15269037 10 hours ago ago

        That implies US jurisdiction plus that was made at the time where "hey, let's actually reverse-engineer and copy our competitor" was very expensive and time-consuming, so slapping one company would discourage everyone else to sink money in it. Currently it seems the latter is becoming either automatic (still somewhat expensive) or even free.

        • bitwize 10 hours ago ago

          Committing copyright infringement against software vendors was automatic and free before, when it involved merely copying files, maybe after cracking the copy protect. It was still illegal. What makes you think this time would be any different? The fact that an LLM did it? Whatever an LLM produces humans assume responsibility for.

          • thway15269037 9 hours ago ago

            When you copy a file, it's trivially simple to prove it's a copy and not an original work.

            How can one prove that X is a copy when none of the source code match original? Clean room re-implementations are not illegal after all, if we go that way (yeah, yeah, I know about clean room argument).

            While I personally do not think it would be different, but a lot of LLM folks are, and even some companies (rushing to blatantly clone and de-compile stuff). Kinda strange feeling sitting and looking around in the midst of it, y'know.

            • imtringued 7 hours ago ago

              Doing a clean room implementation would require you to build a test suite using an AI against the binary and then you build a test server against which other people, not you, can run their own implementations and you just report how many tests pass.

  • mawadev 8 hours ago ago

    It is wild how AI ends up in the trenches of plagiarism, going from art to text, from software to games...

  • Fizz43 4 hours ago ago

    the 80% is a meaningless AI generated % by the way.

  • xnx 11 hours ago ago

    A previous version of the page said the game was Call of Duty: Modern Warfare 2 (2009).

  • vivzkestrel 11 hours ago ago

    - i want to very very badly see a post of this using LM studio and one of the open source models

    - please someone do it

  • georgemcbay 11 hours ago ago

    In case anyone is curious about the obvious question of which game they are talking about... based on months-old reddit posts (which seem like links to prior progress reports of the same project) the game in question appears to be Call of Duty: Modern Warfare 2 (the original 2009 version).

    https://www.reddit.com/r/ReverseEngineering/comments/1vxig19...

  • Uptrenda 5 hours ago ago

    I'm more interested in how the OP managed to convince claude to do what was probably against Activison's TOS. I am assuming they didn't hand specifics of the game but just a blank binary. Because if you had of given specifics claude would have checked the TOS of the IP holder, found that the company doesn't want people to decompile their software, and refused to work on it. Or am I missing something here?

  • tasubotadas 6 hours ago ago

    Interesting read and a thread here. I am working on decompiling two obscure abandoware games from my childhood and it's been a difficult and quirky process...

    Even after building whole infra for that and not pursuing identical byte output it's been pretty token hungry process

  • esafak 11 hours ago ago

    Can anyone think of any lessons to draw from this for normal development, where we don't have oracles to serve as guardrails? I write specs but the agents still find ways to insert bugs between the lines. Oh well, job security.

    • ashdnazg 8 hours ago ago

      It's not really comparable. When decompiling you only care about a one time result. Code quality doesn't matter much. Stability doesn't matter much. Only whether the result is accurate. Once you have the matching decompiled code, all tools used until that point will be usually thrown away.

      Normal development isn't so lucky, you do care about code quality. So you have to review the code the agents write and handle the testing yourself. The agents will not do what you want, they'll do what they think you want, and you're the only person who knows what you think.

  • WillAdams 12 hours ago ago

    To save folks looking this up:

    >At current 2026 API rates, 500 billion AI tokens would cost roughly $100,000–$750,000 depending on the model, with most flagship models in the $150–$400 per million input tokens range

    • Gigachad 12 hours ago ago

      This stuff is all done on highly subsidized subscription plans. I suspect Antropic is quite happy to sell these people $100k of compute for $4k because it boosts their growth numbers and they can tell investors once they stop subsidizing, this will grow to $100k. Despite the fact that most of this stuff simply wouldn't be done without the subsidization.

      • chii 11 hours ago ago

        the exact same arguments were said about uber's business at the start.

        Yet, it is now profitable.

        The bet is that people realize how valuable these services are, and despite complaining, they still would pay the higher price. This realization would not happen without this initial subsidy from investors.

        It isn't too different from drug dealer's first sample free...

        • thway15269037 11 hours ago ago

          Uber business model wasn't a subscription "for 10 bucks you can travel 900 lightyears a week"

        • Gigachad 11 hours ago ago

          There’s also examples where this didn’t work. Moviepass for example.

          Taxis were an established profitable business model and the uber subsidisation wasn’t anywhere near as much as AI subsidies.

          • 0x457 9 hours ago ago

            Moviepass failed because they thought it's going to be like Gym membership - people buy and don't go, but guess what? People love going to movie theaters. On its own its not bad, see AMC Stubs, but AMC owns the theater, they sell you popcorn and soda. OpenAI an Anthropic is closer to AMC than Moviepass.

            • xienze 5 hours ago ago

              > Moviepass failed because they thought it's going to be like Gym membership - people buy and don't go, but guess what? People love going to movie theaters.

              You think developers won't/don't stretch subscriptions to the absolute limit? The AI subscription model is like the gym model except a large percentage of the customers work out 24/7/365.

          • isubkhankulov 11 hours ago ago

            Disagree with your second paragraph. Uber/lyft subsidized into deep negative margin territory around ~2015 or so. Anthropic (and OpenAI) are subsidizing but not losing money on these consumer plans.

            • edg5000 10 hours ago ago

              > subsidizing but not losing money

              ????

              • HWR_14 7 hours ago ago

                The claim is that uber and lyft lost money on each ride but that openai and anthropic make money on inference. Just not enough to pay the cost of developing the models. The difference is that uber and lyft had to change their pricing (or payment) models to make money, where anthropic or openai could just sell enough inference (at some level of sales).

              • isubkhankulov 6 hours ago ago

                to clarify, I mean that the big model companies are charging consumers much less than equivalent API pricing but they're not actually losing money so its more of a steep at-cost discount for inference. It likely does not fully cover amortized R&D just like the other reply stated but it does still cover marginal inference cost (GPU/power/etc)

                Uber and Lyft were paying drivers $X but charging users way less than $X so they were literally burning investor money to get market share.

                • xienze 5 hours ago ago

                  > It likely does not fully cover amortized R&D just like the other reply stated but it does still cover marginal inference cost (GPU/power/etc)

                  Why does this point come up over and over again, pretending that you can truly separate training and inference costs. Yes, they are separate things but the value OpenAI and Anthropic are presenting to the world is "we're the absolute best, no one else comes close." Well, to keep that up you can't just not train for extended periods of time. You have to keep that engine going non-stop when there's free Chinese models nipping at your heels. You can be profitable on "just inference" all you want but if training expenses dwarf that, you're not going to be profitable overall, and that's the bottom line.

                  • edg5000 5 hours ago ago

                    > does still cover marginal inference cost Simply comparing to the larger models on OpenRouter implies that the pure hosting costs (equipment + power + minimal overhead) still exceed plan pricing if we assume all users always use their full weekly allowance.

                    So my conclusion is that it only works because the majority of users doen't fully utilize their allowance. Last month I used almost nothing of my Claude 20x plan (did use Codex though).

              • Gigachad 9 hours ago ago

                It's profitable if you don't count the expenses.

                • oblio 8 hours ago ago

                  I think you're joking but that's Amodei & co are claiming with a straight face.

        • nmfisher 9 hours ago ago

          Last I checked, Uber underperformed the S&P index since its IPO. Even Softbank didn't make a great return on its investment. The growth story was definitely oversold.

          Also, Uber was peak ZIRP + COVID, money was cheap and growth was easy. I think the mountain to climb is a lot steeper now.

    • TomatoCo 12 hours ago ago

      Where are you getting 150-400 per million input?

      https://developers.openai.com/api/docs/pricing https://platform.claude.com/docs/en/about-claude/pricing

      OpenAI and Anthropic are both 10/mil in.

      https://openrouter.ai/z-ai/glm-5.3#providers https://openrouter.ai/moonshotai/kimi-k3#providers Other frontier models are like 1-3/mil in.

      Also, 500 billion is 500,000 millions. At the lower end of your 140/mil estimate that's 70 million dollars. Even at my 2/mil lookup for Chinese frontier models that's one million. Show your math for 100k-750k, please.

      • WillAdams 11 hours ago ago

        It was a search for "average cost of 500 billion ai tokens" which may or may not have been the correct phrasing, but seemed straight-forward enough that at first blush, accepting the AI-generated answer seemed reasonable.

    • j2kun 12 hours ago ago

      I struggle to believe that someone would find it worth that much money to have a decompiled version of a game they also commit to keeping private.

      • girvo 12 hours ago ago

        > they also commit to keeping private

        Reading between the lines:

        "The avid reader of my blog might have noticed that I had previously written two posts that have since been removed. Everyone else might now be wondering which game I am talking about. To both of you I can only say that corporate America was here to ruin our fun."

        Keeping it private likely wasn't the original plan...

    • GaggiX 11 hours ago ago

      The vast majority of these tokens are cache input tokens so even at API price the cost would be much lower.

      • thway15269037 11 hours ago ago

        Even with 95% cache hit, and on cheapest chinese model, it would still be in tens of thousands of dollars (I assume that of 500B tokens at least 10B would be in "output"). If we switch gears to Kimi K* then if would very quickly escalate in hundred thousands USD.