HP ZGX Fury Is Now Orderable: GB300 Superchip, 748GB Unified Memory

(storagereview.com)

39 points | by rbanffy 15 hours ago ago

57 comments

  • fc417fc802 15 hours ago ago

    > a Kensington slot

    I think you might need a bit more than that at this price point ...

    • alexfoo 14 hours ago ago

      At the expected price point it would be nice if it came with a free pied-à-terre in Kensington (https://en.wikipedia.org/wiki/Kensington). Just a small studio apartment, nothing special.

      • noir_lord 14 hours ago ago

        You'd need to spent ~£500K+ to get a studio in Kensington these days that isn't a shoebox.

        One listing is 410K for 280 sq/ft coming out at £1464 per square foot, almost exactly 10x the price per sq/ft we paid for our house a few years ago.

        So $100K would be £74K which would get you ~50 sq/ft.

        • yread 13 hours ago ago

          A square (or cubic?) foot would be enough to house it

    • Retr0id 14 hours ago ago

      A kensington lock in this context is like a sign that says "do not move", which is still kinda useful

      • fortran77 13 hours ago ago

        I have a $50,000 Tektronix 'scope on my bench that's Secured with a kensington lock. I am in a small office in a building I own and I rent out some extra space to another 3 person engineering company. One weekend they tried to remove my 'scope from the bench because they had something "important" to work on. They quickly discovered the lock. (The whole thing was caught on camera.)

        They weren't trying to steal it, they just wanted to bring it over to their side. But the lock did its job! This is the perfect situation for a Kensington lock.

        I yelled at them the next day telling them if they need a Series 5 'scope they could go rent or buy one.

  • pritambarhate 14 hours ago ago
  • dsrtslnd23 15 hours ago ago

    252GB of HBM3e at 100k vs. a multi A6000 96GB setup. GB300 seem expensive in comparison at ~$100k. Am I missing something?

    • _diyar 15 hours ago ago

      Without looking it up and doing the math, I bet GB300 has higher Tflops and memory bandwidth, especially when used with e.g. NVFP4.

      • hgoel 14 hours ago ago

        Plus lower peak power draw

        • rbanffy 14 hours ago ago

          How long until the higher power draw nullifies the higher acquisition price of the GB300 machine?

          • cyanydeez 14 hours ago ago

            Are you in a world where energy prices are goimg down?

            • rbanffy 14 hours ago ago

              Quite possibly, depending on grid-scale renewable deployment. I also am about to install a set of solar panels, so a base level of power will cost me only the depreciation of the PV hardware.

              We could condition datacenter installs to providing power to the grid - you want to build a datacenter, you also need to build a wind or solar farm that can fully power its peak load.

              • cyanydeez 10 hours ago ago

                America has decided the answer is no.

            • nutjob2 14 hours ago ago

              Yes, which planet are you on? Australia has recently lowered power prices due to renewables, probably other places will too for similar reasons as the rollout continues.

              • cyanydeez 10 hours ago ago

                Country: America, where we pay people not to build windmills.

    • cmrdporcupine 14 hours ago ago

      There's a lot more going on here because the host machine has a boatload (496GB) of expensive LPDDR5x (which is stupidly expensive) that can also be used as unified (but slower) memory to the GPU, 72 ARM64 cores, stupid fast NVlink/QSFP networking, the PSU to support all that etc. etc.

      Basically it's the same as a tray in a GB300 NVL72, but in workstation/desktop form. Niche would be AI researchers.

      They will... not sell a lot of these. But what a beast.

      I am too lazy to price out 496GB of DDR5 but, um, mostly because it is terrifying to see what today's prices look like.

      You can't build such a machine out yourself but if you could I suspect the price point would come out about the same.

      • rbanffy 14 hours ago ago

        > You can't build such a machine out yourself but if you could I suspect the price point would come out about the same.

        That's kind of the nature of capitalism and price elasticity - the manufacture price only limited the minimum sale price, and sale price usually reflects how much is the market willing to pay for the good.

        Where you might save a lot of money is on building something that's very targeted to your needs that matches them better than a GB300 workstation, for a lower price.

        • cmrdporcupine 14 hours ago ago

          But the point of such a machine generally is to have the equivalent of a GB300 NVL72 tray on your desk. It's so that you can do the work that belongs in a production DC eventually. Or at least that's now NVIDIA would like you to use it.

          You could build out a complicated multi GPU setup of your own, but the work you do on inference tuning for your kernels etc would not necessarily translate to the real world.

          But yes, if you just want to run GLM 5.3 on your own machine, that's a whole other story.

    • segmondy 11 hours ago ago

      have you priced ddr5 memory?

    • LargoLasskhyfv 14 hours ago ago

      HP’s spec sheet fills in details the platform announcements skipped. The CPU memory is four 128GB SOCAMM modules delivering 396GB/s, and the Grace CPU is soldered to the host processor module rather than socketed. The two pools add up to the 748GB coherent space that lets the GPU address CPU memory directly, which is what makes trillion-parameter inference and fine-tuning of models in the 100 billion parameter class possible on a single box. HP’s footnote on those model sizes is that the harness quantizes at FP4.

  • theplumber 14 hours ago ago

    Someone should edit the title to make it clear it only has 252GB of “AI” memory

    • edg5000 13 hours ago ago

      Agree, the CPU memory is about 400GB/s vs. Apple M5 Ultra at 1.2TB/s. But the M5 only has that, not the 7TB/s portion. So this box could* still be faster when using the full 748GB if both RAM types can be used without the slower part bogging everything down. *That comes down to implementation. Maybe somebody can chime in on this.

  • Retr0id 15 hours ago ago

    How many arms and legs does one of these cost?

    • shmoil 15 hours ago ago

      3.14 kidneys.

  • dyzone 15 hours ago ago

    But can it run crysis?

  • karmakaze 10 hours ago ago

    This is a bundle where you end up paying for parts that aren't as useful:

        - 252 GB HBM3e VRAM
        - 496 GB LPDDR5X RAM
    
    I would rather have a system using 2x Instinct MI350P GPUs (288GB total) for much less.
  • einpoklum 15 hours ago ago

    Can I get 32 GB of RAM for a sane price instead?

    • m4rtink 13 hours ago ago

      Yeah, all these overpriced toys will be obsolete soon anyway, but regular memory for PC for actually doing some productive stuff is still overpriced. Pure madness.

  • bentt 13 hours ago ago

    It's actually kind of nostalgic to see off the shelf computers have such a sky high price again. Back in the Silicon Graphics days the prices were aspirational! Give me something to lust over again... makes me feel young.

    • znpy 12 hours ago ago

      aye! 'twas so jolly in ye ol' days when only the master could afford a large computer! /s

      • bentt 5 hours ago ago

        /u/znpy gets it! :)

  • self_awareness 15 hours ago ago

    I guess I'm too poor to even know the price

    • wolttam 15 hours ago ago

      MSI’s DGX Station (what this is) was listed for $99K

    • chedabob 15 hours ago ago

      Dell's equivalent starts at £180k.

    • NKosmatos 13 hours ago ago

      These are for billionaires/multi millionaires/cryptobros/stupidly rich people who can afford such “toys”. I’d really love to see a geek/tech freak/gadget lover/IT/homelab billionaire buying large number of things we mere mortals can’t afford, but only dream of, and then he blogs about them ;-)

      • Gigachad 13 hours ago ago

        I’m gonna bet in 5 years the price will be in range for the average tech bro to get their hands on one.

        It just doesn’t make sense to buy this stuff at peak prices.

      • fortran77 13 hours ago ago

        Since I'm a "retro computing" guy, I look forward to buying one of these on eBay or at a Vintage Computing Fair for a few hundred bucks 5 years from now.

  • theplumber 15 hours ago ago

    >> 252GB of HBM3e at 7.1TB/s,

    So it has only 252GB of actual ”AI” memory making it “useless”/toy for actual real world AI workloads(I.e it can’t replace something like opus 5)

    • theplumber 14 hours ago ago

      I find it just misleading by advertising 700GB RAM as AI headline. I could plug a 32GB GPU to my 1.5TB ram server and call it “AI station with 1.5TB+ RAM” just so that you find it actually useless compared to the headline

    • kees99 15 hours ago ago

      Should work just fine for MoE models where active set fits into 252GB.

      • theplumber 14 hours ago ago

        Can’t you do that already more or less with a Mac Studio with 256 or better 512fb of ram?

        • fc417fc802 14 hours ago ago

          Where are you getting 7 TB/s of memory bandwidth?

          • theplumber 12 hours ago ago

            Fair point but you get that only for 252GB of ram so it’s not even “totally better” than a Mac Studio ultra with 512GB RAM. It still the old “smaller but faster than a Mac” stuff nvidia sells.Maybe in 1-2 years they will match the Mac 512 but then Apple will release an even bigger Mac

    • fc417fc802 15 hours ago ago

      I don't know about "useless" (it seems quite useful to me) but I do feel mislead. It's unified memory in the same way that my current dGPU has unified memory. I guess nvlink-c2c probably (?) doesn't introduce a bottleneck but it's still two distinct arenas with very different performance characteristics.

    • cmrdporcupine 14 hours ago ago

      7.1TB/s of HBM is not "useless" -- that's 30 times the memory bandwidth of my DGX Spark -- and nobody is expecting such a machine to run "Opus 5" on its own. For such large models even datacentre GB300 NVL72 are multiple trays linked together via NVlink etc. This machine has QSFP ports and ConnectX for linking up for larger models.

      It's a workstation, not a rack. It's for AI researchers. I'd love to have one on (err, under) my desk.

      What even is this comment?

      • theplumber 12 hours ago ago

        My point is this is “useless” because you can’t run large models. It’s like having a very fast and expensive SSD with a very small capacity.

        It’s not useless but it is for “real work”. Apple has been providing 512GB ram machines for several years so to say you need a rack to run a large model for personal usage it’s missing the point.

        You needed a rack of nvidia cards to match an 128GB ram Mac as well a while ago so it’s more of the same.

        • cmrdporcupine 12 hours ago ago

          People who buy Macs to run local inference are not the target audience for this machine.

          It's an AI research workstation for people whose ultimate work goes on production GB300 NVL72 data centre racks (and for large models you link them together.)

          It's also about 15x the memory bandwidth of a Mac. For models that fit you'd be looking at hundreds of tokens per second on decode and prefill many times that.

          And has CUDA, which is (likely) what your production system will use.

          And runs a real server operating system.

          Also by the time you spec'd out a Mac with the same total memory capacity and computation you'd also be as expensive. And still not have as many cores, nor have the ConnectX RDMA networking speeds.

    • LargoLasskhyfv 14 hours ago ago

      [dead]

  • kstrauser 13 hours ago ago

    Can you imagine a Beowulf cluster of these?

  • cagenut 14 hours ago ago

    I'm just barely starting to wrap my head around mapping model sizes and quants to hardware components and constraints.

    I don't get what this device is for.

    I a million percent understand wanting 252GB-VRAM, that would get me a 280-320B model like glm-5.3 or deepseek-v4-flash, which would be a massive improvement over my gpt-oss:20b 16GB toy. I would gladly pay a grand for this, I would never pay ten grand for this, and it seems to be priced around a hundred grand.

    So obviously the customer is commercial not consumer.

    Can anyone planning a project around one of these at work share what their workload is shaped like and how they're modeling price/performance?

    For instance, I don't get the 512GB of system memory, I'd gladly drop that to 128 to save money. Am I missing something about commercial workloads? Is a 1T parameter model at 20 tok/s more important to your workload than a 300B one at 60? Is it simply a co-dependency of not the model but the other software you're running on the machine thats using/interacting-with/being-driven-by the model?

    Whats your napkin math to justify 100k? Actually thats not even really the question, its more like - whats your napkin math to determine between the "dual linked GB10" use case vs this product's use case vs an 8U supermicro with 4 cards use case.

  • varispeed 15 hours ago ago

    Furious

    What is it with those stupid names?

    No price, so of course this is not for the smelly working class.

  • Egg-Man 15 hours ago ago

    Imagine a beowolf cluster of these!

  • undefined 15 hours ago ago
    [deleted]
  • surcap526 14 hours ago ago

    [dead]

  • TechDebtDevin 14 hours ago ago

    [dead]