3 comments

  • schopra909 9 hours ago ago

    The approach is cool; but the results feel ~sus~.

    GenEval and GenEval2 prompts are very terse (~10 tokens) while these models are primarily trained on long-dense captions. So, it's pretty common to upsample prompts before running them through benchmarks. That way you're testing the model as it was trained, rather than testing it on prompts that are out of distribution.

    If you look at the Qwen-Image Report before they enhanced it for their 12-25 release, upsampled prompts score 0.87 on GenEval* In this paper they take it from 0.74 to 0.81.

    To me that reads to me that they're basically getting the model to adapt to the benchmarks' terse prompts. And it's unclear to me if simply finetuning the model on shorter prompts would work just as well as the RL solution.

    * https://arxiv.org/pdf/2508.02324v1#page=21

  • Lerc 10 hours ago ago

    Interesting, I had been wondering if you could cycle distillation and then splitting weights to turn a saturated model into a non-saturated with the same parameter count for further training.

    Train until you stop getting sidnificant improvements. distill to a quarter size model. Expand back up, and continue training.

    Split the weights so W1+W2 = W with W1 = (W + Randoffset)/2, W2 = (W - RandOffset)/2. Double the width of the layers using the split weights, you get the same result from a 4x size network. My hypothesis is that this has way more scope to train that the model you distilled from.

    • calebkaiser 8 hours ago ago

      This is a productive line of research with lots to explore, if you're interested. There are different flavors based on what you're describing that you might like. Maybe you already know all of this, but just in case any other reader is curious :)

      A classic paper on "expanding" a neural network (the second part of your distil-then-grow loop) is Net2Net. Their goal is basically to get a smaller model's knowledge transferred into a larger one to bootstrap the larger model training: https://arxiv.org/abs/1511.05641 . Bert2Bert is basically the same idea but language models: https://arxiv.org/abs/2110.07143

      There's also stuff on "growing" a pretrained LLM that is somewhat similar to what you're suggesting: https://arxiv.org/abs/2303.00980

      And then there's a bunch of stuff on "refreshing" weights in the network as well, sort of like pruning but without reducing capacity.

      One of the things to explore are assumptions around how information really compresses down in distillation, and what is really happening when we deem a model saturated. There's often an intuition that when you distill the model down, you're finding a sort of "truer", more essential representation of the functions you're modeling. And so if you distill it down and then expand the weights, you've added additional capacity for it to learn more, since the distilled core now contains the important information from the original larger model. But very often that's not actually how it plays out, and the problem is thornier than it seems on its face. Plateaus in training don't necessarily mean the model has run out of representational capacity. And distillation is not guaranteed to preserve the most valuable information. It's entirely possible that the representations the original model has learned for performing on the tasks it is evaluated against are more generalizable than the compressed representations you've distilled.

      Initialization is also tricky when you're expanding the weights like this. Non-linearities make it such that splitting the weights directly like you're describing doesn't preserve the function. Not that this means the model can't recover from further training, but it's just not where you want to start training from if you can avoid it. There's also a strong likelihood that the added parameters are redundant relative to the distilled core, which will make learning more challenging. A lot of work in this space as a result focuses on stuff like symmetry breaking. A cool paper you might find interesting that isn't about LLMs, but has a similar "Progressively give the model more degrees of freedom" flavor is ARC-AGI Without Pretraining: https://iliao2345.github.io/blog_posts/arc_agi_without_pretr...

      But there's lots of active research around this general space. You should test out your ideas and share them! I can't recall now exactly but I feel like I've seen MoE stuff that touches on some similar ideas.