DeepSeek-V4-Flash-Vision-Exp

(huggingface.co)

20 points | by amulyabaral 17 hours ago ago

5 comments

  • BrucecarlL 5 hours ago ago

    I'm not that impressed with this version. When using it, it often acts as if it has no visual capability and refuses to recognize images unless I remind it.

  • syntaxing 17 hours ago ago

    I’m honestly surprised this is better benchmark wise than the text only model. I figured the addition of vision would take away from some of the text capabilities.

    • mcbuilder 17 hours ago ago

      One of the early results from multimodal training is that it kinda works like cross training. Training vision helps with text tasks and visa versa.

    • Llamamoe 14 hours ago ago

      I believe that multimodal training increases the robustness of latent representations regardless of which modality is being processed.

      • nateb2022 13 hours ago ago

        Tangentially this makes me wonder how large Opus really is. Perhaps Opus is a lot smaller than most of the 1T+ assumptions, just a lot more post-training/ finetuning on a 300-400B sized MoE model.