• Asafum@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      It’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16

      • Brett@programming.dev
        link
        fedilink
        English
        arrow-up
        0
        ·
        2 months ago

        What are you using to run the model? Llama.cpp will automatically split the model between your system ram and graphics card’s vram.

        Qwen 3.6 is a mixture of experts model with only 3B parameters active at a time. Even without quantization your card could easily run that.