Some_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 3 months agoGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comexternal-linkmessage-square193linkfedilinkarrow-up11arrow-down10
arrow-up11arrow-down1external-linkGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comSome_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 3 months agomessage-square193linkfedilink
minus-squareBrett@programming.devlinkfedilinkEnglisharrow-up0·3 months agoIs that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
minus-squareAsafum@lemmy.worldlinkfedilinkEnglisharrow-up0·3 months agoIt’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
minus-squareDamage@feddit.itlinkfedilinkEnglisharrow-up0·3 months agoMy framework 13 with shared RAM runs qwen quite well
minus-squareBrett@programming.devlinkfedilinkEnglisharrow-up0·2 months agoWhat are you using to run the model? Llama.cpp will automatically split the model between your system ram and graphics card’s vram. Qwen 3.6 is a mixture of experts model with only 3B parameters active at a time. Even without quantization your card could easily run that.
Is that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
It’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
My framework 13 with shared RAM runs qwen quite well
What are you using to run the model? Llama.cpp will automatically split the model between your system ram and graphics card’s vram.
Qwen 3.6 is a mixture of experts model with only 3B parameters active at a time. Even without quantization your card could easily run that.