• leanleft@lemmy.ml
    link
    fedilink
    English
    arrow-up
    0
    ·
    16 hours ago

    if you start excluding the 1000+ B param models …
    using smaller models, would initially ease hardware demand by 60% .
    OFC you cant… and probably shouldnt, ignore and disrespect SOTA flagship models

    • brucethemoose@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      15 hours ago

      Even “big” open source models like DSV4 and Ling/Ring are very efficient. They’re big, but (seemingly) sparser than US models, so they’re cheap.

      They run surprisingly well with hybrid CPU+GPU inference on desktops. And thats not even getting into the efficient attention mechanisms.

      I can run DSV4 Flash, barely quantized, with ~1M context on my Ryzen desktop at ~11 tokens/s. If you told me that two years ago, I would not have believed you.