• De Lancre@lemmy.worlddeleted by creator
    link
    fedilink
    English
    arrow-up
    0
    ·
    4 months ago

    You don’t need 170+ GB of VRAM. Whole model can be run at around 1 token/second on a modern hardware from an ssd. Which is slow, don’t get me wrong, but it still somewhat useable.

    • placebo@lemmy.zip
      link
      fedilink
      English
      arrow-up
      0
      ·
      4 months ago

      “Somewhat” is doing a lot of heavy lifting there 😂 How much time does it take to process your average request?