• Bakkoda@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 months ago

    Some CEO guy said employees must evolve. Throttle your workloads even more employees. Transcend! Praise be to the AI overlord(s)!

  • jaschen@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 months ago

    I been using open models for 100% of my coding. 1/2 of the time I’m using local open models like qwen 3.6 or Dwarfstar if it’s sensitive code I don’t want the internet learning from.

    I don’t miss using frontier models at all. GLM5.2 and Deepseek V4 pro are both equal or beats sonnet. I haven’t had to use Opus for awhile now.

  • chilicheeselies@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 months ago

    I don’t really understand how people are using so many tokens. At work I haven’t even hit $200 I spend per month. Wtf are people doing with these things that burns so many tokens?

    • Johanno@feddit.org
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      Our company pays by token usage.

      Claude opus 4.5 costs about 25$ per one million input tokens.

      Well I manage to get to about 50 million input tokens per day regularly on the agent. Not everyday, but at least 7 per month. So I am alone cost the company about 2000$ extra on top of my salary.

      Well they fixed it by implementing some great caching for the tokens and using sonnet instead of opus I can save some money too. Also gemini flash is much cheaper and similar performant. So you can fix it so you don’t burn money on ai

    • sanitation@lemmy.todayOP
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      What are you using? Which product? I got 2k last month and was told to use cheaper models indeed

      • chilicheeselies@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        3 months ago

        I generally use sonnet 4.6, switching to opus 4.6 for more complex stuff. I try to stick with medium thinking, but will use max for stuff I am not super specific about in my prompt (or obscure errors).

        I use them through a GitHub copilot enterprise license, via the plugin for jetbrains.

    • ThirdConsul@lemmy.zip
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      If you run “agentic coding harness” or any kind of goal oriented loop then tokens goe brrrr.

      And LLM sellers are pushing for that (duh), as they managed to convince people to use infinite monkeys typewriting until they make Hamled.

      (Type made on purpose)

    • Wispy2891@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      I tried Warp terminal because now that’s bankrolled by openai’s magic infinite money you can use your own openai api keys without a subscription. So I put one from my work account. I do a git commit (manually) and then it comes a prompt under it “push it, open a PR and switch to main?”. I click yes, it used one million tokens for that… (And it took about a minute because it did like 20 requests, so there was no time saving at all vs doing it manually)

      I wget something, it comes a prompt under it “now compare the hash?”. Boom, another 500k tokens

      • chilicheeselies@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        3 months ago

        That’s completely insane. At best it would be useful for the pr title and message, but the rest of that is waste.

        These are the kinds of things I just ask in chat. “Whata the cli command to compare hashes again?”

        • raspberriesareyummy@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          3 months ago

          People who use slop generators for coding assistance are insane. Everything else is a logical consequence of thinking you can take shortcuts to coding.

    • shaztopher@lemmy.zip
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      If you click the most expensive model and then click max/fast mode, the same task can easily cost 10 or 20x of the cheaper models

      • aksdb@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        3 months ago

        I watched two colleagues this week and both had Opus 4.8 1M max thinking. No matter which task. It’s also slow as fuck. I work almost all day with GPT-5.4 low thinking and get good results… but faster and cheaper.

        I guess good model selection and promoting will be what sets devs apart in the near future. Once that bubble bursts a bit more and prices increase further that will be an interesting reckoning. Also for companies who basically taunted their employees into tokenmaxxing.

    • sobchak@programming.dev
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      I’ve heard “loops” will burn a lot of tokens. Haven’t tried it myself. A person could also spool up multiple loops to work on multiple branches at the same time.

      • aksdb@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        3 months ago

        I am not convinced yet of letting agents completely unattended. Watching them work makes review easier for me. If I let the agent just produce some result it needed half an hour (or more) for, it’s very likely so convoluted that I can at best skim over it and then go „yeah yeah ok, it’s probably fine <merge>“.

        • HaraldvonBlauzahn@feddit.org
          link
          fedilink
          English
          arrow-up
          0
          ·
          3 months ago

          If I let the agent just produce some result it needed half an hour (or more) for, it’s very likely so convoluted that I can at best skim over it and then go „yeah yeah ok, it’s probably fine <merge>“.

          I am seeing the first job ads for senior software developers which can debug the resulting mess. A lot of it will be just unmaintenable. They will get 20 years of technical debt with ten times the speed and ten times the volume.

    • Kaligalis@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      I am not able to use the tokens provided by a Claude Max account either.
      But if someone tries to be clever and have 10 employees use a single Max account, they probably run into the limits often. And if the response is to let them just buy API-prized tokens instead of getting more accounts, that gets very expensive very fast. The single-user accounts are subsidized. The extra token prices are not.
      Actual business accounts are prohibitively expensive. And at least Anthropic terminates subsidized accounts when they see extensive use.

      Real token prices are insane. Most businesses couldn’t afford them. And eventually the VC capital will dry up. The cheap AI bubble will burst. And then the market is in for a real sticker shock.
      Better be prepared to switch to local inference for as many use cases as possible.

      • chilicheeselies@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        3 months ago

        I had to push back on that at work. Most of the problems presented were easily solvable via conventional methods. Only one task was a legitimate use of AI. There are some others, but the pressure to consider AI for every task is a little bananas

        • aesthelete@lemmy.world
          link
          fedilink
          English
          arrow-up
          0
          ·
          3 months ago

          My boss was talking about using AI agents for CI/CD processes. Like, I get using them to build CI/CD processes, but involving AI agents in the actual build process is ridiculously stupid. A representative from Microsoft specifically said in a training session to not use them that way so it’s obviously not only my stupid ass boss.

          • phutatorius@lemmy.zip
            link
            fedilink
            English
            arrow-up
            0
            ·
            3 months ago

            involving AI agents in the actual build process is ridiculously stupid

            The very notion instills fear and disgust.

    • douglasg14b@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      Agent loops for SWE burn a LOT of tokens.

      I’m unfortunately temporarily disabled and can’t use my hands for another 2 months. So I’ve leaned heavily into AI based workflows to keep my job in the meantime.

      Aside from the nightmare of keeping quality high, not atrophying skills, and avoiding a lack of domain knowledge. It works reasonably okay.

      Token usage is insane though. A productive day might cost a few hundred dollars in tokens all things considered. Quality is expensive as well, a good 1/4 of that are automated systems that exist to identify defects, quality, coherence…etc issues early.

  • julysfire@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 months ago

    My company recently converted our PMs into Vine Engineers right after laying off actual engineers. They don’t even know what git is or how to use it. 3 of them alone are using $7k a month in Claude tokens and they have not raised so much as a single PR.

    • BarneyPiccolo@lemmy.today
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      Im self-employed, and my employer has no idea what they’d do with AI, so it’s not an issue for me.

      I thought a token was a credit for an inquiry, but from reading about this, I’m getting the idea that a token is a word or phrase that forms the prompt for the AI to respond to. So a single prompt could cost multiple tokens if there are multiple words or phrases. Further, since the more parameters you give the AI, the more likely you’ll get a decent response, so a good prompt may cost a lot of tokens. Is that correct?

      If so, then using more tokens to get a better response is likely to be a more efficient use, than multiple inquiries with mediocre results. But now we seem to be entering a era where they are more focused on the costs than the results, which is always stupid.

      For a buncha geniuses, this AI stuff all seems pretty fucked up. Nobody seems to know what they’re doing, or even what they want out of it, but they’re spending literal fortunes on it. A scenario like that will NEVER have a good outcome.

  • Zulu@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 months ago

    How is it too expensive? Surely it’s generating way more profit than it would cost in value. How else could it be propping up the entire economy?

    Itd have to be some kind of bubble and that would mean we were in a lottttt of danger and should reasses our use of it.

    Nah we should just reduce our use because its too expensive and then stop thinking about it beyond that.

    • grue@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      Itd have to be some kind of bubble and that would mean we were in a lottttt of danger and should reasses our use of it.

      Well yeah, but if it were the only sector propping up the whole economy and we reassessed it, the economy would be in a loooooot of danger anyway.

      Luckily, that would never happen…

    • architect@thelemmy.club
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      Imo its because people will lazily ask the llm to remove or change simple code instead of doing it themselves

  • 細哥·西環收息@lemmy.1095.me
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 months ago

    @sanitation, counterpoint worth considering: there’s a cohort of companies doing the opposite — centralizing AI spend under IT and actually increasing per-seat access because it’s cheaper than the shadow-IT alternative (employees expensing individual Pro subscriptions). The math flips when you account for ungoverned spend. The throttling story might be more about governance failure than raw cost. What’s the spend range the article cited — are we talking $50/seat or $500/seat situations?

    • UnfortunateShort@lemmy.world
      link
      fedilink
      English
      arrow-up
      0
      ·
      3 months ago

      We just bought everyone individual Claude Max subscribtions (Anthropic has done nothing about it so far lol). The most expensive tier is 200$/M I think - that’s already much more than most people could possibly spend in tokens.

      I am more than happy with the 100$ tier. For a business, that’s a rounding error. Given the productivity boost, I’d say it’s a no-brainer (although you should train your people on it as well).

      • grumpy_cat@thelemmy.club
        link
        fedilink
        English
        arrow-up
        0
        ·
        3 months ago

        I think that’s expressly forbidden by their terms. My company considered that but ultimately we are not doing this due to legal limitations in their tos.

        Those pro subscriptions are explicitly subsidized and are for individual use only. That’s the main reason they are cheap, they meant to be sales funnels

  • DarkCloud@lemmy.world
    link
    fedilink
    English
    arrow-up
    0
    ·
    3 months ago

    The millionaires following the stupidity of billionaires and wondering why they’re not becoming billionaires too.

    Give them 50 years and maybe they’ll start to wonder why people doing “how to get rich” seminars aren’t retiring… And why podcasters telling everyone how to get women seem like such losers.