3
3 Comments

Anyone else watching their LLM API bill climb for no good reason?

If you’re shipping anything on GPT / Claude / Gemini, you’ve probably felt this:

Same product. Same quality. Bill goes up because prompts got longer — more instructions, more context, more “please carefully…”.

A big chunk of that spend is often pure filler tokens. Not smarter prompts. Just more words.

Curious for people actually feeling this in production:

  1. Roughly what % of your monthly AI budget is input tokens vs output?

  2. Have you tried any prompt compression / cleanup — or is it still “just write shorter”?

  3. What’s the highest monthly LLM bill you’ve hit so far?

Building CuToken (cutoken.in) around this exact pain: compress the prompt, keep the intent, cut the waste.

Drop your numbers / stack in the comments — useful to see how common this is beyond hobby usage.

posted toAvatar for product CuToken
CuToken
  1. 1
    Definitely relatable. Input tokens can quietly become a huge cost once prompts start accumulating context and instructions. I think the tricky part is reducing tokens without removing information the model actually needs. Curious to see how CuToken handles that trade-off in practice.
    1. 1

      Spot on—that’s the ultimate tightrope. Saving pennies isn’t worth it if the app starts hallucinating or missing edge cases.

      We built CuToken to strip out the conversational fluff humans use (but models don't need) while keeping core instructions 100% intact.

      You can actually test it out yourself-Just by signup you will get 5 free PRO Prompt optimiztion .

      If you drop a bloated prompt in there, let me know how the before/after looks!

  2. 1

    The cat-and-room concept gives the habit tracking a very different feel from the usual productivity apps. The 10-day challenge origin story is interesting too.