Welcome to Prompt, Tinker, Innovate—my AI playground. Each edition gives you a hands-on experiment that shows how AI can sharpen your thinking, streamline your process, and power up your creative work.
This week's playground: Three settings determine how fast you burn through your AI limits
Every AI subscription—Claude, ChatGPT, Gemini, Copilot—has usage limits, including business plans. Yet I hear this all the time: people on ChatGPT Plus or Claude Pro assume limits do not apply because they "are not on a token plan."
If you pay a flat monthly fee, it is easy to assume the model dropdown only affects answer quality. With a flagship model available, why not choose the best model for everything? It is all included in the same monthly subscription, after all.
But your plan is not unlimited. Three settings determine how quickly you use your allowance: which model you choose, how hard you ask it to think, and how quickly you need the answer.
Here's the shorthand: model sets the rate, effort drives the volume, and speed adds the surcharge.
Those are three different levers, and each can use more of your allowance than you intended.
One caveat: these levers do not work identically across tools. Some apps do not expose all three, while others bundle them into a single setting.
Another caveat: the way you interact with AI can also affect usage. That is next week's newsletter.
Model—the rate
This is the most direct lever. Every model included in your subscription draws differently against your usage limit, even if the app never shows you a price. The gap is large. Across the market, models can range from roughly $0.10 to $168 per million tokens processed—more than a 1,000x difference.
Your subscription does not show that math, but the effect is similar: choose a flagship model for work a lighter model could handle, and you spend a larger share of your allowance on the same conversation.
One thing that quietly makes this worse: output tokens—what the model writes back to you—typically cost more than input tokens (by about 4x). And the "smartest" models often write longer answers. So a flagship model can hit your limit twice: pricier tokens and more of them.
The takeaway: the model picker isn't just a quality dial. It's one of the biggest things deciding how many messages you get before your limit resets.
Effort—the volume
This is the one nobody sees coming, because it doesn't look like a cost at all. A lot of tools now let you set how hard the model "thinks" before answering: low, medium, high, sometimes labeled reasoning or thinking mode. Raising that setting does not necessarily change the model. It changes how much work happens before the answer arrives—and that work can count against your usage limit even though you never see it.
The gap can be much bigger than the visible response suggests. A simple factual question might produce 50 visible tokens in the answer but trigger 500 to 2,000 tokens of hidden reasoning behind it—10 to 40x what you actually see.
So two people can ask the same question, get similarly short answers, and use very different amounts of their monthly limit—simply because one left the reasoning setting on high.
The takeaway: if you're blowing through your limit faster than expected and your answers aren't getting noticeably better, check your reasoning/thinking setting before you blame the app.
Speed—the surcharge
Speed is the easiest dial to misunderstand because tools handle it differently.
In subscription apps, faster or priority access may come from a smaller usage pool. Use it for work that did not need an immediate answer, and it may not be available when speed actually matters.
On the API side, the tradeoff is even more explicit: a slower, best-effort queue runs roughly 50% cheaper, while priority processing costs roughly 2x the standard rate, in exchange for responses up to 2.5x faster.
The takeaway: speed isn't free just because there's no price next to it. If your tool gives you a faster or priority mode, save it for work where the wait actually matters.
One more trap: some consumer chat apps also switch to a lighter model when you choose "fast." A toggle that looks like a speed setting may be changing the model, too. Check what it actually does in the tool you use.
Your AI experiment: Go find your dials
👉 Time to tinker: Open whatever AI tool you use most. Look for a model picker, a reasoning or thinking level, and, if your tool offers one, a speed or priority mode.
📝 Prompt:
Pick one task you run regularly. Test the setting you suspect costs you the most. Compare the answer to your usual setup: did the extra power materially improve the result?
💡 Pro tip: Check your usage meter before and after, if your tool shows one. That is the clearest evidence of which dial is costing you the most.
What did you discover?
Which dial made the biggest difference for you: model, effort, or speed? Did any of them surprise you? Tell me where the real cost was hiding.
Until next time—keep tinkering, keep prompting, keep innovating.
📩 Not subscribed yet? Hit the button at the top. You won't want to miss what's next.



