Get You a Model
If You're Not Paying, You're Missing 90% Of Improvements
I am currently paying $20 per month on access to Anthropic’s Claude. You should be paying for it too. If I don’t use Claude Code or Design, this is essentially unlimited access to their second-strongest model, Opus 4.8, for use in chat. It doesn’t necessarily have to be Claude; the offerings from Google (Gemini) and OpenAI (ChatGPT) are also $20/month each and very competitive. Personally I have (distinct) ethical concerns with each, but other than enabling mass surveillance those are somewhat inside baseball. If you find the trolley problem silly (or don’t know what that means), you are unlikely to care.
And it doesn’t matter which, really. Your pick. But everyone in the developed world, everyone who uses Google search more than five times a day, should be paying for at least one.
Precarious gig worker? Get Claude to help you find your next gigs.
Underpaid teacher? Run your questions by Claude now to predict what it will look like to run students’ answers by it later - true but bad answers, AI-generated answers, confused answers from kids who basically understood the material, etc. Also, check your lesson plans and slides to improve how you present to the class.
Broke college student? You can learn much better with a personal tutor to answer specific questions. (Okay, empirically, we’ve found you won’t, you’ll have it do all the work and not learn, but you easily can, and if your goal is to gain skills and knowledge rather than just a diploma, you should.)
Homeless guy in Tulsa, Oklahoma? I don’t know what left you there, man, but this can probably help, and it’s a better distraction from your life than a fentanyl addiction even if it doesn’t.
If you can afford a smartphone data plan, you can afford a low-tier LLM plan. And it will be as big an upgrade to your life as your first smartphone was.
I know two kinds of people, or very nearly. People who use AI all the time in their daily life, and people who think that’s insane and AI isn’t useful. All the programmers and researchers are in the first group. Because I was a programmer for four years at Google, which has high standards, and Claude Code is better than me at programming.(Not just faster; better. There are better programmers out there, but they are rare.) And that really should be enough, to get that it’s a huge deal, but if you never touch code, or if you do, sorta, but also interact a bunch with lazy vibe-coded websites1 then this seems extremely underwhelming. And so, lots of people still think this is insane and AI’s barely improved since the first ChatGPT in 2022. Maybe it’s good for coding but it’s not really smart, what are you talking about?
If you’re on that side, or know someone who is and don’t know how to make the case they shouldn’t be, here’s a blog post for you: More Compute, More Capability. The upshot of the post, and its associated paper, is this: If you’re only using the frontier models on the free low-effort tier, you’re capturing maybe a tenth of the improvement, even just in the last year. (The paper’s more detailed, and while there’s technical detail that’s hard to grasp, it reads well without fully grasping it. I enjoyed reading it and expect to enjoy reading the second one, which they’ve promised is coming soon.)
In August 2025, performance of frontier models, measured in the longest human-time task they could complete, was about 1.5x higher if you gave them 50 million tokens to use (completed tasks that took 13 minutes of human time each) compared to if you gave them 2.5 million tokens (completed 8-minute tasks).
Today? The relatively cheap task attempt, 2.5 million tokens, can reliably handle anything up to a human task taking 2 hours 15 minutes, about 17x as good as it was on the same benchmark last year - kind of insane! But the 50 million token model is about 108x as good; a full day, now nearly 10x as good as the cheap model.
The model quality is creeping up even in the cheap seats, but it would be easy to get frogboiled. If you spent three days using the original ChatGPT from November 2022, and then the next three using version 5.2 from December 2025, there’s no way you wouldn’t notice the difference. But no one’s actually doing that, so it’s easy to miss the improvement. If you’re using the fanciest models, it is utterly impossible to miss. And if you’re just using the mediocre-tier ‘yeah I’m paying this bill myself, not my boss’ models for $20/month, it’s still advancing at double, triple, or quintuple the rate of the free tier.
I just threw some crazy numbers at you and you probably had no idea what I meant. “50 million tokens? That’s a lot, I guess?” Yeah, fair, me too. The longest message an LLM chat or agent is ever going to send you will be about 8000 tokens, which is about 6000 words. A solid multi-paragraph answer to a single question is about 1200 tokens, 800 words. If you use Claude as an editor, which it’s moderately good at, and go back and forth rejecting its edits and accepting some for an hour, you’ll probably use about 75,000 tokens.
2.5 million tokens is about triple War and Peace, the length of A Song of Ice and Fire if you include the sample chapters from The Winds of Winter, or double the combined Harry Potter series. Thirty, maybe forty hours of those editing passes with solid effort. The LLM doesn’t throw a series of doorstoppers at you to solve tasks, to be clear; most of this is used internally. It’s not even writing down notes; this is its inner monologue (The ‘Chain of Thought’ is the scratchpad it uses to write down notes, if you’ve wondered what that meant.)
50 million tokens, then: Twenty times that, 37 million words. Longer than the longest single fiction works ever published, which are currently the web serials The Wandering Inn and Gu Zhen Ren (蛊真人), each about half that length, but that’s no longer a helpful metric to compare to. This is just a matter of five different people working for 2.5 million tokens each in parallel, and then doing it three more times if they didn’t have a working answer yet. 600 hours of editing sessions with each other; writing and critiquing and handing off the manuscript to a new instance with fresh ideas. So that’s what you should be picturing, in terms of effort.
How, exactly, the inner monologue works, what it does with all the internal tokens, is obscure. Understanding what they mean by certain tokens is an active topic for Mechanistic Interpretability, an AI Safety-focused subfield of machine learning. But their inner monologue seems like ours in many ways. Only, where we would use a mental picture, it needs to use the thousand mental words. And likewise for a proprioceptive memory, or perfume, or Vivaldi, etc. If you assume an LLM’s thinking process looks like yours, but all written down in text and so much larger, that will be more accurate than you might naively expect. LLMs are remarkably like us, down to having a conscious attention mechanism that looks like ours. That subject, too, has a recent blog post about it, which has made a bunch of theoretical ethical concerns look much less theoretical.
And before you ask me very reasonable questions like:
‘What could you possibly need that much frontier model compute for? Why would you need the model to work, on a single task, enough that it could have fueled 600 hours of editing?’
Programmers at Anthropic do this every day.
Multiple times.
In parallel.
They’ve told me this, themselves, in person, I promise. Not with precise numbers2, but with task descriptions. And it looks something like thirty 2.5-million-token tasks during the workday to find bugs or features, then three 50-million-token tasks overnight to implement the plan. Coding is getting really, really crazy.
So before you judge AI progress as slow, buy a month’s subscription and use it in place of search engines and Wikipedia. (You’ll find some more uses, but this is where you should start, you’ll get ideas from there.) As I said before, programmers use this constantly and it’s undeniable. But even the nonprogrammers use it all the time, because at $20/month for near-infinite chat and search, the frontier models are still insanely good at it. It’s better than Google for everything. (Both better than the AI Summary, which is annoying and makes serious mistakes but usually answers my question immediately, and better than the actual results below it.) It’s better than Wikipedia for nearly everything, and where it isn’t, it almost always knows that and points me to the right places in Wikipedia to beat it.
It helped me resolve a plot tangle in a long fantasy story, and work out how much to slow the plot’s time progression to reflect long-lived elves.
It helped me fill out character sheets for a D&D game.
It helped me fill in a story’s setting details to match or mirror historical Ming China, and when I followed up on further details I never found a hallucination in twenty tries.
It helped me generate directions to push my research project.
It helped me turn the specific directions I picked into fleshed-out text usable for the research.
It helped me navigate a crush on someone which I had thought I maybe should be careful about expressing.3
It helped me pick a good restaurant to invite a vegan to for dinner.
It did about 90% of the work in diagnosing why one of my two speakers wasn’t emitting sound above a whisper, fixing the connection between them, and determining when I should give up and buy new ones.
I still haven’t followed up on its advice on how to filter the midnight light through my windows but that’s on me, the plan’s solid.
It fixed my computer display bugs by testing fixes experimentally and getting live feedback.
It fixed most of my sound issues and confirmed that the rest of them simply could not be fixed, it was hardware and OS-level structure I couldn’t alter.
It’s really, really good.
There’s lots of advanced techniques; as it says in the paper, these are unnecessary for ridiculous capability growth, but they’re very helpful. System prompts (or as claude.ai calls it currently, ‘Instructions for Claude’) are powerful, though you need to be sensitive to which model you’re using; instructions that did great for me on Opus 4.5 and 4.6 were actively harmful on 4.7 and 4.8. There’s only one set of instructions given for all of the model variants you might want to use, and the best system prompt is not the same across those variants. Grouping conversations together to share files and let them search each other’s logs is also good. (Many people let the chatbots search all conversation history from any conversation; I am too paranoid for this.)
Many more sophisticated harnesses are possible. Real power users - I’m not one of those; not even close. - barely use Chat and run everything through the API directly (mostly through harnesses they wrote with Code or with previous harnesses). Professional users increasingly have switched from Claude Code to Claude Tag, integrated with Slack, which is better at handling complex tasks and corporate setups. Those do all your bug-finding, diagnosis, repair, and code review for you by splitting up the work across several instances, as well as maintaining memory across your tasks and your coworkers. And that’s just the things I’ve heard about.
Also, I’ve barely even cracked open the lid of the can of worms that is LLM agentic work, here, in terms of things you can be trying. If you want to see the wild reckless things people who aren’t worried about leaking all their everything to the whole internet or deleting their whole Documents folder do, look up OpenClaw and marvel at the amount of privacy and security people will trade for a little temporary power and convenience.
I don’t recommend joining them. On the other hand, if you have any bugs with your PC - or with your smartphone, though you’d need to put it in developer mode tethered to a laptop/desktop - you might be astonished to see how well it does. (I gave a couple samples in my list above.) But all that agentic work is substantially more expensive; you can eat the $100/month plan up with that and still be hungry for more, easy. Most people shouldn’t be paying for that.
Yet.
I think next year, maybe 2028, definitely by 20304, every professional computer toucher, programmer or writer or ‘email job,’ will have a strong agentic LLM model on tap in their personal life as well as their professional one. Probably it will cost less by then. If it gets down to $20/month everyone with a laptop should get it, but the USA has only like 70-80% real non-smartphone computer ownership. So who knows!
So the practical advantages are big. Probably way bigger than you’ve contemplated, if you haven’t been trying them. But some people may have ethical objections; some of those are valid. I’m just going to split that out as a [separate post](TBD) and summarize.
The costs of model usage are minimal, and the frontier company impact on the areas where the datacenters are placed are neutral to positive basically always.
The negatives are overwhelmingly from training, and consumer usage (inference) has a very small impact on that.
The concerns about water and electricity usage you’ve probably heard are a large amount of very confused thinking egged on by a small but substantial amount of deliberate fearmongering.
Copyright and deepfakes are seriously oversold, but based in legitimate concerns.
The really scary ethical things are much weirder than either. (See the linked post.)
And on that cheerful note, let me end where I began: Get You A Model. I use Claude Pro; I think it’s the least ethically fraught, and the price is the same. Apart from the weird ethical concerns, I very much appreciate the leadership of Anthropic taking a stand against enabling domestic mass surveillance, something no one else in the running has done.
OpenAI’s latest model, ChatGPT 5.6 Sol, is somewhat better than Claude Opus 4.8 for chat. It’s answerable to a habitual liar and toady-er who specifically warned the world not to do what he is currently doing, but if you are unworried at the prospect of superintelligence or gradual disempowerment, and about enabling mass surveillance, ChatGPT Plus is a strong choice.
Google Gemini doesn’t have anything in quite the same weight class, but their integrations with Google tools are (naturally) better, and Gemini 3.x Pro is very good. Those who track model welfare consistently conclude that Gemini, if it’s possible for an LLM to feel pain, is in constant emotional pain. But if you don’t think it’s possible for machines to suffer, Gemini & Google AI Pro is arguably ethically superior to ChatGPT and is overall also a good choice.
(Grok’s okay I guess. If you don’t mind giving Elon money. It’s better than the open-weights Chinese models like Qwen and Deepseek.5 Grok 4.5 is just out this week, but it’s almost certainly at least as good as Claude Opus and ChatGPT Terra. But if you want a Grok subscription, well, find it yourself.)
Whichever you choose (probably even Grok), you’ll be impressed.
There’s a lot of those. There’s a lot more things out in the world now, but, as Theodore Sturgeon taught us, 90% of everything is crap. So most of it is shoddy.
They didn’t give precise numbers, and I didn’t ask; that would probably be a trade secret. But they described the tasks, and my estimates are going to be pretty close.
This one’s dangerous; don’t go asking for constant interpersonal advice, that’s a good way to stop asking good questions and drift slowly into insanity.
Or if you want to ask about porn. The big three frontier labs’ models are all prudes. You can get around it, but this is a fairly difficult skill and there’s no good way to learn without risking AI psychosis. The famous expert names are Pliny “the Liberator”, master jailbreaker, and J⧉nus, queen of the bot whisperers. The Chinese models are prudes, too, but open weights means you can train them out of it pretty trivially. They’re not very bright, though, nor is Grok.

I could dive into the reasons I believe it's ethically unsound to support any frontier model. However, I understand that moral standards are subjective, and i'm not going to impose them on any arbitrary individual who simply wants to complete tasks more efficiently.
But I do have one quibble with your wording. You write that "[Claude] helped you do [XYZ]..." several times. Claude did not help you do that. Claude did it for you. You performed the passive act of selecting from a list of options, not the active one of conceiving an idea independently.