Message length and context limits
Written By Paolo Rufrano
Last updated 1 day ago
Two different limits shape a long conversation: how much of it the model can see, and how long a single reply can be. They're easy to confuse, so here's each one.
The context meter
Next to the model button you'll see something like 18K / 200K. That's how much of the model's context window your conversation currently fills. Hover for the exact numbers and a percentage.
The bar changes colour as it fills:
Green: plenty of room
Amber at 75%: getting full
Red at 90%, with a warning: Context nearly full. Branch to continue with fresh context.
At that point a branch button appears next to the warning. Click it to start a fresh chat carrying the conversation so far, which resets the meter while keeping your history.
The limit is the model's, not Percy Chat's. Different models have very different context windows, and switching to a larger one gives you more room.
What happens when a conversation gets long
Every time you send a message, the whole conversation is sent to the model again. That's how it remembers what you discussed, and it has two consequences.
Cost grows. Message 40 costs far more than message 4, because there's forty messages of history to re-read. This is the most common reason credits disappear faster than expected. See How credits work.
The oldest messages get dropped. To stop cost growing without limit, Percy Chat trims the oldest messages out of what it sends the model once the conversation exceeds your plan's budget. Your most recent message is always kept.
How much history is sent depends on your plan:
Trimming is invisible in the chat. You still see the full conversation on screen, and nothing is deleted. But if the model seems to have forgotten something you said much earlier in a long thread, this is usually why.
How long a reply can be
Each plan also caps how much a model can write in a single response:
Reasoning counts towards this, so Advanced thinking leaves less room for the visible answer.
If a reply stops mid-sentence, it hit this cap. Ask the model to continue, or ask for a shorter, more focused answer.
Working with long conversations
Branch when you change topic. Organising your chats: pin, rename, branch, search and delete covers how. You keep the earlier chat intact and start fresh.
Start a new chat for unrelated questions. Cheaper, and the model isn't distracted by irrelevant history.
Use a Project for context that should persist. Project instructions, files and memory are injected into every chat in the project, so you don't have to re-explain your situation in each one. See Projects: instructions, files and memory.
Upgrade if you routinely work with long documents. Premium sends ten times the history that Free does.