Context and compaction
Every model has a limit on how much text it can read at once. Raw tracks the size of what the model will receive, shows you that size, and can summarize older turns so a long session keeps working. This page explains the estimates, compaction and cache behavior.
Context estimate
Section titled “Context estimate”The context estimate covers three things: the system prompt, the tool definitions the agent selects, and the messages that are still active. It is an estimate of size, not a count from a tokenizer. It is marked with ~ in the terminal footer.
- If the last response reported token counts, Raw uses that count as a baseline and adds an estimate for what changed since.
- Otherwise, the whole context is estimated.
- A percentage appears only when the model sets
context_window_tokens. Without it, Raw has no window to measure against.
To get a percentage, set the window in the model entry:
{ "models": { "flash": { "provider": "deepseek", "method": "openai-chat-completions", "model_id": "deepseek-flash", "base_url": "https://api.deepseek.com", "api_key_env": "DEEPSEEK_API_KEY", "context_window_tokens": 1048576 } }}context_window_tokens is a value you supply. Raw does not look it up from the provider.
Compaction
Section titled “Compaction”Compaction replaces the model’s older context with a summary. The visible history stays as it was, so you can still scroll back through the conversation.
Compaction runs in one of two ways:
- Manual. Run
/compactin the REPL or the dashboard. This is the default when no trigger is set. - Automatic. Set
compact.trigger_tokenson the agent. Before a model request would exceed the threshold, Raw compacts first and shows its progress.
{ "agents": { "deepseek": { "model": "flash", "tools": { "use": ["builtin/read_file", "builtin/write_file", "builtin/bash"] }, "compact": { "trigger_tokens": 800000, "keep_recent_turns": 2, "max_output_tokens": 512 } } }}The compact settings are:
| Field | Default | Purpose |
|---|---|---|
trigger_tokens |
Not set (manual only) | The estimated input size that starts automatic compaction. Requires context_window_tokens, and must leave room for the output reserve and a safety margin. |
keep_recent_turns |
2 | The number of most recent complete turns kept as they are. |
max_output_tokens |
512 | The cap on the summary response. |
What a summary contains
Section titled “What a summary contains”Compaction keeps the original task once and the most recent complete turns, then asks the selected model to summarize the older turns. The summary request has no tools. Older images are described by file path, type and size, never by their data.
A summary is accepted only when it is nonempty, complete and shorter than the transcript it replaces. If the summary request fails, is cancelled, or does not shrink the context, the previous context is kept unchanged.
Summaries are lossy. Decisions, changed files and open problems are usually kept, but a summary can omit details. If a detail matters, state it again in your next message.
When the history is large and the model’s window is known, Raw splits the older turns into chronological chunks, and summarizes each chunk separately.
After compaction
Section titled “After compaction”Compaction rewrites the start of the model’s context once, so the next request may not reuse the cached prefix from before the compaction. Later turns build on the new prefix.
/clear removes the conversation history while keeping your agent configuration and the cumulative usage counts.
Prompt caching
Section titled “Prompt caching”Some providers can reuse the start of a prompt when it repeats. Raw helps by keeping that start stable: the system prompt, the tool definitions and the committed messages are sent in the same order on every turn. Raw does not add padding or warm-up requests to chase cache hits.
Whether a cache hit happens is decided by the provider. Raw reports what the provider returns and does not claim a hit without evidence.
| Method and service | What Raw sends |
|---|---|
| OpenAI Chat Completions and Responses | A stable prompt cache key for the session, unless you set an explicit cache.key. An explicit key takes precedence. |
| Anthropic Messages | A cache_control marker in auto mode. The retention can be 5m or 1h. |
| Google Generate Content | Stable content, relying on Google’s implicit caching. |
| Other services | No guessed cache field. |
The cache key changes when the effective runtime changes, such as when the selected tools change. A change that only edits a selected skill keeps the key and may add a reminder to reload the skill. Changing terminal appearance never changes the cache key.
Usage and unknown values
Section titled “Usage and unknown values”Raw shows token counts only when the provider reports them. A missing value is shown as unknown, never as zero. Cache-read and cache-write counts each have their own coverage, because a provider might report one and not the other.
The cache-read ratio is cache reads divided by input tokens. It is computed only for requests where both values are reported. A ratio does not prove that a particular backend served a cached prefix.
Use /stats in the REPL, or the dashboard’s Details section, to see the counts and how many requests reported each one.
Practical advice
Section titled “Practical advice”- Set
context_window_tokensfor every model you use, so the footer and dashboard can show a percentage. - Set
compact.trigger_tokensbelow the context window, leaving room for the response. A value close to the window leaves little headroom. - Keep your system prompt and tool selection stable during a long session. Each change to them alters the cached prefix.
- Use
/compactbefore a long task when the footer shows a high percentage, instead of waiting for automatic compaction.