# Context and compaction

> Read Raw's context estimates, summarize long conversations with manual or automatic compaction, and understand how prompt caching and token usage are reported.

Every model has a limit on how much text it can read at once. Raw tracks the size of what the model will receive, shows you that size, and can summarize older turns so a long session keeps working. This page explains the estimates, compaction and cache behavior.

## Context estimate

The context estimate covers three things: the system prompt, the tool definitions the agent selects, and the messages that are still active. It is an estimate of size, not a count from a tokenizer. It is marked with `~` in the terminal footer.

- If the last response reported token counts, Raw uses that count as a baseline and adds an estimate for what changed since.
- Otherwise, the whole context is estimated.
- A percentage appears only when the model sets `context_window_tokens`. Without it, Raw has no window to measure against.

To get a percentage, set the window in the model entry:

```json
{
  "models": {
    "flash": {
      "provider": "deepseek",
      "method": "openai-chat-completions",
      "model_id": "deepseek-flash",
      "base_url": "https://api.deepseek.com",
      "api_key_env": "DEEPSEEK_API_KEY",
      "context_window_tokens": 1048576
    }
  }
}
```

`context_window_tokens` is a value you supply. Raw does not look it up from the provider.

## Compaction

Compaction replaces the model’s older context with a summary. The visible history stays as it was, so you can still scroll back through the conversation.

Compaction runs in one of two ways:

- **Manual.** Run `/compact` in the REPL or the dashboard. This is the default when no trigger is set.
- **Automatic.** Set `compact.trigger_tokens` on the agent. Before a model request would exceed the threshold, Raw compacts first and shows its progress.

```json
{
  "agents": {
    "deepseek": {
      "model": "flash",
      "tools": { "use": ["builtin/read_file", "builtin/write_file", "builtin/bash"] },
      "compact": {
        "trigger_tokens": 800000,
        "keep_recent_turns": 2,
        "max_output_tokens": 512
      }
    }
  }
}
```

The compact settings are:

| Field               | Default               | Purpose                                                                                                                                                      |
| ------------------- | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| `trigger_tokens`    | Not set (manual only) | The estimated input size that starts automatic compaction. Requires `context_window_tokens`, and must leave room for the output reserve and a safety margin. |
| `keep_recent_turns` | 2                     | The number of most recent complete turns kept as they are.                                                                                                   |
| `max_output_tokens` | 512                   | The cap on the summary response.                                                                                                                             |

### What a summary contains

Compaction keeps the original task once and the most recent complete turns, then asks the selected model to summarize the older turns. The summary request has no tools. Older images are described by file path, type and size, never by their data.

A summary is accepted only when it is nonempty, complete and shorter than the transcript it replaces. If the summary request fails, is cancelled, or does not shrink the context, the previous context is kept unchanged.

Summaries are lossy. Decisions, changed files and open problems are usually kept, but a summary can omit details. If a detail matters, state it again in your next message.

When the history is large and the model’s window is known, Raw splits the older turns into chronological chunks, and summarizes each chunk separately.

> **Note:**
>
> Automatic compaction runs at most once for each turn. If the next request is still too large after compaction, Raw returns an error instead of sending it.

### After compaction

Compaction rewrites the start of the model’s context once, so the next request may not reuse the cached prefix from before the compaction. Later turns build on the new prefix.

`/clear` removes the conversation history while keeping your agent configuration and the cumulative usage counts.

## Prompt caching

Some providers can reuse the start of a prompt when it repeats. Raw helps by keeping that start stable: the system prompt, the tool definitions and the committed messages are sent in the same order on every turn. Raw does not add padding or warm-up requests to chase cache hits.

Whether a cache hit happens is decided by the provider. Raw reports what the provider returns and does not claim a hit without evidence.

| Method and service                    | What Raw sends                                                                                                       |
| ------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| OpenAI Chat Completions and Responses | A stable prompt cache key for the session, unless you set an explicit `cache.key`. An explicit key takes precedence. |
| Anthropic Messages                    | A `cache_control` marker in auto mode. The retention can be `5m` or `1h`.                                            |
| Google Generate Content               | Stable content, relying on Google’s implicit caching.                                                                |
| Other services                        | No guessed cache field.                                                                                              |

The cache key changes when the effective runtime changes, such as when the selected tools change. A change that only edits a selected skill keeps the key and may add a reminder to reload the skill. Changing terminal appearance never changes the cache key.

## Usage and unknown values

Raw shows token counts only when the provider reports them. A missing value is shown as unknown, never as zero. Cache-read and cache-write counts each have their own coverage, because a provider might report one and not the other.

The cache-read ratio is cache reads divided by input tokens. It is computed only for requests where both values are reported. A ratio does not prove that a particular backend served a cached prefix.

Use `/stats` in the REPL, or the dashboard’s Details section, to see the counts and how many requests reported each one.

## Practical advice

- Set `context_window_tokens` for every model you use, so the footer and dashboard can show a percentage.
- Set `compact.trigger_tokens` below the context window, leaving room for the response. A value close to the window leaves little headroom.
- Keep your system prompt and tool selection stable during a long session. Each change to them alters the cached prefix.
- Use `/compact` before a long task when the footer shows a high percentage, instead of waiting for automatic compaction.
