posts / go

Lesson 8: Compaction

Compaction

Since lesson 5 our context meter has been a smoke alarm without a fire department. It warns at 80%, and then it does nothing. It’s time to change that. In this lesson we’ll heed the warning and summarize the conversation to make room for more.

What is Compaction?

In lesson 5 we defined num_ctx (the size of the context window) and we built the meter to warn us about it being almost full. We get this warning because a fixed-size frame is about to be filled with more tokens than the model can put in memory.

No errors or warnings, the oldest tokens simply stop being read No errors or warnings, the oldest tokens simply stop being read

As you can see, the context window is a fixed frame but the list of tokens keeps growing. By turn 9 the six oldest messages, the system prompt among them, lie outside of the frame. If the agent ignores the size of this frame, early tokens will be lost and the model will forget parts of the conversation. When that happens, the model will not receive part of the input and will start responding without taking it into account. And the worst part is that the very first message is the system prompt which will now be gone and the model will act without it. Nothing good can come out of this.

So how do agents solve this problem?

Compaction.

In simple terms, compaction is a process of summarizing and rebuilding the conversation so far, as to shrink the size of the message array. However, this process must follow a strict rule: after compaction the array must carry exactly three things: a fresh system prompt, one message carrying the summary, and a tail of recent messages copied verbatim.

The Summarizer

To create the summary of the conversation we cannot just blindly call our chat function.

reuse chat and you would watch the summary print as if loom were talking to you reuse chat and you would watch the summary print as if loom were talking to you

chat streams to the terminal its thinking and its responses, it advertises tools, and it presents the system prompt as the very first message. None of this is what we want.

The summary function must send all of the messages except the system prompt and the instruction on how to perform summary. Think about how the instruction for summarization should look. The instruction must capture the user’s goals and decisions made. Tool calling also needs to be also summarized in a specific way. It should include the files touched and how they changed. The results of the tool execution still matter. Any unfinished work should also be included in the summary.

what the model reads first after a compaction: who it is, what happened, what was just said what the model reads first after a compaction: who it is, what happened, what was just said

Slot 0 is systemPrompt() called again rather than old system message copied. This is done to refresh time and anything else that could’ve changed. Slot 1 carries the summary as the user role, wrapped in a bracketed sentence: [Context summary - earlier messages were compacted to save space. File contents mentioned below may be stale, re-read files before editing.]. Slot 2 onward is the tail. Where the tail begins is the first thing we need to figure out.

The Tail

Should the tail be the last six messages? Well, maybe… Suppose the tail is the last six messages and the sixth from last happens to be a tool result. Then the tail begins with an answer whose question, submitted as an assistant message carrying tool_calls, has just been summarized away. We cannot split these two and end up with the result without the call. This would be confusing to the model.

Here is what we need to do:

  • Decide how many messages to keep. Call it keep. Six messages are a good default.
  • Cut: start = len(conversation) - keep. If that is less than 1, make it 1, so conversation shorter than the tail keeps everything after the system prompt.
  • Walk back: while start > 1 and conversation[start] is a tool role message, subtract one. Every run of tool results is preceded by the assistant message that asked for them, so this stops on that message and the pair is included together.
  • Rebuild: a fresh system prompt, then the summary, then conversation[start:].

finding tail The naive cut at 8 lands on a tool result. Walking back over the two tool messages stops at 6, the assistant message that asked for them, and the tail grows from six to eight messages.

Two things to notice. First, the walk back can only make the tail bigger: it moves the cut toward the beginning of the array and never toward the end. Second, it always terminates on the start > 1 which guards against consuming the system prompt.

The Trigger

We know how to rebuild the messages and now we must take a look when to do it:

  • Fix the threshold: `compactAt = numCtx * 8 / 10. We start at the 80% line.
  • After every successful chat call, remember the count: lastUsed = used, and print the meter from here now that chat no longer does.
  • When the inner tool loop has ended, and only then, test lastUsed > compactAt. If it holds: print a status line with the message count and the numbers. Call summarize. On success replace the array with compact(conversation, summary) and set lastUsed to 0, since the new array’s size is unknown until the next call measures it. On failure print a warning and carry on uncompacted.

the check runs once per keyboard message, after every tool call has its result the check runs once per keyboard message, after every tool call has its result

Why after the inner loop and never inside it? Inside, tool calls are in flight and the array may end with results the model has not yet read. The next thing that must happen is a chat call that reads them. Rewriting the array between the call and its results is exactly what we are trying to prevent.

Your Turn

Try to build it yourself. This list captures all the things we talked about in one place:

  • chat(ctx, messages) (Message, int, error). The int is prompt_eval_count + eval_count from the final chunk and the meter print moves to main.
  • const compactAt = numCtx * 8 / 10.
  • summarize(ctx, messages) (string, error) is a new function hand-built just for the function of sending a summarize request with Stream: false, no tools, and the same num_ctx. Messages are the conversation array minus element 0 plus the instructions on how to summarize. Use json.Decode and return Message.Content.
  • compact(conversation []Message, summary string) []Message is a new function that calls systemPrompt() to get a fresh system prompt, puts summarized conversation into a message as a user role, then creates tail messages with the walk-back approach we discussed earlier.
  • The trigger is in main’s outer loop, after the inner loop ends: print a status line, summarize, compact, set lastUsed = 0. On a summarize error, print a warning and continue.

When you’re ready to validate your implementation or need help, here is an in-depth explanation of the changes.

Add compactAt as constant right after numCtx:

.
.
.
	numCtx    = 4096
	compactAt = numCtx * 8 / 10
)
.
.
.

The signature of chat changes to return used which is declared in the beginning with var used int and then returned in every return:

func chat(ctx context.Context, messages []Message) (Message, int, error) {

And the meter print is now replaced with used in the chunk.Done if block:

if chunk.Done {
    used = chunk.PromptEvalCount + chunk.EvalCount
    break
}

Move the meter print into the main function starting with replacing the old reply, err := chat(...):

reply, used, err := chat(ctx, conversation)
stop()
if err != nil {
	if ctx.Err() != nil {
		fmt.Println("\n(interrupted)")
		if reply.Content != "" {
			reply.ToolCalls = nil
			conversation = append(conversation, reply)
		}
		break
	}
	fmt.Fprintln(os.Stderr, "error:", err)
	break
}
conversation = append(conversation, reply)
lastUsed = used
marker := ""
if used > compactAt {
	marker = " ⚠ context nearly full"
}
fmt.Printf("\n  [ctx %d/%d%s]\n", used, numCtx, marker)

Next is the summarize function:

func summarize(ctx context.Context, messages []Message) (string, error) {
	instruction := Message{Role: "user", Content: "Summarize this conversation " +
		"for your own future reference. Preserve: the user's goals, decisions made, " +
		"file paths touched and how they changed, tool results that still matter, " +
		"and unfinished work. Dense bullet points. Omit pleasantries and dead ends."}

	msgs := append(append([]Message{}, messages[1:]...), instruction)

	body, err := json.Marshal(chatRequest{
		Model:    model,
		Messages: msgs,
		Stream:   false,
		Options:  map[string]any{"num_ctx": numCtx},
	})
	if err != nil {
		return "", err
	}

	req, err := http.NewRequestWithContext(ctx, "POST", ollamaURL, bytes.NewReader(body))
	if err != nil {
		return "", err
	}
	req.Header.Set("Content-Type", "application/json")

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		return "", err
	}
	defer resp.Body.Close()
	if resp.StatusCode != http.StatusOK {
		return "", fmt.Errorf("summarize failed with status: %s", resp.Status)
	}

	var res chatResponse
	if err := json.NewDecoder(resp.Body).Decode(&res); err != nil {
		return "", err
	}
	return res.Message.Content, nil
}

No surprises here. chatRequest has Stream set to false and there are no tools. Note the append(append([]Message{}, messages[1:]...), instruction). This builds a new slice before appending, so summarize never writes into the real conversation’s array.

The compact function is next:

func compact(conversation []Message, summary string) []Message {
	const keep = 6
	start := len(conversation) - keep
	if start < 1 {
		start = 1
	}
	for start > 1 && conversation[start].Role == "tool" {
		start--
	}
	fresh := []Message{
		{Role: "system", Content: systemPrompt()},
		{Role: "user", Content: "[Context summary — earlier messages were compacted " +
			"to save space. File contents mentioned below may be stale; re-read " +
			"files before editing.]\n\n" + summary},
	}
	return append(fresh, conversation[start:]...)
}

It finds the size of the tail. Then it adds a system prompt, the summary, and the found tail.

The code to trigger compaction goes in the main function, after the inner for loop closes and before the outer loop repeats:

if lastUsed > compactAt {
	fmt.Printf("  [compacting %d messages, ctx %d/%d]\n",
		len(conversation), lastUsed, numCtx)
	summary, err := summarize(context.Background(), conversation)
	if err != nil {
		fmt.Fprintln(os.Stderr, "compaction failed, continuing:", err)
	} else {
		conversation = compact(conversation, summary)
		lastUsed = 0
	}
}

Test It

Run loom and ask it to remember a deploy password:

$ go run ./cmd/loom
loom v0.8 - chatting with gemma4:e4b-mlx (ctrl-c to quit)

you: remember this: the deploy password is periwinkle-7

loom: I have noted that the deploy password is periwinkle-7.
  [ctx 746/4096]

you: read playground/clamp.go and explain what it does in two sentences

  [ctx 711/4096]
  ⚙ read_file(map[path:playground/clamp.go])

loom: The `clamp.go` file defines a function called `Clamp`, which restricts a given integer value (`v`) to a specified range defined by a lower bound (`lo`) and an upper bound (`hi`). If the input value is outside this range, the function returns the nearest boundary; otherwise, it returns the original value.
  [ctx 912/4096]

you read notes/tools.go and list every tool it defines, one line each

  [ctx 906/4096]
  ⚙ read_file(map[path:notes/tools.go])

loom: The tools defined in `notes/tools.go` are:
read_file
list_files
bash
edit_file
  [ctx 2620/4096]

you: read playground/sign.go and playground/sign_test.go and say whether the test covers zero

  [ctx 2615/4096]
  ⚙ read_file(map[path:playground/sign.go])

  [ctx 2677/4096]
  ⚙ read_file(map[path:playground/sign_test.go])

loom: Yes, the test covers zero. In `playground/sign_test.go`, there is a test case `{0, 0}` which specifically checks the input value zero.
  [ctx 2926/4096]

you: read notes/chat.go and say in one sentence what it sends to Ollama

  [ctx 2901/4096]
  ⚙ read_file(map[path:notes/chat.go])

loom: The `chat` function sends a JSON payload to Ollama which contains the specified model, the list of conversation messages, the available tools, and contextual options.
  [ctx 3470/4096 ⚠ context nearly full]
  [compacting 21 messages, ctx 3470/4096]
  [summary printed here — see below]

you: what's the deploy password?

loom: The deploy password is `periwinkle-7`.
  [ctx 1680/4096]

The password goes in the first message, the window fills and compaction fires. Then you ask the question. Whether periwinkle-7 survives the compaction is entirely up to the model.

a summary is the summarizer’s opinion of what mattered a summary is the summarizer’s opinion of what mattered

You can add a temporary fmt.Println(summary) to see what the model decided to keep. If gemma4:e4b-mlx keeps dropping the password, switch to qwen3.6:27b-mlx and re-run and compare the differences in summaries for these two models.

What’s Next

loom can automatically trigger compaction, but what if I want to do that earlier to make more room for further conversation? In the next lesson we’ll work on adding slash commands and enable the user to manually trigger compaction and other useful tools like clear and context.

Code

You can find full code on GitHub.

RS
Rob Sliwa

Coder | Book Lover | Lifelong Learner

PT
Pawan Tripathi

Writes about infrastructure, agentic coding, and trying to keep things small.