posts / go

Lesson 5: The System Prompt

The System Prompt

loom sends its messages to the model and gathers responses. However, any self-respecting agent also sends a system prompt which is a message with "role": "system" that contains a set of standing orders for the model. While user prompts tell the model what to do, a system prompt tells the model how to do it. It defines the model’s persona, formatting rules, safety guardrails and overall capabilities.

What do real coding agents put in there? Three kinds of content:

Type What it contains Why the model can’t do without it
Identity “You are loom, a coding agent. You complete tasks by calling tools.” Frames every decision that follows. Without it, the model defaults to chat assistant behavior describing changes rather than making them.
Environment Working directory, OS/arch, today’s date The model knows only what’s in its context. It cannot see your clock or your shell.
Working rules Read before editing. Verify after changing. Don’t repeat a failed call unchanged. This fills the gap between having tools and using them well.

Before we start working on adding a system prompt, let’s establish the baseline behavior of loom without it. This way we’ll be able to observe the difference it makes.

Run loom with go run ./cmd/loom and ask it: “Add a playground/sign.go with a Sign function returning -1, 0, 1.”, “What directory are you working in?” and “what’s today’s date?”. Write down the responses, we’ll compare them to the ones generated after adding a system prompt.

❯ go run ./cmd/loom
loom v0.4 — chatting with gemma4:e4b-mlx (ctrl-c to quit)

you: Add a playground/sign.go with a Sign function returning -1, 0, 1.                        
  ⚙ bash(map[command:mkdir -p playground])
  ⚙ edit_file(map[new_string:package playground

// Sign returns -1, 0, or 1 based on the value of x.
// -1 if x is negative, 0 if x is zero, and 1 if x is positive.
func Sign(x int) int {
        if x < 0 {
                return -1
        } else if x > 0 {
                return 1
        }
        return 0
} old_string: path:playground/sign.go])

loom: The file `playground/sign.go` has been created successfully.

you: What directory are you working in?
  ⚙ bash(map[command:pwd])

loom: I am currently working in the directory `/Users/robertsliwa/code/robjsliwa/coding-agent-articles/lesson-0005/loom`.

you: what's today's date?

loom: I do not have a tool that provides me with the current date.

you: 

Notice your model may respond differently.

Step 1: Your Turn: give loom standing orders

Here is what you need to implement:

  • Create a systemPrompt() string function that composes entire prompt. Pass in some values about the environment so loom knows time, OS, and folder it operates in. Environment values must be looked up: working directory with os.Getwd(), platform with runtime.GOOS / runtime.GOARCH, date with time.Now().Format("2006-01-02").
  • Seed the conversation with it: conversation starts as []Message{{Role: "system", Content: systemPrompt()}}. It needs to be at index 0. System prompt will always be at index 0.
  • Write your rules block. It should cover at least:
    • Read a file before editing it, and quote old_string exactly.
    • After any change, verify with bash (build or test) and never claim a success you haven’t seen.
    • On a tool error, adjust and don’t repeat the identical call.
    • Prefer small edits over rewriting files.
    • Finish with a brief summary of what changed and how it was verified.
    • Keep it under ~250 tokens. Every token contributes to filling in the context window.

When you’re ready to validate your implementation or need help, here is the finished code:

func systemPrompt() string {
	cwd, _ := os.Getwd()
	return fmt.Sprintf(`You are orb, a coding agent. You complete tasks by calling tools,
not by describing what could be done.

Environment:
- Working directory: %s
- Platform: %s/%s
- Today's date: %s

Rules:
- Before editing a file, read it first. Quote old_string exactly, including whitespace.
- After any code change, verify it: build or run tests with bash. Never claim a success you have not seen.
- If a tool returns an error, read it and change your approach; never repeat the same call unchanged.
- Prefer small, targeted edits over rewriting whole files.
- When done, summarize what you changed and how you verified it, in a sentence or two.`,
		cwd, runtime.GOOS, runtime.GOARCH, time.Now().Format("2006-01-02"))
}

And in main, replace var conversation []Message with:

conversation := []Message{{Role: "system", Content: systemPrompt()}}

Step 2: Verify it

Now run loom and type the same commands as before:

❯ go run ./cmd/loom
loom v0.4 — chatting with gemma4:e4b-mlx (ctrl-c to quit)

you: Add a playground/sign.go with a Sign function returning -1, 0, 1.
  ⚙ list_files(map[path:/Users/robertsliwa/code/robjsliwa/coding-agent-articles/lesson-0005/loom])
  ⚙ edit_file(map[new_string:package playground

// Sign returns the sign of x: -1 if x < 0, 0 if x == 0, and 1 if x > 0.
func Sign(x float64) int {
        if x < 0 {
                return -1
        } else if x > 0 {
                return 1
        }
        return 0
}
 old_string: path:playground/sign.go])
  ⚙ list_files(map[path:playground])
  ⚙ bash(map[command:go build ./...])

loom: I added `playground/sign.go` containing the `Sign` function and verified the addition by running `go build ./...`.

you: What directory are you working in?

loom: I am working in the `/Users/robertsliwa/code/robjsliwa/coding-agent-articles/lesson-0005/loom` directory.

you: what's today's date?

loom: Today's date is 2026-08-25.

you: 

Notice that loom now follows the system prompt instructions and verifies the code it added. It also answers directory and time immediately based on the knowledge provided in the system prompt rather than running the bash tool or saying I don’t know.

Context Length

We added a system prompt and we are happily stuffing prompts and responses into the conversation slice, but how many can we actually put in there?

In order to answer this, we need to understand how many tokens the model has access to in memory. Context length is what defines this. Ollama sets default values for the models based on their VRAM size requirement.

You don’t have to accept the defaults and you can increase or decrease the size. Setting a larger context length will increase memory usage, so be aware of this when setting the context length on your machine. On macOS with 32GB I could run both qwen3.6:27b-mlx and gemma4:e4b-mlx with 128k context length.

There are two ways to control the context length. Ollama allows you to overwrite its defaults either from the Settings dialog UI or from the command line: OLLAMA_CONTEXT_LENGTH=64000 ollama serve. You can also send the context length as part of the request via the API, and this is the approach we’ll use.

Note, while the API allows for requesting the context length, Ollama will only respect it on s fresh model load. That is, if the model was already loaded with a 64k context length and the API requests 128k, Ollama will ignore it and use what is in memory. It will start using 128k only when the model is reloaded. Ollama keeps the model in memory for a default of 5 minutes. So if you don’t talk to it for 5 or more minutes, the next request will use the requested context length. This is not a problem for us as we’ll always load the model with loom’s requests, but just FYI.

So what happens when the conversation array no longer fits in the context length? Well, part of the conversation is lost because the model never sees it. The way we can solve this problem is through compaction, which we’ll work on in lesson 8. However, before we do that, we need a way to know what the context length we work with is and how much of it was consumed. For the remainder of this lesson, we’ll work on showing these statistics the user and warning them when context is 80% full. For now, we’ll show the warning but we’ll take no action. We’ll take care of compaction in future lesson.

Step 1: Your Turn: write context meter

Here’s what you need to do:

  • chatResponse adds two new fields PromptEvalCount and EvalCount that correspond to Ollama’s API response fields prompt_eval_count / eval_count.
  • chatRequest adds an Options map[string]any field (JSON tag options, with omitempty), and chat sends {"num_ctx": numCtx} where numCtx is a const we set to 16384 for now. You can adjust it based on your hardware and needs.
  • chat needs to surface the counts to the loop.
  • After each reply, print a meter line like [ctx 2105/16384] using PromptEvalCount + EvalCount. When usage crosses 80% of numCtx, print warning.

When you’re ready to validate your implementation or need help, here is the finished code:

Add a numCtx constant and update structures with new fields:

const numCtx = 16384

type chatRequest struct {
	Model    string         `json:"model"`
	Messages []Message      `json:"messages"`
	Stream   bool           `json:"stream"`
	Tools    []Tool         `json:"tools,omitempty"`
	Options  map[string]any `json:"options,omitempty"`
}

type chatResponse struct {
	Message         Message `json:"message"`
	PromptEvalCount int     `json:"prompt_eval_count"`
	EvalCount       int     `json:"eval_count"`
}

Update chat to return (Message, int, error) where int is the out.PromptEvalCount + out.EvalCount:

func chat(messages []Message) (Message, int, error) {
	var used int
	body, err := json.Marshal(chatRequest{
...
	used = out.PromptEvalCount + out.EvalCount

	return out.Message, used, nil
}

Also make sure to add used to all error returns.

In the main function print the statistics:

...
			conversation = append(conversation, reply)

			marker := ""
			if used > numCtx*8/10 {
				marker = " ⚠ context nearly full"
			}
			fmt.Printf("  [ctx %d/%d%s]\n", used, numCtx, marker)
...

Step 2: Test it

Run loom with numCtx set to 16384. You can verify that Ollama accepted your context by running:

❯ ollama ps
NAME              ID              SIZE      PROCESSOR    CONTEXT    UNTIL              
gemma4:e4b-mlx    aa6f2058d5dc    8.4 GB    100% GPU     16384      4 minutes from now

Notice that Ollama accepts num_ctx and seemingly honors it. Why only seemingly? Because if you keep chatting and go over the context length, Ollama, instead of forgetting older messages, keeps accepting them. The proof is in the value returned by used = out.PromptEvalCount + out.EvalCount.

So what is happening?

It trusts our payload where we set numCtx to 16384 to allocate a memory baseline. But because loom is not truncating the conversation, it eventually sends more than the requested context length. Ollama, instead of truncating it, tries to accommodate it by allocating more memory. Keep chatting and occasionally running ollama ps and you will see how the SIZE grows. This is happening because loom is not keeping its conversation within the context length.

Another thing you might have noticed is that your token meter will sometimes go down between turns. Obviously this is not the conversation shrinking, this is caused by KV cache. Ollama reuses the computed state for the unchanged prefix of the prompt, and prompt_eval_count reports only the tokens it actually evaluated during this call. So the meter right now is slightly inaccurate as far as measuring the total size of messages.

We’ll address both of these issues in lesson 8 as we work on the message compaction feature.

What’s Next

The next lesson is about streaming and UX. Right now we have to wait for responses without any indication of what is happening. We’ll fix this by setting stream: true. We’ll also add a handler for CTRL-C so loom will be more like a real coding agent and allow user to pause execution and try something else.

Code

You can find full code on GitHub.

RS
Rob Sliwa

Coder | Book Lover | Lifelong Learner

PT
Pawan Tripathi

Writes about infrastructure, agentic coding, and trying to keep things small.