posts / go

Lesson 6: Streaming and UX

Streaming and UX

loom is shaping into a very capable agent. It also is really good at creating anticipation. When you type in the command and press enter, nothing happens for a couple seconds and then boom, the answer just appears and scrolls through the terminal. This isn’t the best experience. In this lesson we’ll address it by turning on streaming.

Streaming

Since lesson 1, we’ve had "stream": false. loom sent a request, and waited for however long it took to receive the response. From a user’s perspective, watching tokens arrive as the model thinks makes a difference between an agent that is working vs. an agent that appears hung.

Before we go running to code, let’s run some curl from the terminal to observe what happens when we send stream true to Ollama. We need to understand what happens on the wire before we implement the changes.

curl -N http://localhost:11434/api/chat -d '{
  "model": "gemma4:e4b-mlx",
  "messages": [{"role": "user", "content": "Count to five, one number per line."}],
  "stream": true
}'

What comes back is not one JSON object. It’s a sequence of them, newline delimited (NDJSON), each arriving the moment the model produces it:

{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:13.992715Z","message":{"role":"assistant","content":"1"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.006998Z","message":{"role":"assistant","content":"\n"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.022367Z","message":{"role":"assistant","content":"2"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.027773Z","message":{"role":"assistant","content":"\n"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.048087Z","message":{"role":"assistant","content":"3"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.048096Z","message":{"role":"assistant","content":"\n"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.097257Z","message":{"role":"assistant","content":"4"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.097264Z","message":{"role":"assistant","content":"\n"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.115006Z","message":{"role":"assistant","content":"5"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-27T20:19:14.115011Z","message":{"role":"assistant","content":""},"done":true,"done_reason":"stop","total_duration":13566344083,"load_duration":12678591000,"prompt_eval_count":25,"prompt_eval_duration":598520500,"eval_count":9,"eval_duration":288061166}

Every chunk carries a message whose content is a fragment of a message, often a single token. Chunks carry "done": false and are sent by the model as it produces them. To indicate the end of the message, chunk carries "done": true as well as prompt_eval_count and eval_count.

Tool calls stream too and their chunks can carry message.tool_calls, usually with empty content.

Step 1: Your turn: the streaming chat

Here is what you need to do to support streaming mode:

  • Add Done bool (tag done) to chatResponse and flip the request to Stream: true.
  • Replace the single response Decode with a loop. dec := json.NewDecoder(resp.Body) is perfect for decoding chunks, as each dec.Decode(&chunk) call consumes exactly one JSON object from the stream and blocks until the next one arrives. Loop until a chunk has Done == true but make sure to also handle io.EOF as end-of-stream too.
  • Do three things per chunk:
    • Print the content fragment immediately.
    • Accumulate it into a strings.Builder.
    • Collect append(assembled.ToolCalls, chunk.Message.ToolCalls...).
  • Print loom on the first non empty content fragment and not before the loop. A turn that’s all tool calls should show ⚙ lines and not an orphaned prefix.
  • On the final chunk, read the metrics and print your meter.
  • chat returns the assembled Message with role assistant.
  • In main, remove fmt.Println("\nloom:", reply.Content). By the time chat returns, the user has already read the answer.

When you’re ready to validate your implementation or need help, here is the finished code:

.
.
.
type chatResponse struct {
	Message         Message `json:"message"`
	PromptEvalCount int     `json:"prompt_eval_count"`
	EvalCount       int     `json:"eval_count"`
	Done            bool    `json:"done"`
}
.
.
.
func chat(messages []Message) (Message, error) {
	var used int
	body, err := json.Marshal(chatRequest{
		Model:    model,
		Messages: messages,
		Tools: func() []Tool {
			tools := make([]Tool, len(registry))
			for i, def := range registry {
				tools[i] = def.Tool
			}
			return tools
		}(),
		Stream:  true,
		Options: map[string]any{"num_ctx": numCtx},
	})
	if err != nil {
		return Message{}, err
	}

	resp, err := http.Post(
		ollamaURL,
		"application/json",
		bytes.NewReader(body),
	)
	if err != nil {
		return Message{}, err
	}
	defer resp.Body.Close()

	if resp.StatusCode != http.StatusOK {
		return Message{}, fmt.Errorf("ollama returned %s", resp.Status)
	}

	dec := json.NewDecoder(resp.Body)
	var (
		assembled = Message{Role: "assistant"}
		content   strings.Builder
		prefixed  bool
	)

	for {
		var chunk chatResponse
		if err := dec.Decode(&chunk); err != nil {
			if errors.Is(err, io.EOF) {
				break
			}
			assembled.Content = content.String()
			return assembled, err // return partial result with error
		}

		if chunk.Message.Content != "" {
			if !prefixed {
				fmt.Print("\nloom: ")
				prefixed = true
			}
			fmt.Print(chunk.Message.Content)
			content.WriteString(chunk.Message.Content)
		}

		assembled.ToolCalls = append(assembled.ToolCalls, chunk.Message.ToolCalls...)
		if chunk.Done {
			used = chunk.PromptEvalCount + chunk.EvalCount
			marker := ""
			if used > numCtx*8/10 {
				marker = " ⚠ context nearly full"
			}
			fmt.Printf("\n  [ctx %d/%d%s]\n", used, numCtx, marker)
			break
		}
	}

	assembled.Content = content.String()
	return assembled, nil
}
.
.
.
func main() {
	fmt.Printf("loom v0.6 — chatting with %s (ctrl-c to quit)\n", model)
	scanner := bufio.NewScanner(os.Stdin)
	conversation := []Message{{Role: "system", Content: systemPrompt()}}

	for {
		fmt.Print("\nyou: ")
		if !scanner.Scan() {
			break
		}
		conversation = append(conversation,
			Message{Role: "user", Content: scanner.Text()})

		for {
			reply, err := chat(conversation)
			if err != nil {
				fmt.Fprintln(os.Stderr, "error:", err)
				break
			}
			conversation = append(conversation, reply)

			if len(reply.ToolCalls) == 0 {
				break
			}

			for _, tc := range reply.ToolCalls {
				fmt.Printf("  ⚙ %s(%v)\n", tc.Function.Name, tc.Function.Arguments)
				var result string
				var toolDef ToolDef
				var toolFound bool
				for _, def := range registry {
					if def.Tool.Function.Name == tc.Function.Name {
						toolDef = def
						toolFound = true
					}
				}

				if toolFound {
					result = toolDef.Run(tc.Function.Arguments)
				} else {
					result = "error: unknown tool " + tc.Function.Name
				}

				conversation = append(conversation, Message{
					Role:     "tool",
					ToolName: tc.Function.Name,
					Content:  result,
				})
			}
		}
	}

	if err := scanner.Err(); err != nil {
		fmt.Fprintln(os.Stderr, "\nreading standard input:", err)
	}
}

Step 2: Test it

Now build loom and run a prompt:

❯ go run ./cmd/loom
loom v0.6 — chatting with gemma4:e4b-mlx (ctrl-c to quit)

you: Verify that tests in playground folder are passing

  [ctx 656/16384]
  ⚙ bash(map[command:ls -R])

  [ctx 741/16384]
  ⚙ bash(map[command:go test ./playground])

loom: The tests in the `playground` folder are passing, as confirmed by the output: `ok example.com/loom/playground (cached)`.
  [ctx 784/16384]

You should see the answers from loom stream as tokens and start arriving as the model sends them.

This is great! But wait! There is more!

You might have noticed that while the tokens streamed just fine, there were still some “frozen” UI moments between the tool calls and messages. The reason for this is that our models are thinking models and that is exactly what they were doing: thinking.

Let’s execute a curl command as we did before but this time with a harder request:

❯ curl -N http://localhost:11434/api/chat -d '{                                                                                                    
  "model": "gemma4:e4b-mlx",
  "messages": [{"role": "user", "content": "Count how many files are in playground folder."}],
  "stream": true
}'
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:20.120147Z","message":{"role":"assistant","content":"","thinking":"Thinking"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:20.138923Z","message":{"role":"assistant","content":"","thinking":" Process"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:20.138929Z","message":{"role":"assistant","content":"","thinking":":"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:20.157094Z","message":{"role":"assistant","content":"","thinking":"\n\n1"},"done":false}
.
.
.
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:23.314957Z","message":{"role":"assistant","content":" me"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:23.314965Z","message":{"role":"assistant","content":" the"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:23.314985Z","message":{"role":"assistant","content":" output"},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:23.336702Z","message":{"role":"assistant","content":"."},"done":false}
{"model":"gemma4:e4b-mlx","created_at":"2026-08-31T16:33:23.336709Z","message":{"role":"assistant","content":""},"done":true,"done_reason":"stop","total_duration":10862546958,"load_duration":7444165208,"prompt_eval_count":25,"prompt_eval_duration":136689958,"eval_count":328,"eval_duration":3281176208}

Notice that some messages contain {"role":"assistant","content":"","thinking":" Process"}. Content is empty but thinking has tokens describing the model’s thinking process. Ah, now we know how to update our code to capture the thinking process.

To differentiate thinking tokens from answer tokens, let’s display them with a grey color.

First add Thinking to Message struct:

type Message struct {
	Role      string     `json:"role"`
	Content   string     `json:"content"`
	Thinking  string     `json:"thinking,omitempty"`
	ToolCalls []ToolCall `json:"tool_calls,omitempty"`
	ToolName  string     `json:"tool_name,omitempty"`
}

Then update the chat function just above if chunk.Message.Content != "" {:

.
.
.
		if chunk.Message.Thinking != "" {
			// \033[90m sets the text color to grey
			// \033[0m resets it back to default
			fmt.Printf("\033[90m%s\033[0m", chunk.Message.Thinking)
		}

		if chunk.Message.Content != "" {
.
.
.

We simply check if Message.Thinking is not empty and write it out in grey color.

Try it out again, and this time you should see output stream immediately and see thinking process as well as final answer:

❯ go run ./cmd/loom
loom v0.6 — chatting with gemma4:e4b-mlx (ctrl-c to quit)

you: Verify that tests in playground folder are passing
I need to check the contents of the directory to locate the "playground" folder and then figure out how to run tests within it. I will start by listing the files in the current directory.
  [ctx 627/16384]
  ⚙ list_files(map[])
The user wants to verify that tests in the `playground` folder are passing.
I see files like `playground/abs.go`, `playground/abs_test.go`, `playground/clamp.go`, `playground/clamp_test.go`, `playground/sign.go`.

I should run tests specifically for the `playground` directory. The `go test` command is the appropriate tool. I will run `go test ./playground`.
  [ctx 758/16384]
  ⚙ bash(map[command:go test ./playground])
The user wants to verify that tests in the `playground` folder are passing.
I have already used `go test ./playground` and the output was `ok example.com/loom/playground (cached)`, which indicates that all tests passed.
I will report this finding to the user.
loom: Tests in the `playground` folder are passing.
  [ctx 761/16384]

Ctrl-C: interruption feature

Sometimes you may a ask coding agent to do something and watch it work only to realize you asked it the wrong thing or it is doing the wrong thing. Instead of having to suffer agony watching tokens getting wasted while waiting for the agent to finish and then start over, it would be very nice to be able to interrupt the current work and give the agent new instructions.

Most coding agents allow this via either the escape key or Ctrl-C. In this section we’ll teach loom to handle Ctrl-C and instead of immediately dying it will present a prompt so you can steer it in a different direction.

Step 3: Your turn: interruption algorithm

Go has signal.NotifyContext which gives us a context that cancels on SIGINT. We can use http.NewRequestWithContext to wire this cancellation into the request. When the context gets a signal, the connection will drop and stop generating tokens. Your GPU will thank you and the Decode loop already handles errors by returning partial answers. I love it when the plan comes together!

The changes:

  • Update chat to take a ctx context.Context as a first parameter. Swap http.Post for http.NewRequestWithContext + http.Default.Client.Do. You will also need to set the header yourself: Content-Type: application/json, Post was doing that for you.
  • In main’s inner loop, start each iteration with the sequence ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt), chat(ctx, conversation), stop(). Notice that this handler exists only while a request is in flight, so Ctrl-C at the prompt still quits loom.
  • Handle ctx.Err() != nil. Print (interrupted) on its own line, append the partial message but strip its ToolCalls. Promised but not executed tool calls would confuse the next turn. Finally, break to the prompt. Any other error, report and break as before.

When you’re ready to validate your implementation or need help, here is the finished code:

func chat(ctx context.Context, messages []Message) (Message, error) {
	chatReq := chatRequest{
		Model:    model,
		Messages: messages,
		Tools: func() []Tool {
			tools := make([]Tool, len(registry))
			for i, def := range registry {
				tools[i] = def.Tool
			}
			return tools
		}(),
		Stream:  true,
		Options: map[string]any{"num_ctx": numCtx},
	}

	body, err := json.Marshal(chatReq)
	if err != nil {
		return Message{}, err
	}

	req, err := http.NewRequestWithContext(
		ctx,
		"POST",
		ollamaURL,
		bytes.NewReader(body),
	)
	if err != nil {
		return Message{}, err
	}
	req.Header.Set("Content-Type", "application/json")

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		return Message{}, err
	}

	if resp.StatusCode != http.StatusOK {
		return Message{}, fmt.Errorf("request failed with status %s", resp.Status)
	}

	dec := json.NewDecoder(resp.Body)
	var (
		assembled = Message{Role: "assistant"}
		content   strings.Builder
		prefixed  bool
	)

	for {
		var chunk chatResponse
		if err := dec.Decode(&chunk); err != nil {
			if errors.Is(err, io.EOF) {
				break
			}
			assembled.Content = content.String()
			return assembled, err // return partial result with error
		}

		if chunk.Message.Thinking != "" {
			// \033[90m sets the text color to grey
			// \033[0m resets it back to default
			fmt.Printf("\033[90m%s\033[0m", chunk.Message.Thinking)
		}

		if chunk.Message.Content != "" {
			if !prefixed {
				fmt.Print("\nloom: ")
				prefixed = true
			}
			fmt.Print(chunk.Message.Content)
			content.WriteString(chunk.Message.Content)
		}

		assembled.ToolCalls = append(assembled.ToolCalls, chunk.Message.ToolCalls...)
		if chunk.Done {
			used := chunk.PromptEvalCount + chunk.EvalCount
			marker := ""
			if used > numCtx*8/10 {
				marker = " ⚠ context nearly full"
			}
			fmt.Printf("\n  [ctx %d/%d%s]\n", used, numCtx, marker)
			break
		}
	}

	assembled.Content = content.String()
	return assembled, nil
}

func main() {
	fmt.Printf("loom v0.6 — chatting with %s (ctrl-c to quit)\n", model)
	scanner := bufio.NewScanner(os.Stdin)
	conversation := []Message{{Role: "system", Content: systemPrompt()}}

	for {
		fmt.Print("\nyou: ")
		if !scanner.Scan() {
			break
		}
		conversation = append(conversation,
			Message{Role: "user", Content: scanner.Text()})

		for {
			ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt)
			reply, err := chat(ctx, conversation)
			stop()
			if err != nil {
				if ctx.Err() != nil {
					fmt.Println("\n(interrupted)")
					if reply.Content != "" {
						reply.ToolCalls = nil
						conversation = append(conversation, reply)
					}
					break
				}
				fmt.Fprintln(os.Stderr, "error:", err)
				break
			}
			conversation = append(conversation, reply)

			if len(reply.ToolCalls) == 0 {
				break
			}

			for _, tc := range reply.ToolCalls {
				fmt.Printf("  ⚙ %s(%v)\n", tc.Function.Name, tc.Function.Arguments)
				var result string
				var toolDef ToolDef
				var toolFound bool
				for _, def := range registry {
					if def.Tool.Function.Name == tc.Function.Name {
						toolDef = def
						toolFound = true
					}
				}

				if toolFound {
					result = toolDef.Run(tc.Function.Arguments)
				} else {
					result = "error: unknown tool " + tc.Function.Name
				}

				conversation = append(conversation, Message{
					Role:     "tool",
					ToolName: tc.Function.Name,
					Content:  result,
				})
			}
		}
	}

	if err := scanner.Err(); err != nil {
		fmt.Fprintln(os.Stderr, "\nreading standard input:", err)
	}
}

Step 4: Test it

Ask loom to do something long and then interrupt in the mid-stream, confirm that you’re back at the prompt and you saw (interrupted), and then give it a different prompt.

❯ go run ./cmd/loom
loom v0.6 — chatting with gemma4:e4b-mlx (ctrl-c to quit)

you: Tell me what the code in each file in playground folder does.       
I need to list all files in the `playground` folder and then read the content of each file to explain what the code does. I'll start by listing the files.
  [ctx 635/16384]
  ⚙ list_files(map[path:playground])
^[The user wants to know what each file in the `playground` folder does.
I have a list of files:
- `playground/abs.go`
- `playground/abs_test.go`
- `playground/clamp.go`
- `playground/clamp_test.go`
- `playground/sign.go`

I will read each file to understand its content and then summarize its purpose.

Start with `playground/abs.go`.
  [ctx 763/16384]
  ⚙ read_file(map[path:playground/abs.go])
The user wants to know what each file in the `playground` folder^[ does.
I have already listed the files:
- `playground/abs.go`
- `playground/abs_test.go`
- `playground/clamp.go`
- `playground/clamp_test.go`
- `playground/sign.go`

I have read `playground/abs.go`. Now I need to read the rest^[ of the files to describe their content.

Next file: `playground/abs_test.go`
Next file: `playground/clamp.go`
Next file: `playground/clamp_test.go`
Next file: `playground/sign.go`^[
  [ctx 889/16384]
  ⚙ read_file(map[path:playground/abs_test.go])
^[^[The user wants to know what the code in each file in the `playground` folder does. I have listed the files and read the content of `abs.go` and `abs_test.go`. I need to continue reading the remaining files: `playground/clamp.go`, `playground/clamp_test.go`, and `playground/sign.go`. After reading them, I will summarize the purpose of all the files.^[
  [ctx 1027/16384]
  ⚙ read_file(map[path:playground/clamp.go])
The user wants a summary of what the code in each file in the `playground` folder does. I have already read `playground/abs.go`, `playground/abs_test.go`, `playground/clamp.go`. I still need to read `playground/clamp_test.go`, `playground/sign.go`, and^[ then I need to read `playground/sign_test.go` (which I assume exists based on the list, but I need to check the file list again or if I should assume a pair).

Let's re-check the file list:
`playground/abs.^[go`
`playground/abs_test.go`
`playground/clamp.go`
`playground/clamp_^[test.go`
`playground/sign.go`

It seems `sign_test.^[go` might not exist, or I missed it. I should read `playground/clamp_test.go` and `playground/sign.go`.

First, I'll read `playground/clamp_test.go`.
  [ctx 1254/16384]
  ⚙ read_file(map[path:playground/clamp_test.go])
The user wants to know what the code in each file in the `playground` folder does. I have already read all the files in this directory (`abs.go`, `abs_test.go`, `clamp.go`, `clamp_test.go`, `sign.go`).

I need to read `playground/sign.go` before summarizing.

After reading all files, I will provide the summary.
  [ctx 1309/16384]
  ⚙ read_file(map[path:playground/sign.go])
The user wants a summary of what the code in each file within the `playground` folder does.
I have successfully read all files in the `playground` folder.
I will now compile this information into a single response.
- `playground/abs.go`: Defines the `Abs` function, which calculates the absolute value of an integer.
- `playground/abs_test.go`: Contains unit tests for the `Abs` function.
- `^C
(interrupted)

you: just tell me how many files are in playground folder
The user wants to know the total number of files in the `playground` folder.
I previously called `list_files` and got a list:
`playground/abs.go`
`playground/abs_test.go`
`playground/clamp.go`
`playground/clamp_test.go`
`playground/sign.go`

I will count them. There are 5 files. I do not need to call a tool, I can answer based on the previous results.
loom: There are 5 files in the `playground` folder.
  [ctx 1460/16384]

What’s Next

With the great power comes great responsibility. loom is growing in its capabilities but it has also been dangerous. bash tool can execute anything, including rm -rf. In the next lesson we’ll address safety and permissions so that loom asks you for permission before it acts.

Code

You can find full code on GitHub.

RS
Rob Sliwa

Coder | Book Lover | Lifelong Learner

PT
Pawan Tripathi

Writes about infrastructure, agentic coding, and trying to keep things small.