What is a Coding Agent?
Before we build a coding agent, we should define what it actually is.
While a standard chat model expects you to type in your text or code and then ask a question, a coding agent interacts with your environment directly. At the basic level it is an LLM with hands that can interact with your file system, see files, read them and even edit or create new ones. It can also write and execute shell scripts or run a test suite. And most importantly it can see the result of these actions, success or failure, and figure out the next steps.
The basic architecture is simple: a loop that repeatedly sends a conversation to a model, acts on the response, appends the result. What makes an agent feel intelligent is that loop plus a handful of tools.
The Agent Loop
This lesson starts development of loom, our coding agent, and is all about building this basic loop. Every lesson that follows builds on this foundation.
A Conversation is just an Array
When talking with a model, you POST a list of messages. Each message has a role (who said it) and content (what they said). Here is the sample message loom is going to send to Ollama’s /api/chat endpoint:
{
"model": "gemma4:e4b-mlx",
"messages": [
{ "role": "user", "content": "My name is Rob." },
{ "role": "assistant", "content": "Nice to meet you, Rob!" },
{ "role": "user", "content": "What's my name?" }
],
"stream": false
}
And the shape that comes back:
{
"message": { "role": "assistant", "content": "Your name is Rob." },
"done": true
}
There are four roles that exist:
| Role | Author | When |
|---|---|---|
user |
The human | This Lesson |
assistant |
The model | This Lesson |
system |
You, the author of harness | Lesson 5 |
tool |
Your tool code | Lesson 2 |
The model is stateless. Every time you call it, it has total amnesia and you have to send it the entire conversation. It knows only the tokens in that request and it is up to the agent to have the “memory.” This memory of a chat is an array of messages that the agent keeps and resends, in full, every single turn. The long conversation can fill up the context window and messages can get lost to the LLM when they fall outside that window. loom will need to implement context management to handle this.
Build loom v0.1
loom v0.1 is a terminal REPL which:
- reads a line
- appends it to the conversation
- sends the whole conversation to Ollama
- prints the reply
- rinse and repeat
Step 1: Pull the course model
As explained in the introduction we are going to use two models, so let’s download them:
ollama pull qwen3.6:27b
ollama pull gemma4:e4b
or if you are on macOS and Apple silicon use -mlx versions of these:
ollama pull qwen3.6:27b-mlx
ollama pull gemma4:e4b-mlx
Step 2: Initialize the module
Initialize the module in the repo root:
go mod init example.com/loom
Step 3: Define the Data Structures
Create main.go and define structs that represent the message as well as the request and response. Notice they look just like the JSON in the conversation array we just talked about.
package main
import (
"bufio"
"fmt"
"os"
)
const (
ollamaURL = "http://localhost:11434/api/chat"
// model = "qwen3.6:27b-mlx"
model = "gemma4:e4b-mlx"
)
type Message struct {
Role string `json:"role"`
Content string `json:"content"`
}
type chatRequest struct {
Model string `json:"model"`
Messages []Message `json:"messages"`
Stream bool `json:"stream"`
}
type chatResponse struct {
Message Message `json:"message"`
}
ollamaURL points to the local Ollama server. Change that if yours is running on a different machine. Also if you prefer to use qwen3.6 or your other favorite model, just update the model variable.
The main function is just a simple loop that takes user input and packages it as a Message that then is appended to the conversation slice. The slice is the entire conversation that gets sent to the model on every turn using the chat function. The chat function returns the reply which is also a Message type. Response gets appended to the conversation and then displayed back to the user. The loop moves to the next turn.
func main() {
fmt.Printf("loom v0.1 โ chatting with %s (ctrl-c to quit)\n", model)
scanner := bufio.NewScanner(os.Stdin)
var conversation []Message
for {
fmt.Print("\nyou: ")
if !scanner.Scan() {
break
}
conversation = append(conversation,
Message{Role: "user", Content: scanner.Text()})
reply, err := chat(conversation)
if err != nil {
fmt.Fprintln(os.Stderr, "error:", err)
continue
}
conversation = append(conversation, reply)
fmt.Println("\nloom:", reply.Content)
}
}
Step 4: Your turn
Now your turn. Try to implement the chat function:
func chat(messages []Message) (Message, error) {
// YOUR TURN:
// 1. json.Marshal a chatRequest (Stream must be false).
// 2. http.Post it to ollamaURL as "application/json".
// Check resp.StatusCode. Anything but 200 is an error.
// 3. Decode the body into chatResponse. Return its Message.
return Message{}, fmt.Errorf("not implemented")
}
When you’re ready to validate your implementation or need help, here is the finished implementation.
func chat(messages []Message) (Message, error) {
body, err := json.Marshal(chatRequest{
Model: model,
Messages: messages,
Stream: false,
})
if err != nil {
return Message{}, err
}
resp, err := http.Post(
ollamaURL,
"application/json",
bytes.NewReader(body),
)
if err != nil {
return Message{}, err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return Message{}, fmt.Errorf("ollama returned %s", resp.Status)
}
var out chatResponse
if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
return Message{}, err
}
return out.Message, nil
}
Step 5: Test it
Run loom to verify it works:
โฏ go run ./cmd/loom/
loom v0.1 โ chatting with gemma4:e4b-mlx (ctrl-c to quit)
you: My name is Rob and my favorite language is Go.
loom: That's interesting! Go is a fascinating language.
What would you like to talk about regarding Go, or is there anything else I can help you with? ๐
you: What's my name and what language do I like?
loom: Your name is **Rob**, and the language you like is **Go**.
you:
Great! It answered correctly, but why? It’s not because the model remembered but rather because loom resent the whole slice. Let’s prove this by breaking the code a bit. Change this line to send only the latest message instead of the whole conversation:
reply, err := chat(conversation[len(conversation)-1:])
Run the same conversation as before.
โฏ go run ./cmd/loom/
loom v0.1 โ chatting with gemma4:e4b-mlx (ctrl-c to quit)
you: My name is Rob and my favorite language is Go.
loom: It's nice to meet you, Rob!
Go is an incredibly fascinating and deep game. Do you enjoy strategic play, or do you prefer the more casual, social aspect of it?
you: What's my name and what language do I like?
loom: I do not have access to personal information about you, so I do not know your name or your language preferences.
If you tell me your name or the languages you like, I would be happy to know!
you:
The model has no idea who you are. All of the memory is in the slice. This is the key takeaway of this lesson.
What’s Next
In the next lesson we’ll work on the tool calling handshake. We’ll teach the model how to answer with tool calls instead of just text, and loom will get its first tool: read_file.
Code
You can find full code on GitHub.