Projects

Halden

A local operations assistant that brings persistent memory, voice, coding agents, and engineering tools into one workspace.

  • Software
  • Hardware

Halden connects my day-to-day agent work with a workspace for engineering projects. It combines conversational memory and voice with coding-agent orchestration, project tools, and an Engineering Studio for working through CAD, electronics, and bench tasks.

When
July 2026
Where
Independent project
Role
Designer and developer
Runs on
My own GPU machine, with local models through Ollama and vLLM
Memory
SQLite with keyword and vector search under a token budget
Delegated work
Supervised coding-agent runs, checked by the server's own builds and tests
Studio
Schematic, PCB, and CAD outputs, each checked before review
Tools
Python, FastAPI, SQLite, React, TypeScript, Tailwind, Ollama, vLLM, faster-whisper, Docker, KiCad

Persistent context: memory is kept in SQLite as distilled facts. Retrieval combines keyword and vector search. Each turn gets a context assembled under a token budget from core memory, a rolling summary, and the retrieved memories, and that context feeds the conversation loop.

Conversation: typed chat and voice reach the same conversation loop. Voice goes through speech-to-text, and the reply is streamed to text-to-speech. Barge-in, speaking over the reply, cancels generation.

Policy gate: between the conversation loop and every tool. Each tool action has a risk tier, the session has a permission mode, and physical actions need confirmation. From the gate, work goes to two lanes.

Delegated work: heavy coding tasks go to an external coding agent that runs supervised and can be paused or stopped. The server runs build, test, and lint checks itself and keeps the output as the record. Light tasks go to local workers with no tool access.

Engineering workspace: schematic code is compiled in a no-network sandbox and checked with ERC and SPICE. A first-pass PCB goes through placement, autorouting, and DRC, then is flagged for human review. CAD part code is measured against a written spec. A bench test plan and script are generated. Hardware flashing is only for registered boards, with matching firmware and checksum and a fresh confirmation for each action.

How Halden is organized. Context feeds one conversation loop, and every tool call passes the policy gate before it reaches delegated work or the engineering workspace.

Problem

Agent work on an engineering project involves more than one conversation. The assistant has to remember what the project is, hand work to other tools without losing track of it, and check results before they reach real hardware. Halden treats conversation, long-term context, delegated work, and engineering tasks as separate parts with defined handoffs.

My role

My role covered design and implementation of the whole system: the FastAPI backend, React and TypeScript dashboard, memory and retrieval layer, voice pipeline, delegation and verification flow, and Engineering Studio.

How it is organized

Halden keeps four concerns apart.

  • Conversation. One agent loop serves typed chat and voice. Voice is another front end to the same loop: speech-to-text with faster-whisper, the agent’s reply streamed sentence by sentence to a local text-to-speech model, and barge-in so I can interrupt and the server cancels generation.
  • Persistent context. Memory lives in SQLite as distilled facts rather than raw transcripts. Each turn gets a context assembled under a fixed token budget from core memory, a rolling summary, and retrieved memories. Retrieval combines full-text keyword search with vector search and merges the two rankings. If the vector extension is missing, it falls back to keyword search instead of failing.
  • Delegated work. Heavy coding tasks go to an external coding agent that Halden starts as a supervised process and can pause or stop. Light tasks go to local workers that have no tool access. The server runs the task’s build, test, and lint commands itself and keeps the output, so a run is judged by its checks rather than by the agent’s own summary. A run with no checks is reported as having nothing to verify.
  • Engineering workspace. The Studio handles circuits, boards, and mechanical parts (below).

Between the agent loop and any tool sits a policy layer. Every tool action is scored for risk, the session has a permission mode, and anything that touches the physical world asks for confirmation.

Engineering Studio

  • Schematics. A model writes circuit code in SKiDL or atopile. It compiles inside a Docker sandbox with no network access and capped resources, and the Studio stops if the sandbox is unavailable. The netlist renders to an SVG schematic, and each revision is graded pass, warn, or fail from electrical rules checks and ngspice simulation.
  • PCBs. Placement, KiCad board synthesis, autorouting with FreeRouting, and design rules checks produce a first-pass layout and fabrication files. Every output carries a note that the board needs human routing review before it is ordered.
  • CAD. A model writes build123d part code against a written spec. The sandbox exports the part, measures its bounding box, volume, and holes, and a grader compares those measurements to the spec. The loop repeats until the checks pass or it hands the problem back to me.
  • Bench handoff. Simulations that pass become steps in a bench test plan, with a PyVISA script generated for a lab computer.

The Studio builds circuits from known reference blocks and produces first-pass boards and simple parametric parts for review.

Hardware in the loop

Halden can only flash boards I have registered, and registering a board is itself a confirmed action. A firmware image has to match the board’s target, and its checksum is recorded and checked again by a small agent that runs next to the hardware and exposes only a few fixed commands. Each flash, power cycle, or test needs a fresh confirmation tied to one board and one image.

Decisions

  • Local first. Models, memory, speech, and retrieval run on my own machine. Only delegated coding runs leave it.
  • The server checks the work. A coding agent’s report is treated as a claim. The server runs the checks and keeps their output as the record of what happened.
  • Fail closed. A missing sandbox, missing checks, or an unregistered board stops the action instead of falling back to something less safe.