A loop that plans, runs tools and reads the results, and that writes the tools it is missing. It runs in plain OTP processes, with no agent framework.
This page walks through the moving parts, replays a run step by step and lists the limits as they are in the code today. Every number on it is taken from the source.
Each page adds one idea to the previous one. They share the same tools and processes.
The same goal goes to two agents that are identical except for their brain. One uses keyword rules (ClassicPlan), the other a model (AiPlan, mock or Claude). The results are shown side by side.
Describe a tool, and a model writes the Elixir module. You can watch it pass the sanitizer, compile, load under a versioned name and answer a probe call. The registry lists every tool generated so far.
Each conversation is a GenServer. A worker runs the loop, and progress, approvals, the trace and remembered facts reach the LiveView over PubSub as they happen.
Pick a task and step through it. The loop diagram lights up the part that is working. This is a replay: the tool names and texts are examples, but every step is one the code actually takes.
The only part that decides anything is the planner, and it is swappable. Everything around it is plain OTP and Phoenix: the loop, the tools, the processes and the checks.
The model remembers nothing between calls. Before every round, TaskAgent assembles the planner's context afresh: six parts, each from a different place, each living for a different span of time. Everything the agent "knows" has to be in it.
The rules: use existing tools, invent new ones in snake_case, plan again after results, remember lasting facts.
The catalog plus every generated tool in the registry, with name, required parameters and description. A tool written in round 1 is offered in round 2.
The facts from the JSON file, at most 50 of 300 characters each. Read again every round, so a fact remembered just now already counts.
The last 4 turns, each reduced to the task and its answer (or its step results). That is what resolves "and now in Kelvin".
The message you just sent, unchanged.
Every step of this task with its params and result or error, each cut at 2000 characters. This is how the loop observes.
system: You plan the tools of a software agent. … (about 40 lines of rules) user: Available tools: - ai_plan(goal) – AI assisted task planning - celsius_to_fahrenheit(celsius) – Converts Celsius to Fahrenheit - celsius_to_kelvin(celsius) – Converts Celsius to Kelvin - classic_plan(goal) – Rule based task planning without AI - forget(id) – Forgets the remembered fact with the given id - remember(fact) – Remembers a short, self-contained fact … Things you remember: [1] The user lives in Hamburg Conversation so far: - User: Convert 30 degrees Celsius to Fahrenheit Result: 30 °C is 86 °F. Task: and now in Kelvin Results so far: - celsius_to_kelvin({"celsius":30}) -> %{kelvin: 303.15}
celsius_to_kelvin was generated in round 1 and is already in the tool list, and its result sits under "Results so far".This is the planner's context with the Claude backend. The mock planner gets the same map but only looks at the task and the results.
If the planner names a tool nobody has written yet, Dynamic.ensure_action/2 takes over. Anything the sanitizer, the compiler, the schema check or the probe rejects goes back to the model with the reason, up to three attempts in total. A rejection by you is final.
The limits are tagged. by design means a deliberate choice for a demo, shortcut something a real system would need, risk a security caveat, and model behaviour seen in eval runs with Claude.
The planner sees the results of every step and plans the next round, so it can chain values from one step into the next, retry a failed step or answer.
An unknown tool name triggers code generation. The module is persisted, compiled, loaded under a versioned name and reused in later turns.
An AST walk with a deny list (shell, eval, spawn, atoms) and an allow list of modules. File writes only go below the sandbox directory, network access is GET only through one function.
New code waits for a click before its first run: approve or reject, shown with the full source. It is on by default and times out after 10 minutes.
Every step runs once, in its own process with a time limit and a heap limit. Hitting either becomes a readable error the planner can react to.
Claude API calls are retried on 408, 429, 5xx and 529 and on dropped connections, honouring retry-after. Tools are never retried, so a page is never fetched twice.
"Remember that ..." stores a fact, "forget fact 3" removes it. Facts go into every planner prompt and show up live in every open chat.
Every planner call, code generation, tool step and wait for approval is listed with its duration and tokens, with totals below. The same events feed LiveDashboard metrics.
One GenServer per conversation, found through a Registry. It survives page reloads and stops after 15 idle minutes. If it crashes, the LiveView starts a fresh one, and other conversations are not affected.
The planner can be the mock or Claude, the generator the mock, Claude or a local Ollama model. You pick per run in the UI, and the mocks run without an API key.
mix agent.eval runs fixed cases several times and reports a pass rate with rounds, tokens and time. The cases check the way the agent took, not only the answer.
Tools are a four-callback behaviour validated with NimbleOptions. Everything else is GenServer, Registry, DynamicSupervisor, Task.Supervisor and PubSub.
The sanitizer only sees the AST. Code that passes it runs with the full privileges of the BEAM node. OS-level isolation is the missing layer.
riskThe URL comes from the model. A GET can hit internal services and metadata endpoints, or leak data in a query string. Where that matters, the restriction belongs in the network.
riskCatalog tools and reused generated tools run without asking. There are no per-tool permissions, and the eval harness approves automatically.
riskThe planner sees the last 4 turns, each summarized, and step results cut at 2000 characters. Nothing condenses older context.
shortcutMessages exist only in the agent process. They are gone after 15 idle minutes or a restart. Only facts and generated tools are persisted.
shortcutFacts are one global list of at most 50, with 300 characters each. That is fine for a single user; with several, each would need their own.
shortcutThe Registry, the ETS tool registry and the loaded modules are local to one node. A cluster would need distributed lookup and code loading on every node.
shortcutThe trace counts tokens but prints no prices, because prices in code go stale. There is no budget that stops a run.
shortcutThe planner returns steps as forced JSON instead of using the API's tool calling. That is what lets the keyword mock drive the same loop.
by designSteps run sequentially, so each keeps its own result. There are no parallel tool calls, and each conversation runs one task at a time; a message sent meanwhile is ignored.
by designAfter 5 rounds of tool calls the loop gives up with :max_rounds and returns what it has.
Generated tools compute, process text or read a page. They get no clock, no random numbers, no processes and no writes outside the sandbox directory.
by designThe mock planner matches keywords and ignores history and facts. The mock generator only echoes its parameters. They show the mechanics, not intelligence.
by designProgress arrives per planned round and per finished step. The model's text itself is not streamed.
by designWith Claude, temperature conversions came back in 0 rounds: the right answer without any tool. That is sensible, but it skips the tools the case meant to test.
modelAsked to count the words of a page title, Claude generated one combined tool instead of chaining a page tool and a text tool.
modelIn 3 of 30 runs, parts of the step leaked into the tool name. The pipeline rejected them and the planner recovered. The output schema now restricts names by pattern; no Claude run has checked that yet.
modelOnce, Claude generated prioritize_todo_list although classic_plan sat in the catalog. It cost about 23 seconds.
These are the defaults in the code. Most can be changed via config :agent_demo, ....
| What | Value | Where |
|---|---|---|
| Rounds of tool calls per task | 5 | TaskAgent @max_rounds |
| Earlier turns the planner sees | 4 | ChatAgent @history_turns |
| Step result passed to the planner | 2000 chars | ToolPlanner.Anthropic |
| Time per tool step / probe | 60 s / 20 s | Dynamic.Limits, Dynamic |
| Heap per tool step | 64 MB | Dynamic.Limits |
| HTTP GET timeout / body size | 10 s / 2 MB | Dynamic.Http |
| Generation attempts (1 + repairs) | 3 | Dynamic generation_attempts |
| Tool name | ^[a-z][a-z0-9_]{2,49}$ | Dynamic.name_pattern/0 |
| Wait for approval | 10 min | ChatAgent.Worker |
| Agent stops when nobody is watching | 15 min | ChatAgent @idle_timeout |
| Facts / length per fact | 50 / 300 chars | Memory |
| Retries of a Claude API call | 3 | AI.Retry |
| Retries of a tool step | 0 | TaskAgent |
This is mix agent.eval --planner anthropic --generator anthropic --runs 3 with the first version of the cases: 27 of 30 runs passed. The rounds column showed that a pass did not always mean the agent took the intended way.
celsius, celsius_reuse and use_memory were answered without a tool, so nothing was generated and nothing reused.min_rounds and nothing_generated. No Claude run of the new cases exists yet.