Elixir BEAM OTP Phoenix LiveView

What this agent
demo can do, and
where it stops.

A loop that plans, runs tools and reads the results, and that writes the tools it is missing. It runs in plain OTP processes, with no agent framework.

This page walks through the moving parts, replays a run step by step and lists the limits as they are in the code today. Every number on it is taken from the source.

Three pages

One app, three ways to look at an agent

Each page adds one idea to the previous one. They share the same tools and processes.

/

Classic agent vs. AI agent

The same goal goes to two agents that are identical except for their brain. One uses keyword rules (ClassicPlan), the other a model (AiPlan, mock or Claude). The results are shown side by side.

/dynamic

Tools written at runtime

Describe a tool, and a model writes the Elixir module. You can watch it pass the sanitizer, compile, load under a versioned name and answer a probe call. The registry lists every tool generated so far.

/chat/:id

An agent that picks its own tools

Each conversation is a GenServer. A worker runs the loop, and progress, approvals, the trace and remembered facts reach the LiveView over PubSub as they happen.

Step through a run

What happens when you send a message

Pick a task and step through it. The loop diagram lights up the part that is working. This is a replay: the tool names and texts are examples, but every step is one the code actually takes.

    results go back as observations, next round no steps left: the reasoning is the answer Planner ToolPlanner.plan/2 Pipeline only if the tool is new Tool Limits.run + Tool.run Answer
    Architecture

    Brain, loop, body

    The only part that decides anything is the planner, and it is swappable. Everything around it is plain OTP and Phoenix: the loop, the tools, the processes and the checks.

    BROWSER PROCESSES BRAIN STATE ChatLive /chat/:id · approve · trace ChatAgent GenServer per conversation cast PubSub Worker Task.Supervisor, one per task progress TaskAgent the loop, max 5 rounds ToolPlanner Mock (keywords) · Claude (forced JSON) sees tools, history, results, facts Generator Mock · Claude · Ollama, writes tools Tools catalog + generated new tool you approve new code Memory facts, JSON file Storage ETS registry + .ex files Trace :telemetry spans, tokens Registry finds the agent by id
    Context

    What the planner gets to see

    The model remembers nothing between calls. Before every round, TaskAgent assembles the planner's context afresh: six parts, each from a different place, each living for a different span of time. Everything the agent "knows" has to be in it.

    System prompt fixed

    The rules: use existing tools, invent new ones in snake_case, plan again after results, remember lasting facts.

    ToolPlanner.Anthropic @system

    Available tools refreshed every round

    The catalog plus every generated tool in the registry, with name, required parameters and description. A tool written in round 1 is offered in round 2.

    TaskAgent.available_tools/0

    Things you remember across conversations

    The facts from the JSON file, at most 50 of 300 characters each. Read again every round, so a fact remembered just now already counts.

    Memory.list/0

    Conversation so far one conversation

    The last 4 turns, each reduced to the task and its answer (or its step results). That is what resolves "and now in Kelvin".

    ChatAgent.history/1

    Task one task

    The message you just sent, unchanged.

    ChatAgent.send_message/2

    Results so far one task, grows per round

    Every step of this task with its params and result or error, each cut at 2000 characters. This is how the loop observes.

    TaskAgent observations
    system: You plan the tools of a software agent. …
            (about 40 lines of rules)
    
    user:
    Available tools:
      - ai_plan(goal) – AI assisted task planning
      - celsius_to_fahrenheit(celsius) – Converts Celsius to Fahrenheit
      - celsius_to_kelvin(celsius) – Converts Celsius to Kelvin
      - classic_plan(goal) – Rule based task planning without AI
      - forget(id) – Forgets the remembered fact with the given id
      - remember(fact) – Remembers a short, self-contained fact …
    
    Things you remember:
      [1] The user lives in Hamburg
    
    Conversation so far:
      - User: Convert 30 degrees Celsius to Fahrenheit
        Result: 30 °C is 86 °F.
    
    Task:
    and now in Kelvin
    
    Results so far:
      - celsius_to_kelvin({"celsius":30}) -> %{kelvin: 303.15}
    Built from scratch, every callThe second planner call of the task above: celsius_to_kelvin was generated in round 1 and is already in the tool list, and its result sits under "Results so far".
    One message, not a transcriptEarlier turns are not replayed as user and assistant messages. They are lines of text in a single user message, each a task and its answer.
    Bounded, except for the toolsHistory, facts and results have limits. The tool list grows with every generated tool and nothing trims it. Nothing is cached either: the system prompt and tool list are sent again with every call.

    This is the planner's context with the Claude backend. The mock planner gets the same map but only looks at the task and the results.

    Dynamic tools

    From "this tool is missing" to a loaded module

    If the planner names a tool nobody has written yet, Dynamic.ensure_action/2 takes over. Anything the sanitizer, the compiler, the schema check or the probe rejects goes back to the model with the reason, up to three attempts in total. A rejection by you is final.

    Capabilities and limits

    What it can do, and what it can't

    The limits are tagged. by design means a deliberate choice for a demo, shortcut something a real system would need, risk a security caveat, and model behaviour seen in eval runs with Claude.

    ↻Plan, act, observe

    The planner sees the results of every step and plans the next round, so it can chain values from one step into the next, retry a failed step or answer.

    Agents.TaskAgent

    ✎Write missing tools

    An unknown tool name triggers code generation. The module is persisted, compiled, loaded under a versioned name and reused in later turns.

    Dynamic · Dynamic.Generator.*

    ☑Static safety gate

    An AST walk with a deny list (shell, eval, spawn, atoms) and an allow list of modules. File writes only go below the sandbox directory, network access is GET only through one function.

    Dynamic.Sanitizer · Dynamic.Http

    ✋Human approval

    New code waits for a click before its first run: approve or reject, shown with the full source. It is on by default and times out after 10 minutes.

    ChatAgent.Worker

    ⌛Bounded execution

    Every step runs once, in its own process with a time limit and a heap limit. Hitting either becomes a readable error the planner can react to.

    Dynamic.Limits

    ↺Retries where they belong

    Claude API calls are retried on 408, 429, 5xx and 529 and on dropped connections, honouring retry-after. Tools are never retried, so a page is never fetched twice.

    AI.Retry

    ★Memory across conversations

    "Remember that ..." stores a fact, "forget fact 3" removes it. Facts go into every planner prompt and show up live in every open chat.

    Memory · Tools.Remember/Forget

    ⏱A trace per answer

    Every planner call, code generation, tool step and wait for approval is listed with its duration and tokens, with totals below. The same events feed LiveDashboard metrics.

    Trace · :telemetry

    ⚙A process per conversation

    One GenServer per conversation, found through a Registry. It survives page reloads and stops after 15 idle minutes. If it crashes, the LiveView starts a fresh one, and other conversations are not affected.

    Agents.ChatAgent

    ⇆Swappable brains

    The planner can be the mock or Claude, the generator the mock, Claude or a local Ollama model. You pick per run in the UI, and the mocks run without an API key.

    ToolPlanner · Dynamic.Generator

    ⚖Eval harness

    mix agent.eval runs fixed cases several times and reports a pass rate with rounds, tokens and time. The cases check the way the agent took, not only the answer.

    Eval · eval/cases.exs

    ◯No framework

    Tools are a four-callback behaviour validated with NimbleOptions. Everything else is GenServer, Registry, DynamicSupervisor, Task.Supervisor and PubSub.

    Tool · application.ex

    Not a sandbox

    The sanitizer only sees the AST. Code that passes it runs with the full privileges of the BEAM node. OS-level isolation is the missing layer.

    risk

    GET reaches everything the node reaches

    The URL comes from the model. A GET can hit internal services and metadata endpoints, or leak data in a query string. Where that matters, the restriction belongs in the network.

    risk

    Approval covers only new code

    Catalog tools and reused generated tools run without asking. There are no per-tool permissions, and the eval harness approves automatically.

    risk

    Short memory of the conversation

    The planner sees the last 4 turns, each summarized, and step results cut at 2000 characters. Nothing condenses older context.

    shortcut

    Conversations live in RAM

    Messages exist only in the agent process. They are gone after 15 idle minutes or a restart. Only facts and generated tools are persisted.

    shortcut

    One memory for everyone

    Facts are one global list of at most 50, with 300 characters each. That is fine for a single user; with several, each would need their own.

    shortcut

    Single node

    The Registry, the ETS tool registry and the loaded modules are local to one node. A cluster would need distributed lookup and code loading on every node.

    shortcut

    Tokens, not costs

    The trace counts tokens but prints no prices, because prices in code go stale. There is no budget that stops a run.

    shortcut

    JSON loop, not native tool use

    The planner returns steps as forced JSON instead of using the API's tool calling. That is what lets the keyword mock drive the same loop.

    by design

    One step at a time

    Steps run sequentially, so each keeps its own result. There are no parallel tool calls, and each conversation runs one task at a time; a message sent meanwhile is ignored.

    by design

    At most 5 rounds

    After 5 rounds of tool calls the loop gives up with :max_rounds and returns what it has.

    by design

    Small tools only

    Generated tools compute, process text or read a page. They get no clock, no random numbers, no processes and no writes outside the sandbox directory.

    by design

    Mocks are deliberately dumb

    The mock planner matches keywords and ignores history and facts. The mock generator only echoes its parameters. They show the mechanics, not intelligence.

    by design

    No streaming

    Progress arrives per planned round and per finished step. The model's text itself is not streamed.

    by design

    Computes in its head

    With Claude, temperature conversions came back in 0 rounds: the right answer without any tool. That is sensible, but it skips the tools the case meant to test.

    model

    Combines instead of chaining

    Asked to count the words of a page title, Claude generated one combined tool instead of chaining a page tool and a text tool.

    model

    Garbled tool names

    In 3 of 30 runs, parts of the step leaked into the tool name. The pipeline rejected them and the planner recovered. The output schema now restricts names by pattern; no Claude run has checked that yet.

    model

    Writes a tool it already has

    Once, Claude generated prioritize_todo_list although classic_plan sat in the catalog. It cost about 23 seconds.

    model
    Numbers

    The limits in figures

    These are the defaults in the code. Most can be changed via config :agent_demo, ....

    WhatValueWhere
    Rounds of tool calls per task5TaskAgent @max_rounds
    Earlier turns the planner sees4ChatAgent @history_turns
    Step result passed to the planner2000 charsToolPlanner.Anthropic
    Time per tool step / probe60 s / 20 sDynamic.Limits, Dynamic
    Heap per tool step64 MBDynamic.Limits
    HTTP GET timeout / body size10 s / 2 MBDynamic.Http
    Generation attempts (1 + repairs)3Dynamic generation_attempts
    Tool name^[a-z][a-z0-9_]{2,49}$Dynamic.name_pattern/0
    Wait for approval10 minChatAgent.Worker
    Agent stops when nobody is watching15 minChatAgent @idle_timeout
    Facts / length per fact50 / 300 charsMemory
    Retries of a Claude API call3AI.Retry
    Retries of a tool step0TaskAgent
    Measured, not claimed

    The first eval run with Claude

    This is mix agent.eval --planner anthropic --generator anthropic --runs 3 with the first version of the cases: 27 of 30 runs passed. The rounds column showed that a pass did not always mean the agent took the intended way.

    All three failures: garbled tool namesName check caught them, the planner recovered in the next round. The schema now forbids such names.
    0 rounds, still greencelsius, celsius_reuse and use_memory were answered without a tool, so nothing was generated and nothing reused.
    What changed sinceThe cases now check the way: numbers too long to compute mentally, min_rounds and nothing_generated. No Claude run of the new cases exists yet.