Six days ago there were three capability packs in a private repository. Today there are thirteen, nine of them forming one framework for running a content website: an inventory ledger, a demand radar, a competitor corpus, an article pipeline, an evidence library, a link planner, a publishing transaction, an operator guide and a cockpit. 133 lanes run on one model. 317 unit tests and 149 render checks pass. Not one line of that was typed by a person. It was prompted, into Claude Code and Codex, through the MCP the runtime exposes, and the runtime is what made that possible without the result being a pile of scripts nobody can inspect.
This article is about the runtime, not the website. The website is a use case. What it demonstrates would hold for a support desk, a compliance review, a trading rule, a research programme: a long process, dozens of steps, some of them judgement calls a model should make, some of them measurements a script must make, and a person who has to be able to see, at any moment, what happened and why.

What Agentic-Nets is
Agentic-Nets is a governed, event-sourced runtime built on bipartite graphs, Petri nets. Its own README calls it a graph harness for AI agents. Two kinds of node:
- Places hold tokens. A token is a JSON document. Places are persistent, event-sourced and queryable with a small query language (ArcQL). A place is where a requirement, a decision, a measurement, an error, a question or a draft lives, durably, with its history.
- Transitions (the runtime calls them lanes) consume tokens from their input places when the required state arrives, act, and emit tokens into their output places. A transition has a kind:
pass,map,http,llm,command,agent,link. Seven kinds, and they cover everything the framework needed.
Three of those kinds do the work in this framework:
- A map lane is deterministic. It reshapes a token, routes on a condition, packs a command. No model, no I/O.
- A command lane hands a command-shaped token to an executor, a separate process that polls the runtime over an egress-only link, runs the script, and returns stdout as a token. The framework’s Python lives here: one bundled runtime file per pack, shipped inside the package. Secrets never sit in the token; a vault injects them into the executor’s environment at fire time, and the executor scrubs its own environment before every run. A command lane that needs a model’s judgement calls a headless CLI (Claude Code or Codex) with a JSON schema for the answer.
- An agent lane is a model with tools in a loop. It can be given external MCP servers to call, including the runtime’s own MCP, fenced by an explicit allow-list of tools and a role string of eleven capability flags. Its reasoning is bounded by
maxIterations; its write access is bounded by the net: it can only emit into the places its arcs point at.
Every fire, every token and every emission is an event. The runtime keeps the event trail, a durable model history, and token lineage. That is the property everything else in this article rests on.
The operator does not build any of this
The operator’s tools are a chat window and a browser. Claude Code and Codex connect to the runtime through its MCP server. The runtime’s MCP carries its own operational documentation (docs/index, docs/emit, docs/recipes, docs/mcp-servers, limits), so the coding agent reads how the platform works from the platform, not from training data. The agent then builds: add_place, add_transition, set_transition_credentials, attach_mcp_server, hub_publish. It verifies: diagnose_transition, dry_run_transition, event_trail, query_tokens. It packages: nets, inscriptions, scripts, seeds, application and manifest travel together as a versioned capability into a private NetHub repository, and are installed from there into a model with one call.
The conversation is the design document. Over six days the operator described how they work: a numbered procedure, its gates, the rules of evidence, what may never be adopted from third-party material without approval. The agent turned that into nets. When the operator wanted to know why a step was blocked, the agent read the model. When the operator wanted an application changed, the agent changed it, republished, reinstalled. The operator never opened the net editor. They opened the applications.
Applications: the net’s human face
A Net Application is a projection over live places plus a set of guarded actions that write tokens back. It is a trusted web component, shipped inside the capability, reading the model through a runtime bridge: readStore, watchStore, readBlob, invoke. It is not a reporting database beside the process. It is the process, rendered.

The Operator Guide above renders the operator’s own written procedure (sixteen steps for a new article) and a run inside it. The left rail is the run’s steps map, read from a token. The right panel is the question the run is waiting on, as an Interview prompt with options. The buttons invoke actions; each action’s input schema is declared in the application’s manifest, validated by the runtime before any lane sees it, and lands as a token in a requests place that one lane consumes. Nothing in the page has state of its own.

The cockpit reads 33 stores from nine packs and answers one question: what needs a person. 37 items when this was captured: gates, questions, approvals, cases. 10 stopped stages, each with the reason its lane stopped. Seven schedules that were deployed and never started, classified from a runtime audit token. Data sources probed with a verdict per adapter.

One process, end to end, as the model recorded it
On 2026-09-21 the framework ran a complete article pipeline on a keyword. What follows is not a description of the design; it is the run’s own event trail, read back the next morning: the work log, the decisions place, the artifacts. This is the point of the article, so it is worth being precise.

18:25:59, one operator action. start-keyword. Nothing had run before it. The request lane, a command lane, took the order and made three measurements before any search: it read the site’s 87 published pages live from the CMS API and judged the keyword new (0 covering pages, 0 close); it read the operator’s own search data from the ledger, 236 queries, 0 matching; it recorded that no revenue measurement exists, as a named gap rather than an estimate. Each of those became a token. Decision inventory-checked, decision keyword-selected.
18:26 to 18:30, four lane pairs, no operator. Search (keyless, one engine). Five competitor pages fetched and measured: 1,037 · 1,076 · 1,473 · 1,672 · 1,843 words. A gap analysis that read all five and named what none of them provides: 14 gaps. A judge that scored the resulting editorial promise at 80.2, above the gate, and wrote one reservation into its verdict: the promise assumed manufacturers publish figures they might not.
18:30 to 18:51, sources and brief. 30 sources proposed, each re-fetched directly and matched against its quotation: 24 verified, 6 rejected, 22 snapshots and 30 claims handed to the evidence library as its own tokens. The brief written as a document (word range 2,600 to 3,400, a 17-section outline, 20 sources), and then the lane stopped at wait-brief. A fire never waits for a person; it writes the gate and finishes.
18:52:09, the one human gate crossed. approve-brief, bound to that exact brief revision. Queued write.
18:56 to 19:37, 22 lane fires, five revisions. Write (headless Claude, schema output): 4,289 words. Context: the reviewer’s packet, draft plus brief plus the verified quotations. Review (headless Codex, a fixed eight-criterion rubric): 43, REJECT, 23 problems. Revise. 65. Revise. 66; the 42 deterministic checks joined: 83.33, seven failing. Revise. Revision 4: review 89 PUBLISH, checks 100.0. Only then did the third check run, the one that tests every claim against the quotations: 37 unsupported claims, three critical. The readiness function (deterministic, no model) routed revise, because any unsupported claim blocks release regardless of the other two scores. Revision 5 cut 1,460 words and fell to 55. The revision cap parked the run at wait-human with reasonCode: qa-cap.
What was never reached: the release gate, the staged draft, publish, submission, post-publish scans.
None of that was reconstructed from logs or memory. It was read from places: p-ap-work-completed, p-ap-decisions, p-ap-gpt-review, p-ap-qa-check, p-ap-qa-llm, p-ap-readiness, the draft blobs. Every number in the paragraphs above has a token behind it, and every token has the fire that produced it.
The strongest claim, stated plainly
A process built by an AI is debuggable in this runtime in a way that an application built by an AI is not.
When a coding agent writes a service, the result is code. To find out why it did something you read the code, add logging, reproduce the input, and hope the failure is deterministic. When a coding agent writes a net, the result is a graph of small modules, each lane a bounded unit with declared inputs, declared outputs and one kind, and the runtime records what every one of them consumed, produced and decided. The debugging surface is not the source; it is the event trail and the places. You do not reproduce a failure; you read it.
Three things in the run above were discovered only because of that:
- The revise loop made the best revision worse and kept the worst. Revision 4 passed two of three checks; the forced revision 5 scored 55 and lost three checks; the cap parked revision 5. Visible in one query over the review place, sorted by revision. In a conventional pipeline that is a print statement someone has to think to add.
- Two reviewers disagreed completely about the same text: 89 PUBLISH against 37 unsupported claims. Both verdicts are tokens with the draft revision on them. Diffing what each was given (the context tokens) showed the rubric never tested claims against evidence. That is a finding about the process, made from the process’s own records.
- The gap judge warned an hour before the failure. Its reservation at 18:30 named the exact failure mode of 19:28. The warning was a token nobody consumed. Now it can be wired to a gate, one arc.
The same property runs through the week’s other finds. An application counted two token properties as strings and rendered 10 where one had stopped, found by reading the run token it summarised. A lane’s health record described a version that was no longer installed, found by comparing the health place with the install. A start action that skipped four gates of the operator’s procedure, found by asking, for one article, whether the places those gates write into held anything. They were empty. Every one of those is a query, not a debugging session.
Best of both worlds, made concrete
The framework decides with models where judgement is required and with code where it is not, and the net is where that line is drawn, visibly, per lane.
- AI decides: search intent, which competitor pages matter, what the gaps are, whether the editorial promise is worth pursuing (with a score), how to write, how to review against a rubric, whether each claim is supported by its quotation.
- Code decides: whether a page already covers a keyword (stems against 87 live slugs), whether the operator’s own data shows demand, whether 42 formal checks pass, whether a source’s quotation actually appears on the fetched page, and, the one that mattered, whether an article may be released, from the three verdicts combined.
readiness()is 30 lines of Python. No model can talk it into a release. - A person decides: the brief, the human gate, the release. A fire never waits for them; it writes a gate token and stops. The person’s answer arrives as another token. “An answer is information, not authorization”: the runtime’s own documentation says it, and the framework treats it as law.
The result is a process where the expensive, fallible calls are surrounded by cheap, certain ones, and where each kind of decision is a different kind of lane, so anyone reading the net knows which is which without reading a prompt.
What it cost, and what it did not
The whole framework runs without a paid API. Search is keyless. The models are the operator’s Claude Code and Codex subscriptions, invoked as headless CLIs from command lanes; the one agent lane in the guide runs on a local provider through Ollama. That was a constraint from the operator, and the runtime accommodated it because the runtime does not care where intelligence comes from, server model, local model, or the MCP client already open.
What it did cost is honesty. The week produced a 60-item audit, and the first line of its verdict is the machinery works; the process has never been completed once. No article has reached the release gate. Half the pipeline (publish, submit, measure, learn) is code that has not met reality. That sentence could only be written because p-ap-published and p-ap-wait-release are places, and they are empty. A framework that cannot tell you what it has never done is not one you can trust with what it has.
Where this is
- Runtime: github.com/alexejsailer/agentic-nets: gateway, executor, vault, CLI, chat, MCP server, blob store; BSL 1.1. Desktop Lite in the releases.
- Documentation and a live net to try without installing: agentic-nets.com.
- Guided tour, eight minutes: youtu.be/hgW11A_7vWY.
- Forum: forum.agentic-nets.com.
- The framework described here lives in a private NetHub repository (thirteen capability packs, 104 versions) and installs into any Agentic-Nets model with
hub_install. The public repository ships the capability tooling (capabilities/tools/pack.mjs) that builds, packages and publishes such packs.
The run ended at a gate. The record of it did not.