Free book The Agentic OS: how to run your business with a team of AI agents. Get it free
Internal build · AI operations

Devian OS: the agent operating system I built to run my own business

Fourteen AI agents, seventy tools, one board and an approval gate: the system behind my SEO and automation work, with the numbers from its first two weeks.

14agents on one roster, each with a role, tools and permissions
70tools, every one behind a named permission
811agent runs in the first two weeks, each one traced
1,300+automated tests that run before anything counts as done

The starting point

Like most people using AI for work, I was the glue. One chat for research, another for writing, a third for code, and me copying results between them, remembering what each one knew and checking everything by hand. The AI was fast. The system around it was me, and I was the bottleneck.

So I built the system instead: one place where AI agents take work, do it with tools they are allowed to use, remember what they learn, and stop and ask before anything reaches the outside world. I call it Devian OS, and it is where every idea in my client work gets tried first.

What I built

A roster of agents, hired by role. Fourteen agents, from a generalist operator to a planner, a reviewer, coding agents and a local model that never leaves the laptop. Each has a job, a set of tools and permissions, and a model that suits the work. Swapping the model behind an agent doesn’t change what it is allowed to do.

A board, not a chat. Work lives on cards that agents pick up, split, finish and hand back for review, so nothing depends on one conversation staying open. Ten saved missions start common jobs, such as a research brief or tidying the memory, with one click.

An approval gate. Anything that leaves the system, such as a published post, a sent message or an outbound request, stops for a person’s yes. In the first two weeks the gate stopped ten actions: four were approved, two rejected, and three failed after approval and were reported as failures, not quietly retried.

A memory that compounds. A Markdown vault of 79 notes that every agent reads from and writes to, drawn as a galaxy so I can see what connects to what. When an agent uses a note, the system records which agent and when.

Models from everywhere, chosen per job. 52 model profiles across 12 providers: subscription tools I already pay for, free gateways, and a model that runs locally. Each agent uses the one that fits its job, and changing it is a setting, not a rebuild.

Seventy tools behind permissions. Search, files, the web, the browser, memory, delegation, SEO research and content studios. An agent only gets a tool if it holds that tool’s permission, and every action lands in an audit log, now more than 8,600 entries long.

A voice. Jarvis, a butler-style assistant I talk to, runs on the same system: he answers from my notes and the live state of the board, opens any screen by name, and when I leave he reaches me on Telegram only if something can’t wait.

063125188250 12 Sep162125 Sep Board and pipelines live 36
Agent runs per day, 12 to 25 September 2026From the system's run log

What didn’t work

Of those 811 runs, 241 failed. That number stays in the log on purpose: a system that hides its failures teaches you nothing. The failures were the most useful part of the build.

  • Green is not the same as done. Early on, runs finished “successfully” with nothing produced. The rule now is to judge a run by what it produced, not by whether it errored.
  • A fast demo can hide a slow brain. A small local model took over 90 seconds to answer with the full context, which is fine for a batch job and useless for a voice. Measuring each model on the real task, not a toy prompt, decided which brain does what.
  • Credits run out. One provider’s prepaid credits ran dry mid-build, which stopped Jarvis’s vision. It now falls back to a free vision model and says which model looked, instead of failing silently.

The results

  • One board where fourteen agents take real work, with a person approving everything that leaves.
  • 811 traced runs in two weeks, each with its model, tools, cost and outcome on record.
  • More than 1,300 automated tests that run before a change counts as finished.
  • The same system runs my SEO research pipeline and drafts content, each piece stopping at the gate before it goes anywhere.

Lessons worth stealing

Hire by role, not by model. Models change monthly. A role with clear tools and permissions survives every change in the market.

Put the gate in first. Once nothing can leave without approval, you can let agents do everything else. The gate is what makes the speed safe.

Keep the failures visible. A log that shows 241 failures out of 811 runs is worth more than a dashboard of green lights, because it tells you exactly what to fix next.

Want a system like this for your business, sized to the work you actually do? That is what I build for clients.

Want results like these? Let's map yours.

Book a free 30-minute strategy session. You'll leave with a 90-day plan built around your market.

Popular

↑↓ to move Enter to open Esc to close