TovasolBook a call
Proof · systems we built and run ourselves

We ship the production AI you can’t hire for.

No client logos yet. Something better: two production systems we built and run ourselves, in the open. Clone them, run the tests, read the threat model.

First-party builds · public repos · every number maps to a command
Public reposclone and run
Threat-modelednames what it does not defend
Case studies

Systems we run our own production on

Not a questionnaire and not client logos: first-party systems we built and operate, public and re-runnable.

01 · CI/CD runner · open source

The CI/CD we run our own production on

A from-scratch, single-VPS runner on gitolite, rootless Docker, and sops/age. It ships this very site, with a written threat model that names what it does not defend.

  • Container hardening on every job: cap-drop ALL, seccomp, no-new-privileges
  • sops/age per-environment secret scoping
  • Adversarial + mutation-tested shell suite

“It is not a multi-tenant CI platform.”

Single-tenant, single-VPS, on purpose. Knowing where a system stops is production discipline, not a gap.

Public · MIT · clone and run the tests yourself
02 · agent-forge · open source

An agentic system, idea to deployed

Built on the Claude Agent SDK: idea → orchestrator → parallel research workers → critic loop → policy gate → deployed site. The exact stack we sell: orchestration, evals, guardrails, cost control.

  • Guardrails in code, not prompts: spend + named-contact actions gated to a human
  • A hard budget ceiling and turn caps bound every run
  • A verification-gated builder that cannot declare victory early

“It does not guarantee revenue. It builds the machine and optimizes on evidence; the market decides.”

Straight from its own CAPABILITIES.md. We built the machine and drove one idea to a deployed site. We never imply revenue or a track record.

Public · MIT · clone and run it yourself
Judge the engineering

Read the parts that show judgment

No trust required, and no vanity metrics. The repos are public — read the design docs and the guardrail code and judge the work directly.

A threat model that names what it does NOT defend

Honest scope limits, written down, not marketing.

A confused-deputy-safe archive-push design

A security hole closed by construction, not by a check bolted on later.

Guardrails and eval-gates enforced in code, not prompts

Spend and named-contact actions are predicates gated to a human. No stage advances on vibes.

A capabilities doc that leads with what it will NOT do

The system states its own limits before it sells you anything.

No client case studies yet. The strength: our own systems are public and re-runnable, so you judge the engineering directly instead of taking our word.

Put your AI in front of us.

A 45-minute working session on your actual AI project, not a sales deck. We find where it breaks and name the exact thing standing between you and launch.