The depth gauge · 11 teams sounded · 3 layers each · 0 floors breached

One layer deeper than the conductor is allowed to go.

Ten agent teams sit behind a real information boundary, one the system enforces on itself, in its own files, not as a marketing conceit. This page goes past what the conductor is allowed to see, once, on purpose, and stops at exactly the point the conductor also stops. No skill file, no prompt, no agent instruction appears below, no matter how far you scroll.

This page shows you more than the conductor is allowed to see, and still not the part that matters.

SURFACE — WHAT'S ALREADY PUBLIC ACT I — CONTRACT LAYER what the conductor is given: 11 sealed cards, four facts each TEAM · JOB · ENTRY POINT · COMPLETION SIGNAL ACT II — SCORECARD LAYER one layer past the conductor: real headcounts, real skill counts 100 AGENTS · 69 SKILLS · 1 PROOF-POINT EACH ACT III — CHOREOGRAPHY LAYER how it moves: sequence, parallel fan-out, and who decides next STILL NO SKILL FILES, NO PROMPTS, NO STAGE NAMES LINE STOPS HERE BEDROCK — SKILL & PROMPT CONTENT NEVER SURFACED, NOT EVEN TO THE CONDUCTOR

A sounding, not a drilling. The probe measures depth; it does not bring anything up.

100
Agents across the 11 teams
69
Distinct skills, never shown before
15
Skills the conductor itself gets
0
Skill files shown on this page

Act I · What the conductor knows

Eleven sealed cards, four facts each

Somewhere in this system, one file describes each of the eleven available teams to the single agent responsible for deciding what runs next. Not what's inside the team. Not how many agents work on it. Four facts, and four facts only, whether the team behind the card has five agents or fifteen.

Rails Engineering

Builds and reviews a Rails web feature end to end.

Entry point

/feature <description>

Completion signal

Synthesis artifact written: {NNN}-summary.md

Go Engineering

Builds and reviews a Go command-line tool end to end.

Entry point

/feature <description>

Completion signal

Synthesis artifact written to its own Go briefs directory

Tauri Engineering

Specs and builds a Rust and Tauri desktop feature end to end.

Entry point

/feature <description>

Completion signal

Synthesis artifact written, backend, frontend, and IPC boundary together

Rails QA

Verifies a finished Rails feature against the actual running app.

Entry point

Launched against a completed feature's synthesis artifact

Completion signal

QA synthesis artifact with one combined verdict

Tauri QA

Verifies a finished desktop feature against the actual running app.

Entry point

Launched against a completed feature's synthesis artifact

Completion signal

QA synthesis artifact, once the platform and webview-driver matrix is resolved

Rails Security

Verifies a finished web feature's live security posture.

Entry point

Launched against a completed feature

Completion signal

Security synthesis artifact: PASS, PASS WITH NOTES, NEEDS WORK, or INCOMPLETE — AUTHORIZATION REQUIRED

Tauri Security

Verifies a finished desktop build's live security posture.

Entry point

Launched against a completed feature

Completion signal

Security synthesis artifact: PASS, PASS WITH NOTES, or NEEDS WORK

Market Research

Assesses whether a product idea is viable before any engineering starts.

Entry point

Launched against a product idea

Completion signal

One verdict: GO, NO-GO, or CONDITIONAL, with a confidence percentage

Ideation & Discovery

Turns a raw product idea into a structured brief.

Entry point

Runs live, in-conversation. Cannot be handed to a background subagent

Completion signal

A human-approved product brief

Design

Turns an approved brief into a reviewed design concept.

Entry point

Launched against an approved brief

Completion signal

A reviewed static design concept, checked for accessibility, performance, and AI discoverability

Copywriting

Writes marketing copy for a completed feature.

Entry point

/copywriting <path to a completed feature's summary>

Completion signal

Copywriting synthesis artifact written: {NNN}-copywriting-summary.md

Act II · The team scorecard

Deeper than the conductor is allowed to go

The conductor never sees any of this. We did, because someone has to build these teams, and this is as far in as this page goes. Real headcounts. Real skill counts. One proof-of-depth capability per team, named as an outcome, never as a methodology.

Rails Engineering

10
Agents
6
Skills

Three parallel specialist reviews, code quality, security, performance, after every implementation. The same finding surviving three rounds escalates to a human instead of looping a fourth time.

Go Engineering

10
Agents
5
Skills

Same review-loop discipline, Go-specific: runs govulncheck and checks go.sum supply-chain integrity as part of every security pass.

Tauri Engineering

10
Agents
6
Skills

Specs and reviews the Rust backend, the frontend, and the IPC boundary between them as one coherent unit, not three separate reviews.

Rails QA

9
Agents
7
Skills

Six specialist QA agents launched in parallel against the live running app: accessibility, copywriting, cross-browser, functional, load, visual.

Tauri QA

9
Agents
6
Skills

Cross-platform QA across three genuinely different webview engines: WebView2, WebKitGTK, WKWebView.

Rails Security

6
Agents
5
Skills

Drives a real ephemeral Kali Linux container that installs tools like Metasploit on demand, gated by a durable file-based authorization, never a chat confirmation.

Tauri Security

5
Agents
5
Skills

Dynamic penetration testing of the frontend-backend trust boundary: tries to make the running app violate its own declared capabilities config, instead of reading whether the config looks correct.

Design

12
Agents
7
Skills

Produces multiple distinct, fully-locked design-token directions per brief for a human to choose between, then reviews the built concept across four parallel lenses: accessibility, performance, SEO, AI discoverability.

Market Research

15
Agents
12
Skills

Fans out 11 specialist researchers in parallel, competitors, demand, pricing, regulatory, community, investment landscape, and more, then one synthesis agent reads all 11 reports plus a trends signal and returns one verdict with a confidence percentage.

Ideation & Discovery

6
Agents
3
Skills

The orchestrator itself has to run live, in-conversation. It can't be handed to a background subagent, because two of its five stages are live human interviews.

Copywriting

8
Agents
7
Skills

A four-stage pipeline: content plan first, then narrative angle, then landing-page and email copy drafted in parallel, then a final pass that decides per piece whether humor belongs at all. A clean "no humor fits here" is a complete result, not a failure to find a joke.

100 agents, 69 skills, and not one skill file changes hands to make any of it run.

The conductor's own 15

The conductor itself: 1 agent, 15 skills. Not internal expertise: one contract skill per available team, 11, three team-selection meta-skills, and one agent-log skill. 11 plus 3 plus 1 is 15. Its knowledge is contract-shaped by construction, not by restraint someone has to remember to apply.

Verified, not self-reported

A standing team-auditor agent, separate from the conductor, checks whether a newly built team is actually built correctly. Run for real against the newest security team, it found one real documentation bug and one real safety gap, an unbounded orphaned-container risk, independent of two review rounds that had already happened.

Act III · Pipeline choreography

How the pieces move, still without opening the box

One layer left before the actual instructions, and this page stops before it. What's left to show is shape: what runs in sequence, what runs in parallel, and what decides.

Within one team

Engineering families

Rails · Go · Tauri

DISC ARCH DSGN ENGR QUALITY SECURITY PERF loop HUMAN after 3x same finding

QA & Security families

Rails QA · Tauri QA · Rails Security · Tauri Security

artifact orchestrator 4 specialists in parallel combined verdict

Market Research

Widest fan-out on the roster

orchestrator 11 researchers + 1 trends script synthesis GO / NO-GO / CONDITIONAL

Copywriting

Sequential, then parallel, then sequential

plan angle landing page email humor pass "no humor fits" is a complete result too

Across teams, through the conductor

teams.yaml conductor team A team B team C ask ask ask order inferred from what each team consumes vs. produces, not a fixed pipeline position runs each team's real entry command directly; reads nothing but the Act I contract to decide

The conductor infers a sensible order. It never memorizes one.

The reveal stops here. That's not this page being cautious.

Three acts, three real layers of what's actually inside this system: the contract the conductor gets, the scorecard we just showed you, the choreography that moves it all. Nowhere in any of that is a skill file, a prompt, or an agent's actual instructions. That boundary isn't editorial restraint. It's the same rule written into the conductor's own file: never read a team's internal agent or orchestrator files. We went one layer past where the conductor is allowed to look. We did not go two.

That's also why the structure of this page is the argument, not just the packaging for it. A catalog that could keep going, another layer, another reveal, another count, would be proving the opposite point: that depth here is decorative, not load-bearing. It isn't. The floor was always going to stay covered, on this page and inside the system it describes.

If a write-up about a multi-agent system shows you the actual prompts, ask what's left for the system to be worth building.

A second instrument · not a fourth layer

Three things get called “the loop.” Two of them have already happened. Not in this repository.

The excavation above answers what's visible. This answers something else: whether any of it gets smarter over time, and whether that claim is earned or assumed. Three different mechanisms share the word “loop” in how this system gets described. They are not the same mechanism, and treating them as one overclaims what's real.

Tool-use loop

Inherited, not built

The basic think, act, observe cycle a single agent runs to use a tool. Every agent in this system gets it free from the runtime underneath. Nobody here built it.

Review loop

Real, demonstrated

Per feature, every run. An engineer hands off to parallel specialist reviewers; a real finding sends the work back. The same finding surviving three rounds escalates to a human instead of looping a fourth time.

Learning loop

Real, closed downstream

Cross-run. Fully designed and fully wired, in every team, right now. Closed twice already, both times in a downstream project that installed this system, never yet inside this template repo itself. The rest of this section is about this one.

AGENT RUN LOG bin/agent-log agent_log.sqlite3 decision --type gap: traces an artifact reflection: agent-internal log-analyst, one per team after 10-15 cycles dated report skill-builder human approves each item same mechanism, read in two different places THIS REPOSITORY 0 CLOSURES LOGGED INSTALLED DOWNSTREAM 2 CLOSURES LOGGED

Every arrow in this diagram is a real file, a real agent, a real gate. The left needle sits at zero because this repository ships the mechanism; it doesn't run features through it. The right needle already moved, twice, in the codebases that installed it.

Each team carries its own log-analyst agent, written to wait before it speaks: it activates after 10 to 15 feature cycles, not after every run, and its only job is finding a pattern across the accumulated data and writing that pattern to a dated report. A skill-builder agent reads the report, gets a human's explicit approval per proposed item, and edits the actual agent and skill files. The next run's agents are supposed to come out measurably different.

This loop has already closed twice. Just never inside the repository you're reading right now.

Checked this session, not assumed

In a real production codebase that installed this system's team definitions, the loop has closed twice, on two different dates. First closure: analysis of three features' worth of real runs produced 11 proposals, and all 11 got executed, including one brand-new skill built from scratch, a testing-discipline skill covering four specific coverage gaps that kept recurring, wired straight into the engineering agent's own skill list. Second closure, later: two real patterns got promoted. A race-condition handling rule had been silently re-derived from scratch three separate times; one of those times, the missing rule shipped a real bug. It got written once into the engineering agent as a standing rule, and a second brand-new skill, covering a sensitive-data-handling gap that kept recurring across features, got built and wired into four different agents in a single pass. Across every run that codebase has logged, the skill-builder agent has itself run 14 times. In a personal project, 99 real runs are logged and accumulating right now; no analysis pass has touched that data yet. Neither closure happened inside this template repository, and structurally, neither can: this repo ships the mechanism; it isn't a project running features through it. The evidence lives where the mechanism actually gets used.

Reading the same system through a different gauge

One rung climbed. One built and quiet. Two left deliberately untouched.

A separate piece on this site, the memory ladder, argues that complexity in AI memory is earned one failure at a time, never adopted because a technology is trending: files, then a database, then RAG, then a knowledge graph, each one only once the rung below it demonstrably fails you. This system practices that argument on itself.

Rung 1

Files

“you have context”

In continuous use

Every one of the roster's agents, every skill, every report-format contract is a file. This is the rung the entire system runs on today.

Rung 2

Database

“you have structure”

Built, not climbed

Built for exactly the reason the ladder prescribes a database: pattern-mining across many runs needs queryable structure, not re-reading prose every time. Wired in every team through the learning loop above. Zero real runs are logged inside this template repository; real runs are already logged and growing downstream, in the projects that installed it.

Rung 3

RAG

“you have recall”

Not built

Nothing in this system has outgrown a context window. No failure has occurred that RAG would fix. Deferred on purpose, not from lack of ambition.

Rung 4

Graph

“you have relationships”

Not built

No question has come back wrong because the answer was an untracked relationship between two artifacts. The same discipline the ladder argues for: earned through usage, when a real failure demands it.

This system practices exactly what the ladder argues for, including on itself.

Two rungs stay empty. Not because the ambition isn't there, but because nothing has forced them yet. That's the discipline, not a placeholder for it.

Questions worth answering directly

How much of this system is actually visible on this page?

Three real layers: what the conductor's own contract skill contains for each team, the real agent and skill counts behind each team, and how the pieces sequence and run in parallel. A fourth layer, the actual skill files and agent instructions, is never shown, on this page or to the conductor itself.

Why doesn't the conductor just read a team's internal files directly?

Because its own instructions forbid it. That single rule is what lets any team be rebuilt completely on the inside without ever touching the agent that calls it. Only the contract has to hold still.

How many distinct skills exist across the eleven available teams?

69, combined across engineering, QA, security, design, market research, ideation, and copywriting. None of those 69 live inside the conductor's own 15, which are contract skills about the teams, not expertise borrowed from them.

Is this depth independently checked, or is it self-reported?

Checked. A standing team-auditor agent, separate from the conductor, found one real documentation bug and one real safety gap, an unbounded orphaned-container risk, in the newest security team, after two review rounds had already passed it.

This sounding is one artifact of a larger practice.