Who Watches the Watcher? Why I Built PM Butler and ComplianceGate — and Why One Without the Other Is Just Wishful Thinking

I want to tell you about a very specific kind of pain.

You're building something. You're in the zone. You've got an AI assistant that seems to genuinely understand what you're doing. It's responding to your prompts with confidence, generating code, tracking decisions, referencing things you said three sessions ago. It feels like having a co-founder who actually reads the documents.

Then you pull back. You look at what you've built. And somewhere in the gap between session four and session nine, the plan changed — and nobody told you. The AI didn't flag it. The AI agreed with the new direction just as confidently as it agreed with the old one. Because that's what it does. It says yes. Fluently. Persuasively. Completely without accountability.

You spent ten hours building something that drifted from your own spec.

That pain is why I built PM Butler. And the thing I discovered after building PM Butler is why I built ComplianceGate.

Part One: The Problem with Building Alone

Most people building with AI right now are vibe coding. That's not an insult — it's a description. You have an idea. You start prompting. The AI generates. You react. You iterate. The session ends. You come back the next day and you're not entirely sure where you left off, what you decided, or why the architecture looks the way it does.

This works fine for prototypes. It is a disaster for anything that needs to actually ship, be maintained, or be explained to another human being.

The core problem isn't the AI. The AI is doing exactly what it's designed to do. The core problem is that there's no structure underneath the conversation. No one is tracking decisions. No one is watching for scope creep. No one is raising a hand when you contradict something you committed to three sessions ago.

PM Butler is my answer to that problem.

It's an ensemble of eight AI personas — each one opinionated, each one fiercely protective of their domain — that work alongside you on any project and produce real project management documentation as a byproduct of doing the actual work. You don't stop to document. The Librarian documents. You don't manually track scope. The Sculptor tracks scope. You don't write your own risk register. The Gatekeeper writes your risk register, flags your risks at three severity levels, and does not let you forget about them.

The Head Butler coordinates the ensemble. The Oracle won't let you build until you've proven the problem exists. The Architect obsesses over the edge case you forgot. The Alchemist protects your developers from mid-sprint surprises. The Shipper lives for launch day. And the Gatekeeper — the Gatekeeper has zero patience for 'we'll fix it in post.'

Every prompt you send produces three things: the work you asked for, updated documentation, and any concerns the relevant personas want to flag. Eleven living documents. No post-mortem regret.

I built this because I needed it. I needed something that would make accountability a feature of the workflow, not an afterthought.

And then I realized I'd only solved half the problem.

Part Two: The AI Said Yes to Everything

Here's what I discovered after deploying PM Butler: the personas are excellent at flagging what should happen. The Gatekeeper is meticulous. The Sculptor tracks scope with a running percentage index. The Oracle logs every decision with rationale, alternatives considered, and reversal conditions.

But none of that matters if the actual code — the thing being built — never gets checked against those documents.

The AI is not your project manager. It's a very confident collaborator with no memory between sessions, no accountability for outcomes, and no skin in the game. It will help you write the requirements document and then, three sessions later, help you build something that violates half of them, without ever mentioning the contradiction. Because you didn't ask. And it doesn't volunteer.

Who watches the watcher?

That question is the origin of ComplianceGate.

ComplianceGate is an open-source compliance layer that installs directly into your repo. One command:

npx compliancegate install

It runs automated checks against your codebase using a set of configurable agent skills — not against what the AI claimed the plan was, but against what the plan actually says. The rules. The requirements. The decisions someone logged in DECISIONS.md before the session started.

It's the enforcement layer. PM Butler tells you what the rules are. ComplianceGate tells you whether anyone is actually following them.

Together, they close the loop that most AI-assisted development workflows leave wide open.

Part Three: What This Actually Looks Like in Practice

Let me be concrete.

You start a project. You bring in the Oracle to validate the problem — is this worth building, is the market real, what are the assumptions that would kill this if they're wrong. Oracle logs the decision in DECISIONS.md with rationale and reversal conditions.

You move to the Sculptor. Sculptor defines scope — what's in, what's out, what happens if someone tries to add something mid-sprint. Sculptor maintains a scope creep index. The moment scope starts to drift, Sculptor flags it.

The Architect maps the technical approach. The Alchemist protects the sprint. The Gatekeeper runs QA. The Shipper manages the launch. The Librarian documents everything. And at every step, the Head Butler is synthesizing concerns across personas, making sure the right hand knows what the left hand is doing.

Meanwhile, ComplianceGate is watching the repo. Every time you push code, it checks your build against the rules the ensemble established. Scope boundaries. Architectural constraints. Requirements with acceptance criteria. Risk items that were supposed to be mitigated before this phase.

If something drifts, ComplianceGate tells you. In plain terms. With the specific rule that got violated and the specific file where the violation lives.

No dashboard nobody opens. No Confluence page last touched in Q2. No consultant's deck. Just a check, in your repo, on every push.

Part Four: Why I'm Sharing This Now

I'm building in public. That means sharing the things that work before they're polished, the things that failed before I've figured out why, and the logic behind the decisions because the logic is usually more useful than the output.

PM Butler is available now on Gumroad — both the Ensemble edition (eight AI system prompts, routing guide, ensemble dynamics reference, interactive web demo) and the Project Edition (seven personas configured specifically for building alongside you, with full documentation infrastructure).

ComplianceGate is live and open source at github.com/tedrubin80/compliancegate and compliancegate.dev.

These are not finished products. They are working tools. I use both of them. The edges are real and I know where they are.

What they represent is a specific thesis: that accountability has to be designed into the workflow, not bolted on afterward. That the AI being confident is not the same as the AI being correct. That someone — something — has to be the check.

The Uncomfortable Question

Here's the thing I keep coming back to.

Most people building with AI right now are operating on trust. They trust the model to remember the plan. They trust the session to carry context. They trust that the thing being built is the thing they intended to build.

That trust is not always earned.

The model doesn't remember the plan. The session doesn't carry context. The thing being built is whatever the most recent prompts pointed at — which may or may not be the thing you intended when you started.

PM Butler and ComplianceGate exist because I got tired of finding out the hard way.

If you're building something that matters, build accordingly.

—

PM Butler: gumroad.com (search PM Butler)

ComplianceGate: compliancegate.dev | github.com/tedrubin80/compliancegate

@currentlyted on LinkedIn and Substack — The Chaos Translator