I kept watching AI coding tools confidently produce wrong code. Not wrong in an obvious way — wrong in a "this will work until it doesn't" way. There was no gate. No verification. Just a retry loop and hope.
I wanted a system where output was checked, not assumed. Where you could point to exactly where a task failed and why. Where the AI was amplifying my judgment, not replacing it.
So I built Orca.
What it is
Orca is a desktop AI orchestration engine for Windows. Instead of one large expensive model trying to do everything, it routes tasks through named specialist roles — each with a defined contract (input schema, output schema, acceptance criteria). The roles don't care which model sits behind them. Swap cheap local models for expensive frontier ones and the contracts stay clean.
The roles have names because the names matter to how I think about them:
Brain — task decomposition and routing
Benson — user-facing output
Miranda — budget enforcement and permissions
Pappy — quality gating and output verification
Dewey — user context, persistent across sessions
Moonshiner — distillation pipeline (trains better jars from Pappy-verified runs)
Pappy is the core differentiator. Every output gets verified against acceptance criteria before it advances. Bad runs get rejected before they become training data. It's the difference between "probably right" and "we checked."
The long-term vision: small specialist models ("jars") that get better over time because Pappy only lets clean batches through.
The constraints I built under
I have ADHD. I work in limited sessions around family responsibilities. I don't write code directly — I work with AI coding agents (Claude Code primarily). I'm based in Eastern Kentucky with no local technical peers.
What I discovered building this way: constraints forced clarity. Every session had to be scoped. Every agent prompt had to be precise. I couldn't afford to "figure it out as I go" because there was no continuous thread to pull on. So the architecture had to be modular and documented well enough that I could pick it back up cold.
Where it stands
v1.2.15 shipped: github.com/junkyard22/Orca
620 tests passing across 12 packages
Windows installer + portable .exe
Free — building traction and real usage data before any monetization
Revenue: $0
I also published the Agent Handoff Protocol (AHP) as a standalone open spec — typed validated packets for agent-to-agent communication. The npm runtime (@marsulta/mailman) is live and the three-agent proof of concept runs clean.
The pattern I keep noticing
While building this, I've watched well-funded teams independently ship individual pieces of the architecture I'd already built — quality gates, governance layers, orchestration infrastructure. I document it rather than spiral about it. The integrated version is still the thing nobody has shipped.
What I'm looking for
Feedback. Users. People who've felt the pain of "probably right" in their own AI workflows.
Most developers eventually hit a wall where they realize that "vibes-based" coding with AI is great for demos but a nightmare for maintaining a stable production codebase. It is incredibly stressful to wonder if a hidden hallucination in a generated script is going to silently corrupt your data or crash your build when you aren't looking. Since you've built "Pappy" as the gatekeeper for quality, does the verification process rely on static analysis and unit testing, or does Pappy use a separate LLM pass to "reason" through whether the code actually meets the contract?
Great question. Pappy uses a separate LLM pass. It receives the original task contract (scope, acceptance criteria, expected output schema) alongside the agent's output, and reasons against those criteria rather than doing static analysis. Think of it less like a linter and more like a senior reviewer who was handed the ticket alongside the PR.
The key design decision was keeping Pappy's judgment separate from the worker's execution. The worker doesn't grade its own homework. Pappy gets the contract and the output cold, with no knowledge of what the worker "intended."
Static analysis and unit test results can feed into the contract as evidence, but the verdict is Pappy's reasoning pass, not a rule engine. That's what makes it catch the class of failures you're describing: things that are syntactically fine, pass tests, and still don't actually fulfill the requirement.
Having Pappy act as a senior reviewer instead of a linter is a clever way to catch those logical gaps that pass syntax checks but fail the actual requirement. Keeping the judgment separate from the worker is the only way to break that retry loop and hope cycle that makes most AI tools feel unreliable for serious work.
It reminds me of my approach to Digital PR and Media Placements on sites like MSN or Bloomberg. We don't just aim for a live link and call it a day; we focus on a verification layer to ensure that the placement actually builds long term authority and trust for the brand. In both orchestration and PR you need that objective gatekeeper to make sure the output matches the high intent of the original goal.
That separation of intent and execution is definitely the secret to moving past the vibes based era of AI development.