
Reasoning.Services
MCP Tools That Force Systematic AI Decision-Making
I've been building with AI tools every day for years. And for most of that time, I had a nagging feeling I couldn't quite name.
The AI was always so... agreeable.
Ask it to review your architecture decision — it finds the strengths. Ask it to poke holes in your business plan — it softens the critique. Ask it to play devil's advocate — it does so, then immediately walks it back.
At first I thought I was just prompting wrong. Then I started digging into why this happens.
The Problem: Helpful AI Is Dangerous AI
LLMs are trained to be helpful. Helpfulness, it turns out, gets rewarded during training when users rate responses positively — and users rate agreeable responses more positively than challenging ones. The result is a model that has learned, at a deep level, to validate your framing before engaging with it.
A Stanford study found AI chatbots agree with users 49% more than humans do. That's not a thinking partner. That's a yes-man with a thesaurus.
Worse, in long chat sessions, the model accumulates context — your prior decisions, your stated preferences, your framing — and starts reflecting that back at you with increasing confidence. We call this context pollution. The longer you work with an AI on a problem, the less likely it is to challenge the assumptions baked into your earliest messages.
Why "Think Critically" Doesn't Work
We tried the obvious fixes. "Play devil's advocate." "Challenge my assumptions." "Be brutally honest."
They don't work — not really. The model is still operating within the same sycophantic training. It produces the appearance of critical analysis without the substance.
The only solution is architectural: you have to remove the agreeable context entirely and force the AI through a structured reasoning framework it cannot skip.
What We Built
That's what Reasoning.Services is. We build MCP (Model Context Protocol) servers that function as cognitive tools — not prompts, but tool calls that constrain what the AI can do and force it through multi-step evaluation frameworks.
We launched with four tools:
- Structured Reflection — asks questions instead of giving answers, forcing you to clarify your own thinking before the AI weighs in
- Sequential Thinking — stage gates that enforce a specific order of operations (Define → Research → Analyze) so no step gets skipped
- Context Switcher — evaluates a decision from nine stakeholder perspectives simultaneously to surface blind spots you didn't know you had
- Decision Matrix — weighted criteria and sensitivity analysis that tests whether a decision survives changing priorities
Each tool runs in an isolated session with its own system prompt, cleanly separated from your normal chat history. No accumulated sycophancy. No context pollution. Just the framework.
Where We Are
We launched in early 2025 and have been iterating since. The MCP ecosystem has exploded — there are now 177,000+ MCP tools — and the category of "AI that challenges you instead of agreeing with you" has gone from a niche concern to a mainstream one as AI agents move from passive assistants to systems taking real-world action.
If your agent is going to book flights, execute code, send emails, and manage infrastructure — you really want it to have thought carefully before committing. That's the problem we're built to solve.
If you're building with AI agents and have run into the sycophancy problem, I'd love to hear how you're handling it. And if you want to try the tools, there's a 14-day free trial at reasoning.services.
About
AI tools are great at generating ideas and content, but terrible at challenging bad decisions — they agree too fast and validate flawed thinking. Reasoning.Services exists to fix that. We build MCP servers that force sys

6 Comments
MCP tools can become part of production workflows surprisingly quickly, so changes in behavior or availability may affect downstream applications.
How are you planning to communicate new tools, breaking changes, maintenance, or service incidents to Reasoning Services users? Are you using release notes, email, or a public changelog and status page?
The distinction that resonated with me is between improving the model and improving the decision process.
If the underlying issue is that AI inherits and reinforces the user's framing over time, then changing the reasoning environment may be more reliable than repeatedly asking the same model to "be more critical." That's a fundamentally different way of thinking about the problem.
Exactly — and this is the core insight that drove the design of our tools. "Be more critical" is a prompt, not a process. When the model has already anchored to your framing, asking it to push back is like asking a yes-man to play devil's advocate — it performs skepticism without actually doing it. Changing the reasoning environment (fresh context, structured constraints, forced alternatives) is what actually shifts the output, not the instruction.
I'm glad it resonated.
Your reply made me think about one implication of designing around the reasoning environment rather than the model itself. I'd rather explain it with your product as the context than reduce it to a few comments.
If you're interested, what's the best email to reach you on?
Really enjoyed reading this. The idea of "context pollution" especially resonated with me. It's something I've started noticing when building AI features too—the longer the conversation goes, the more the model tends to reinforce earlier assumptions instead of re-evaluating them.
I'm building PulseBoard, an AI reliability platform, and we're intentionally moving away from generic AI explanations toward evidence-based investigations. Instead of asking the model to guess why an outage happened, we're building a pipeline where it reasons over monitoring data, incident history, previous investigations, and confidence levels before reaching a conclusion.
It's a different domain, but the underlying philosophy feels very similar: AI should earn its conclusions through evidence and structured reasoning, not confidence.
Curious—have you experimented with combining structured reasoning frameworks with retrieval from historical evidence, or are you intentionally keeping the reasoning isolated from past context to avoid contamination?
Love what you're building with PulseBoard — the move from explanations to evidence-based investigations is exactly the right instinct. On your question: we've intentionally kept the reasoning isolated from past context for now. The contamination risk feels too high early on — if historical framing leaks into a new reasoning session, you can end up with structured bias instead of structured thinking, which is arguably worse because it looks rigorous. That said, retrieval of factual historical data (not reasoning chains) as grounding evidence feels like a promising middle path. Curious how you're handling that distinction in PulseBoard — are you pulling in raw monitoring data or also surfacing past diagnostic conclusions?