1
0 Comments

Rebuilt my AI article generator from one monolithic prompt to a per-section multi-agent pipeline. Quality scores jumped from 6.5 to 8.5+

Quick build-in-public update on Articfly (AI content engine for SEO articles, mostly WordPress).

For the first few months the pipeline was one big "Writer" agent. User picks a topic → one giant prompt → article comes back. It worked. Sort of. Quality was inconsistent — sometimes 8/10, sometimes 5/10, no idea why. Hallucinations were the worst part. The model would confidently invent stats, dates, even sources.

I tried the usual fixes — better prompts, longer prompts, examples in context, temperature tweaks. Marginal gains. The real problem was that one agent was trying to do six jobs at once: research, plan, write intro, write body, write conclusion, format HTML. So it did all of them okay-ish.

So I rebuilt the whole thing. Now the pipeline is:

Analyze Raport → Planner → per-section AI Writer (running per H2) → Aggregate → Clean HTML → Callback

Each agent does one thing. The Analyze Raport agent only researches (Brave Search + Wikipedia + Perplexity). The Planner only builds the section-level brief. Then a Writer agent runs once per section with just that section's brief in context — not the whole article plan, not the research dump, just what it needs to write those 200-400 words.

Three things I didn't expect:

1. Token cost stayed roughly the same. More agent calls, but each one has way less context, so total tokens are flat. Quality jumped from ~6.5 to 8.5-9.0 on my internal scoring rubric.

2. Hallucinations dropped massively. Not because the models got better at being honest — because the Writer agent now sees verified research notes from the Analyze step instead of having to "remember" facts mid-generation. The hallucination surface area shrank.

3. Debugging got easy. When an article is bad now I can see which agent failed. Before, every bad article was a mystery prompt-engineering session.

Cost of doing it this way: it's slower (parallel section writes help but it's still 60-90s per article vs 30s before), and the orchestration layer in n8n is way more complex. If a single agent fails the article fails. So you need retry logic, fallback models, the whole reliability stack.

Anyone else running multi-agent pipelines in production? Curious how you're handling failures mid-pipeline — I'm using Claude Sonnet as primary with a Gemini fallback but it feels fragile.

articfly.com if curious.

posted toAvatar for product Articfly
Articfly