1
0 Comments

I scored my AI app 64/100. Here's the architectural mistake I fixed.

My AI fitness app (Pelaris - https://pelaris.io) had a problem I didn't see for weeks.

I'd built a program generator that looked good on the surface. Users entered their goals, the system produced personalised training programs. Reasonable structure, correct terminology, sensible periodization phases.

Then I ran a proper quality assessment. Scored the output against 10 criteria. Got 64 out of 100.

The specific failure that made it obvious: a user training for a major endurance swimming event got a balanced plan with roughly equal splits across swimming, running, and cycling. Three goals were all swim-focused. The plan treated them like a triathlete.

My first instinct was to fix the symptoms. Tighten the prompt, add better rules, inject more instructions.

I started building those fixes. Then I deleted them.

The prompts weren't the problem. The architecture was.

What was actually wrong

The pipeline classified what sport someone did. It never reasoned about what achieving their goal actually required.

There's a big difference. Classifying "this person swims" takes one lookup. Answering "what does this goal physiologically demand, and what training structure produces that adaptation" requires a reasoning step that simply didn't exist.

Every fix I was building was patching the output of a system that had already reasoned incorrectly. The output looked like the problem. The problem was one layer earlier.

What I found in the research

While figuring out how to fix this properly, I dug into how the established platforms in this space actually work.

None of them use LLMs to architect programs from scratch. The ones that have been doing this for years use expert systems, mathematical models, and rule engines. AI, when it appears at all, handles personalisation within a structure that's already been validated.

That distinction - the LLM as the coach filling in a correct skeleton, not the architect building it - reframed everything.

The fix

The pipeline now has a layer that didn't exist before: a focused AI call that reasons about what a goal actually requires before anything else is generated. What energy system. What adaptations. Whether the initially suggested training methodology even fits the event. (Some methodologies are designed for short race-pace efforts. They're wrong for ultra-endurance. The new layer catches this.)

Everything downstream - program structure, weekly targets, session generation - works from that physiological foundation rather than category labels.

Two other findings from the research that changed the implementation:

Forcing JSON output directly in a generation prompt degrades reasoning quality. The model splits attention between format compliance and content. Separating reasoning from extraction - one pass in natural language, one pass to pull structured data from it - produces materially better output. It also makes the reasoning legible, which turns out to be useful for debugging.

The existing validator caught structural failures. It couldn't catch semantic ones. Wrong exercise for the goal. Generic coaching cues. Vague intensity. Added a separate evaluation step using a different model as the judge, with binary criteria, to catch what code can't.

Where things stand

The pipeline changes are built and validated. The before/after assessment is next.

Building in public at pelaris.io. Will post the scores when I have them.

posted toAvatar for product Pelaris
Pelaris