3
5 Comments

After 10 Years Running a Shopify Agency, I Think Small Development Tasks Need a Different Business Model

For years, I treated small Shopify development requests as the least interesting part of agency work. They were the tasks that sat between larger projects: fixing a Liquid bug, changing a product page component, cleaning up structured data, removing a script that was slowing down the storefront, adjusting a collection template, or implementing a small CRO idea that had already been approved internally.

Most of these changes were not technically difficult. An experienced Shopify developer could often complete the actual implementation in a few hours. What took me longer to understand was that the technical work was rarely the real problem. The larger issue was that the traditional agency model is not particularly efficient when the client already knows exactly what needs to be changed.

At Shugert, we still work on the kind of projects that genuinely benefit from an agency structure: migrations, complex integrations, technical SEO, custom development, performance work, and anything where there is real ambiguity around architecture or implementation. Those projects justify senior involvement because the cost of making the wrong decision can be much higher than the cost of writing the code.

Small tasks behave differently.

If a merchant wants a component moved above the buy button, a field added to Product schema, or an approved design implemented in a theme, the outcome is already clear. The work may still pass through intake, review, estimation, scheduling, QA, and client communication, but the surrounding process can easily become larger than the implementation itself.

That is when a simple technical request starts to look like a business-model problem.

Retainers solve some of this, especially for merchants with a steady stream of SEO, CRO, performance, and engineering work. But there is another type of Shopify merchant that does not need a full agency team every month and still has a meaningful backlog of changes that should be shipped.

These are usually not emergencies. They are small improvements that sit in the uncomfortable middle between “important enough to matter” and “urgent enough to force action.” A schema issue can wait. A small performance problem can wait. A PDP improvement can wait. A tracking fix can wait. A collection template can wait.

The store keeps working, so the backlog keeps growing.

Over time, those tasks stop looking small in aggregate. They create technical debt, slow down merchandising, leave CRO ideas unimplemented, and make the storefront harder to improve. The merchant may not need another large project, but they do need a better way to get routine engineering work out of the queue.

That was the point where I stopped thinking of small Shopify development as simply cheaper agency work and started thinking of it as a separate category.

Freelancers are an obvious answer, and for many merchants they are the right one. A good freelancer can be fast, experienced, and economical. The tradeoff is that the merchant often becomes the coordination layer. They have to find the right person, explain the context, manage access, review the result, and decide whether the change is safe.

The quality of the process depends heavily on the individual developer and on how much technical judgment exists inside the merchant’s own team. For sophisticated ecommerce companies, that may be manageable. For everyone else, coordination becomes part of the cost.

Then AI changed the economics of the implementation layer.

Liquid, JavaScript, CSS, JSON templates, schema changes, and common Shopify theme patterns are increasingly easy for capable models to generate. When we started building TaskerArmy, that made code generation look like the obvious opportunity. If a merchant could describe a change and the system could produce the necessary Shopify code, it seemed like much of the friction might disappear.

It did not.

The merchant does not actually want code. The merchant wants the store changed safely.

That distinction changes the product entirely.

The hard questions are not whether a model can write a Liquid snippet. They are whether the request is clear enough to execute, whether the right theme and files are being modified, whether the change conflicts with existing customizations, how the result should be validated, what happens if publishing fails, and when the system should stop and escalate instead of continuing automatically.

Those questions pushed us toward what we now think of as bounded engineering.

A bounded task has a clear objective, a limited scope, and an outcome that can be validated without forcing the system to make a broad business decision. Adding a field to Product structured data is relatively bounded. Implementing an approved PDP component can be bounded. Removing an unused script is usually understandable.

“Improve our conversion rate” is not bounded. Neither is “fix our SEO,” “migrate our ERP,” or “redesign our product experience.”

Those are judgment problems before they are implementation problems.

The more we work on this, the more I think the boundary matters more than the automation itself. There is a tendency in AI products to treat increasing autonomy as the goal, but an engineering system may become more useful by becoming better at recognizing what it should refuse to do.

If a Shopify task is clear, contained, and testable, software can potentially handle much of the workflow. If the task becomes architectural, ambiguous, or commercially risky, escalation should be part of the design.

That does not make products like TaskerArmy replacements for agencies. A better framing may be that they remove a class of work that never really needed an agency-style process in the first place.

There is still enormous value in senior engineering for migrations, B2B implementations, complex integrations, technical SEO strategy, international architecture, and projects where the client does not yet know what the right solution should be. Using senior engineers for predictable storefront changes because there is no more efficient operating model is a different issue.

The business model around this is still something we are figuring out.

Hourly billing is familiar, but it keeps the customer focused on developer time. Fixed project pricing makes little sense when the “project” is a small theme change. Unlimited development subscriptions sound attractive until usage becomes unpredictable. Credits or engineering runs are easier to understand, but engineering work is still variable, so the product needs clear boundaries without recreating the same estimating machinery it was supposed to replace.

That is what makes this category interesting to me.

The opportunity may not be simply to make development cheaper. It may be to redesign the transaction itself.

If a merchant can submit a clear Shopify task and trust that it will be implemented, validated, and shipped safely without turning the request into a miniature project, that begins to look like a different kind of service.

It is not conventional SaaS because the customer is buying an outcome more than a tool. It is not really an agency because many tasks should not need a project team. It is not a freelancer marketplace because the merchant should not have to coordinate individual developers. It is not fully autonomous software because there are still cases where human engineering judgment is essential.

I am not sure the category name matters much yet.

What matters is whether the economics work and whether merchants trust the workflow enough to use it.

After years of watching Shopify backlogs grow for reasons that have very little to do with the difficulty of writing code, I am increasingly convinced there is a real problem here. The interesting question is no longer whether AI can generate Shopify code. That part is becoming ordinary.

The harder question is whether we can build an operating model around that capability that is cheaper than traditional development, safer than blindly generating code, and simple enough that merchants actually use it.

That is what we are trying to figure out.

If you run an agency, manage a Shopify store, or have tried productizing technical services, I would be interested in where you draw the line between work that can be standardized and work that still needs senior human judgment.

on September 7, 2026
  1. 1

    I would measure the human touches after a task is accepted, alongside implementation time. A narrow change can still become expensive when the brief misses an approval or someone changes the acceptance criteria. That seems like a useful way to test per-task pricing before committing to an unlimited plan.

    I am working on the intake side of this with AiZipZap (my product: https://aizipzap.com/agent?utm_source=indiehackers&utm_medium=community&utm_campaign=intake_launch). It produces source excerpts, questions and a reviewed handoff; it does not touch the store or send anything. An early test let instructions inside the customer request contaminate the draft, so the application now controls the quoted source references. That still does not make the model’s interpretation reliable enough to skip review.

    No paid-demand result to report yet. The experiment is whether a better brief reduces clarification work enough to justify a scoped service. In your backlog, which creates more rework: missing context at intake, or changes after the merchant sees the result?

    Disclosure: I used AI assistance to build the tool and prepare this comment.

  2. 1

    The pricing question and the category question are the same question, and I learned that the expensive way.

    My product sits in a space where the nearest mental category is dominated by free tools, and as long as buyers filed it under that category, every price looked outrageous. Nothing about the product changed when the price finally started making sense to people. The description did.

    Bounded Shopify tasks have the same exposure. If the unit is "a small dev task", the anchor is a freelancer's hour, and you get squeezed toward the bottom no matter how good the operating model is. But what the client buys from you is not the schema change. It is the guarantee that the change ships safely on a live store that makes money, with someone accountable if it does not. Agencies price judgment, freelancers price hours, and the gap you are describing prices certainty: a flat rate per bounded task, validated, warrantied, with a clear line on what falls outside.

    When a client already knows exactly what needs changing, what do they say they are afraid of? And what did your riskiest small task cost when it went wrong?

  3. 1

    Your stated question (where's the line between standardizable and judgment work) and your unsolved one (the business model) are the same question. The line IS the pricing model, and you're stuck because every option you listed prices the work — hourly, credits, subs all pay for execution — when you should price the boundary itself.

    Here's the resolve: "bounded" isn't only a safety property, it's a cost property. A task with clear scope and validatable output is cost-predictable by definition. Unbounded-in-scope and unbounded-in-cost are the same thing. So the refusal mechanism you built for engineering safety is also your margin control, for free — same line, two jobs.

    Which kills the pricing paradox. "Unlimited subs get unpredictable" only if you let unbounded work in. A sub covering only bounded tasks is predictable by construction. So: flat price for anything that passes the bounded test, honest handoff to agency pricing for anything that fails. You don't price the task, you price which side of the line it's on.

    And the same refusal does a third job — trust, your other open question. A system that says "too risky to standardize, this needs a human" is more trustworthy than one that attempts everything, which is the only reason a merchant lets it touch a live store. Safety, pricing, and trust are all one line.

    So the real question: can you write the bounded test crisply enough that a merchant predicts, before submitting, which side their task lands on? That predictability is the product. What percent of your actual backlog sits cleanly on the bounded side?

  4. 1

    The line I would draw is not task size or whether AI can generate the code. It is whether the request has a testable outcome and a named decision owner.

    A request like ‘move this component’ can still be unbounded when device states, theme variants, rollback conditions or approval criteria are missing. Before standardising it, I would require five things: the exact outcome, affected surfaces, constraints and access, the acceptance check, and the person who can approve it.

    If any of those are unknown, the job still needs discovery or senior judgment. Once they are explicit, per-task pricing becomes much more credible because the uncertainty is no longer hidden inside the implementation fee.

  5. 1

    The bounded-task distinction seems stronger than the AI-code-generation angle. Have you seen merchants naturally prefer a particular transaction model—per task, credits, subscription, etc.—or is pricing still being shaped around how predictable their backlog actually is?