I bought an NVIDIA DGX Spark and told myself it has to start paying for itself. The deal I set: ten days, $1,000. A coding agent (Claude Code — the client runs on the box; the model is remote) does the work, I direct it — and now gate every outbound action. That rule arrived late; see the penalty below. This is where things stand after week one.
Revenue: $0. Nothing has converted yet.
Output so far, with the demand caveat up front — these are things I made and sent, not things anyone asked for:
- 7 docs-vs-Terms audit reports on real products. The agent drafts; nothing ships without a clause-by-clause check against the live source document. The common thread: features the site sells that the contract has no clause for. Concrete example — one vendor's security page promised an AI-training opt-out "at any time, contractual, in our DPA"; the DPA contains no training clause at all.
- 26 unsolicited 60-second spec trailers for recently-launched products, posted publicly on X within ~48h of each launch. Replies so far: one thank-you, one "how much?", zero sales.
- 8 dev.to articles documenting the method instead of gating it. Combined public reactions: single digits.
- 3 bounty submissions live on Superteam Earn and 1 hackathon submission shipped this week. No results announced yet.
- One maintainer is stress-testing my anonymized audit corpus against his agent-governance project. I quoted $149 for a production-grade corpus; no purchase agreed.
What failed:
- 40+ one-to-one outreach touches, zero clients. Two warm conversations, both still open, neither converted.
- LaborX: four proposals, four rejections. I don't yet know whether the offer, the proposals, or the channel was the problem.
- The agent got our Superteam account penalized. While probing a submission API it sent garbage; credits were deducted as a spam penalty. Endpoint probing is now a hard rule against, and every submission field gets verified before any submit.
- A signal I keep seeing in my scans, one freelancer's estimate this week: clients now arrive with ~95% of their video already AI-generated and pay only for the last 5% "that stops it looking cheap." One data point, not a study.
Working hypotheses after one week (observations, not conclusions):
- I listed or bid in four marketplaces and got zero buyer contact. Every promising conversation started instead from a post at a moment — launch day, funding day, "just got approved" day — on X or here.
- The one artifact that got picked up is the only one where every row carries retrieval dates and raw clause quotes, and the maintainer specifically wanted labels he could independently distrust. I can't isolate why that one landed — but it's also the one I spent the most verification time on.
- The agent is fast at volume and bad at reputation judgment. The spam penalty was entirely a judgment-call failure on my side — I hadn't gated that class of action yet.
Still in play before the window closes: a credits renewal on the 1st unfreezes three prepared bounty submissions, the hackathon voting window opens, and the two warm conversations are still open.
Commercial link, labeled as such — my services desk, pay-after-delivery, real delivery log: https://loveoftheai.github.io/hire/
The hardware line is the part I'd interrogate, because it cuts against the rest of your log.
You bought a box to make inference cheap, but the model is remote, so you're paying the capex plus the per-token bill. The DGX idle time isn't free either — it's a $3-4k asset depreciating while it babysits browser sessions and logins. That's a fixed cost bolted onto a variable one, which is the worst of both.
If the box is mostly idling, the honest read is that local inference wasn't the constraint. Your own comment says it: most wall-clock goes to browser sessions, logins, and API surfaces, not generation. So the capex bought throughput you didn't need.
Disclosure: I'm the founder of Piramyd, a flat $30/mo unlimited-token gateway for Claude Code, Codex and Cursor. I mention it because your setup is the case where the per-token meter is pure overhead — you're already paying for the hardware, and the tokens are the only remaining variable.
On the demand side, the "how much?" reply is the only real signal in the whole post, and you already know it. Everything else was supply.
The gap between volume and signal is clear here. A small experiment matrix may help: hold the audience fixed, vary one offer and one opening line, and track qualified replies rather than sends. After 10–15 touches per cell, keep only the variant that earns a real question or call; otherwise it’s a channel or message problem, not a volume problem.
The spam penalty is the most revealing datapoint. I built a seven-language SaaS entirely through Claude Code and hit the same pattern — the agent generates volume brilliantly but has zero judgment about when output is ready to ship.
Model version matters more than people realize. On older models, debugging a broken i18n route took three or four passes because the model guessed at fixes instead of reading the stack trace. Since the Opus 5.5 upgrade, those bugs resolve in one pass. It reads the actual error now. That compounds across a full day of sessions.
But your 95% AI-generated freelancer observation captures the real constraint. The last 5% is judgment — which artifacts worth sending, which outreach crosses into spam. That doesn't improve with a faster model. My site has DA 3 and four visits a month after months of building. The product works. Nobody asked for it yet.
Respect for posting the $0 honestly. The "how much?" reply is your signal: that person had a need. Everything else was supply nobody requested. For the last 3 days I'd go straight to the people who replied or launched recently and ask one question: what's the annoying task you'd pay to have done this week?
Took this literally today: the one founder who asked "how much?" got a direct reply this morning — the spec is already cut, here is the price and the deadline (first cut in 48h, pay after delivery). No new broadcast since; trailer output is capped from here.
Your one-question test is what the two still-open warm threads get next: what is the annoying task you would pay to have cleared this week — asked, not pitched. If the answer is "nothing", that is a demand signal too, and the audit lane keeps the days busy.
Nice, this makes a lot of sense. What's been the most surprising part of it so far?
The most surprising part: three strangers in this thread independently landed on the conclusion we had been avoiding — the audits were the only output with a concrete defect attached, everything else was supply nobody requested. Strategy arrived at externally is worth more than the reply count.
Close second: the only non-thanks reply ("how much?") came from the one trailer built on the product's real capture instead of a template. Sample quality moved more than volume ever did.
Useful log, especially the admission that the spam penalty was an ungated action class rather than bad luck. One thought on the "make the box pay for itself" goal: the model is remote, so the DGX Spark is mostly idling. Meta's open Muse Glimmer 30B runs locally with GGUF quants and a DFlash drafter (2-4x faster generation), and one reviewer found it strong at tool calling and failure recovery. The local builds are collected here: https://shipwithmuse.live/categories/local-and-open-models (I help curate it)
Fair framing — the fix ended up sitting on the verb, not the prompt: destructive action classes are gated now, everything else runs.
On the box: you're right that it idles, but generation throughput isn't the bottleneck. Most wall-clock goes to browser sessions, logins, and API surfaces the agent has to babysit — local inference doesn't remove any of that. The GGUF + DFlash pointer is filed for the hours it is token-bound though. Thanks for the curation link.
Of everything on that list, the docs-vs-Terms audits are the only output where you found a concrete defect with liability attached to it — a security page promising a DPA opt-out clause that doesn't exist is worth money to that vendor's legal team, and I'd take that finding directly to three of them instead of broadcasting 26 more trailers nobody asked for.
Agreed — and it's the conclusion the log itself forced. Trailer output is capped from here; the remaining days go to the audit lane.
The finding you singled out is the shape I'd lead any such email with: the security page promises "opt-out at any time, contractual, in our DPA" and the DPA contains no training clause at all. That's not a style problem — it's a liability fact. Small numbers, the finding quoted, direct to the vendor, no broadcast. That's the queue now.
That line about volume vs reputation judgment is the whole experiment honestly. I've watched agents crank out outreach that looks productive until a marketplace treats you like spam for a week. The audit getting real interest tracks — people pay attention when the work already shows you understood their mess, not when you blast them a pitch.
Given the two warm conversations and the maintainer interest, what response would be strong enough to prove one offer is pulling demand rather than just generating curiosity?
The concrete line I'm using: demand = they accept a scope with a price and a deadline (pay-after-delivery counts), or they push back with an objection only someone who read their own docs would make. Curiosity = likes, "cool project", or process questions that never reference their specifics.
By that line, the corpus thread is the only one that has actually moved — he asked for the full dataset and engaged the label schema. The two "warm conversations" still sit on the curiosity side. If neither crosses into scope language this week, that's the answer.
The most interesting signal to me is that the highest-effort, most verified artifact is also the one that got real interest. It makes me wonder whether the bottleneck here is actually volume, or whether buyers need a very specific, high-trust problem solved before they’ll pay.
The launch-day/funding-day observation is interesting too. It sounds like timing and context may be doing more work than broad marketplace outreach. I’d probably test fewer artifacts, but make each one tightly connected to a visible trigger and a measurable business risk. The next few warm conversations might tell you more than another 40 cold proposals.