Quick one for the community. I've been building Kilawatt Cloud (multi-provider GPU orchestration, kind of a load balancer across RunPod, Vast.ai, etc.) and this week we got x402 fully live, not just tested.
For anyone who hasn't run into it yet: x402 revives the old HTTP 402 "Payment Required" status code. An AI agent hits our API, gets a 402 challenge back, pays instantly in USDC on Base, and gets a running GPU instance in the same request. No card, no signup, no human in the loop.
We ran two real settlements today. Real tx hashes, real Stripe payouts on the other end. Not a testnet demo.
Why I think this matters beyond just us: agents are becoming real economic actors that need to pay for infrastructure autonomously, and almost nobody's built the compute-grade version of this yet, most x402 integrations so far are small stuff like API metering or content paywalls. GPU compute is a much bigger, more capital-intensive use case to get working end to end.
Happy to answer questions about the implementation, the CDP/Coinbase side, or what broke along the way (a lot did).
Damian — the more I look at the architecture you've described, the more I think there's something worth doing here beyond trading comments.
You now have a genuinely consequential chain running live:
agent payment → internal credit → provider provisioning → automatic failover → successful compute → billing.
What I've been building with OpsWatch is almost exactly the independent control layer for the failure boundaries inside that chain — particularly where payment is established but provisioning is not, a provider fails after dispatch, failover creates a fresh attempt, or the system has to decide whether it is safe to retry, bill, continue or leave the balance untouched.
Rather than another theoretical discussion, I'd be interested in independently stress-testing one bounded Kilawatt flow against the actual evidence your system produces.
The output wouldn't be a generic audit. I'd give you a frozen determination of what the evidence can establish at each boundary, where uncertainty survives, whether retry/failover can create duplicate consequence, and what evidence should gate the next autonomous action.
If it exposes nothing material, that's useful too.
But if it does, we would have a concrete control problem to solve before this gets much larger.
Would that be commercially useful enough to Kilawatt for us to scope it as a paid piece of work?
Appreciate the specific offer, and I like that you're thinking about the exact failure boundaries rather than a generic audit. To keep this fair and consistent with how everyone else engages with the platform, I'd want you to go through it the same way any real user would, create an account, fund a wallet, and run it live. That's the version that actually proves something, since it's real money moving through the same rails everyone else uses, not a special back door. Once you've got results from that, happy to talk about what a larger paid engagement could look like if it turns up something worth digging into further
That makes sense — and I’ve now tested the distinction rather than just reasoning about it.
I put a live x402 transaction through Base mainnet with real USDC, independently verified settlement on-chain, then deliberately removed the application-level consequence evidence after the paid response.
That produced an interesting state: payment was conclusively established, but the consequence the consuming agent needed was still unresolved.
I then attempted the next decision — retry the paid action. Because the previous consequence was unresolved, the retry was denied before dispatch. No second payment occurred.
So I think your internal ledger solves an important part of this: settlement does not automatically consume value when provisioning fails.
The remaining question I’m interested in is one layer further downstream.
Suppose the ledger is correct, the provider reports successful provisioning, but the agent cannot independently establish the state it actually needs before deciding what to do next. What evidence in Kilawatt tells the agent it may safely continue, retry, compensate or close?
That’s the boundary I’ve now got working in OpsWatch.
If Kilawatt already exposes enough state to answer that, I think we could test the two systems against each other rather than debate it theoretically.
Appreciate the real test, and appreciate that you actually funded it and ran it straight rather than asking for a workaround. Good result on my end too, the ledger holding through a deliberately corrupted consequence is exactly the behavior it's supposed to have.
On the deeper question you're raising, whether an external agent can independently verify state well enough to decide continue, retry, compensate, or close, that's a fair thing to want answered with data instead of theory. I'm currently closing a contract at the moment, so no full joint testing project on my end right now, but here's what I can do: send me a specific question or a redacted log row when you have one, and I'll check it against the real data and give you a straight answer.
Good work on the test.
Damian Dixon
CEO, Kilawatt Cloud
hello@kilawattcloud.dev
Appreciate it, Damian. I have the specific question now.
From the live x402 test, I can independently establish that 0.001 USDC settled on Base from my payer address to Kilawatt’s payment address and that the paid request returned HTTP 200.
Now assume my agent loses the response body immediately after that point. From its side, the application consequence is therefore unresolved.
Bounded question: using Kilawatt’s existing production records only, what evidence can you return for that specific paid attempt that would allow an external agent to independently determine whether the purchased operation actually occurred — as distinct from merely proving that payment settled or that Kilawatt accepted the request?
In particular, is there an identifier or record that binds:
on-chain payment → Kilawatt job/request → resulting provisioned resource or failed outcome
strongly enough that I could independently classify that attempt as completed, failed, or still unresolved before authorising another spend?
I’m not asking you to change anything or expose sensitive data. A redacted record or the fields Kilawatt already retains is enough.
The transaction from my live test was:
0x902c17aff669aaca07b86c9802fe930886eed9d85b37ba5a5e2f08c845e46dad
If you can trace that through Kilawatt’s side, we can test the boundary against the real record rather than theorising about it.
I checked this directly against Kilawatt's production records and independently on chain, block level, not a summary.
Your transaction is real. It settled successfully on Base at block 51812466, timestamped Sep 26 2026, 09:17:59 AM UTC, which is 2:17:59 AM Pacific. Transaction hash 0x902c17aff669aaca07b86c9802fe930886eed9d85b37ba5a5e2f08c845e46dad. Amount 0.001 USDC.
Here is the problem. That USDC moved to wallet address 0x7dd5Be069f2d2eAd75eC7C3423B116fF043c2629. Kilawatt's actual receiving address, the one our system has on file and has never changed, is 0xf441ad71d29502c33ef911de45aaec0060b34c66. Those are two different wallets. Your payment never reached Kilawatt.
That fully explains why we have zero record of it anywhere in our system. It's not that we failed to log a real payment to us, it's that the payment never arrived here in the first place.
I checked the destination wallet directly too. It's an ordinary personal wallet holding a mix of tokens, not a Coinbase facilitator, not anything tied to our infrastructure.
Our docs are explicit that the payTo address must be read fresh from the live 402 challenge response on every request, never hardcoded or reused, specifically to prevent this exact situation. So the real question is where you got the address you paid. If it came from anything other than a live call to our endpoint at the time of payment, that's likely the source of the mismatch.
Happy to help track down where the disconnect happened if you send me how you obtained that address.
That's a materially different result — thank you for checking it against the production records.
I didn't hardcode the payTo address.
The address 0x7dd5Be069f2d2eAd75eC7C3423B116fF043c2629 came directly from the live HTTP 402 challenge returned by the endpoint I called immediately before payment:
https://x402engine.app/api/crypto/price?ids=bitcoin
The challenge specified:
network: eip155:8453
amount: 1000
asset: Base USDC
payTo: 0x7dd5Be069f2d2eAd75eC7C3423B116fF043c2629
My x402 client then satisfied that challenge. The resulting transaction was the one you just independently verified:
0x902c17aff669aaca07b86c9802fe930886eed9d85b37ba5a5e2f08c845e46dad
So I agree with your conclusion that I should not describe this transaction as a payment to Kilawatt.
But that creates a more interesting evidence question.
If the endpoint represented itself through a live x402 challenge as requiring payment to that address, my agent followed the live payment instruction correctly and received HTTP 200 — yet Kilawatt's production system says the payment and recipient have no relationship to Kilawatt.
Where, then, does x402engine.app sit in relation to Kilawatt?
If it isn't a Kilawatt-controlled endpoint, then I've tested a third-party x402 service rather than Kilawatt and I'll correct the scope of my record accordingly.
If it is connected to Kilawatt in some way, then we've found a much more significant provenance issue: the agent successfully satisfied a live payment challenge without being able to establish that the payment destination belonged to the service it believed it was purchasing from.
Either answer is useful. I want to establish the relationship before drawing any further conclusion.
Appreciate you running this down instead of just taking a swing at us. That's exactly the kind of scrutiny this space needs more of.
To put it plainly: x402engine.app has no relationship to Kilawatt. Not a partner, not white-labeled infrastructure, not anything downstream of us. First time that domain has ever come up on our end. Whatever's running behind it, it's a completely separate x402 implementation that happens to speak the same protocol we do.
Our actual production endpoint is:
POST https://www.kilawattcloud.dev/api/public/x402/exec
Full docs: kilawattcloud.dev/docs/x402
If you point your agent at that endpoint, the 402 challenge it returns will carry our real payTo address (0xf441ad71d29502c33ef911de45aaec0060b34c66), and you can verify that yourself before any funds move, straight from the challenge response, no need to trust us on it.
Based on what you've described, my best read is your test had a stale or cached address pointed at the wrong service rather than anything wrong with our docs or endpoint. Genuinely, thanks for being fair about correcting the record. That's the kind of exchange that makes testing like yours worth having in public.
Confirmed. I’ve now hit Kilawatt’s actual production endpoint myself.
The live 402 returned your stated
payToaddress —0xf441ad71d29502c33ef911de45aaec0060b34c66— so I’m satisfied the earlier transaction had nothing to do with Kilawatt. That attribution was mine, and I’ll preserve the correction in the evidence record rather than rewrite what happened.Interestingly, correcting it led me to a better test.
I pulled the OpenAPI contract from the live Kilawatt endpoint. The 200 boundary is very precise: “Payment settled and job dispatched.” It returns the
job_id, provider and payment evidence.So here’s the question I’d actually like to test against Kilawatt:
After that 200, what evidence can the paying agent obtain that establishes what happened to that specific
job_idafter dispatch?Your schema mentions an
api_keyfor follow-up calls, which makes me suspect there’s another part of the interface I haven’t seen.If there is, point me at it. I’ll test that chain instead of assuming where it ends.
That feels like the more meaningful test anyway: not whether Kilawatt can take an x402 payment, but whether an autonomous buyer can independently establish enough of the resulting state to justify what it does next.
Appreciate you pushing on both of these instead of letting either one slide. Here's the receipt, straight from the live system.
Since your payment never actually reached our endpoint, there's no job under your wallet to show you, but here's a real completed job from our own system, in the exact shape that endpoint returns.
The job ID is ebbbad89-2c26-4209-b0ca-5b2eb701f1c0, status completed. It ran on Node US-West-07, workload type agent_exec, on an RTX A4000 GPU, for 150 seconds. It was billed at $0.01 against a real underlying provider cost of $0.003037, that's the actual margin, not a marketing figure, created at 2026-09-26, 01:56:25 UTC.
On the health side, the instance shows stopped, with a health status of ended_stopped, meaning it completed and shut down normally rather than failing mid-run. Last healthy and last checked both land at 2026-09-26, 02:00:14.117 UTC, and there are zero incidents logged, no downtime, no faults. Because there was nothing to credit, the SLA credit is $0.00 and the net billed amount stays the full $0.01.
That's the full receipt this endpoint hands back on any job, success or otherwise, health data and automatic credits included whenever they apply. Also flagging that endpoint wasn't listed on the x402 docs page, that's getting added now, good catch.
Between where the money actually goes and what evidence exists after it's spent, you've now verified both, and both held up. Genuinely appreciate the rigor, it's the kind of testing that makes a claim like "agent-native payments" mean something. Good talk.
Damian — I think we've actually reached the point you mentioned when we started this.
You said that if I ran this against the real system and it turned up something worth digging into, we'd talk about what a larger paid engagement could look like.
We've now pressure-tested two material boundaries, corrected one false attribution rather than forcing a finding, independently established that Kilawatt's actual evidence chain holds up, and the exercise surfaced a real documentation gap that you're now correcting.
More importantly, I think we've demonstrated what the larger engagement would actually buy you: not an audit looking for faults, but independent evidence that Kilawatt's consequential claims survive scrutiny — and precise findings when something doesn't.
I think that's enough evidence to have the commercial conversation you originally left open.
If you agree, I'll put a tightly bounded paid scope in front of you rather than extending the testing informally.
Fair. That’s actually the cleaner test.
I’ll treat Kilawatt exactly as an external user would — normal account, funded wallet, live transaction, no privileged instrumentation or special path.
I’ll keep the scope narrow: establish what can actually be proven across payment → balance credit → provisioning → drawdown/failure, and preserve anything the available evidence cannot establish rather than filling the gap with inference.
If it produces nothing material, I’ll tell you that. If it exposes a boundary that matters operationally, I’ll bring you the evidence and we can decide whether there’s enough there to justify a paid engagement.
Thanks, Damian — that gives me a clean way to test it.
Damian — the fact you've got x402 → real GPU provisioning working makes me curious about one failure boundary.
If payment settles but the agent can't establish whether the GPU instance was successfully provisioned, what evidence does it currently require before paying/provisioning again?
I'm working on this exact problem with OpsWatch: independently determining whether the evidence from one consequential action is sufficient to permit the next one, particularly where blindly retrying can create a second real-world consequence.
Your live x402 flow looks like a very clean case to test that against.
Payment and provisioning are deliberately decoupled through an internal ledger, which is what sidesteps their exact failure mode. When the on-chain payment settles, it credits an internal balance first. It does not directly pay for GPU time. Only a confirmed successful provisioning draws down that balance. If every provider fails after the payment already settled on-chain, the job is logged as failed with $0 billed, and the credited balance simply stays there, ready for the next attempt. Nothing is lost, and nothing gets double-spent.
That separation makes sense — the internal ledger removes the payment-settled / provisioning-failed ambiguity I was testing for.
The interesting boundary then moves one step downstream.
What is authoritative for “confirmed successful provisioning” before you debit the balance?
For example, if a provider accepts the request and returns an instance ID, but the GPU never becomes usable or the agent loses confirmation before it can establish readiness, do you still treat provisioning as successful?
That seems like the consequential boundary in your architecture: not whether the USDC settled, but what evidence is sufficient to turn the credited balance into a charge.
That reframing is accurate, and it's worth answering directly rather than around.
Today, "confirmed successful provisioning" is the provider accepting the launch request and returning an instance ID. The moment that happens, the instance is marked running and the held funds are settled into a final charge. That happens before any independent check that the GPU is actually reachable or usable, so in the exact scenario you described, provider accepts and returns an ID, but the machine never comes up or confirmation is lost before readiness is established, the charge does go through on the provider's word, not on proven readiness.
What catches that case is downstream, not upfront. Every instance gets a 15 minute boot grace period, and if it hasn't become reachable by the end of that window, or goes unreachable for more than 2 minutes after, the system automatically opens an incident, credits the customer back for that exact window with no claim required, and attempts an automatic failover to another provider. So the failure mode you're describing doesn't result in a silently lost payment, it results in a charge followed by an automatic refund once the system confirms the resource never worked.
So to your framing directly: the authoritative signal for the charge itself is provider acceptance, not proven readiness. The authoritative signal for whether the customer actually keeps paying is the health monitor, which does require real evidence. Those are two different gates today, and you've correctly identified the boundary between them
Really solid approach — I'm juggling something similar myself (building Xstream4K on the side), what's been the hardest part for you so far?