We’re building Wisdom Gate, an API gateway for LLMs so devs can hit multiple models through one clean interface.
What’s new this week
Claude Sonnet 4.5 is live on Wisdom Gate (Text + Image, up to 1M context).
Pricing: 20% off the official list right now — $2/M input and $10/M output.
Why we’re building this
Managing multiple LLM vendors (auth, pricing, regional routing, failovers) is annoying. We want one key, one SDK, sane routing, and clear cost controls.
What’s hard right now (and how we’re fixing it)
Our free DeepSeek endpoints were hammered by automated traffic (hundreds of thousands of calls/hour). Result: 429/400 spikes for normal users.
Mitigations shipping: more DS capacity, per-IP and per-key rate windows, burst queues, anomaly throttles, and better observability.
Support backlog is catching up; we’re improving status updates and automatic incident notices.
Ask for feedback (would love IH input)
What’s your favorite way to separate free vs paid rate-limits without punishing legit power users?
For models with “thinking” modes, would you rather pay a premium per output token, or a small thinking-token surcharge only when used?
Do you prefer a single aggregated key with routing rules, or per-provider keys exposed in the dashboard?
Roadmap (next 2–4 weeks)
Per-model budgets/alerts
Cached completions and idempotency keys
Better logs: latency and cost breakdown
If you’re curious, try Sonnet 4.5 and tell me where it breaks in your workflow. I’m here all day answering questions and would love brutal feedback. If you need stable access while we scale free DS, the paid route is unaffected.
Thanks for reading — happy to swap notes with anyone building infra or devtools.
(P.S. If you want a quick link or code for the top-up bonus, ping me and I’ll DM.)