01
Cost
$190kper month
Median client arrives spending $190k/month on inference with 31% of it serving tokens nobody asked for — retries, oversized contexts, uncached prefixes, unprompted 4k completions.
Kusukang—LLM inference scaling · Los Angeles · Jakarta · Seoul
Kusukang is a consultancy that specialises in LLM inference at scale — serving architecture, autoscaling, quantisation and cost engineering for teams whose usage outgrew their first deployment.
We do one thing: the last 90% of the gap between a model that works in a notebook and a model that answers 4,000 requests a second on a Tuesday.
01—Why we exist
Three failures show up in every stack we open. None of them are model failures, and none of them get fixed by a bigger cluster.
01
$190kper month
Median client arrives spending $190k/month on inference with 31% of it serving tokens nobody asked for — retries, oversized contexts, uncached prefixes, unprompted 4k completions.
02
1.4–3.2s
Time-to-first-token is the user-visible metric. It is usually 1.4–3.2s in a stack that could sit under 300ms, because batching, KV cache reuse and prefill/decode separation were never tuned together.
03
4×evening surge
Load is spiky and GPU capacity is not. Quotas, cold starts and a single-region dependency turn a 4× evening surge into a 40-minute outage.
Every engagement starts with the same artefact: a profile of your stack that says where the milliseconds and the dollars actually go.
02—Services
Every engagement starts with the profile. What happens after that is your call, priced against a named scope.
01
Kernel-to-gateway profile: token accounting, GPU occupancy, batch efficiency, cache hit rate, queue depth, per-endpoint latency distribution. Deliverable: a ranked list of wins with expected savings and an effort estimate. Typically 2 weeks.
02
Continuous batching, paged KV cache, prefill/decode disaggregation, speculative decoding, multi-LoRA serving, scheduler tuning. We write the config and the load test, and we prove the throughput claim on your traffic shape.
03
Predictive scale-up on queue-depth and token-rate signals, spot orchestration with graceful drain, warm pools for cold-start elimination, multi-region spill, quota-aware admission control.
04
Model router with quality/cost tiers, deterministic fallbacks, per-team quotas, budget caps, request coalescing, semantic and prefix caching, cost attribution down to the feature flag.
05
FP8/INT4/AWQ evaluation against your eval set, speculative decoding with a drafter you own, distilled task-specific models, and a quality gate that blocks a regression from shipping.
06
Token budgets per feature, prompt compaction, cache strategy, committed-use versus on-demand mix, per-request cost telemetry, and an alerting policy that fires before the invoice does.
03—Approach
Four steps, fixed windows, no discovery phase. You pick the portfolio; we pull the levers and prove them.
01
Week 1
Read-only access, replay of your traffic, full latency and cost decomposition. No recommendations yet — just the map.
02
Week 2
A capacity model of your stack: what each lever buys in milliseconds and dollars, and what it costs to pull. You pick the portfolio.
03
Weeks 3–6
Ship the top three levers to a shadow deployment, then 10% of traffic, then 100%. Every change carries a quality gate and a rollback.
04
Weeks 7–12
Dashboards, alerts, runbooks, load tests in CI, and two weeks of paired on-call so your team owns it. We leave; the numbers stay.
04—Proof
Medians are across completed programmes. Nothing on this page comes from a lab box or a vendor deck.
Challenge
Consumer search agent, p95 time-to-first-token of 1.9s, 41% of queries timed out during the 18:00–20:00 peak.
Approach
Prefix caching across the agent loop, prefill/decode disaggregation on 24 H100s, queue-depth autoscaling, semantic cache for the top 8% of queries.
Result
| Metric | Measured change |
|---|---|
| p95 time-to-first-token | 1.9s → 410 ms |
| Peak-window timeouts | 41% → 0.3% |
| Cost per query | −61% |
| Delivery | 7 weeks |
Results verified with the client under NDA.Client names and references are available under NDA.
Challenge
A HIPAA workload that could not leave a single on-prem cluster; 4.3 requests/s ceiling and a queue that grew every Monday morning.
Approach
FP8 on the frozen weights after a 12k-pair clinical eval gate, continuous batching with priority scheduling, KV cache sizing per specialty, capacity headroom model tied to the appointment calendar.
Result
| Metric | Measured change |
|---|---|
| Throughput | 4.3× at the same H100 count |
| Availability | 99.98% over 11 months |
| Clinical eval delta | +0.4 points — within noise, gate passed |
Results verified with the client under NDA.Client names and references are available under NDA.
Challenge
IDE autocomplete with a 250ms interactive budget, priced itself out of the free tier at 340k daily users.
Approach
Speculative decoding with a distilled drafter, prefix caching on the repo context, a two-tier router that sends 71% of requests to a small model, per-repo token budgets.
Result
| Metric | Measured change |
|---|---|
| Tokens per second per user | 3.1× |
| Spend at identical quality | −74% |
| Requests routed to the small model | 71% |
| Free tier | reinstated |
Results verified with the client under NDA.Client names and references are available under NDA.
Identical hardware — 8×H100, Llama-class 70B, 1k input / 256 output tokens. Every row is the row above plus one lever.
| Configuration | Tokens/s | p50 TTFT | Cost per 1M tokens |
|---|---|---|---|
| Naive HF serve | 310 | 1,240 ms | $18.40 |
| Continuous batching | 1,180 | 420 ms | $6.90 |
| + Paged KV, tuned scheduler | 2,040 | 260 ms | $4.10 |
| + FP8 + speculative decoding | 3,410 | 140 ms | $2.60 |
| + Prefix cache, traffic-shaped | 4,890 | 96 ms | $1.70 |
05—Engagement models
Fixed-fee, no change orders, no seat licences. If the Diagnostic does not name at least two changes worth more than its fee, we refund it.
01
Diagnostic
from$24,000
2 weeks, fixed scope, one engineer
You need the map: where the latency and the money are going, ranked.
Commission a Diagnostic02
Programme Most chosen
from$85,000
6–12 weeks, two engineers, shipped changes
You know the problem and want the throughput and the bill moved.
Scope a programme03
Fractional
from$14,000/month
Retained, 2 days/week, rolling
You have the roadmap and want a scaling engineer who has seen the failure modes before.
Retain a scaling engineer06—Team
The engineer who profiles your stack is the engineer who ships the fix. Founded 2024; Los Angeles, Jakarta and Seoul, remote-first, 11 people.
Founder / Principal Inference Engineer
Founder + Engineering
Fifteen years in low-latency systems, six of them on GPU serving. Writes the profile, runs the first pilot, still does the diffs.
Distributed Systems Lead
Autoscaling + failure modes
Built multi-region capacity planning for a 40,000-node fleet at Helio Systems. Owns autoscaling and the failure-mode catalogue.
Performance Engineering Lead
Quantisation + quality gates
Kernels, schedulers, profiler output. Has read more nsys traces than is healthy. Owns quantisation and the quality gates.
07—FAQ
No. We are two to twelve weeks of focused pressure on one problem, then we hand over dashboards, runbooks and the diffs. Most clients keep us on fractional after that because they like the numbers staying put.
vLLM, SGLang, TensorRT-LLM, Hugging Face TGI, LMDeploy and managed endpoints (Together, Fireworks, Modal, Bedrock, Vertex AI, OpenAI-compatible gateways). Roughly half our work is on managed APIs, where the levers are routing, caching and prompt shape rather than kernels.
No. Self-hosted buys you kernel-level wins; managed buys you architectural ones. The Diagnostic tells you which side of that line you are on.
Fixed fee against a named scope. No hourly, no seats, no revenue share. Travel is included for the two on-site weeks every programme starts with.
Nothing ships without an eval gate on your own data — a held-out set, an LLM judge, or both, sized before the work starts. If the gate fails, the change does not ship and we tell you in writing what it would have cost.
You keep the load tests in CI, the dashboards, the alert policies and the runbooks. We stay reachable for six months at no charge. About half of our clients move to a fractional retainer.
08—Contact
Send us your traffic shape and your invoice. We will tell you the three changes we would make first, and roughly what each is worth. No pitch deck, no discovery call before the substance.
The firm
Fixed-fee, no change orders, no seat licences.
Request sent
Your request left this page as a form submission and is in the shared inbox. We answer the address above with the three changes we would make first.
One more step to send
This copy of the site has no delivery endpoint configured, so nothing has left your browser yet. Everything you typed is below, and in the draft link — send it from your own mail client and we will pick it up at hello@kusukang.com.