Kusukang—LLM inference scaling · Los Angeles · Jakarta · Seoul

Inference is where
your AI budget goes to die.

Kusukang is a consultancy that specialises in LLM inference at scale — serving architecture, autoscaling, quantisation and cost engineering for teams whose usage outgrew their first deployment.

We do one thing: the last 90% of the gap between a model that works in a notebook and a model that answers 4,000 requests a second on a Tuesday.

Founded
2024
Offices
Los Angeles · Jakarta · Seoul
People
11, remote-first

Serving stacks we work in

  • vLLM
  • SGLang
  • TensorRT-LLM
  • Hugging Face TGI
  • LMDeploy
  • Together AI
  • Fireworks
  • Modal
  • Amazon Bedrock
  • Vertex AI

01—Why we exist

Where the milliseconds and the dollars actually go.

Three failures show up in every stack we open. None of them are model failures, and none of them get fixed by a bigger cluster.

01

Cost

$190kper month

Median client arrives spending $190k/month on inference with 31% of it serving tokens nobody asked for — retries, oversized contexts, uncached prefixes, unprompted 4k completions.

02

Latency

1.4–3.2s

Time-to-first-token is the user-visible metric. It is usually 1.4–3.2s in a stack that could sit under 300ms, because batching, KV cache reuse and prefill/decode separation were never tuned together.

03

Peak

4×evening surge

Load is spiky and GPU capacity is not. Quotas, cold starts and a single-region dependency turn a 4× evening surge into a 40-minute outage.

Every engagement starts with the same artefact: a profile of your stack that says where the milliseconds and the dollars actually go.

02—Services

Six services. One problem: the serving path.

Every engagement starts with the profile. What happens after that is your call, priced against a named scope.

01

Inference Audit

Kernel-to-gateway profile: token accounting, GPU occupancy, batch efficiency, cache hit rate, queue depth, per-endpoint latency distribution. Deliverable: a ranked list of wins with expected savings and an effort estimate. Typically 2 weeks.

02

Serving Architecture

Continuous batching, paged KV cache, prefill/decode disaggregation, speculative decoding, multi-LoRA serving, scheduler tuning. We write the config and the load test, and we prove the throughput claim on your traffic shape.

03

Autoscaling & Capacity

Predictive scale-up on queue-depth and token-rate signals, spot orchestration with graceful drain, warm pools for cold-start elimination, multi-region spill, quota-aware admission control.

04

Gateway & Routing

Model router with quality/cost tiers, deterministic fallbacks, per-team quotas, budget caps, request coalescing, semantic and prefix caching, cost attribution down to the feature flag.

05

Quantisation & Distillation

FP8/INT4/AWQ evaluation against your eval set, speculative decoding with a drafter you own, distilled task-specific models, and a quality gate that blocks a regression from shipping.

06

Cost Engineering

Token budgets per feature, prompt compaction, cache strategy, committed-use versus on-demand mix, per-request cost telemetry, and an alerting policy that fires before the invoice does.

03—Approach

Twelve weeks from profile to handover.

Four steps, fixed windows, no discovery phase. You pick the portfolio; we pull the levers and prove them.

  1. 01

    Profile

    Week 1

    Read-only access, replay of your traffic, full latency and cost decomposition. No recommendations yet — just the map.

  2. 02

    Model

    Week 2

    A capacity model of your stack: what each lever buys in milliseconds and dollars, and what it costs to pull. You pick the portfolio.

  3. 03

    Pilot

    Weeks 3–6

    Ship the top three levers to a shadow deployment, then 10% of traffic, then 100%. Every change carries a quality gate and a rollback.

  4. 04

    Hardening

    Weeks 7–12

    Dashboards, alerts, runbooks, load tests in CI, and two weeks of paired on-call so your team owns it. We leave; the numbers stay.

04—Proof

Numbers, measured on production traffic.

Medians are across completed programmes. Nothing on this page comes from a lab box or a vendor deck.

Across 140+ GPU-node audits

Median p50 latency reduction
62%
Median cost per 1M tokens reduction
47%
Best time-to-first-token achieved
96msfrom 380 ms
Peak sustained throughput observed
9.1Mtokens/min
Gateway availability, trailing 12 months
99.95%
GPU-node audits completed
140+

Case studies

Case 01 — Consumer agentic search

EU · 2025

Challenge

Consumer search agent, p95 time-to-first-token of 1.9s, 41% of queries timed out during the 18:00–20:00 peak.

Approach

Prefix caching across the agent loop, prefill/decode disaggregation on 24 H100s, queue-depth autoscaling, semantic cache for the top 8% of queries.

Result

MetricMeasured change
p95 time-to-first-token 1.9s → 410 ms
Peak-window timeouts 41% → 0.3%
Cost per query −61%
Delivery 7 weeks

Results verified with the client under NDA.Client names and references are available under NDA.

Case 02 — Clinical triage assistant

US · 2025

Challenge

A HIPAA workload that could not leave a single on-prem cluster; 4.3 requests/s ceiling and a queue that grew every Monday morning.

Approach

FP8 on the frozen weights after a 12k-pair clinical eval gate, continuous batching with priority scheduling, KV cache sizing per specialty, capacity headroom model tied to the appointment calendar.

Result

MetricMeasured change
Throughput 4.3× at the same H100 count
Availability 99.98% over 11 months
Clinical eval delta +0.4 points — within noise, gate passed

Results verified with the client under NDA.Client names and references are available under NDA.

Case 03 — IDE code assistant

APAC · 2024

Challenge

IDE autocomplete with a 250ms interactive budget, priced itself out of the free tier at 340k daily users.

Approach

Speculative decoding with a distilled drafter, prefix caching on the repo context, a two-tier router that sends 71% of requests to a small model, per-repo token budgets.

Result

MetricMeasured change
Tokens per second per user 3.1×
Spend at identical quality −74%
Requests routed to the small model 71%
Free tier reinstated

Results verified with the client under NDA.Client names and references are available under NDA.

Benchmark: one box, five configurations

Identical hardware — 8×H100, Llama-class 70B, 1k input / 256 output tokens. Every row is the row above plus one lever.

Configuration Tokens/s p50 TTFT Cost per 1M tokens
Naive HF serve 310 1,240 ms $18.40
Continuous batching 1,180 420 ms $6.90
+ Paged KV, tuned scheduler 2,040 260 ms $4.10
+ FP8 + speculative decoding 3,410 140 ms $2.60
+ Prefix cache, traffic-shaped 4,890 96 ms $1.70
Bar chart of throughput for five configurations, rising from 310 to 4,890 tokens per second Five horizontal bars on a shared scale, each configuration adding one serving lever to the row above it. Naive HF serve Continuous batching Paged KV, tuned FP8 + speculative Prefix cache 310 1,180 2,040 3,410 4,890
Throughput by configuration, tokens per second. Each bar adds one serving lever to the configuration above it, on identical hardware.

05—Engagement models

Three ways to start.

Fixed-fee, no change orders, no seat licences. If the Diagnostic does not name at least two changes worth more than its fee, we refund it.

01

Diagnostic

from$24,000

2 weeks, fixed scope, one engineer

You need the map: where the latency and the money are going, ranked.

Commission a Diagnostic

02

Programme Most chosen

from$85,000

6–12 weeks, two engineers, shipped changes

You know the problem and want the throughput and the bill moved.

Scope a programme

03

Fractional

from$14,000/month

Retained, 2 days/week, rolling

You have the roadmap and want a scaling engineer who has seen the failure modes before.

Retain a scaling engineer

06—Team

Eleven people, no account managers, no bench.

The engineer who profiles your stack is the engineer who ships the fix. Founded 2024; Los Angeles, Jakarta and Seoul, remote-first, 11 people.

PK PK

Phillip Kang

Founder / Principal Inference Engineer

Founder + Engineering

Fifteen years in low-latency systems, six of them on GPU serving. Writes the profile, runs the first pilot, still does the diffs.

PR PR

Priya Raghunathan

Distributed Systems Lead

Autoscaling + failure modes

Built multi-region capacity planning for a 40,000-node fleet at Helio Systems. Owns autoscaling and the failure-mode catalogue.

TB TB

Tomás Beaulieu

Performance Engineering Lead

Quantisation + quality gates

Kernels, schedulers, profiler output. Has read more nsys traces than is healthy. Owns quantisation and the quality gates.

07—FAQ

Six questions, answered before the contract.

Do you replace our platform team?

No. We are two to twelve weeks of focused pressure on one problem, then we hand over dashboards, runbooks and the diffs. Most clients keep us on fractional after that because they like the numbers staying put.

Which stacks do you work in?

vLLM, SGLang, TensorRT-LLM, Hugging Face TGI, LMDeploy and managed endpoints (Together, Fireworks, Modal, Bedrock, Vertex AI, OpenAI-compatible gateways). Roughly half our work is on managed APIs, where the levers are routing, caching and prompt shape rather than kernels.

Do we need to be self-hosted?

No. Self-hosted buys you kernel-level wins; managed buys you architectural ones. The Diagnostic tells you which side of that line you are on.

How do you price?

Fixed fee against a named scope. No hourly, no seats, no revenue share. Travel is included for the two on-site weeks every programme starts with.

What if a change hurts quality?

Nothing ships without an eval gate on your own data — a held-out set, an LLM judge, or both, sized before the work starts. If the gate fails, the change does not ship and we tell you in writing what it would have cost.

What happens after the engagement?

You keep the load tests in CI, the dashboards, the alert policies and the runbooks. We stay reachable for six months at no charge. About half of our clients move to a fractional retainer.

08—Contact

Book a 30-minute scaling review

Send us your traffic shape and your invoice. We will tell you the three changes we would make first, and roughly what each is worth. No pitch deck, no discovery call before the substance.

The firm

  • Emailhello@kusukang.com
  • OfficesLos Angeles · Jakarta · Seoul
  • Founded2024, remote-first, 11 people

Fixed-fee, no change orders, no seat licences.

Scaling review request

Traffic shape, model, stack, current p95 — whatever you have.

Posted to hello@kusukang.com through one form processor and nothing else — no newsletter, no trackers. See the privacy policy.