Onboarding the first endpoints.Talk to us

Recursive Self-Improving Intelligence.
Starting with inference.

A dedicated endpoint built around your traffic that keeps getting faster after it ships, run by a system that improves with every customer.

What a customer sees: an endpoint that keeps converging
Speedup vs. the default serving config, same model, same card.
1.0x1.5x2.0x2.5x1.0xBaselinevLLM default1.4xEngine tunedweek 11.9xKernels fusedweek 12.6xTraffic re-tunedcontinuous
Illustrative target curve. Every published multiple will carry its hardware, quantization, workload shape, and baseline.
Vision

We are building a real Recursive Self-Improving Intelligence. The first place we are pointing it is the hardest and most profitable problem in AI: inference acceleration. We will prove that RSI is the basic organizational form of intelligence, then take it into every field. But it starts with inference.

Inference first, then every field with a measurable outcome
Problem

Every AI application company runs on rented intelligence, with no privileges.

Who

Any company putting intelligence into its product or service. In a few years, almost every company.

Choice one

Build their own

A seven-figure payroll for expertise that is scarce, and a setup that breaks every time the model updates.

seven-figure annual bill / slow progress / uncertainty
Choice two

Buy from the general model market

A ceiling. They hand over the core of their product with no control over its speed, quality, or price.

control of your product, held by a provider that does not care much about you
GPU performance engineers
4 : 1openings per candidate
Supply-to-demand ratio fell to 0.26 in spring 2026. The people who fix this cannot be hired.
Public MaaS gross margin, a leading token seller
-119%
Per its 2026 IPO filing. Selling generic tokens loses money. The provider has every reason to cut corners on you.
Open-model API output prices, 2026
2-4x up
Three years of 99% price decline reversed. Renting got more expensive, not less.
Dedicated

A third choice: a dedicated inference endpoint, built around your model, your traffic, and your SLO, that keeps getting better after it ships.

Custom in a week, not years.

Bring the model, a sample of real traffic, and a target. Get a private endpoint with performance written into the SLA.

model + traffic + SLO → endpoint + SLA

Full-stack, on real traffic.

All you see is an endpoint. Kernels, engine, quantization, scheduling, parallelism, and the card itself are searched together against your actual request mix.

one endpoint, whole stack

Evolves with your product.

Traffic drifts, models update, drivers change. We adapt as it happens: endless optimization, tailored only to your production needs.

never finished
Custom in a week: the target onboarding
From first call to a private endpoint with an SLA
Day 0Model,traffic,targetDay 1Baseline +profileDay 4VerifiedwinnerDay 7Endpointlive, SLAsignedOngoingRe-tuned astrafficchanges
Target timeline for a standard open model on stocked hardware. New hardware or an unsupported architecture adds time. We say so before we quote.
Technology

An RSI system that optimizes inference across the whole stack, and gets faster at it with every customer.

The fastest setup is never a single switch. Kernels, serving engine, batching, quantization, hardware, and traffic shape all push on each other. Real performance is a combination, and combinations are what the system searches.

Find the bottlenecksTry many pathsShip the measured winnerRSIminutes per cycle, on real trafficonly verified gains shipverified gainsbecome priors
  1. Find the bottlenecks

    Profiles the customer's stack on real traffic to see where latency or throughput is lost: scheduling, decoding, kernels, memory pressure, or hardware fit.

  2. Try many paths

    Generates candidate configurations across batching, decoding, quantization, engines, kernels, and hardware, and measures every one. Thousands of runs, verified in hours.

  3. Ship the measured winner

    Deploys the fastest verified configuration, then keeps profiling production traffic as load, models, and hardware evolve.

  4. Only verified, reproducible gains ship.

    That is what makes the loop trustworthy, and what makes its memory worth something. Every solved endpoint shortens the next search.

Real performance is not won at a single layer.

Heterogeneous by design

Model, engine, kernels, and hardware optimized together against the outcome you care about.

NVIDIA BlackwellAMD MI300 / MI355Ascend

Kernels per traffic shape

Fused ops, attention paths, GEMM variants, and decode kernels for the model shape and hardware target.

CUDAHIPAscendCTriton

Engine configs per workload

Serving engine auto-tuned to the model, traffic shape, memory pressure, and latency target.

vLLMSGLangMindIEschedulerKV cache

Decode strategy search

Speculative decoding, FP8 / FP4 / W8A8 formats, batching, and expert sharding for MoE.

Spec decodequantizationbatchingexpert sharding

The card-hour gap is 15x. The per-token gap is a tie. The difference is software.

Card-hour price, September 2026
US$ per GPU-hour, rental. Ascend is the cheapest card in the world by a wide margin.
Ascend 910B$0.42AMD MI355X$2.95NVIDIA H100$3.25NVIDIA B200$6.44
Public rental list prices (TensorWave, major NVIDIA clouds) and mainland China cloud pricing for 910B, converted at market rates.
Cost per million output tokens, same cards
DeepSeek-class MoE, 20-50 tok/s per user. The 15x card-hour gap collapses to a tie.
$0.25$0.50$0.75$1.00$1.25Ascend 910B$0.15-$0.33AMD MI355X$0.23-$0.34NVIDIA H100$0.32-$1.22NVIDIA B200$0.23-$0.36
Range = default engine config (high end) to well-tuned config (low end), from public benchmarks and our own estimates. On 910B the 500 to 1,900 tok/s per card span is software. That span is the business.
Why now

Open models won. Every company is becoming an AI company. And they all want to own their intelligence.

Open models are taking over.

Companies are moving off closed APIs and onto open weights they can run themselves. Open-model router volume is up 5x in six months.

Every company is now an AI company.

Businesses in every sector are putting intelligence into their own product, and discovering that the hard part is not the model. It is running it well.

The market wants to own its intelligence.

Product companies have figured out that the model is the moat, and renting it from a big lab is not one. Nobody has made owning it affordable yet.

Inference share of AI compute
Percent of AI compute spent on inference rather than training
25%50%75%33%202375%2026 (est.)
Industry estimates; 2026 is the midpoint of a 70-80% range.
Where open-model traffic runs
Share of OpenRouter open-model volume by model origin, mid-2026
72%Chinese open modelsAll others · 28%
OpenRouter public stats. The labs shipping the volume are the ones we work with.
Open-model router volume, last six months
5x
25 trillion tokens a week and climbing. The demand side already moved. The supply side is still generic.

Three ways to pay. All of them sit on the bill you already have.

Dedicated endpoint

Per card-hour plus reserved capacity, with performance targets written into the SLA. At least 30% below your current cost per token at equal or better latency.

Voice, coding, and agent products with an SLO

Savings share

20-30% of the measured reduction against your current bill. The easiest contract to sign, and the one that proves our numbers.

Platforms where inference is COGS

Public fast tier

The fastest endpoint for the hottest open models, listed on OpenRouter at a speed premium. The router sells; inbound converts to dedicated.

Try the public models
Distribution, not the main line
Team
Kyrie Cai
Founder & CEO

CS, Sun Yat-sen University (national top-talent program, HCP Lab); MSc Venture Creation, NUS. Self-funded two years delivering private multi-agent systems to enterprise clients in exoskeletons, game advertising, and e-commerce. Founded Launchpad S1 and SG Hacker House; ran a 200-person agent hackathon.

Bring the model, a sample of real traffic, and a target.

Get a private endpoint with performance written into the SLA. Custom in a week, not years.

Talk to us