Model-as-a-Service with OpenEverest: Your Own OpenAI-Compatible API on Kubernetes

Model-as-a-Service with OpenEverest: Your Own OpenAI-Compatible API on Kubernetes

• By Sergey Pronin Sergey Pronin

Almost every platform team I talk to these days runs models and inference workloads locally. Sometimes it is data sovereignty, sometimes cost, sometimes the GPUs are already there and nobody uses them well. The pattern has a name - Model-as-a-Service (MaaS) - and it looks a lot like Database-as-a-Service did ten years ago.

And just like DBaaS back then, building it yourself is a puzzle. KServe and vLLM to run the model. Gateway API and Envoy to expose it. An AI Gateway to route by model name. cert-manager for TLS. API keys per team, token quotas, scaling the model that does not fit into one server, Prometheus for metrics. Every piece is great on its own. Gluing them together is a project.

In this post I will show how to get there with OpenEverest and the KServe provider for it. By the end you will have:

  • a model deployed with one click in the UI (or ~20 lines of YAML),
  • served over HTTPS on a single OpenAI-compatible endpoint,
  • protected by a per-model API key and an hourly token quota,
  • and, as a bonus, a model split across three GPU nodes.

Models are just another Instance

If you read OpenEverest v2: Why is it a big deal, you know that v2 moved every technology into a Provider. OpenEverest does not really care if the thing behind an Instance is MongoDB, PostgreSQL, or a large language model. Same UI, same API, same RBAC, same Presets.

The KServe provider takes one Instance with topology.type: llm and turns it into everything that is needed to serve a model behind an authenticated gateway. Click on the boxes to see what each object does:

One Instance in, a full serving stack out

Instancetopology: llm1 click or ~20 lines of YAMLprovider-kservereconcilesLLMInferenceServiceKServe · vLLMAIGatewayRoutemodel routingSecurityPolicyAPI keyBackendTrafficPolicytoken quotaPodMonitormetricsSecret <name>-connconnection details

The provider never manages pods directly. KServe does the heavy lifting of running vLLM, Envoy AI Gateway handles the edge, and OpenEverest gives you one place to drive it all from.

How a request travels

All models share one Gateway with one public address. Clients choose the model with the standard OpenAI model field - nothing custom in the client. Every request passes three checks before it ever touches a GPU. Pick a request and send it through:

Follow a request through the AI Gateway

YOUR KUBERNETES CLUSTEREnvoy AI GatewayClientOpenAI SDKLoad BalancerHTTPS :443API keyModelQuota401403429InferencePoolendpoint pickervLLMGPU pods
→
Pick a requestChoose one of the buttons above to send it through the Gateway.

Because the Gateway reads the OpenAI request, it also meters tokens per key. That is what makes it a service and not just a load balancer in front of vLLM: you know who is using which model and how much.

What you need

ComponentRequiredInstalled by
GPU nodes + NVIDIA GPU Operator (or device plugin)yesyou
A LoadBalancer implementation with a public IPyesyour cloud, or MetalLB on-prem
Gateway API CRDs and cert-manageryesstep 1
OpenEverest coreyesstep 2
KServe, Envoy Gateway, Envoy AI Gateway, LeaderWorkerSetyesprovider chart, step 3
A domain nameyes (or nip.io for tests)step 4
Redis / Valkeyonly for token quotasyou

I tested this end to end on a Linode LKE cluster with three NVIDIA RTX 4000 Ada GPUs (20 GB each), using nip.io and Let’s Encrypt. Any managed Kubernetes with GPU nodes or a bare-metal cluster with MetalLB works the same way - see the deployment guide for EKS, GKE, AKS, CoreWeave and friends.

Check that your GPUs are visible to Kubernetes:

$ kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\\.com/gpu
NAME                     GPU
lke664704-985021-bd2pn   1
lke664704-985021-mwh5d   1
lke664704-985021-vqpf4   1

Step 1: Gateway API and cert-manager

Two things need to be in place before anything else:

  • Gateway API CRDs. Gateway, HTTPRoute and friends are not built into Kubernetes - they are CRDs. The AI Gateway is a Gateway object, so the cluster needs these types.
  • cert-manager with Gateway API support. It issues the HTTPS certificate for the Gateway (and the webhook certificates for KServe). Gateway API support is off by default, and without it cert-manager ignores the Gateway - no certificate, no HTTPS.

Order matters: cert-manager only checks for the Gateway API CRDs when it starts. So first install the CRDs - the exact set bundled with the Envoy Gateway version the provider chart pins, so the versions match - and then cert-manager with Gateway API support switched on:

helm pull oci://docker.io/envoyproxy/gateway-helm --version v1.5.9 --untar -d /tmp/eg
kubectl apply --server-side -f /tmp/eg/gateway-helm/crds/gatewayapi-crds.yaml

helm upgrade -i cert-manager oci://quay.io/jetstack/charts/cert-manager \
  -n cert-manager --create-namespace \
  --version v1.21.2 \
  --set crds.enabled=true \
  --set config.gatewayAPI.enabled=true \
  --wait

If cert-manager already runs in your cluster, upgrade it with config.gatewayAPI.enabled=true and restart it (kubectl -n cert-manager rollout restart deployment cert-manager).

Step 2: OpenEverest v2.0.0-dev.3

helm repo add openeverest https://openeverest.github.io/helm-charts/
helm upgrade -i everest openeverest/openeverest -n everest-system --create-namespace \
  --devel --version 2.0.0-dev.3 \
  --set server.initialAdminPassword=<admin-password> \
  --wait

Open the UI with a port-forward and log in as admin:

kubectl port-forward svc/everest 8080:8080 -n everest-system

Step 3: The KServe provider

One Helm chart installs the provider together with the KServe controllers, Envoy Gateway, Envoy AI Gateway, and LeaderWorkerSet. If you already have those running - disable them below. Put the settings into a values file:

# provider-values.yaml
envoy-gateway:
  enabled: true          # set to false if Envoy Gateway already runs in the cluster
envoy-ai-gateway:
  enabled: true
aiGateway:
  enabled: true
  gatewayService:
    type: LoadBalancer
lws:
  enabled: true          # LeaderWorkerSet - needed for multi-node models
storageInitializer:
  resources:
    limits:
      memory: 4Gi        # the default 1Gi gets OOM-killed on multi-GB model downloads
      cpu: "2"
NS=provider-kserve
CHART=oci://ghcr.io/openeverest/charts/provider-kserve

helm upgrade -i provider-kserve $CHART -n $NS --create-namespace -f provider-values.yaml \
  --set aiGateway.auth.allowInsecureHTTP=true   # temporary, until step 4

It takes 2-3 minutes: hook Jobs fetch the KServe CRDs and apply the vLLM runtimes and LLM presets once the controllers are ready. Then grab the public address of the shared Gateway:

kubectl -n $NS get gateway provider-kserve-ai-gateway \
  -o jsonpath='{.status.addresses[0].value}'

The provider registers itself in OpenEverest and shows up in the UI right away, together with two presets it ships out of the box: smollm-cpu for a quick smoke test without a GPU, and qwen-gpu with Qwen3-4B on a single GPU.

Solanica - Blog - OpenEverest model creation with Kserve

Step 4: HTTPS

API keys over plain HTTP are a bad idea, so the chart refuses to enable authentication without TLS (that is why we needed allowInsecureHTTP above). For a test, nip.io gives you a hostname without touching DNS: if the Gateway got 203.0.113.10, use 203-0-113-10.nip.io.

Create a Let’s Encrypt issuer that solves HTTP-01 challenges through the Gateway itself:

apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-prod-ai-gateway
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    privateKeySecretRef:
      name: letsencrypt-prod-ai-gateway-account
    solvers:
      - http01:
          gatewayHTTPRoute:
            parentRefs:
              - name: provider-kserve-ai-gateway
                namespace: provider-kserve
                kind: Gateway
                sectionName: acme-http01

A ClusterIssuer is cluster-wide, so it has no namespace - just kubectl apply -f it. The only namespace that matters is in parentRefs: it must point to where the Gateway lives (the provider release namespace, provider-kserve here).

Then turn TLS on in provider-values.yaml and upgrade - this time without allowInsecureHTTP:

aiGateway:
  enabled: true
  gatewayService:
    type: LoadBalancer
  tls:
    enabled: true
    hostname: 203-0-113-10.nip.io
    acmeHTTP01: true       # port-80 listener for ACME challenges only
    issuerRef:
      name: letsencrypt-prod-ai-gateway
      kind: ClusterIssuer
helm upgrade provider-kserve $CHART -n $NS -f provider-values.yaml
kubectl -n $NS wait certificate/provider-kserve-ai-gateway-tls --for=condition=Ready --timeout=10m

curl -s -o /dev/null -w '%{http_code}\n' https://203-0-113-10.nip.io/v1/models
401

401 is exactly what we want: the endpoint is up, and it does not talk to strangers. For wildcard certificates or clusters without public port 80, use DNS-01 - see ai-gateway-tls.md.

Step 5: Deploy a model

Now the fun part. In the UI, pick the qwen-gpu preset, or start from scratch and choose a model from the catalog. The form is generated from the provider definition and is split into sections: Basic Information, Scaling & Parallelism, Runtime, Routing, and a few advanced ones. In Routing, set External access to Envoy AI Gateway and, optionally, a token budget per user per hour.

OpenEverest

Solanica - Blog - OpenEverest choose the model
Solanica - Blog - OpenEverest expose with AI Gateway

If you prefer YAML (or GitOps), this is the whole thing:

apiVersion: core.openeverest.io/v1alpha1
kind: Instance
metadata:
  name: qwen-small
  namespace: default
spec:
  providerRef:
    name: provider-kserve
  topology:
    type: llm
    parameters:
      externalAccess: EnvoyAIGateway
      tokenLimitPerHour: 3000      # enforced only with Redis/Valkey configured
  components:
    llmEngine:
      type: vllm
      replicas: 1
      resources:
        requests: { cpu: "3", memory: 12Gi }
        limits: { memory: 12Gi }
      parameters:
        modelURI: hf://Qwen/Qwen3-4B
        modelName: qwen3-4b        # must be unique across the Gateway
        gpuCount: 1
$ kubectl -n default get instance qwen-small -w

The first start takes a few minutes: the vLLM image is large and the model is downloaded from Hugging Face. After that the Instance turns Ready.

Step 6: Call it

The provider generates a random API key per Instance and keeps it for the life of the Instance - just like a database password. Everything a client needs is in the connection details:

Solanica - Blog - OpenEverest ready model overview

FieldContent
urihttps://<domain>/v1 - use as the OpenAI base_url
passwordthe API key
usernamethe key ID (not secret, appears in metrics)
modelthe name to send in the model field

From the terminal, the same values live in the <instance>-conn Secret:

c() { kubectl -n default get secret qwen-small-conn -o jsonpath="{.data.$1}" | base64 -d; }

curl "$(c uri)/chat/completions" \
  -H 'Content-Type: application/json' \
  -H "Authorization: Bearer $(c password)" \
  -d "{\"model\":\"$(c model)\",\"messages\":[{\"role\":\"user\",\"content\":\"Hello\"}]}"

And since it is OpenAI-compatible, any SDK or tool that speaks OpenAI just works:

from openai import OpenAI

client = OpenAI(base_url="https://203-0-113-10.nip.io/v1", api_key="<password>")
resp = client.chat.completions.create(
    model="qwen3-4b",
    messages=[{"role": "user", "content": "Explain Kubernetes in one sentence."}],
)
print(resp.choices[0].message.content)

Anthropic-style clients can send the key in the x-api-key header instead.

Multi-tenancy for free

Deploy a second model - say llama-3.1-8b for another team - and it lands on the same Gateway, the same hostname, with its own key. Try calling it with the Qwen key and you get a 403. That is the whole multi-tenancy model: one endpoint, one key per model, and namespaces plus OpenEverest RBAC decide who can see which key.

A few operational things that come with it:

  • Rotate a key: kubectl -n team-a delete secret qwen-small-ai-gateway-key. A new key is generated within seconds, connection details update, and the old key stops working.
  • Token quotas: point the bundled Envoy Gateway at a Redis or Valkey (envoy-gateway.config.envoyGateway.rateLimit.backend), and tokenLimitPerHour is enforced per key and model. Exhausted budgets return 429.
  • Metrics: Gateway GenAI metrics carry an openeverest_key_id label, so you can answer “who burned the GPU budget this week” with a single PromQL query.

When a model does not fit on one node

Small models are easy. The interesting part starts when one copy of a model is bigger than a single server - even after you split it across all GPUs in that server. It happens sooner than you think: Qwen3 14B is ~30 GB of bf16 weights, which already does not fit on a node with a single 20 GB card. Llama 4 Scout is ~220 GB - more than a whole 8x 24 GB box.

vLLM solves this with two kinds of parallelism, and the KServe provider exposes both:

You setWhat it meansTypical case
tensorParallelSizeSplit each layer across GPUs in one nodeFits on one 8-GPU box
workerCount + pipelineParallelSizeSplit the layers across several nodes (head + workers)Still too big for one box
replicasHow many full copies of that layoutMore traffic

Under the hood KServe runs a head + workers layout on a LeaderWorkerSet. The head is pipeline stage 0 and serves the API, workers hold the remaining layers and join the head over the pod network. Clients see nothing new - same URL, same model, same key.

Play with the planner to see how the pieces fit together:

Multi-node planner - one model, several GPU nodes

Model (bf16)
Worker nodes
GPUs per node
Replicas

The default state of the planner is the setup I actually ran on the demo cluster: Qwen3 14B from the model catalog across three nodes with one 20 GB RTX 4000 Ada each. Its ~30 GB of bf16 weights do not fit on a single card. Split over two cards, the weights would fit, but there would be almost no memory left for the KV cache at Qwen3’s 40k-token context. Three nodes it is - head plus two workers, one pipeline stage each:

apiVersion: core.openeverest.io/v1alpha1
kind: Instance
metadata:
  name: qwen3-14b
spec:
  providerRef:
    name: provider-kserve
  topology:
    type: llm
    parameters:
      externalAccess: EnvoyAIGateway
  components:
    llmEngine:
      type: vllm
      replicas: 1
      resources:
        requests: { cpu: "3", memory: 16Gi }
        limits: { memory: 16Gi }
      parameters:
        modelURI: hf://Qwen/Qwen3-14B
        modelName: qwen3-14b
        gpuCount: 1               # per pod
        tensorParallelSize: 1
        pipelineParallelSize: 3   # must be workerCount + 1
        workerCount: 2

LeaderWorkerSet spreads the group across the GPU nodes - the head (-mn-0) and two workers (-mn-0-1, -mn-0-2):

$ kubectl get pods -o wide
NAME                                  READY   STATUS    AGE     NODE
qwen3-14b-kserve-mn-0                 1/1     Running   9m34s   lke664704-985021-mwh5d
qwen3-14b-kserve-mn-0-1               1/1     Running   9m34s   lke664704-985021-bd2pn
qwen3-14b-kserve-mn-0-2               1/1     Running   9m34s   lke664704-985021-vqpf4
qwen3-14b-kserve-router-scheduler-…   1/1     Running   9m34s   lke664704-985021-mwh5d

From kubectl apply to Ready took under ten minutes; most of it was every pod pulling the vLLM image and downloading the ~30 GB model in parallel. The vLLM logs show how the model got sliced - each GPU holds about a third of the weights, and the rest goes to the KV cache:

(Worker_PP0) Model loading took 9.46 GiB
(Worker_PP1) Model loading took 8.62 GiB
(Worker_PP2) Model loading took 9.46 GiB
(Worker_PP0) Available KV cache memory: 7.1 GiB
(EngineCore) GPU KV cache size: 143,104 tokens
(EngineCore) Maximum concurrency for 40,960 tokens per request: 3.49x

And for the client it is just another model behind the same Gateway, with its own API key. The only hint that three machines are answering is in the response fingerprint:

$ curl "$(c uri)/chat/completions" -H "Authorization: Bearer $(c password)" ...
{
  "model": "qwen3-14b",
  "choices": [{ "message": { "content": "Pipeline parallelism is a technique in distributed computing where different stages of a computational task are executed in parallel across multiple processors or devices..." } }],
  "system_fingerprint": "vllm-0.26.0-pp3-6fb39616",
  "usage": { "prompt_tokens": 22, "completion_tokens": 41, "total_tokens": 63 }
}

In the UI the same fields live in Scaling & Parallelism: Worker count, Pipeline Parallel Size, and optional worker CPU and memory.

Solanica - Blog - OpenEverest multi-node flags

A few things worth knowing before you go big:

  • workerCount is the switch. pipelineParallelSize alone only sets a vLLM flag - it does not create worker pods. The provider rejects an Instance where pipelineParallelSize is not workerCount + 1, or where the GPU count is smaller than tensorParallelSize.
  • One replica is the whole ring. replicas: 2 means two heads and two sets of workers, not one extra GPU.
  • Every pod downloads the full model. Size storageInitializer memory accordingly, or serve the weights from a pvc:// URI.
  • The ring lives and dies together. If any pod restarts, LeaderWorkerSet recreates the whole group. Pods also need to reach each other on the pod network (port 29501 and the NCCL ports).

For the full set of rules, including disaggregated prefill/decode with its own worker ring, see llm-multi-node.md.

What is next

What we built here is the core of a Model-as-a-Service platform: models on your own GPUs, one OpenAI-compatible endpoint, a key per model, quotas, metrics, and models bigger than a single server. All of it driven by the same Instance API and UI that OpenEverest users already use for databases.

The KServe provider is young and moving fast. Autoscaling with the Workload Variant Autoscaler, disaggregated prefill/decode, LoRA adapters and predictive models (sklearn, XGBoost, ONNX, Triton) are already in the provider - each of them deserves a separate post. If you want to try it, break it, or help shape it - the provider-kserve repository is the place to start, and we hang out in #openeverest-users on CNCF Slack.

See Also