Configure the AI Gateway
Enable the AI Gateway through the platform chart, then apply its custom resources in the order described below. See AI Gateway for ongoing configuration.
Prerequisites
- A corporate identity provider, configured once at
global.stacklok.primaryIdp. The gateway ties every request to an identity from that provider. See Configure identity. - PostgreSQL, which the platform already requires. Budgets, pricing, and recorded spend live there.
- Redis or Valkey, only if you intend to enable the detection result cache. It is optional and off by default. See PCI/PII controls.
Enable it
Set the install toggle in your platform values and upgrade:
global:
stacklok:
aiGateway:
enabled: true
This installs the AI Gateway operator and custom resource definitions. Apply an
AIGateway resource to create a gateway instance.
Bring it up in this order
Complete the following sequence before sending production traffic:
-
Create budgets that cover every caller, before you enable the budget webhook target. Set an organization default for resolved directory users, or create budgets for individual users and groups. A caller with no applicable budget is refused. See Budgets and pricing.
-
Enable gateway-level budget webhooks. Add the webhook target and its receiver configuration to your platform values, then upgrade the release:
values.yamlglobal:webhooks:issuerRef:name: <WEBHOOK_CLUSTER_ISSUER>kind: ClusterIssuercaBundleSecret: <WEBHOOK_CA_BUNDLE_SECRET>enterprise-manager:webhookTLS:enabled: trueport: 443webhookAuth:audience: <BUDGET_WEBHOOK_AUDIENCE>enterprise-ai-gateway-operator:upstream:budgetsWebhook:serviceName: <ENTERPRISE_MANAGER_SERVICE>port: 443audience: <BUDGET_WEBHOOK_AUDIENCE>Set
serviceNameto the Enterprise Manager Service in the same namespace as the gateway. The twoaudiencevalues must match exactly. The operator adds admission and usage webhooks to every OIDC-enabled gateway it manages. -
Apply an
AIGatewayresource with at least one provider and one route. See Connect model providers. -
Verify. Confirm the gateway reports its providers ready and that budget enforcement probed successfully:
kubectl get aigw -n <NAMESPACE>kubectl get aigw <NAME> -n <NAMESPACE> \-o jsonpath='{.status.webhooks}' | jq .
Publish and verify the endpoints
The operator creates a Gateway API Gateway with the same name and namespace as
your AIGateway. Envoy Gateway creates the proxy Service that receives model
inference requests. Configure publishing through the AIGateway resource so the
operator preserves your settings during reconciliation.
Publish the inference listener
For direct HTTPS exposure through a load balancer, merge this configuration into
your existing AIGateway, keeping its authentication, providers, and routes:
spec:
gateway:
listeners:
- port: 443
protocol: HTTPS
tls:
certificateRef:
name: ai-gateway-tls
tls:
issuerRef:
name: <CERTIFICATE_ISSUER>
kind: ClusterIssuer
dnsNames:
- '<AI_GATEWAY_HOST>'
proxy:
serviceType: LoadBalancer
This example requires cert-manager, a ready ClusterIssuer that can issue a
certificate for <AI_GATEWAY_HOST>, and a cluster with load balancer support.
The operator creates the certificate Secret in the AIGateway namespace and
references it from the HTTPS listener. If you supply your own TLS Secret in that
namespace, set listeners[].tls.certificateRef.name to its name and omit
gateway.tls.
Apply your resource and inspect the listener, routes, and proxy Service:
kubectl apply -f aigateway.yaml
kubectl get gateway <AIGATEWAY_NAME> -n <NAMESPACE> -o yaml
kubectl get httproute -n <NAMESPACE>
kubectl get svc -A \
-l gateway.envoyproxy.io/owning-gateway-name=<AIGATEWAY_NAME>,gateway.envoyproxy.io/owning-gateway-namespace=<NAMESPACE>
The proxy Service name is generated by Envoy Gateway. Use the selected Service's
namespace, ports, and load balancer address rather than assuming the operator's
Service receives inference traffic. Check that the Gateway's Programmed
condition and its listener's Accepted condition are True, and that the
inference routes have Accepted and ResolvedRefs set to True. Point
<AI_GATEWAY_HOST> in DNS at the load balancer address.
Clients use https://<AI_GATEWAY_HOST>/v1 as the OpenAI-compatible base URL.
Preserve paths such as /v1/chat/completions and the authentication headers if
you place another proxy in front of Envoy. Configure that proxy to pass streamed
responses without buffering and allow your longest model requests. The gateway's
spec.gateway.timeouts sets a total request deadline, including the streamed
response; leave it unset to avoid imposing that deadline.
If your cluster uses an existing edge proxy, keep
spec.gateway.proxy.serviceType: ClusterIP and route to the discovered Envoy
proxy Service on its configured listener port. An HTTPS listener requires an
HTTPS backend connection with certificate validation at the edge proxy.
Keep the management API reachable by the console
When spec.auth.virtualAPIKeys.enabled: true, the operator also creates
<AIGATEWAY_NAME>-api-key-service in the AIGateway namespace. Its HTTP port
8080 serves the management API, including /v1/me and model catalogs. The
console can reach it inside the cluster. For external automation, publish a
separate HTTPS hostname routed to that Service and port, preserving /v1 paths.
See the management API reference for the
base URL, authentication, and a port-forward alternative.
Verify authenticated use
Send a request from a client machine using a model configured in your routes and a token accepted by your gateway's authentication configuration. The caller must have a directory identity, model access, and an applicable budget:
curl --fail-with-body --silent --show-error \
https://<AI_GATEWAY_HOST>/v1/chat/completions \
-H 'Authorization: Bearer <ACCESS_TOKEN>' \
-H 'Content-Type: application/json' \
-d '{"model":"<MODEL_NAME>","messages":[{"role":"user","content":"Reply with pong"}]}'
A successful completion verifies DNS, TLS, authentication, admission, and the
provider connection. For a private CA, use --cacert <CA_FILE> and configure
your clients to trust that CA. Repeat with "stream": true in the request body
and curl --no-buffer to confirm chunks arrive before the response completes.
If you enabled the management API, verify it separately through its port-forward:
curl --fail --silent --show-error http://localhost:8080/v1/me \
-H 'Authorization: Bearer <ACCESS_TOKEN>' | jq .
Use a token accepted by that API's authentication configuration. Then follow Roll out gateway clients to distribute the verified endpoint.
Content screening posture
Detection failures deny requests by default. An experimental waiver can allow traffic during a rollout or incident, but it is unavailable on the stable release channel.
Next steps
- AI Gateway for providers, routing, screening, and budgets.
- Roll out gateway clients to distribute deployment-specific setup instructions.