Skip to main content
Version: Next

LLMInferenceService with agentgateway

This guide walks through using agentgateway as an Inference Gateway for KServe LLMInferenceService. agentgateway supports the Gateway API Inference Extension, so KServe's generated HTTPRoute can route directly to the standard InferencePool backend. An AgentgatewayBackend is only needed when you want to apply AI policies that require LLM-aware processing, such as token-based rate limiting. It can wrap the generated InferencePool so that endpoint selection remains available.

agentgateway Overview

agentgateway is a Rust-based proxy under the AI Agent Infrastructure Foundation (AAIF) at the Linux Foundation. It implements the Kubernetes Gateway API but is LLM-aware: it parses OpenAI chat completion requests and responses, extracts token usage, emits OpenTelemetry GenAI semantic conventions, and enforces token-based rate limits and policies. Key custom resources:

  • InferencePool: Standard Gateway API Inference Extension backend supported directly by agentgateway.
  • AgentgatewayBackend: Optional backend that declares an LLM provider so the gateway can apply LLM-aware processing and AI policies.
  • AgentgatewayPolicy: Attaches governance policies such as token-based rate limiting.
  • HTTPRoute: Standard Gateway API routing that can reference either an InferencePool or an AgentgatewayBackend.

For more information, see the agentgateway KServe integration guide, the llm-d agentgateway guide, and the Gateway API Inference Extension implementation list.

Choose a Backend Type

When the managed scheduler is enabled, KServe generates HTTPRoute resources with an InferencePool backend. agentgateway supports this standard backend without any route override. Use AgentgatewayBackend only when an AI policy needs agentgateway to parse the LLM request or response.

BackendUse it forRoute override required
InferencePoolStandard inference routing through the Gateway API Inference ExtensionNo
AgentgatewayBackendAI policies such as token-based rate limiting, plus LLM-aware telemetry and model trackingYes
note

KServe supports distributed tracing natively via spec.tracing, which provides request-level spans and traces. The LLM-aware telemetry available through AgentgatewayBackend is complementary — it adds LLM-specific attributes such as token counts, model name, and operation type at the gateway level.

Prerequisites

Before you begin, ensure you have the following components installed and configured:

Configure the KServe LLMInferenceService controller to attach generated routes to the shared agentgateway Gateway. During the GAIE v1 migration, KServe also installs transitional CRDs that its controller uses for compatibility:

export KSERVE_VERSION=v0.20.0-rc0

helm upgrade -i kserve-llmisvc-resources \
oci://ghcr.io/kserve/charts/kserve-llmisvc-resources \
--version $KSERVE_VERSION \
--namespace kserve \
--set kserve.controller.deploymentMode=Standard \
--set kserve.controller.gateway.ingressGateway.enableGatewayApi=true \
--set kserve.controller.gateway.ingressGateway.createGateway=false \
--set kserve.controller.gateway.ingressGateway.kserveGateway=kserve/kserve-ingress-gateway \
--set kserve.controller.gateway.ingressGateway.className=agentgateway \
--set kserve.controller.gateway.disableIstioVirtualHost=true \
--set kserve.controller.gateway.disableIngressCreation=false \
--set kserve.controller.knativeAddressableResolver.enabled=false \
--set kserve.controller.gateway.localGateway.gateway="" \
--set kserve.controller.gateway.localGateway.gatewayService=""

Wait for the updated controller, then apply the final GAIE v1.5.0 CRD bundle. Applying the bundle after the KServe chart updates the stable API definitions and retains KServe's transitional CRDs:

kubectl rollout status deployment/llmisvc-controller-manager \
--namespace kserve \
--timeout=240s

kubectl apply --server-side -f \
https://github.com/kubernetes-sigs/gateway-api-inference-extension/releases/download/v1.5.0/manifests.yaml

Install or upgrade agentgateway after the GAIE CRDs so that its controller discovers InferencePool, then install the matching KServe runtime configuration:

export AGENTGATEWAY_VERSION=v1.4.1

helm upgrade -i agentgateway-crds \
oci://cr.agentgateway.dev/charts/agentgateway-crds \
--version $AGENTGATEWAY_VERSION \
--namespace agentgateway-system \
--create-namespace

helm upgrade -i agentgateway \
oci://cr.agentgateway.dev/charts/agentgateway \
--version $AGENTGATEWAY_VERSION \
--namespace agentgateway-system \
--set inferenceExtension.enabled=true

helm upgrade -i kserve-runtime-configs \
oci://ghcr.io/kserve/charts/kserve-runtime-configs \
--version $KSERVE_VERSION \
--namespace kserve \
--set kserve.llmisvcConfigs.enabled=true

KServe creates the InferencePool and deploys the llm-d Router endpoint picker from its runtime configuration. Do not install the llm-d Router Helm chart separately for this workflow.

Deploy LLMInferenceService

Create Namespace

kubectl create namespace kserve-test

Create Gateway

Create a shared agentgateway Gateway resource in the kserve namespace. Routes from model namespaces can attach to this Gateway:

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: kserve-ingress-gateway
namespace: kserve
spec:
gatewayClassName: agentgateway
listeners:
- name: http
protocol: HTTP
port: 80
allowedRoutes:
namespaces:
from: All
infrastructure:
labels:
serving.kserve.io/gateway: kserve-ingress-gateway

Deploy Your Model

Deploy an LLMInferenceService. This example uses a small model for demonstration; replace with your model of choice:

apiVersion: serving.kserve.io/v1alpha2
kind: LLMInferenceService
metadata:
name: my-model
namespace: kserve-test
spec:
model:
uri: "hf://Qwen/Qwen2.5-0.5B-Instruct"
name: Qwen/Qwen2.5-0.5B-Instruct
replicas: 1
router:
route: {}
scheduler: {}
template:
containers:
- name: main
resources:
limits:
nvidia.com/gpu: 1
requests:
nvidia.com/gpu: 1

Wait for the LLMInferenceService to be ready:

kubectl wait --for=condition=Ready llminferenceservice/my-model \
-n kserve-test --timeout=300s

Use the Standard InferencePool Backend

The managed scheduler and default LLMInferenceServiceConfig route template generate an InferencePool and an HTTPRoute that references it. agentgateway supports this backend directly, so no AgentgatewayBackend or route override is required for standard inference routing.

Verify the generated backend reference:

kubectl get httproute my-model-kserve-route \
-n kserve-test \
-o jsonpath='{.spec.rules[?(@.name=="v1-chat-completions-path")].backendRefs[0]}'

If you do not need AI policies, continue to Configure the gateway URL.

Optional: Configure an AgentgatewayBackend for AI Policies

To use AI policies that require LLM-aware request or response processing, create an AgentgatewayBackend and override the generated route to reference it.

Step 1: Create AgentgatewayBackend

Create an AgentgatewayBackend that wraps the InferencePool generated by KServe. This tells agentgateway to activate its LLM pipeline while retaining the pool's endpoint selection:

apiVersion: agentgateway.dev/v1alpha1
kind: AgentgatewayBackend
metadata:
name: my-model-backend
namespace: kserve-test
spec:
ai:
provider:
custom:
backendRef:
group: inference.networking.k8s.io
kind: InferencePool
name: my-model-inference-pool
model: Qwen/Qwen2.5-0.5B-Instruct
formats:
- type: Completions
path: /v1/chat/completions
tip

The generated pool name is {llminferenceservice-name}-inference-pool. Verify it with:

kubectl get inferencepool -n kserve-test

Step 2: Override HTTPRoute backendRef

You have two options to override the backendRef in KServe's auto-generated HTTPRoutes.

Override the route configuration on an individual LLMInferenceService using spec.router.route.http:

apiVersion: serving.kserve.io/v1alpha2
kind: LLMInferenceService
metadata:
name: my-model
namespace: kserve-test
spec:
model:
uri: "hf://Qwen/Qwen2.5-0.5B-Instruct"
name: Qwen/Qwen2.5-0.5B-Instruct
replicas: 1
router:
scheduler: {}
route:
http:
spec:
parentRefs:
- group: gateway.networking.k8s.io
kind: Gateway
name: kserve-ingress-gateway
namespace: kserve
rules:
- backendRefs:
- group: agentgateway.dev
kind: AgentgatewayBackend
name: my-model-backend
matches:
- path:
type: PathPrefix
value: /v1/chat/completions
timeouts:
backendRequest: 0s
request: 0s
- backendRefs:
- group: agentgateway.dev
kind: AgentgatewayBackend
name: my-model-backend
matches:
- path:
type: PathPrefix
value: /v1/completions
timeouts:
backendRequest: 0s
request: 0s
template:
containers:
- name: main
resources:
limits:
nvidia.com/gpu: 1
requests:
nvidia.com/gpu: 1

Step 3: Attach Token-Based Rate Limiting

Apply an AgentgatewayPolicy to enforce token-based rate limits on the route:

apiVersion: agentgateway.dev/v1alpha1
kind: AgentgatewayPolicy
metadata:
name: token-ratelimit
namespace: kserve-test
spec:
targetRefs:
- group: gateway.networking.k8s.io
kind: HTTPRoute
name: my-model-kserve-route
traffic:
rateLimit:
local:
- tokens: 10000
unit: Hours

Configure $GATEWAY_URL

Check if your Gateway has an external IP address assigned:

kubectl get svc -n kserve \
-l gateway.networking.k8s.io/gateway-name=kserve-ingress-gateway

If the EXTERNAL-IP shows an actual IP address (not <pending>):

export GATEWAY_URL="http://$(kubectl get gateway -n kserve kserve-ingress-gateway \
-o jsonpath='{.status.addresses[0].value}')"

Testing the Integration

Set the request path for the backend you selected:

export GATEWAY_PATH="/kserve-test/my-model/v1/chat/completions"

Send a test request:

curl -s "$GATEWAY_URL$GATEWAY_PATH" \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-0.5B-Instruct",
"messages": [{"role": "user", "content": "Hello"}]
}' | jq .

Verify AI Policy Processing

If you configured an AgentgatewayBackend, check the agentgateway logs to confirm the LLM pipeline is active:

kubectl logs -n kserve deploy/kserve-ingress-gateway --tail=10

With AgentgatewayBackend, you should see GenAI fields in the log:

route=kserve-test/my-model-kserve-route
http.status=200
protocol=llm
gen_ai.operation.name=chat
gen_ai.request.model=Qwen/Qwen2.5-0.5B-Instruct
gen_ai.response.model=Qwen/Qwen2.5-0.5B-Instruct
gen_ai.usage.input_tokens=12
gen_ai.usage.output_tokens=15

How It Works

For standard inference routing, KServe generates an HTTPRoute whose backendRef points to InferencePool (kind: InferencePool, group: inference.networking.k8s.io). agentgateway implements the Gateway API Inference Extension and routes this traffic through the pool's endpoint picker.

For AI policies, overriding spec.router.route.http changes the backendRef to AgentgatewayBackend (kind: AgentgatewayBackend, group: agentgateway.dev). The AgentgatewayBackend wraps the generated InferencePool, so agentgateway can parse OpenAI request and response payloads, extract token usage, emit GenAI telemetry, and enforce token-based policies without bypassing the endpoint picker.

Next Steps