Deploy AI Runtime Security using Helm
Deploy AI Runtime Security to a Kubernetes cluster using Helm.
Prerequisites
- Access to a Kubernetes cluster
- kubectl — Kubernetes command-line tool
- Helm — Kubernetes package manager (v3+)
- Resource Requirements — license keys, tools, and scaling guidance
- Hybrid and Disconnected Modes — connection mode details
Create a Helm Values File
Create a values.yaml file to override the chart’s default settings. Application configuration goes inside config.settings.yaml as an embedded YAML string block. Secrets (auth credentials and license) go in the env: block.
Sensitive values in env: should use the secret: prefix followed by the base64-encoded value (e.g., the output of echo -n "my-value" | base64). The chart automatically creates a Kubernetes Secret and injects it via secretKeyRef, so the plaintext value never appears in the pod spec.
Choose the connection mode for your deployment. See Hybrid and Disconnected Modes for details on what data is sent in each mode.
Hybrid
Disconnected
In Hybrid mode, metadata per inference is sent to the HiddenLayer Console to power visualizations and alerting. This requires authentication credentials.
By default, prompts and responses are also sent to the Console so you can review Interactions in context. To disable prompt collection, set log-chat-context to false under aidr-genai.detector.engine.
Deployment
Log In to the HiddenLayer Helm Registry
-
Run the following command in a terminal to log in to the HiddenLayer registry.
- The
usernameis your Registry username. - The
passwordis your License ID. - For more information, see Resource Requirements.
- The
Install
-
Create a
values.yamlfile to customize installation.- See Create a Helm Values File above.
-
Run the following command to deploy Runtime Security.
Verify the Deployment
-
Check that all pods are running:
-
Port-forward the service to your local machine:
-
Verify the health endpoint:
Using the Interactions Endpoint
Once deployed, you can analyze LLM input and output by sending requests to the Interactions endpoint:
For SDK examples and the full response format, see Getting Started with Interactions.
Additional Configuration
The following sections cover additional configuration beyond the baseline. The examples in Create a Helm Values File include working defaults for all of these — adjust as needed after verifying your deployment.
Device Configuration
The aidr-genai.proxy.device.type setting in config.settings.yaml controls which hardware is used for running the ML-based detection models (e.g., prompt injection classifier). It does not affect LLM provider routing.
CPU (Default)
GPU
CPU is the default device type. Allocate 8 CPU units per replica for production workloads.
Horizontal Autoscaling
The chart creates a Horizontal Pod Autoscaler (HPA) when resources.targetUtilization.cpu and resources.requests.cpu are set. Use replicas.min and replicas.max to control scaling bounds.
Allocate 8 CPU per replica and set OMP_NUM_THREADS: 8 to match. Scale replicas to fill node capacity — for example, 4 replicas on a 32-vCPU node. See Resource Requirements for detailed scaling guidance.
Detection Policy
Configure detection policy inside the config.settings.yaml block: