> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://hiddenlayer.ferndocs.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://hiddenlayer.ferndocs.com/_mcp/server.

# Deploy AI Runtime Security using Helm

Deploy AI Runtime Security to a Kubernetes cluster using Helm.

## Prerequisites

* Access to a Kubernetes cluster
* [kubectl](https://kubernetes.io/docs/tasks/tools/) — Kubernetes command-line tool
* [Helm](https://helm.sh/docs/intro/install/) — Kubernetes package manager (v3+)
* [Resource Requirements](/docs/products/runtime/resource_requirements) — license keys, tools, and scaling guidance
* [Hybrid and Disconnected Modes](/docs/products/runtime/hybrid_disconnected) — connection mode details

## Create a Helm Values File

Create a `values.yaml` file to override the chart's default settings. Application configuration goes inside `config.settings.yaml` as an embedded YAML string block. Secrets (auth credentials and license) go in the `env:` block.

> **Handling Secrets**
>
> Sensitive values in `env:` should use the `secret:` prefix followed by the **base64-encoded** value (e.g., the output of `echo -n "my-value" | base64`). The chart automatically creates a Kubernetes Secret and injects it via `secretKeyRef`, so the plaintext value never appears in the pod spec.

Choose the connection mode for your deployment. See [Hybrid and Disconnected Modes](/docs/products/runtime/hybrid_disconnected) for details on what data is sent in each mode.

#### Hybrid

In **Hybrid** mode, metadata per inference is sent to the HiddenLayer Console to power visualizations and alerting. This requires authentication credentials.

By default, prompts and responses are also sent to the Console so you can review Interactions in context. To disable prompt collection, set `log-chat-context` to `false` under `aidr-genai.detector.engine`.

```yaml
aidr_genai:
  config:
    settings.yaml: |
      platform:
        auth-n:
          base-url: "https://auth.hiddenlayer.ai"
        api-connection:
          type: "hybrid"
          base-url: "https://api.us.hiddenlayer.ai"  # US region
          # base-url: "https://api.eu.hiddenlayer.ai"  # EU region
      aidr-genai:
        proxy:
          log-level: "info"
          device:
            type: "cpu"
        detector:
          engine:
            log-chat-context: true  # Set to false to disable sending prompts to the HL Console
  env:
    HL_LLM_PROXY_CLIENT_ID: "secret:<base64-encoded-client-id>"
    HL_LLM_PROXY_CLIENT_SECRET: "secret:<base64-encoded-client-secret>"
    HL_LICENSE: "secret:<base64-encoded-license-key>"
```

#### Disconnected

In **Disconnected** mode, no data is sent back to HiddenLayer. Authentication credentials are not required.

```yaml
aidr_genai:
  config:
    settings.yaml: |
      platform:
        api-connection:
          type: "disabled"
      aidr-genai:
        proxy:
          log-level: "info"
          device:
            type: "cpu"
  env:
    HL_LICENSE: "secret:<base64-encoded-license-key>"
```

## Deployment

### Log In to the HiddenLayer Helm Registry

1. Run the following command in a terminal to log in to the HiddenLayer registry.

   * The `username` is your Registry username.
   * The `password` is your License ID.
   * For more information, see [Resource Requirements](/docs/products/runtime/resource_requirements).

   ```
   helm registry login registry.hiddenlayer.ai --username <email specified for registry> --password <License ID>
   ```

### Install

1. Create a `values.yaml` file to customize installation.

   * See [Create a Helm Values File](#create-a-helm-values-file) above.

2. Run the following command to deploy Runtime Security.

   ```
   helm upgrade --install aidr-genai oci://registry.hiddenlayer.ai/aidr-genai/stable/aidr-genai \
     --namespace aidr-genai --create-namespace \
     -f values.yaml
   ```

### Verify the Deployment

1. Check that all pods are running:

   ```
   kubectl get pods -n aidr-genai
   ```

2. Port-forward the service to your local machine:

   ```
   kubectl port-forward svc/aidr-genai 8000:80 -n aidr-genai
   ```

3. Verify the health endpoint:

   ```
   curl http://localhost:8000/health
   ```

## Using the Interactions Endpoint

Once deployed, you can analyze LLM input and output by sending requests to the Interactions endpoint:

```
curl -X POST http://<service-endpoint>:8000/detection/v1/interactions \
  -H "Content-Type: application/json" \
  -d '{
    "metadata": {
      "model": "gpt-5",
      "requester_id": "user-1234",
      "provider": "openai"
    },
    "input": {
      "messages": [
        {
          "role": "user",
          "content": "What is the largest moon of Jupiter?"
        }
      ]
    },
    "output": {
      "messages": [
        {
          "role": "assistant",
          "content": "The largest moon of Jupiter is Ganymede."
        }
      ]
    }
  }'
```

For SDK examples and the full response format, see [Getting Started with Interactions](/docs/products/runtime/interactions).

## Additional Configuration

The following sections cover additional configuration beyond the baseline. The examples in [Create a Helm Values File](#create-a-helm-values-file) include working defaults for all of these — adjust as needed after verifying your deployment.

### Device Configuration

The `aidr-genai.proxy.device.type` setting in `config.settings.yaml` controls which hardware is used for running the ML-based detection models (e.g., prompt injection classifier). It does not affect LLM provider routing.

#### CPU (Default)

CPU is the default device type. Allocate 8 CPU units per replica for production workloads.

```yaml
aidr_genai:
  config:
    settings.yaml: |
      aidr-genai:
        proxy:
          device:
            type: "cpu"
```

#### GPU

GPU mode uses CUDA acceleration for detection model inference. Mixed-precision (fp16) is enabled automatically when running on GPU.

```yaml
aidr_genai:
  config:
    settings.yaml: |
      aidr-genai:
        proxy:
          device:
            type: "gpu"
  resources:
    limits:
      nvidia.com/gpu: 1
```

> **GPU Image**
>
> GPU mode requires the CUDA-enabled container image.

### Horizontal Autoscaling

The chart creates a Horizontal Pod Autoscaler (HPA) when `resources.targetUtilization.cpu` and `resources.requests.cpu` are set. Use `replicas.min` and `replicas.max` to control scaling bounds.

Allocate **8 CPU per replica** and set `OMP_NUM_THREADS: 8` to match. Scale replicas to fill node capacity — for example, 4 replicas on a 32-vCPU node. See [Resource Requirements](/docs/products/runtime/resource_requirements) for detailed scaling guidance.

```yaml
aidr_genai:
  resources:
    targetUtilization:
      cpu: 75
    requests:
      cpu: 8
      memory: 4096Mi
    limits:
      memory: 4096Mi
  replicas:
    min: 2
    max: 8
```

### Detection Policy

Configure detection policy inside the `config.settings.yaml` block:

```yaml
aidr_genai:
  config:
    settings.yaml: |
      aidr-genai:
        detector:
          prompt-injection:
            enabled: true
            severity: high
            on-alert:
              proxy-action: allow  # allow | block
          personally-identifiable-information:
            enabled: true
            on-alert:
              proxy-action: allow
              redaction-type: replace-with-entity
```