> ## Documentation Index
> Fetch the complete documentation index at: https://docs.getlimina.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Kubernetes Autoscaling Guide

> How to autoscale the Limina container in Kubernetes with the Horizontal Pod Autoscaler, driven by the backlog depth metrics that the container exposes to Prometheus.

This guide describes how to scale the number of Limina Pods automatically with the Kubernetes Horizontal Pod Autoscaler (HPA), using the Prometheus metrics that the container exposes on its `/metrics` route. The example below scales on `text_processing_queue_depth`, the per-pod average backlog of in-flight text processing calls.

## How autoscaling works

The HPA is a built-in Kubernetes controller that adjusts a Deployment's replica count to keep an observed metric near a target value. On its own it can only read CPU and memory. To scale on an application metric such as `text_processing_queue_depth`, the Prometheus Adapter is needed to publish that metric through the Kubernetes custom metrics API.

The steps in this guide build that chain end to end:

```text Metrics Pipeline theme={"theme":"poimandres"}
Limina Pod (/metrics)
  -> ServiceMonitor -> Prometheus                 # scrape
  -> Prometheus Adapter -> custom.metrics.k8s.io  # publish
  -> HPA -> Deployment replicas                   # scale
```

## Available metrics

Limina exposes three backlog depth gauges, one per pipeline layer. Any of them can be used as the scaling signal.

| Metric | Layer | What it counts |
| - | - | - |
| `request_depth` | HTTP | Requests received but not yet completed, including those still waiting for a free worker. |
| `queue_depth` | Worker | De-identification requests currently being handled by a worker. |
| `text_processing_queue_depth` | Text processing | Worker task executing the text processing pipeline. |

The `queue_depth` meter covers every workload type, while `text_processing_queue_depth` counts only the text processing workload which may include the NER, the coreference resolution, the synthetic entity generation and the relation extraction.

<Info>
  This guide uses `text_processing_queue_depth`, which best tracks the bottleneck for text processing workloads. For image or audio files, the inference work (OCR, object detection, audio transcription) is not counted by this metric — use `queue_depth` as the scaling signal instead. For more information on the `/metrics` route, see [Deploying into Production](/installation/deploying-into-production#metrics).
</Info>

## 1. Prerequisites

Before you begin, make sure that your cluster and nodegroup have been created and that Limina Pods are running. See the [Kubernetes Setup Guide](/installation/kubernetes-setup-guide) for the Deployment and Service manifests, and the [AWS EKS Setup Guide](/installation/aws/eks-guide) if you are running on EKS.

The examples in this guide assume a nodegroup defined as follows.

```shell eksctl Command wrap theme={"theme":"poimandres"}
eksctl create nodegroup \
--cluster limina-cluster \
--region us-east-1 \
--name limina-ng \
--node-ami-family AmazonLinux2023 \
--node-type g4dn.4xlarge \
--nodes 2 \
--nodes-min 1 \
--nodes-max 3 \
--ssh-access \
--ssh-public-key test-pk
```

<Warning>
  Set `--nodes-max` to at least the HPA's `maxReplicas`. The Limina GPU container occupies one GPU per Pod, so the number of available GPUs caps the Pod count.
</Warning>

## 2. Install Prometheus

```shell Helm Command wrap theme={"theme":"poimandres"}
kubectl create namespace monitoring
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm install prometheus prometheus-community/kube-prometheus-stack --namespace monitoring
```

expected output

```text Output theme={"theme":"poimandres"}
NAME: prometheus
LAST DEPLOYED: Fri Sep 11 01:07:26 2026
NAMESPACE: monitoring
STATUS: deployed
REVISION: 1
DESCRIPTION: Install complete
TEST SUITE: None
NOTES:
kube-prometheus-stack has been installed. Check its status by running:
  kubectl --namespace monitoring get pods -l "release=prometheus"
...
```

## 3. Configure the Prometheus Adapter

The adapter is what makes Prometheus metrics visible to the HPA. Create a values file as follows.

```yaml Prometheus Adapter Values lines wrap theme={"theme":"poimandres"}
# prometheus-adapter-values.yaml
prometheus:
  url: http://prometheus-kube-prometheus-prometheus.monitoring.svc
  port: 9090

rules:
  default: false # Drop the built-in cpu/memory rules; only custom metrics from Limina.
  custom:
    # Text NER layer depth, smoothed over 2 minutes, exposed per pod so the HPA
    # can use a Pods-type metric and divide by the replica count.
    - seriesQuery: 'text_processing_queue_depth{namespace!="",pod!=""}'
      resources:
        overrides:
          namespace:
            resource: namespace
          pod:
            resource: pod # Required for Pods-type HPA metrics
      name:
        matches: "^text_processing_queue_depth$"
        as: "text_processing_queue_depth_avg2m" # Renamed because the value is transformed
      metricsQuery: "avg_over_time(<<.Series>>{<<.LabelMatchers>>}[2m])"
```

<Note>
  Always set `prometheus.url` explicitly as shown. The chart's default value does not resolve.
</Note>

Now install the adapter with this values file.

```shell Helm Command wrap theme={"theme":"poimandres"}
helm install prometheus-adapter prometheus-community/prometheus-adapter --namespace monitoring -f prometheus-adapter-values.yaml
```

expected output

```text Output theme={"theme":"poimandres"}
NAME: prometheus-adapter
LAST DEPLOYED: Fri Sep 11 01:12:25 2026
NAMESPACE: monitoring
STATUS: deployed
REVISION: 1
DESCRIPTION: Install complete
TEST SUITE: None
NOTES:
prometheus-adapter has been deployed.
In a few minutes you should be able to list metrics using the following command(s):

  kubectl get --raw /apis/custom.metrics.k8s.io/v1beta1
```

## 4. Create an internal ClusterIP Service

This step is required when Limina is exposed through an ELB. It provides a monitoring-only Service for the ServiceMonitor to select. If Limina is only exposed internally, skip this step and see [Without an ELB](#without-an-elb) below.

```yaml Monitoring Service Manifest lines wrap theme={"theme":"poimandres"}
# internal-service-clusterip.yaml
#
# Monitoring-only Service. No application traffic flows through it; it exists
# solely so the ServiceMonitor has a labelled Service to select, which in turn
# populates the Endpoints that Prometheus scrapes.
# Application traffic goes through the LoadBalancer Service (private-ai-service).
apiVersion: v1
kind: Service
metadata:
  name: private-ai-app-internal-svc
  namespace: default # Namespace the app runs in
  labels:
    app: private-ai-app # Selected by the ServiceMonitor's spec.selector
spec:
  type: ClusterIP # Cluster-internal only; no ELB is provisioned
  selector:
    app: private-ai-app # Matches the app's Pod labels
  ports:
    - name: metrics # Referenced by name from the ServiceMonitor's endpoints[].port
      port: 8080 # Numeric only; a port name here is invalid
      targetPort: 8080 # Port the Pod listens on
```

Apply it with this `kubectl` command.

```shell Kubernetes Command theme={"theme":"poimandres"}
kubectl apply -f internal-service-clusterip.yaml
```

## 5. Create the ServiceMonitor

The ServiceMonitor tells Prometheus which Pods to scrape.

```yaml ServiceMonitor Manifest lines wrap theme={"theme":"poimandres"}
# service-monitor.yaml
#
# Tells the Prometheus Operator to scrape the Limina app. Prometheus does not go through
# the Service ClusterIP or the ELB: it reads the Service's Endpoints and scrapes
# the Pod IPs directly.
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: private-ai-app-monitor
  namespace: monitoring
  labels:
    release: prometheus # Required: matches kube-prometheus-stack's serviceMonitorSelector
spec:
  namespaceSelector:
    matchNames:
      - default # Namespace of the target Service
  selector:
    matchLabels:
      app: private-ai-app # Matches the Service's metadata.labels (not the Pod labels)
  endpoints:
    - port: "metrics" # The Service port's *name*, not its number
      path: /metrics
      interval: 15s
```

Apply it with this `kubectl` command.

```shell Kubernetes Command theme={"theme":"poimandres"}
kubectl apply -f service-monitor.yaml
```

expected output

```text Output theme={"theme":"poimandres"}
servicemonitor.monitoring.coreos.com/private-ai-app-monitor created
```

### Without an ELB

If the app already has a ClusterIP Service, step 4 is unnecessary — point the ServiceMonitor at that Service instead. Match `selector` to its labels and `port` to its port name.

```yaml ServiceMonitor Manifest lines wrap theme={"theme":"poimandres"}
# service-monitor-internal.yaml
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
  name: private-ai-app-monitor
  namespace: monitoring
  labels:
    release: prometheus
spec:
  namespaceSelector:
    matchNames:
      - default
  selector:
    matchLabels:
      app: private-ai-app # Labels on the app's existing ClusterIP Service
  endpoints:
    - port: "http" # Whatever that Service names its port
      path: /metrics
      interval: 15s
```

## 6. Create the HPA

```yaml HPA Manifest lines wrap theme={"theme":"poimandres"}
# hpa-manifest.yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: private-ai-app-hpa
  namespace: default # Must match the target Deployment's namespace
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: private-ai-deployment
  minReplicas: 1
  maxReplicas: 3
  metrics:
    # Text NER layer depth. This metric counts in-flight text NER inference calls, so it
    # best tracks text processing workloads (for image/audio workloads, scale on queue_depth).
    - type: Pods # Per-pod metric; the HPA divides the sum by the replica count
      pods:
        metric:
          name: text_processing_queue_depth_avg2m # 2-minute average, via prometheus-adapter
        target:
          type: AverageValue
          averageValue: "4" # Example threshold; tune to your own workload
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 0 # The 2-minute average already absorbs spikes
    scaleDown:
      stabilizationWindowSeconds: 300 # GPU pods are costly to churn
```

Apply it with this `kubectl` command.

```shell Kubernetes Command theme={"theme":"poimandres"}
kubectl apply -f hpa-manifest.yaml
```

expected output

```text Output theme={"theme":"poimandres"}
horizontalpodautoscaler.autoscaling/private-ai-app-hpa created
```

<Tip>
  `averageValue` is the per-pod backlog that the HPA holds the Deployment at, so a lower value scales out earlier. Pick it relative to the number of simultaneous requests each container is expected to handle — see [Concurrency](/configuration-and-operations/container-management/concurrency) for the recommended levels — then tune it against your own traffic patterns.
</Tip>

## 7. Verify the setup

First, check that all Pods have started successfully.

```shell Kubernetes Command theme={"theme":"poimandres"}
kubectl get pods --all-namespaces
```

expected output

```text Output theme={"theme":"poimandres"}
NAMESPACE     NAME                                                     READY   STATUS    RESTARTS   AGE
default       private-ai-deployment-5bb54b4979-bgnvv                   1/1     Running   0          65m
kube-system   aws-node-npt9d                                           2/2     Running   0          74m
kube-system   coredns-6fbcc88588-dlzdc                                 1/1     Running   0          81m
kube-system   kube-proxy-8spzc                                         1/1     Running   0          74m
kube-system   nvidia-device-plugin-daemonset-7hfp9                     1/1     Running   0          73m
monitoring    alertmanager-prometheus-kube-prometheus-alertmanager-0   2/2     Running   0          40m
monitoring    prometheus-adapter-df85459bc-rdmm5                       1/1     Running   0          38m
monitoring    prometheus-grafana-75f9c7f5d4-d9qfh                      3/3     Running   0          40m
monitoring    prometheus-kube-prometheus-operator-798b6cd6b5-4g86t     1/1     Running   0          40m
monitoring    prometheus-kube-state-metrics-bddfd945-8zfkm             1/1     Running   0          40m
monitoring    prometheus-prometheus-kube-prometheus-prometheus-0       2/2     Running   0          40m
monitoring    prometheus-prometheus-node-exporter-5wgd9                1/1     Running   0          40m
```

Then check that Kubernetes recognizes the custom metric.

```shell Kubernetes Command wrap theme={"theme":"poimandres"}
kubectl get --raw "/apis/custom.metrics.k8s.io/v1beta1" | grep text_processing
```

expected output

```text Output theme={"theme":"poimandres"}
{"kind":"APIResourceList","apiVersion":"v1","groupVersion":"custom.metrics.k8s.io/v1beta1","resources":[{"name":"namespaces/text_processing_queue_depth_avg2m","singularName":"","namespaced":false,"kind":"MetricValueList","verbs":["get"]},{"name":"pods/text_processing_queue_depth_avg2m","singularName":"","namespaced":true,"kind":"MetricValueList","verbs":["get"]}]}
```

Once `pods/text_processing_queue_depth_avg2m` is returned, the metric is available to the HPA.

## 8. Test autoscaling

Replace `<your-load-balancer>` in the commands below with your own ELB hostname.

Before submitting any requests, the metric reads zero.

```shell Command wrap theme={"theme":"poimandres"}
curl http://<your-load-balancer>.us-east-1.elb.amazonaws.com/metrics | grep text_processing_queue_depth
```

expected output

```text Output theme={"theme":"poimandres"}
...
# HELP text_processing_queue_depth Number of text NER inference calls currently in-flight (text processing depth).
# TYPE text_processing_queue_depth gauge
text_processing_queue_depth 0.0
```

After submitting 5 concurrent redaction requests, the backlog is reflected in the gauge.

```shell Command wrap theme={"theme":"poimandres"}
curl http://<your-load-balancer>.us-east-1.elb.amazonaws.com/metrics | grep text_processing_queue_depth
```

expected output

```text Output theme={"theme":"poimandres"}
...
# HELP text_processing_queue_depth Number of text NER inference calls currently in-flight (text processing depth).
# TYPE text_processing_queue_depth gauge
text_processing_queue_depth 5.0
```

The `averageValue: "4"` threshold has been exceeded, so the HPA starts a second Pod.

```shell Kubernetes Command theme={"theme":"poimandres"}
kubectl get pods --all-namespaces
```

expected output

```text Output theme={"theme":"poimandres"}
NAMESPACE     NAME                                                     READY   STATUS              RESTARTS   AGE
default       private-ai-deployment-5bb54b4979-bgnvv                   1/1     Running             0          78m
default       private-ai-deployment-5bb54b4979-ksprm                   0/1     ContainerCreating   0          27s
...
```

Once the backlog clears, the replica count returns to `minReplicas`. This takes at least 5 minutes, because `scaleDown.stabilizationWindowSeconds: 300` holds the count steady until the metric has stayed low for that long.

```shell Kubernetes Command theme={"theme":"poimandres"}
kubectl get pods --all-namespaces
```

expected output

```text Output theme={"theme":"poimandres"}
NAMESPACE     NAME                                                     READY   STATUS    RESTARTS   AGE
default       private-ai-deployment-5bb54b4979-bgnvv                   1/1     Running   0          110m
...
```

## Additional Resources

* [Kubernetes Setup Guide](/installation/kubernetes-setup-guide) — the Deployment and Service manifests that this guide scales.
* [Concurrency](/configuration-and-operations/container-management/concurrency) — recommended simultaneous requests per container, for choosing your HPA target.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.