/metrics route. The example below scales on text_processing_queue_depth, the per-pod average backlog of in-flight text processing calls.
How autoscaling works
The HPA is a built-in Kubernetes controller that adjusts a Deployment’s replica count to keep an observed metric near a target value. On its own it can only read CPU and memory. To scale on an application metric such astext_processing_queue_depth, the Prometheus Adapter is needed to publish that metric through the Kubernetes custom metrics API.
The steps in this guide build that chain end to end:
Metrics Pipeline
Available metrics
Limina exposes three backlog depth gauges, one per pipeline layer. Any of them can be used as the scaling signal.
The
queue_depth meter covers every workload type, while text_processing_queue_depth counts only the text processing workload which may include the NER, the coreference resolution, the synthetic entity generation and the relation extraction.
This guide uses
text_processing_queue_depth, which best tracks the bottleneck for text processing workloads. For image or audio files, the inference work (OCR, object detection, audio transcription) is not counted by this metric — use queue_depth as the scaling signal instead. For more information on the /metrics route, see Deploying into Production.1. Prerequisites
Before you begin, make sure that your cluster and nodegroup have been created and that Limina Pods are running. See the Kubernetes Setup Guide for the Deployment and Service manifests, and the AWS EKS Setup Guide if you are running on EKS. The examples in this guide assume a nodegroup defined as follows.eksctl Command
2. Install Prometheus
Helm Command
Output
3. Configure the Prometheus Adapter
The adapter is what makes Prometheus metrics visible to the HPA. Create a values file as follows.Prometheus Adapter Values
Always set
prometheus.url explicitly as shown. The chart’s default value does not resolve.Helm Command
Output
4. Create an internal ClusterIP Service
This step is required when Limina is exposed through an ELB. It provides a monitoring-only Service for the ServiceMonitor to select. If Limina is only exposed internally, skip this step and see Without an ELB below.Monitoring Service Manifest
kubectl command.
Kubernetes Command
5. Create the ServiceMonitor
The ServiceMonitor tells Prometheus which Pods to scrape.ServiceMonitor Manifest
kubectl command.
Kubernetes Command
Output
Without an ELB
If the app already has a ClusterIP Service, step 4 is unnecessary — point the ServiceMonitor at that Service instead. Matchselector to its labels and port to its port name.
ServiceMonitor Manifest
6. Create the HPA
HPA Manifest
kubectl command.
Kubernetes Command
Output
7. Verify the setup
First, check that all Pods have started successfully.Kubernetes Command
Output
Kubernetes Command
Output
pods/text_processing_queue_depth_avg2m is returned, the metric is available to the HPA.
8. Test autoscaling
Replace<your-load-balancer> in the commands below with your own ELB hostname.
Before submitting any requests, the metric reads zero.
Command
Output
Command
Output
averageValue: "4" threshold has been exceeded, so the HPA starts a second Pod.
Kubernetes Command
Output
minReplicas. This takes at least 5 minutes, because scaleDown.stabilizationWindowSeconds: 300 holds the count steady until the metric has stayed low for that long.
Kubernetes Command
Output
Additional Resources
- Kubernetes Setup Guide — the Deployment and Service manifests that this guide scales.
- Concurrency — recommended simultaneous requests per container, for choosing your HPA target.