Skip to main content
The CPU container can be run with the following command:
Docker Command
The command to run the GPU container requires an additional --gpus flag to specify the GPU ID to use. The full GPU container also requires 4GB shared memory, which is set via the --shm-size flag. For the text-only GPU container the shared memory flag is not necessary:
Docker Command
It is recommended to deploy the container on single GPU machines. For multi-GPU machines, please launch a container instance for each GPU and specify the GPU_ID accordingly, as described in Running on Multi-GPU Machines. You can get the GPU_ID using the nvidia-smi command if you have access to runner. You can find more information regarding using GPUs with docker here. For private or public cloud deployment, please see Deployment and the Kubernetes Setup Guide.
crprivateaiprod.azurecr.io is intended for image distribution only. For production use, it is strongly recommended to set up a container registry inside your own compute environment to host the image.

Enhanced Synthetic Container

(new in 4.5) The gpu-synthetic-enhanced image, which includes the enhanced synthetic entity generation system, is deployed exactly like the gpu-synthetic image. It uses the same --gpus and --shm-size flags, and only the image tag differs:
Docker Command
There are, however, three deployment considerations specific to this image:
  • It requires significantly more GPU VRAM than the other GPU flavours. Because of this, we recommend restricting gpu-synthetic-enhanced deployments to instances offering at least 24GB of VRAM on a single GPU. On AWS, we recommend g5.4xlarge (single Nvidia A10G GPU, 24GB VRAM) from the g5 family, or g6e.4xlarge (single Nvidia L40S GPU, 48GB VRAM) from the newer g6e family. On other cloud providers or in a private cloud, choose an equivalent instance or node with a single GPU offering at least 24GB of VRAM.
  • It takes longer to become healthy and start accepting requests than the other images. Allow for a longer startup time before the container is considered unhealthy — if you are deploying on Kubernetes, use the GPU Synthetic Enhanced Configuration manifest in the Kubernetes Setup Guide, which raises the initialDelaySeconds of the readiness and liveness probes accordingly.
  • The VRAM recommendation applies to a single GPU. A container uses only one GPU, so combining several smaller GPUs does not meet the recommendation; for example, two 16GB GPUs are not a substitute for a single 32GB GPU. On multi-GPU machines, you can run one container per GPU as long as each GPU meets the VRAM recommendation on its own.

Running on Multi-GPU Machines

Each GPU container uses a single GPU and does not benefit from additional GPUs. On a machine with multiple GPUs, you can instead run one container per GPU, assigning each container a different GPU ID and host port:
Docker Command
Please keep the following in mind when running multiple GPU containers on the same machine:
  • Assign exactly one GPU to each container. Always specify the GPU ID with the --gpus flag. Depending on the host configuration, a container started without an explicit GPU assignment may be able to see every GPU on the machine and compete with other containers for the same GPU.
  • Do not assign more than one GPU to a container. A container does not get faster with additional GPUs. To increase throughput, add more containers instead.
  • Do not share a GPU between containers. This includes GPU sharing features such as time-slicing. Each container reserves a large portion of its GPU’s memory, so containers sharing a GPU are likely to fail with out-of-memory errors.
  • MIG is not supported. MIG (NVIDIA Multi-Instance GPU) partitions a single physical GPU into multiple isolated GPU instances. MIG is currently not tested or supported, and assigning individual models or features to specific MIG instances is not possible.
  • Size the machine for all containers. The RAM, CPU and shared memory requirements in System Requirements apply to each container, so the machine needs enough resources for all of them combined. Each container checks the available RAM on startup. Finer-grained control over how containers use the CPU is available through advanced environment variables, and a guide on when and how to use them is coming soon. In the meantime, please contact Limina support before changing these values.
  • Use a different host port for each container, and distribute requests across the containers with a load balancer. Do not use --network host, as the containers would then conflict with each other.
  • For the gpu-synthetic-enhanced image, each GPU must meet the VRAM recommendation described in Enhanced Synthetic Container.
For Kubernetes deployments, see the note on multi-GPU nodes in the Kubernetes Setup Guide.

Apple Silicon

It is possible to run the Limina container on Apple Silicon-based Macs, such as the M1 Macbook Pro, even though it is not officially supported. To do this, please make sure you use Docker Desktop 4.25 or later and enable Rosetta2 support. During our testing we didn’t encounter issues with M1 Macs but did encounter some container startup issues on M2 Macs. If this occurs, please try disabling Rosetta2 support. Due to the need to emulate x86 instructions, performance is significantly lower than x86-based machines, let alone GPU-equipped machines. On a M1 Macbook Pro, our tests revealed a throughput of approximately 250 words per second.

Authentication and External Communications

The container makes external communications to Limina’s servers for authentication and usage reporting. To this end, please make sure that the following are reachable:
  • https://verify1.getlimina.ai:443/license-verification/license_status
  • https://verify2.getlimina.ai:443/license-verification/license_status
  • https://app.amberflo.io:443/ingest/
  • https://ingest.amberflo.io:443/ingest/ (failover url for Limina version 4.4.0+)
These communications do not contain any customer data - if training data is required, this must be given to Limina separately. Please see the FAQ for more details on what is sent. An authentication call is made upon the first API call after the Docker image is started, and again at pre-defined intervals based on your subscription.

URI-Based File Support

Running the container with the above commands allows for base64-encoded files to be processed with /process/files/base64. However, to utilize the /process/files/uri route, a volume where input files are stored and PAI_OUTPUT_FILE_DIR must be provided. Note that PAI_OUTPUT_FILE_DIR must reside inside the mounted volume.
Docker Command
For example, if your license file is in your home directory, the input directory you wish to mount is called inputfiles and the output directory is output:
Docker Command