On-Premises Deployment
Lynx AI Agent supports on-premises deployments for organizations that require data to remain within their own infrastructure - whether in a private cloud or a fully air-gapped network.
Both deployment types share the same backend setup. The only difference is the AI engine:
- Private Cloud - AI inference runs via a managed cloud provider, fully within your cloud tenant.
- Air-gapped - AI inference runs on a self-hosted open-weight model with no internet connectivity whatsoever.
AI Engine
The Lynx AI backend requires an OpenAI-compatible completions endpoint. How you source it depends on your deployment type.
Private Cloud
For private cloud deployments, use your cloud provider's managed AI service as the inference endpoint. Popular choices include:
- Amazon Bedrock
- Google Vertex AI
- Microsoft Foundry
All of these expose an OpenAI-compatible API and support Bearer key authentication.
Air Gapped
Warning
A local inference engine is a pre-requisite for an air-gapped environment.
In an air-gapped deployment, Lynx AI Agent relies on the power of local LLMs for Splunk reasoning, tool execution, and guardrail protections (including malicious payload and prompt-injection checks).
We continuously evaluate new checkpoints, publish benchmarks, and update the supported model list as the industry evolves. Learn more about our supported models here.
The inference endpoint must expose an OpenAI-compatible Completions API, be reachable from the backend over either HTTP or HTTPS, and may accept an optional Bearer key.
You can test the connection to the inference endpoint like this:
curl -H "Authorization: Bearer <API_KEY>" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.1","messages":[{"role":"user","content":"What is the meaning of life?"}]}' \
https://example.com/api/v1/chat/completions # (1)!
- Set the base URL to the inference endpoint in your organization.
Lynx AI Backend
The backend ships as a hardened container image, deployable with Docker, Docker Compose, Kubernetes, or OpenShift.
Minimum allocation: 1 CPU core, 1 GB RAM, and a stable network path to the model endpoint. The backend is stateless and can be scaled horizontally behind a load balancer.
Warning
The backend image and the Splunk app are released together, and the app only talks to a backend whose major and minor version match its own. Upgrade both at the same time.
We recommend pinning the image to the matching version tag rather than tracking :latest.
-
Create the following
.envfile:# .env # App ONPREM_MODE=true # (1)! TRUSTED_HOSTS=<host1>[,<host2>] # (2)! PORT=<port> # (3)! SSL_ENABLED=<boolean> # (4)! IPV6=<boolean> # (5)! # Inference INFERENCE_URL=<url> # (6)! INFERENCE_API_KEY=<api_key> INFERENCE_URL2=<url> # (7)! INFERENCE_API_KEY2=<api_key> GUARDRAIL=<boolean> # (8)! COMPACT_MAX_TOKENS=<number> # (9)! AUXILIARY_MODEL=<model> # (10)! REPLAY_REASONING=<boolean> # (11)! PROMPT_CACHING=<boolean> # (12)! VERIFY_SSL=<boolean> # (13)!- This is required to be set to
truewhen deployed in an air-gapped environment. - Hostnames that the backend can accept requests on.
- Port the server listens on. Default:
3000 - Whether to listen for HTTPS requests.
falsemeans HTTP.truerequires SSL certificates to be mounted into the container. Default:false - Bind to
::(IPv6) instead of0.0.0.0(IPv4). Default:false - Full URL of the primary inference endpoint, including the completions path (e.g.
https://example.com/api/v1/chat/completions). You can also useINFERENCE_URL1. - Optional secondary inference endpoint. You can configure multiple endpoints by incrementing the number (e.g.,
INFERENCE_URL3,INFERENCE_URL4). Models are routed to one of them withinference_indexin ai.conf. - Enable request filtering via a lightweight LLM guardrail. Default:
true - Maximum number of tokens to be generated by the compaction process. Default:
2000 - AI model used for guardrail protections and auxiliary tasks like chat title generation and context compaction. Must be a model served by the primary inference endpoint, named exactly as that endpoint serves it - auxiliary requests always go to
INFERENCE_URL, regardless of the model the user selected. Default:deepseek/deepseek-v4-flash-0731 - Convert replayed
reasoning_contenttoreasoning_detailsinstead of stripping it. Usually not required for on-premises deployments. Default:false - Mark the longest messages of a conversation for upstream prompt caching. Disable it if your inference engine rejects the
cache_controlfield. Default:true - Whether to verify the SSL certificate of the inference endpoint. Default:
true
- This is required to be set to
-
Run the container image with the
.envfile:If
SSL_ENABLED=true, mount a certificate chain and private key into the container at/ssl/cert.pemand/ssl/privkey.pem:
Tip
Ensure the certificate files are readable by the container process (e.g., chmod 644).
OpenShift Container Platform (OCP)
On OpenShift, the same configuration is supplied as container env entries rather than an .env file.
Apply the following manifests with oc apply -f, replacing the <YOUR-*> placeholders with your own values:
apiVersion: apps/v1
kind: Deployment
metadata:
name: lynx-ai-backend
namespace: <YOUR-NAMESPACE>
labels:
app.kubernetes.io/name: lynx-ai-backend
app.kubernetes.io/instance: lynx-ai-backend
spec:
replicas: 1 # (1)!
selector:
matchLabels:
app.kubernetes.io/name: lynx-ai-backend
app.kubernetes.io/instance: lynx-ai-backend
template:
metadata:
labels:
app.kubernetes.io/name: lynx-ai-backend
app.kubernetes.io/instance: lynx-ai-backend
spec:
containers:
- name: lynx-ai-backend
image: <YOUR-IMAGE> # (2)!
ports:
- name: http
containerPort: 3000
protocol: TCP
env:
- name: TZ
value: <YOUR-TIMEZONE>
- name: ONPREM_MODE
value: 'true'
- name: TRUSTED_HOSTS
value: <YOUR-HOST> # (3)!
- name: INFERENCE_URL
value: <YOUR-INFERENCE-URL>
- name: INFERENCE_API_KEY
value: <YOUR-API-KEY> # (4)!
- name: GUARDRAIL
value: 'true'
- name: AUXILIARY_MODEL
value: <YOUR-AUXILIARY-MODEL>
resources:
requests:
memory: 512Mi
cpu: 500m
limits:
memory: 1Gi
cpu: '1'
securityContext:
readOnlyRootFilesystem: true # (5)!
volumeMounts:
- name: tmp
mountPath: /tmp
- name: var-tmp
mountPath: /var/tmp
volumes:
- name: tmp
emptyDir:
sizeLimit: 64Mi
- name: var-tmp
emptyDir:
sizeLimit: 64Mi
- The backend is stateless, so you can raise the replica count freely - nothing else needs to be scaled alongside it.
- Full image reference, as pushed to your container registry (e.g.
registry.example.com/lynx-ai-backend:v0.13.1). - Should match the hostname exposed by the route. If you let OpenShift auto-generate that hostname, create the route first and fill this in afterwards.
- Consider storing credentials in a secret and referencing them with
valueFrom.secretKeyRefinstead of writing them directly into the manifest. - The backend writes nothing outside
/tmpand/var/tmp, both of which are mounted below, so its root filesystem can stay read-only (optional).
apiVersion: v1
kind: Service
metadata:
name: lynx-ai-backend
namespace: <YOUR-NAMESPACE>
labels:
app.kubernetes.io/name: lynx-ai-backend
app.kubernetes.io/instance: lynx-ai-backend
spec:
type: ClusterIP
ports:
- name: http
port: 3000
targetPort: http
protocol: TCP
selector:
app.kubernetes.io/name: lynx-ai-backend
app.kubernetes.io/instance: lynx-ai-backend
apiVersion: route.openshift.io/v1
kind: Route
metadata:
name: lynx-ai-backend
namespace: <YOUR-NAMESPACE>
annotations:
haproxy.router.openshift.io/timeout: 300s
labels:
app.kubernetes.io/name: lynx-ai-backend
app.kubernetes.io/instance: lynx-ai-backend
spec:
host: <YOUR-HOST> # (1)!
path: /
to:
kind: Service
name: lynx-ai-backend
weight: 100
port:
targetPort: http
tls:
termination: edge
insecureEdgeTerminationPolicy: Redirect
wildcardPolicy: None
- Omit this field to let OpenShift auto-generate a hostname from the route and namespace names.
Monitoring
The backend exposes two unauthenticated operational endpoints, both outside the versioned API prefix:
/health- Returns{"status": "ok"}while the backend is up./metrics- Prometheus metrics for request rate, latency, in-progress requests, and status codes.
You should keep them reachable only from your monitoring stack.
Splunk App
When deploying the Splunk app, in addition to setting up the details of your backend server, you must also set onprem_mode = true and specify your available on-prem AI models in local/ai.conf.
See ai.conf for more details.
From this point on, setting up the Splunk app is the same as for an internet-connected deployment.
Follow the instructions in Installation and Configuration to complete the setup.