WEBCON Self-hosted AI Proxy installation guide

Run WEBCON AI Proxy with a fully local LLM using Docker, LiteLLM, Ollama and Qwen3:8b. A step-by-step guide for a secure self-hosted AI setup.

WEBCON Self-hosted AI Proxy installation guide

This post provides a complete, step-by-step setup of the self-hosted WEBCON AI Proxy, based on the official AI Proxy Self-hosted | WEBCON documentation and extended with the additional configuration, networking, certificate and troubleshooting steps required to build the environment from scratch.

Running WEBCON AI Proxy with a Local LLM using Docker, LiteLLM, Ollama and Qwen3:8b

WEBCON 2026 R1 introduces the option to run the WEBCON AI Proxy in a self-hosted environment, providing greater control over how and where the AI infrastructure is operated.

WEBCON provides two deployment variants for the AI Proxy: the WEBCON-hosted cloud service and AI Proxy self-hosted. With the self-hosted variant, the AI Proxy runs within the organization’s own infrastructure. It can connect to cloud-based AI services such as Microsoft Foundry (formerly Azure AI Foundry), OpenAI or Google Vertex AI, as well as OpenAI-compatible APIs such as LiteLLM.

This guide focuses on a fully local approach. The WEBCON AI Proxy, LiteLLM, Ollama and Qwen3:8b are hosted within the organization’s own infrastructure, without requiring an external AI provider.

Running the complete AI stack locally can be an important building block for digital sovereignty and data protection. Prompts, documents and generated content remain within the organization’s own infrastructure instead of being sent to an external AI provider. This provides greater control over data, infrastructure and the choice of AI models, which can be particularly relevant in regulated or privacy-sensitive environments.

The architecture used in this guide is:

WEBCON Application Server
│
│ HTTPS :8081 (exposed)
▼
Self-hosted WEBCON AI Proxy
│
│ OpenAI-compatible API
▼
LiteLLM :4000 (not exposed)
│
│ Ollama API
▼
Ollama :11434 (not exposed)
│
▼
Qwen3:8b

Choosing the Docker host

The WEBCON AI Proxy can also be deployed using Docker Desktop on Windows. WEBCON provides an official Docker Desktop example that is particularly useful for evaluation, development and testing.

This guide deliberately uses a dedicated Ubuntu Server VM as the Docker host. For a server-based deployment, Ubuntu provides a lightweight environment with relatively low operating system resource overhead, leaving more CPU and memory available for the AI workloads. It also provides a straightforward Docker and Docker Compose environment without requiring Docker Desktop. Depending on the existing licensing model, using Linux for a dedicated VM may also avoid additional Windows Server licensing costs.

Running the AI stack on a separate host also keeps its resources independent from the WEBCON Application Server. Local LLM inference can consume significant CPU, memory and, when available, GPU resources. Separating these workloads reduces resource contention with WEBCON services and provides clearer boundaries for maintenance, updates, troubleshooting and future scaling.

This is an architectural choice rather than a general requirement. Windows with Docker Desktop, a dedicated Linux VM or another suitable container environment may be the better choice depending on the existing infrastructure, operational requirements, licensing and intended workload. The appropriate architecture should therefore be evaluated individually for each environment.

The AI components run as Docker containers on the same VM and communicate through a dedicated Docker bridge network. Only the HTTPS endpoint of the WEBCON AI Proxy is exposed to the surrounding infrastructure. The AI Proxy acts as the intermediary between WEBCON and the configured AI providers, routing requests to the appropriate provider and model and returning the generated response to WEBCON.

Important:

This guide is based on my own implementation and the knowledge available at the time of writing. It does not claim to be complete or error-free. Feedback, corrections and new findings are very welcome and will be incorporated into future updates whenever possible.

This guide aims to lower the barrier to entry to self-hosted AI by providing detailed step-by-step instructions and additional background for readers with different levels of AI and IT experience.

Sources:

Target architecture

The resulting environment consists of three containers:

Ubuntu VM
│
└── Docker
│
└── Docker Compose: webcon-ai-stack
│
├── webcon-ai-proxy
│ └── HTTPS :8081 (exposed)
│
├── litellm
│ └── :4000 (not exposed)
│
└── ollama
├── :11434 (not exposed)
└── qwen3:8b

The persistent configuration and data will be stored below:

/opt/webcon-ai-stack/
├── docker-compose.yml
├── .env
│
├── webcon-ai-proxy/
│ ├── aiconfiguration.json
│ └── certificates/
│ ├── certificate.pem (used by WEBCON AI Proxy and Designer Studio)
│ ├── certificate.key
│ └── ai-proxy-public-certificate-for-certlm.cer
│
├── litellm/
│ └── config.yaml
│
└── ollama/

Using /opt keeps the configuration of the complete AI stack in one clearly identifiable location. Strictly speaking, variable application data could also be stored below /var/opt, but keeping the configuration and persistent Docker data together below /opt/webcon-ai-stack makes backup, migration and administration of this self-contained stack straightforward.

Virtual Machine - Ubuntu

The chosen operating system in this setup is “Ubuntu Server 26.04 LTS”. The ISO file can be downloaded here: Get Ubuntu Server | Download | Ubuntu

Create a new virtual machine in your hypervisor and attach the downloaded Ubuntu Server ISO as the installation media. The creation and configuration of the virtual machine itself depends on the hypervisor being used and is therefore not covered in this guide.

VM resource sizing

The hardware requirements mainly depend on the model being used and the expected workload. The following values are practical starting points for the Qwen3:8b setup described in this guide and should not be considered official WEBCON or Ollama hardware requirements.

The hardware requirements mainly depend on the model being used and the expected workload. The following values are practical starting points for the Qwen3:8b setup described in this guide and should not be considered official WEBCON or Ollama hardware requirements.

ResourceMinimumRecommendedPerformance
vCPU4816+
RAM12 GB16 GB32 GB+
GPUNot requiredNot requiredDedicated GPU with sufficient VRAM
Disk space25 GB40 GB60 GB+
Qwen3:8b inferenceCPUCPUGPU accelerated
Text analysisSuitable for testingSuitable for regular workloadsSignificantly faster
ReasoningSlow for complex tasksUsableStrongly benefits from GPU acceleration

The Recommended configuration of 8 vCPU and 16 GB RAM corresponds to the development environment used for the setup described in this guide.

A dedicated GPU is not required to run Qwen3:8b. CPU-only inference works well for testing and moderate workloads, but response times depend heavily on the underlying physical CPU, available memory bandwidth, prompt and context size, generated output length and concurrent requests.

As a rough indication from the environment used for this guide, a simple text prompt can take approximately 10–30 seconds when running Qwen3:8b CPU-only. More complex prompts, larger inputs and reasoning tasks can take considerably longer. These values are only practical observations and should not be considered benchmarks.

Qwen3 supports both thinking and non-thinking modes. For simple tasks where extensive reasoning is not required, adding /no_think to the prompt can disable the model’s thinking mode and significantly reduce the amount of generated reasoning and therefore the overall response time. For example:

Summarize the following text in three sentences. /no_think

For tasks that benefit from deeper reasoning, the default thinking behavior can be retained or explicitly requested using /think.

The Recommended configuration is suitable for development, testing and moderate workloads. For production environments, especially with multiple users, concurrent requests or more demanding AI tasks, the Performance configuration should be considered as the preferred starting point. Actual resource requirements should always be validated against the expected workload and the underlying hardware.

Installation of Ubuntu

The installation is straightforward and does not require any special Ubuntu configuration for this setup. Ubuntu Server does not include a desktop environment by default. If required, a GUI can be installed separately. Instructions can be found in How to Install Desktop (GUI) on Ubuntu Server.

Prepare Ubuntu

First update the Ubuntu installation.

sudo apt update
sudo apt upgrade -y

Verify the operating system:

cat /etc/os-release | grep PRETTY_NAME

Expected result:

PRETTY_NAME="Ubuntu ..."

Install some basic tools used throughout this guide:

sudo apt install -y ca-certificates curl openssl jq

Verify OpenSSL:

openssl version

Expected result:

OpenSSL 3.x.x ...

Install Docker Engine and Docker Compose

Docker recommends installing Docker Engine through its official APT repository instead of using the Docker packages shipped by the Linux distribution.

Source:

Install Docker Engine on Ubuntu | Docker Docs

Preparation steps

Follow these steps from the official Docker installation guide:

Remove potentially conflicting packages first:

sudo apt remove $(dpkg --get-selections docker.io docker-compose docker-compose-v2 docker-doc podman-docker containerd runc | cut -f1)

It is not a problem if some of these packages are not installed.

Set up Docker’s apt repository.

# Add Docker's official GPG key:
sudo apt update
sudo apt install ca-certificates curl
sudo install -m 0755 -d /etc/apt/keyrings
sudo curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc
sudo chmod a+r /etc/apt/keyrings/docker.asc
# Add the repository to Apt sources:
sudo tee /etc/apt/sources.list.d/docker.sources <<EOF
Types: deb
URIs: https://download.docker.com/linux/ubuntu
Suites: $(. /etc/os-release && echo "${UBUNTU_CODENAME:-$VERSION_CODENAME}")
Components: stable
Architectures: $(dpkg --print-architecture)
Signed-By: /etc/apt/keyrings/docker.asc
EOF

Important: Make sure that outbound HTTPS traffic from the Ubuntu VM to download.docker.com is allowed.

Update the package index:

sudo apt update

Error handling

If apt update returns the following error, Docker may not yet provide a repository for your Ubuntu release. In this setup, Ubuntu 26.04 uses the codename “resolute”:

Error: The repository 'https://download.docker.com/linux/ubuntu resolute InRelease' is not signed.

You can then fix this by changing the Ubuntu version codename in the Docker repository configuration. Open the file with nano:

sudo nano /etc/apt/sources.list.d/docker.sources

In the nano editor, replace the value in the line starting with “Suites:”. Ubuntu 26.04 uses the codename “resolute”. If Docker does not yet provide a repository for this release, use the Ubuntu 24.04 LTS repository with the codename “noble” instead.

Types: deb
URIs: https://download.docker.com/linux/ubuntu
Suites: noble
Components: stable
Architectures: amd64
Signed-By: /etc/apt/keyrings/docker.asc

Press CTRL+O to save the file and CTRL+X to exit nano. Then retry the update and check if the error is gone:

sudo apt update

Complete installation

Install Docker Engine and the Docker Compose plugin:

sudo apt install -y \
docker-ce \
docker-ce-cli \
containerd.io \
docker-buildx-plugin \
docker-compose-plugin

Verify the Docker installation:

sudo docker --version

Expected result:

Docker version ...

Verify Docker Compose:

sudo docker compose version

Expected result:

Docker Compose version v...

Finally run Docker’s test container:

sudo docker run --rm hello-world

Expected result:

Hello from Docker!
This message shows that your installation appears to be working correctly.

Useful Docker commands

Show the status of all services in the stack:

cd /opt/webcon-ai-stack
sudo docker compose ps

Show logs:

sudo docker compose logs --tail 100

Follow logs:

sudo docker compose logs -f

Restart the stack:

sudo docker compose restart

Stop the stack:

sudo docker compose down

Start it again:

sudo docker compose up -d

Pull the configured container images and recreate the stack if newer images are available:

sudo docker compose pull
sudo docker compose up -d

Important: Be careful when using moving tags such as latest or main-latest in production. docker compose pull may download a newer application version, which will be used when the containers are recreated with docker compose up -d.

Final system updates

You can check if there are any apt updates left by running:

apt list --upgradable

And if there are results you can just upgrade them:

sudo apt upgrade -y

Docker Network Configuration

This chapter covers Docker networking modes and subnets.

Docker bridge network vs. host networking

Docker containers can use different networking modes.

With host networking, a container shares the network stack of the Ubuntu host. A service listening on port 4000, for example, is effectively listening directly through the host network stack. Port mappings are not used in this mode.

With a bridge network, Docker creates an isolated virtual network. Containers can communicate with each other without making their ports accessible outside the Docker network.

Docker Compose also provides DNS resolution within the Docker network. Therefore, the containers can address each other using their service names:

http://litellm:4000
http://ollama:11434

There is no need to assign static container IP addresses.

For this installation, the containers will therefore use a dedicated Docker bridge network.

Source:

Docker – Networking in Compose

Dedicated Docker subnet

Docker supports defining custom IPAM subnets directly in a Compose network. We will explicitly define the Docker subnet:

networks:
webcon-ai:
name: webcon-ai
driver: bridge
ipam:
config:
- subnet: 172.30.0.0/24

Important: 172.30.0.0/24 is only an example.

Before using it, make sure that the subnet is not already used by:

  • your physical LAN,
  • server VLANs,
  • VPN networks,
  • routed corporate networks,
  • or another Docker network on the host.

Overlapping networks can result in difficult-to-diagnose routing problems.

Check the networks and routes already known to the Ubuntu host:

ip route

Check existing Docker networks:

sudo docker network ls

Inspect their configured subnets:

sudo docker network inspect $(sudo docker network ls -q) \
--format '{{.Name}} {{range .IPAM.Config}}{{.Subnet}}{{end}}'

Expected result:

bridge 172.x.x.x/...
...

Make sure the subnet selected for webcon-ai does not overlap with any network found in the host routing table or the existing Docker networks.

If the example subnet overlaps with an existing network, choose another private subnet and use it in the Docker Compose configuration shown later in this guide.

Source:

Docker – Define and manage networks in Docker Compose

Create the directory structure

Create the root directory for the webcon-ai-stack:

sudo mkdir -p /opt/webcon-ai-stack

Create the required subdirectories:

sudo mkdir -p \
/opt/webcon-ai-stack/webcon-ai-proxy/certificates \
/opt/webcon-ai-stack/litellm \
/opt/webcon-ai-stack/ollama

Verify:

find /opt/webcon-ai-stack -maxdepth 5 -type d

Expected result:

/opt/webcon-ai-stack
/opt/webcon-ai-stack/webcon-ai-proxy
/opt/webcon-ai-stack/webcon-ai-proxy/certificates
/opt/webcon-ai-stack/litellm
/opt/webcon-ai-stack/ollama

Change to the stack directory:

cd /opt/webcon-ai-stack

Create the TLS certificate

WEBCON AI Proxy supports certificates in PEM or PFX format. When using PEM, the file used by the AI Proxy must contain both the certificate and its private key. For development environments, WEBCON explicitly allows the use of a self-signed certificate.

Important: For production environments, a certificate issued by a trusted internal or public Certificate Authority (CA) should be used instead of a self-signed certificate.

Source:

WEBCON – Docker Configuration

Change to the certificate directory:

cd /opt/webcon-ai-stack/webcon-ai-proxy/certificates

Replace the FQDN in the following command with the actual DNS name of the AI Proxy.

Generate the private key and self-signed certificate:

sudo openssl req \
-x509 \
-newkey rsa:4096 \
-sha256 \
-days 3650 \
-nodes \
-keyout certificate.key \
-out ai-proxy-public-certificate-for-certlm.cer \
-subj "/CN=webcon-ai-proxy-dev.example.local" \
-addext "subjectAltName=DNS:webcon-ai-proxy-dev.example.local"

This initially creates the private key certificate.key and public certificate ai-proxy-public-certificate-for-certlm.cer.

Now create the combined PEM file containing both the certificate and its private key:

sudo sh -c 'cat certificate.key ai-proxy-public-certificate-for-certlm.cer > certificate.pem'

You now have three certificate files:

certificate.key
certificate.pem
ai-proxy-public-certificate-for-certlm.cer

The files serve the following purposes:

  • certificate.key
    • Private key.
    • Stays only on the Ubuntu AI Proxy host.
  • certificate.pem
    • Combined certificate + private key.
    • Used by the WEBCON AI Proxy and later in WEBCON Designer Studio AI Settings.
    • Its contents will later be copied to the WEBCON Application Server.
  • ai-proxy-public-certificate-for-certlm.cer
    • Public certificate only.
    • Its contents will later be copied to the WEBCON Application Server and imported into the Windows Local Computer → Trusted Root Certification Authorities store.

Verify the certificate:

openssl x509 \
-in ai-proxy-public-certificate-for-certlm.cer \
-noout \
-subject \
-issuer \
-dates

Expected result:

subject=CN = webcon-ai-proxy-dev.example.local
issuer=CN = webcon-ai-proxy-dev.example.local
notBefore=...
notAfter=...

Important: Also note the notAfter value. The certificate must be renewed and replaced in both the AI Proxy and WEBCON before this date.

Check the Subject Alternative Name:

openssl x509 \
-in ai-proxy-public-certificate-for-certlm.cer \
-noout \
-ext subjectAltName

Expected result:

X509v3 Subject Alternative Name:
DNS:webcon-ai-proxy-dev.example.local

Important: The DNS name in the Subject Alternative Name (SAN) must match the hostname used later in the WEBCON AI Proxy network address in Designer Studio.

Verify that the combined PEM file contains the private key and certificate:

sudo grep "BEGIN" certificate.pem

Expected result:

-----BEGIN PRIVATE KEY-----
-----BEGIN CERTIFICATE-----

Important: The certificate configured in WEBCON AI Settings must contain both the RSA private key and the public certificate. Using only the public certificate causes loading the AI model list to fail with No RSA private key found in PEM.

Protect the private key and certificate files:

sudo chmod 600 certificate.key

Verify:

ls -l

Expected result:

-rw-r--r-- ... ai-proxy-public-certificate-for-certlm.cer
-rw------- ... certificate.key
-rw-r--r-- ... certificate.pem

Important: Only certificate.pem and ai-proxy-public-certificate-for-certlm.cer need to be copied to the WEBCON Application Server later.

Configure LiteLLM

WEBCON documents LiteLLM as an OpenAI-compatible intermediary that can connect the AI Proxy to locally hosted models running through Ollama.

Source:

WEBCON – LiteLLM

Create the configuration:

sudo nano /opt/webcon-ai-stack/litellm/config.yaml

Add:

model_list:
- model_name: qwen3-8b
litellm_params:
model: ollama/qwen3:8b
api_base: http://ollama:11434
general_settings:
master_key: sk-change-this-to-a-long-random-key

The important distinction is:

qwen3-8b

is the model name exposed by LiteLLM, while:

qwen3:8b

is the actual Ollama model.

The AI Proxy configuration will therefore use the LiteLLM model name qwen3-8b.

Note about the LiteLLM master_key

Security note: The master_key shown above is only an example. Choose a long, randomly generated secret for your environment and do not reuse the example value. The master_key protects access to LiteLLM and should therefore be treated like any other API credential. The WEBCON AI Proxy can later authenticate against LiteLLM using this master key or a dedicated LiteLLM virtual key.

A random key can, for example, be generated with:

openssl rand -hex 32

Expected result:

9d3c...<64 hexadecimal characters>...a74f

Copy the generated value and use it as the master_key in config.yaml. The same value will also be configured as the API key for the LiteLLM connection in the WEBCON AI Proxy Config Admin UI.

Verify:

sudo grep -E "model_name|model:|api_base|master_key:" \
/opt/webcon-ai-stack/litellm/config.yaml | sed 's/\(master_key:\).*/\1 <configured>/'

Expected result:

model_name: qwen3-8b
model: ollama/qwen3:8b
api_base: http://ollama:11434
master_key: <configured>

Prepare the AI Proxy configuration

The WEBCON AI Proxy stores its provider and model configuration in aiconfiguration.json. Instead of starting with an empty configuration, we can already provide the basic LiteLLM configuration. This configuration can later be reviewed and modified through the Config Admin UI.

Create the file:

sudo nano /opt/webcon-ai-stack/webcon-ai-proxy/aiconfiguration.json

Add the following configuration:

{
"ProviderConnections": {
"LiteLlm": {
"Description": "Local LiteLLM",
"Type": "OpenAiApiCompatible",
"ProviderConfiguration": {
"ApiKey": "YOUR_LITELLM_MASTER_KEY",
"Endpoint": "http://litellm:4000",
"InlineFileAttachments": true,
"InlineFileAttachmentsMaxBytes": 20971520
}
}
},
"ProviderModels": [
{
"Id": "11111111-1111-1111-1111-111111111111",
"ConnectionName": "LiteLlm",
"Priority": 100,
"Name": "Qwen3:8b",
"Description": "Local Qwen3:8b via LiteLLM and Ollama",
"TextModel": {
"ModelName": "qwen3-8b"
},
"Factor": 1.0
}
],
"AiTaskTypesConfiguration": {
"transcribeAudio": [],
"aiAgent": [
"11111111-1111-1111-1111-111111111111"
],
"processBuilder": [
"11111111-1111-1111-1111-111111111111"
],
"concierge": [
"11111111-1111-1111-1111-111111111111"
],
"textSentiment": [
"11111111-1111-1111-1111-111111111111"
],
"textSummary": [
"11111111-1111-1111-1111-111111111111"
],
"prompt": [
"11111111-1111-1111-1111-111111111111"
],
"translation": [
"11111111-1111-1111-1111-111111111111"
],
"imageGeneration": [],
"embeddingGeneration": [],
"conciergeAgent": [
"11111111-1111-1111-1111-111111111111"
]
}
}

Replace:

YOUR_LITELLM_MASTER_KEY

with the master_key generated and configured in the previous LiteLLM section.

Save the file in nano using Ctrl+O, confirm with Enter, and exit with Ctrl+X.

Validate the JSON syntax:

jq . /opt/webcon-ai-stack/webcon-ai-proxy/aiconfiguration.json

Expected result:

jq prints the formatted JSON configuration without reporting a parse error.

The Config Admin UI must be able to modify aiconfiguration.json. Therefore, the file must not be mounted read-only.

For the initial setup, grant write access:

sudo chmod 666 /opt/webcon-ai-stack/webcon-ai-proxy/aiconfiguration.json

Verify:

ls -l /opt/webcon-ai-stack/webcon-ai-proxy/aiconfiguration.json

Expected result:

-rw-rw-rw- ... aiconfiguration.json

Security note: chmod 666 is convenient for the initial setup but unnecessarily permissive for a permanent installation. After determining the UID/GID under which the AI Proxy container runs, ownership should be assigned accordingly and the permissions tightened.

The configuration can later be reviewed and modified through the WEBCON AI Proxy Config Admin UI in the browser:

https://webcon-ai-proxy-dev.example.local:8081/config-ui

The example assigns the local Qwen3 model to the text-based AI task types configured in AiTaskTypesConfiguration, while audio transcription, image generation and embedding generation remain unassigned.

Create the Docker Compose configuration

Return to the stack directory:

cd /opt/webcon-ai-stack

Create:

sudo nano docker-compose.yml

Use the following configuration:

name: webcon-ai-stack
services:
ollama:
image: ollama/ollama:latest
container_name: ollama
restart: unless-stopped
volumes:
- /opt/webcon-ai-stack/ollama:/root/.ollama
networks:
- webcon-ai
litellm:
image: ghcr.io/berriai/litellm:main-latest
container_name: litellm
restart: unless-stopped
command:
- "--config=/app/config.yaml"
volumes:
- /opt/webcon-ai-stack/litellm/config.yaml:/app/config.yaml:ro
depends_on:
- ollama
networks:
- webcon-ai
webcon-ai-proxy:
image: webconbps/aiproxy:2026.2.135.35
container_name: webcon-ai-proxy
restart: unless-stopped
ports:
- "8081:8081"
environment:
ASPNETCORE_ENVIRONMENT: Production
AppConfiguration__SelfHosted__Enabled: true
AppConfiguration__SelfHosted__Certificate__Path: /app/certificates/certificate.pem
# Only if file is PFX, don't place password in plain text here!
#AppConfiguration__SelfHosted__Certificate__Password: <pfx-pwd>
AppConfiguration__SelfHosted__UseAzureKeyVault: false
Logging__LogLevel__Default: Information
Logging__LogLevel__Microsoft: Warning
Kestrel__Endpoints__Https__Protocols: Http1
# For Admin UI Config, don't place password in plain text here!
ConfigAdmin__AccessKey: ${AIPROXY_CONFIG_ADMIN_ACCESS_KEY}
volumes:
- /opt/webcon-ai-stack/webcon-ai-proxy/aiconfiguration.json:/app/aiconfiguration.json:rw
- /opt/webcon-ai-stack/webcon-ai-proxy/certificates/certificate.pem:/app/certificates/certificate.pem:ro
depends_on:
- litellm
networks:
- webcon-ai
networks:
webcon-ai:
name: webcon-ai
driver: bridge
ipam:
config:
- subnet: 172.30.0.0/24

About the image versions

This example uses:

webconbps/aiproxy:2026.2.135.35
ghcr.io/berriai/litellm:main-latest
ollama/ollama:latest

At the time of writing this blog post, the WEBCON AI Proxy version is explicitly pinned to 2026.2.135.35, matching the version used in the official WEBCON AI Proxy Getting Started example.

Important:

main-latest and latest, however, are moving tags. A deployment performed later may therefore pull different versions of LiteLLM or Ollama.

For a reproducible production deployment, replace these tags with the exact LiteLLM and Ollama versions that have been tested in your environment.

Source:

WebCon-AiProxy-GettingStarted/DockerDesktop at main · WEBCON-BPS/WebCon-AiProxy-GettingStarted

Configure secrets

Do not put the Config Admin access key directly into docker-compose.yml.

Create:

sudo nano /opt/webcon-ai-stack/.env

Add:

AIPROXY_CONFIG_ADMIN_ACCESS_KEY=replace-with-a-long-random-secret

Protect the file:

sudo chmod 600 /opt/webcon-ai-stack/.env

Verify:

ls -l /opt/webcon-ai-stack/.env

Expected result:

-rw------- ... /opt/webcon-ai-stack/.env

Do not commit this file to a Git repository.

The aiconfiguration.json file may also contain sensitive provider API credentials. WEBCON explicitly warns that this file should therefore not be committed to source control.

Source:

WEBCON – AIConfiguration Configuration

Validate Docker Compose

Before starting anything, let Docker validate the complete configuration without printing the resolved configuration with the quiet parameter because of the AIPROXY_CONFIG_ADMIN_ACCESS_KEY:

cd /opt/webcon-ai-stack
sudo docker compose config --quiet

Expected result:

The command completes without an error. If the YAML is invalid, Docker reports the affected configuration before any containers are created.

Pull the container images

Pull the container images defined in docker-compose.yml:

sudo docker compose pull

Verify:

sudo docker images

Expected result should include images for:

webconbps/aiproxy
ghcr.io/berriai/litellm
ollama/ollama

Start Ollama first

Start only Ollama:

sudo docker compose up -d ollama

Verify:

sudo docker compose ps ollama

Expected result:

ollama ... Up

Verify that the Ollama API is responding:

sudo docker exec ollama ollama list

Expected result:

NAME ID SIZE MODIFIED

At this point, the model list may be empty. This confirms that the Ollama API is responding correctly.

Check its logs:

sudo docker compose logs --tail 20 ollama

The container should start without fatal errors.

Notice that Ollama does not have a ports: section.

Port 11434 therefore exists inside the Docker network but is not published on the Ubuntu host.

Download Qwen3:8b

Pull the model directly inside the Ollama container:

sudo docker exec ollama ollama pull qwen3:8b

Depending on the network connection, downloading the model may take some time.

Verify:

sudo docker exec ollama ollama list

Expected result should contain a column:

NAME
qwen3:8b

Test the model:

sudo docker exec ollama ollama run qwen3:8b "Reply only with: OK"

Expected result should contain the model-generated response OK.

At this point:

Ollama
│
└── qwen3:8b ✓

is working independently of LiteLLM and WEBCON.

Start LiteLLM

Start LiteLLM:

sudo docker compose up -d litellm

Verify:

sudo docker compose ps litellm

Expected result:

litellm ... Up

Check the logs:

sudo docker compose logs --tail 50 litellm

The logs should show that LiteLLM started successfully and is listening on port 4000 inside the Docker network.

Again, there is deliberately no:

ports:
- "4000:4000"

LiteLLM is reachable from the other containers but not directly from the LAN.

Test LiteLLM from inside the Docker network

Because LiteLLM is intentionally not exposed to the host network, test it from another container attached to the same Docker network.

Replace <your-litellm-master-key> before executing this command:

sudo docker run --rm \
--network webcon-ai \
curlimages/curl \
http://litellm:4000/v1/models \
-H "Authorization: Bearer <your-litellm-master-key>"

Expected result is JSON containing the LiteLLM model:

qwen3-8b

Now test an actual completion and replace <your-litellm-master-key> before executing this command:

sudo docker run --rm \
--network webcon-ai \
curlimages/curl \
-s http://litellm:4000/v1/chat/completions \
-H "Authorization: Bearer <your-litellm-master-key>" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3-8b",
"messages": [
{
"role": "user",
"content": "Reply only with: OK"
}
]
}'

Expected result:

A JSON response containing OK as the assistant response generated by Qwen3.

The request path is now:

curl
↓
LiteLLM
↓
Ollama
↓
Qwen3:8b

Start WEBCON AI Proxy

Start the final service:

sudo docker compose up -d webcon-ai-proxy

Verify all containers:

sudo docker compose ps

Expected result:

NAME STATUS
ollama Up
litellm Up
webcon-ai-proxy Up

Check the AI Proxy logs:

sudo docker compose logs --tail 50 webcon-ai-proxy

Expected result:

The AI Proxy initializes without certificate or configuration errors.

In particular, there should no longer be a warning stating that the Config Admin UI is enabled but ConfigAdmin:AccessKey is not configured.

Verify the Docker network

Inspect the network:

sudo docker network inspect webcon-ai

Expected result:

The output contains all three containers:

ollama
litellm
webcon-ai-proxy

and addresses from:

172.30.0.0/24

The individual container IP addresses do not need to be configured manually because Docker’s internal DNS resolves the service names.

Verify exposed ports

Check the ports published by Docker:

sudo docker compose ps

Expected result:

Only the AI Proxy should contain a published host port:

0.0.0.0:8081->8081/tcp

LiteLLM and Ollama should not have published ports.

You can additionally check the Ubuntu host:

sudo ss -lntp

From this AI stack, only HTTPS port 8081 should be published on the Ubuntu host.

In particular, these should not be published:

4000
11434

This is intentional.

The resulting security boundary is:

LAN
│
│ HTTPS :8081
▼
WEBCON AI Proxy
│
├──────── Docker bridge network ────────┐
│ │
▼ │
LiteLLM :4000 │
│ │
▼ │
Ollama :11434 │
│ │
▼ │
Qwen3:8b │
│
not directly reachable from LAN ┘

For additional security, the network firewall should restrict TCP port 8081 so that only the WEBCON Application Server(s) can reach the AI Proxy.

Configure WEBCON AI Proxy

Open the Config Admin UI of the WEBCON AI Proxy through its HTTPS endpoint (https://webcon-ai-proxy-dev.example.local:8081/config-ui) and authenticate using the Config Admin access key defined in the .env file.

image.png

Provider Connections

Here you can edit the existing LiteLLM provider connection or add a new provider.

image.png

The LiteLLM configuration looks like:

Type:
OpenAiApiCompatible
Endpoint:
http://litellm:4000
API key:
<LiteLLM master or virtual key>

The endpoint deliberately uses litellm instead of localhost, the Ubuntu hostname or an IP address.

Inside the Docker network, litellm resolves directly to the LiteLLM container.

WEBCON documents OpenAiApiCompatible as the provider type for LiteLLM, with the API key sent using the standard Authorization: Bearer <your-litellm-master-key> in the header.

Provider Models

In this tab you can define the models that can be called by the AI Proxy. Each model belongs to a Provider Connection and maps to the provider’s model names (e.g. qwen3-8b in this guide).

image.png

Qwen3:8b Provider Model Configuration

image.png

The model name is using the LiteLLM alias:

qwen3-8b

The complete chain is therefore:

WEBCON AI Proxy
│
│ ModelName: qwen3-8b
▼
LiteLLM
│
│ qwen3-8b → ollama/qwen3:8b
▼
Ollama
│
▼
qwen3:8b

Task Types

For each of the following Task Types you can assign your desired models and their priority order:

  • Transcribe Audio (transcribeAudio)
  • AI Agent (aiAgent)
  • Process Builder (processBuilder)
  • Concierge (concierge)
  • Text Sentiment (textSentiment)
  • Text Summary (textSummary)
  • Prompt (prompt)
  • Translation (translation)
  • Image Generation (imageGeneration)
  • Embedding Generation (embeddingGeneration)
  • Concierge Agent (conciergeAgent)

image.png

Finalize AI Proxy Configuration

Review the configuration in all tabs of the Config Admin UI and save the changes.

Verify that the file was updated in your Ubuntu ssh session:

sudo jq . \
/opt/webcon-ai-stack/webcon-ai-proxy/aiconfiguration.json

Expected result:

Valid JSON containing the configured provider connection and model.

This test is particularly useful because it also confirms that the Config Admin UI has sufficient filesystem permissions to persist its configuration.

Prepare the public certificate for Manage computer certificates

The WEBCON Application Server(s) must trust the self-signed certificate presented by the AI Proxy.

Copy public certificate content

Display the public certificate in your Ubuntu ssh session:

cat \
/opt/webcon-ai-stack/webcon-ai-proxy/certificates/ai-proxy-public-certificate-for-certlm.cer

Expected result:

-----BEGIN CERTIFICATE-----
...
-----END CERTIFICATE-----

Copy the complete block, including both the BEGIN CERTIFICATE and END CERTIFICATE lines.

Do not copy anything containing (if it exists):

BEGIN PRIVATE KEY

On the WEBCON Application Server(s), create the file:

ai-proxy-public-certificate-for-certlm.cer

and paste the copied certificate into the file.

Trust the self-signed certificate on the WEBCON Application Server(s)

The public certificate copied from the Ubuntu AI Proxy host must also be trusted by the Windows server(s) running WEBCON.

The file ai-proxy-public-certificate-for-certlm.cer contains only the public certificate. This certificate can conveniently be imported through the Windows certificate management console.

Open:

certlm.msc

and navigate to in the left pane:

Trusted Root Certification Authorities
└── Certificates

Right click on Certificates and choose All Tasks > Import.

In the Certificate Import Wizard, follow these steps:

  1. Keep “Local Machine” as Store Location and click on Next.
  2. Click on Browse and select the file - ai-proxy-public-certificate-for-certlm.cer and click on Next.
  3. Keep the Certificate Store Trusted Root Certification Authorities and click on Next.
  4. Review your settings and click on Finish to complete the certificate import.

After the import, verify that the certificate appears in the store and that its Subject Alternative Name (SAN) contains the DNS name used for the AI Proxy connection.

The Windows certificate trust can now be tested with PowerShell ISE. This test only verifies that Windows trusts the TLS certificate presented by the AI Proxy; it does not yet test the mTLS authentication or AI Proxy functionality:

Invoke-WebRequest "https://webcon-ai-proxy-dev.example.local:8081" -UseBasicParsing

Before the certificate was trusted, this may fail with an error similar to:

Could not establish trust relationship for the SSL/TLS secure channel.

After importing the certificate into:

Trusted Root Certification Authorities\Certificates

the TLS trust error should no longer occur.

The connection may still fail at this stage because no client certificate is provided for mTLS authentication. The important part of this test is that PowerShell no longer reports a certificate trust error for the certificate presented by the AI Proxy.

After resolving any remaining certificate trust issues, repeat the HTTPS test:

Invoke-WebRequest "https://webcon-ai-proxy-dev.example.local:8081" -UseBasicParsing

Expected result:

The request reaches the WEBCON AI Proxy without a TLS trust error.

The certificate can now be used when configuring the self-hosted AI engine in WEBCON.

Prepare the certificate for WEBCON Designer Studio

WEBCON Designer Studio requires the combined PEM file containing both the certificate and its private key when configuring the self-hosted AI Proxy. The certificate was already created as certificate.pem on the Ubuntu host.

Copy the certificate for WEBCON Designer Studio

Display the combined PEM file in your Ubuntu SSH session:

sudo cat /opt/webcon-ai-stack/webcon-ai-proxy/certificates/certificate.pem

The file must contain both the private key and the certificate:

-----BEGIN PRIVATE KEY-----
...
-----END PRIVATE KEY-----
-----BEGIN CERTIFICATE-----
...
-----END CERTIFICATE-----

Copy the complete content. On the WEBCON Application Server(s), create a file named certificate.pem and paste the complete content into this file.

Configure WEBCON Designer Studio

Open WEBCON Designer Studio on your WEBCON application server and navigate to:
System settings
→ Global parameters
→ AI Settings

Select:

Self-hosted AI

Using the self-hosted AI Proxy requires the AI Proxy Self-hosted license.

Enter the network address of the AI Proxy, for example, in the shown dialog:

https://webcon-ai-proxy-dev.example.local:8081

Provide the combined certificate and private key file:

certificate.pem

WEBCON uses the configured certificate and private key for certificate-based authentication when communicating with the self-hosted AI Proxy.

Restart Designer Studio and WEBCON services

After saving changes in the AI Proxy Config Admin UI, restart WEBCON Designer Studio before continuing with the AI configuration. This ensures that Designer Studio starts with the current AI configuration before the AI Settings are configured.

After changing and saving the AI Settings in Designer Studio, reload the WEBCON service configuration on the Application Server so that the new AI configuration is loaded. This stops the active service tasks and restarts them with the updated configuration. Click “Load configuration” under System settings > Services configuration > Services > .

If the configured self-hosted models do not appear in Designer Studio afterwards, close and restart Designer Studio once more.

A useful sequence after configuration changes is therefore:

Save AI Proxy Config Admin UI configuration
↓
Restart Designer Studio
↓
Configure / save AI Settings
↓
Reload WEBCON service configuration
↓
Restart Designer Studio
↓
Verify the self-hosted models

Verify the self-hosted models

Verify the self-hosted models by calling the following URL in your browser. Replace webcon.example.local with your WEBCON Portal URL and 1 with the ID of the relevant WEBCON database:

https://webcon.example.local/api/studio/db/1/getaimodels?method=Concierge

Expected Result:

[{"id":"11111111-1111-1111-1111-111111111111","name":"[LiteLlm] qwen3-8b (x1)"}]

Conclusion

After a successful connection, the models configured through the AI Proxy become available to the corresponding WEBCON AI functionality.

You can find further guidance on using AI in the following knowledge base article:

Selecting an LLM for AI features in WEBCON

Source:

WEBCON – Global parameters / AI Settings

Communication flow between WEBCON and the self-hosted AI Proxy

The WEBCON Designer Studio does not communicate directly with the AI Proxy. When an AI model is selected or the model list is refreshed, the request is handled by the WEBCON Application Server.

The communication flow is:

WEBCON Designer Studio
│
│ HTTPS
▼
WEBCON Application Server
│
│ HTTPS :8081
│ mTLS / client certificate
▼
WEBCON AI Proxy
│
│ HTTP :4000
▼
LiteLLM
│
│ HTTP :11434
▼
Ollama
│
▼
Qwen3:8b

When Designer Studio requests the available AI models, the following happens:

  1. Designer Studio contacts the WEBCON Application Server

Designer Studio calls the Studio API of the WEBCON Portal. For example:

GET /api/studio/db/<database-id>/getaimodels?method=Concierge

Therefore, Designer Studio itself does not require direct network access to TCP port 8081 of the AI Proxy. 2. The WEBCON Application Server checks AI availability

Before contacting the AI Proxy, WEBCON verifies whether AI functionality is available for the environment, including the corresponding license information. 3. The WEBCON Application Server contacts the AI Proxy

The Application Server uses the AI Proxy endpoint configured under the WEBCON AI settings, for example:

https://webcon-ai-proxy-dev.example.local:8081

Consequently, TCP port 8081 must be reachable from the WEBCON Application Server to the AI Proxy host. 4. The AI Proxy certificate is validated

The Windows server must trust the certificate presented by the AI Proxy.

The public certificate is therefore imported into:

Local Computer
└── Trusted Root Certification Authorities

For a self-signed certificate, the certificate must also contain the AI Proxy FQDN in its Subject Alternative Name (SAN). 5. WEBCON uses the certificate configured in AI Settings

The combined certificate.pem created earlier and configured in WEBCON AI Settings must contain both the private key and the public certificate:

-----BEGIN PRIVATE KEY-----
...
-----END PRIVATE KEY-----
-----BEGIN CERTIFICATE-----
...
-----END CERTIFICATE-----

This is important: providing only the public certificate is not sufficient.

If the private key is missing, requesting the model list fails on the WEBCON Application Server with HTTP 500, while the AI Proxy itself may not show any corresponding request.

The underlying error is:

No RSA private key found in PEM.
  1. AI Proxy forwards AI requests to LiteLLM

After authentication and validation, the AI Proxy uses the provider configuration from aiconfiguration.json and forwards the request to LiteLLM:

http://litellm:4000

Within the Docker bridge network, the service name litellm resolves directly to the LiteLLM container. 7. LiteLLM maps the WEBCON model to Ollama

LiteLLM receives the configured model name:

qwen3-8b

and maps it through config.yaml to the corresponding Ollama model. 8. LiteLLM contacts Ollama

LiteLLM sends the inference request to Ollama:

http://ollama:11434

Ollama then executes the locally installed:

qwen3:8b
  1. The response travels back through the same chain
Qwen3:8b
↓
Ollama
↓
LiteLLM
↓
WEBCON AI Proxy
↓
WEBCON Application Server
↓
Designer Studio / Portal

Network requirements

With the Docker bridge network used in this setup, the relevant communication paths are therefore:

SourceDestinationPortProtocol
Designer StudioWEBCON Application Server443HTTPS
WEBCON Application ServerAI Proxy8081HTTPS
AI ProxyLiteLLM4000HTTP
LiteLLMOllama11434HTTP

Only TCP 8081 needs to cross from the WEBCON infrastructure to the Ubuntu AI host. Ports 4000 and 11434 remain internal to the Docker bridge network and are not exposed on the Ubuntu host.

Important: The WEBCON Designer Studio does not directly connect to the AI Proxy. The WEBCON Application Server acts as the communication endpoint towards the AI Proxy. Therefore, certificate trust, DNS resolution and firewall connectivity to the AI Proxy must be verified from the WEBCON Application Server.

End-to-end verification

At this point the complete request path is:

WEBCON
│
│ HTTPS :8081
▼
WEBCON AI Proxy
│
│ http://litellm:4000
▼
LiteLLM
│
│ http://ollama:11434
▼
Ollama
│
▼
Qwen3:8b

A successful AI request from WEBCON confirms that the complete request path through the AI Proxy, LiteLLM, Ollama and the local model is working.

A simple way to perform an end-to-end test is through the AI Models Configuration in WEBCON Designer Studio. Navigate to System settings > Global parameters > AI Settings, open AI Models Configuration and check if you can find the configured self-hosted model for one of the available AI features.

If WEBCON cannot retrieve the configured AI models, the error:

An error occurred while fetching AI models. Please check your connection to AiProxy and try again.

may appear:

This usually indicates that the WEBCON Application Server cannot establish the required connection to the AI Proxy. Check DNS resolution, network and firewall connectivity to TCP port 8081, certificate trust, and the certificate configured in the WEBCON AI Settings, which must contain both the private key and the public certificate.

If the AI models are retrieved successfully and no connection error appears, select the configured self-hosted model and execute an AI test request. While the request is running, monitor the container logs on the Ubuntu host:

cd /opt/webcon-ai-stack
sudo docker compose logs -f \
webcon-ai-proxy \
litellm \
ollama

Expected result:

The logs should show activity in the AI Proxy, LiteLLM and Ollama while the AI request is processed. If the request completes successfully, the generated response is returned through the same chain to WEBCON.

Stop the log output with Ctrl+C. This only stops log monitoring. It does not stop the containers.

Final Result

The final setup provides a local LLM to WEBCON while keeping Ollama and LiteLLM isolated within the Docker bridge network and exposing only the WEBCON AI Proxy to the corporate network:

Corporate network
WEBCON Application Server
│
│ HTTPS :8081
▼
┌─────────────────────────────────────────────┐
│ Ubuntu VM │
│ │
│ WEBCON AI Proxy │
│ webconbps/aiproxy:2026.2.135.35 │
│ │ │
│ │ OpenAI-compatible API │
│ ▼ │
│ LiteLLM :4000 │
│ Model alias: qwen3-8b │
│ │ │
│ │ Ollama API │
│ ▼ │
│ Ollama :11434 │
│ Model: qwen3:8b │
│ │
│ Docker bridge: webcon-ai │
│ Example: 172.30.0.0/24 │
└─────────────────────────────────────────────┘
Only HTTPS :8081 exposed

This architecture keeps the interfaces between the individual components simple:

  • WEBCON only communicates with WEBCON AI Proxy.
  • AI Proxy sees LiteLLM as an OpenAI-compatible provider.
  • LiteLLM abstracts the locally hosted model.
  • Ollama is responsible for running Qwen3:8b.
  • Docker provides isolation and internal DNS-based service discovery.
  • Only the WEBCON AI Proxy is exposed outside the Docker bridge network through HTTPS port 8081.

The result is a self-hosted AI architecture in which the complete LLM inference path remains within the local infrastructure.

References

Categories

Last updated on

Comments

Comments will be added here later.