Deployment Overview of gpt-oss-20b on Server¶
Prerequisites and Basic Requirements¶
To ensure the successful deployment and operation of the application, the following requirements must be met:
-
Operating System: Ubuntu
-
Privileges: Root or sudo access is required for package installation and service management.
-
Hardware Acceleration: NVIDIA GPU support (CUDA) is required for running the containerized interface with hardware acceleration.
-
Network Ports:
-
443/TCP: External HTTPS traffic. -
8080/TCP: Internal application access.
FQDN of the final panel on the hostkey.in domain¶
The application is accessible via a specific subdomain template based on the server ID:
| Parameter | Value |
|---|---|
| Prefix | gpt-oss |
| Domain | hostkey.in |
| Full template | gpt-oss{Server_ID}.hostkey.in |
Application installation process¶
The installation follows a multi-stage process involving system-level package management and container orchestration:
-
System Dependencies: The system is updated, and required packages such as
curlare installed. -
Ollama Installation:
-
The Ollama service is installed via the official installation script.
-
A dedicated system user named
ollamais created. -
The
ollama.serviceconfiguration is modified to allow remote connections (OLLAMA_HOST=0.0.0.0) and enable Flash Attention (LLAMA_FLASH_ATTENTION=1). -
Model Deployment: The specific model
gpt-oss:20bis pulled into the local Ollama storage. -
Docker Environment: Docker and NVIDIA CUDA drivers are configured to allow containerized applications to utilize the GPU.
-
Container Orchestration:
-
An Nginx proxy container is deployed via a
compose.ymlfile located in/root/nginx. -
The Open WebUI container is launched using
docker runwith specific environment variables and volume mappings.
Docker Containers and Their Deployment¶
The deployment utilizes the following containers:
open-webui¶
This container provides the web interface for interacting with the LLM.
-
Image:
ghcr.io/open-webui/open-webui:cuda -
Ports:
8080:8080 -
Volumes:
open-webui:/app/backend/data -
Environment Variables:
-
ENV=dev -
OLLAMA_BASE_URLS=http://host.docker.internal:11434 -
Restart Policy:
always -
Additional Configuration: Uses
--gpus allfor hardware acceleration and--add-host=host.docker.internal:host-gatewayto communicate with the host's Ollama service.
nginx¶
This container acts as a reverse proxy to handle SSL termination and routing.
-
Image:
jonasal/nginx-certbot:latest -
Network Mode:
host -
Volumes:
-
nginx_secrets:/etc/letsencrypt -
/data/nginx/user_conf.d:/etc/nginx/user_conf.d -
Restart Policy:
unless-stopped
Custom Scripts and Additional Setup¶
The following actions are performed during the setup process:
-
Ollama Service Modification: The system service file for Ollama is backed up to
/etc/systemd/system/ollama.service.bakbefore being updated with specific environment variables (OLLAMA_ORIGINS,OLLAMA_HOST). -
Nginx Configuration Injection: A configuration line is dynamically added to the Nginx user configuration files in
/data/nginx/user_conf.d/to proxy traffic from the external domain to the internal service athttp://127.0.0.1:8080.
Application Update Instructions¶
To update the main application interface, execute the following command in the directory containing the compose file:
Location of configuration files and data¶
| Component | Path / Resource |
|---|---|
| Nginx Compose File | /root/nginx/compose.yml |
| Nginx User Configs | /data/nginx/user_conf.d/ |
| Open WebUI Data Volume | open-webui (Docker managed) |
| Ollama Models | /usr/share/ollama/.ollama/models/ |
Available ports for connection¶
-
External Access:
443(HTTPS via Nginx proxy) -
Internal API:
8080(Open WebUI) -
Local LLM API:
11434(Ollama)