Deployment Overview of gpt-oss-120b on Server¶
Prerequisites and Basic Requirements¶
To ensure the successful deployment and operation of the application, the following requirements must be met:
-
Operating System: Ubuntu (recommended).
-
Privileges: Root or sudo access is required for package installation and service management.
-
Hardware Acceleration: NVIDIA GPU support with CUDA installed is required for containerized acceleration.
-
Network Ports:
-
443(HTTPS) for external web access. -
8080(HTTP) for internal application communication.
FQDN of the final panel on the hostkey.in domain¶
The application is accessible via a specific subdomain template based on the server ID.
| Parameter | Value |
|---|---|
| Prefix | gpt-oss |
| Domain | hostkey.in |
| Full template | gpt-oss{Server_ID}.hostkey.in |
Application installation process¶
The installation involves a multi-stage process including system package management, model retrieval, and container orchestration:
-
System Dependencies: Installation of Docker and NVIDIA CUDA drivers to support GPU acceleration.
-
Ollama Installation:
-
The Ollama service is installed via the official shell script.
-
A dedicated
ollamasystem user is created. -
The systemd service configuration is modified to allow remote connections (
OLLAMA_HOST=0.0.0.0) and enable Flash Attention (LLAMA_FLASH_ATTENTION=1). -
Model Retrieval: The
gpt-oss:20bmodel is pulled via the Ollama CLI. -
Web Interface Deployment: The Open WebUI interface is deployed as a Docker container with GPU support enabled and linked to the local Ollama instance.
-
Reverse Proxy Setup: An Nginx configuration is generated and managed via Docker Compose to handle SSL termination and traffic routing.
Docker Containers and Their Deployment¶
The deployment utilizes two primary containers:
Open WebUI Container¶
-
Image:
ghcr.io/open-webui/open-webui:cuda -
Ports:
8080:8080 -
Volumes:
open-webui:/app/backend/data -
Environment Variables:
-
ENV=dev -
OLLAMA_BASE_URLS=http://host.docker.internal:11434 -
Restart Policy:
always -
Additional Config: Uses
--gpus allfor hardware acceleration and--add-host=host.docker.internal:host-gatewayto communicate with the host's Ollama service.
Nginx Proxy Container¶
-
Image:
jonasal/nginx-certbot:latest -
Network Mode:
host -
Volumes:
-
nginx_secrets:/etc/letsencrypt -
/data/nginx/user_conf.d:/etc/nginx/user_conf.d -
Restart Policy:
unless-stopped
Custom Scripts and Additional Setup¶
The following actions are performed during the setup process:
-
Service Configuration: The Ollama systemd service is updated to include specific environment variables for network accessibility and performance optimization.
-
Nginx Proxy Routing: A configuration file is dynamically generated in
/data/nginx/user_conf.d/that routes incoming traffic from port 443 to the internal application running onhttp://127.0.0.1:8080.
Application Update Instructions¶
To update the main web interface component, execute the following commands in the directory containing the container configuration:
Location of configuration files and data¶
-
Nginx Configuration Directory:
/root/nginx -
Nginx User Configurations:
/data/nginx/user_conf.d -
Open WebUI Data Volume:
open-webui(Docker managed volume) -
Ollama Models: Managed by the Ollama service in the system path.
Available ports for connection¶
| Port | Service | Access Type |
|---|---|---|
443 | Nginx (HTTPS) | External |
8080 | Open WebUI | Internal/Local |