Skip to content

Deployment Overview of gpt-oss-120b on Server

Prerequisites and Basic Requirements

To ensure the successful deployment and operation of the application, the following requirements must be met:

  • Operating System: Ubuntu (recommended).

  • Privileges: Root or sudo access is required for package installation and service management.

  • Hardware Acceleration: NVIDIA GPU support with CUDA installed is required for containerized acceleration.

  • Network Ports:

  • 443 (HTTPS) for external web access.

  • 8080 (HTTP) for internal application communication.

FQDN of the final panel on the hostkey.in domain

The application is accessible via a specific subdomain template based on the server ID.

Parameter Value
Prefix gpt-oss
Domain hostkey.in
Full template gpt-oss{Server_ID}.hostkey.in

Application installation process

The installation involves a multi-stage process including system package management, model retrieval, and container orchestration:

  1. System Dependencies: Installation of Docker and NVIDIA CUDA drivers to support GPU acceleration.

  2. Ollama Installation:

  3. The Ollama service is installed via the official shell script.

  4. A dedicated ollama system user is created.

  5. The systemd service configuration is modified to allow remote connections (OLLAMA_HOST=0.0.0.0) and enable Flash Attention (LLAMA_FLASH_ATTENTION=1).

  6. Model Retrieval: The gpt-oss:20b model is pulled via the Ollama CLI.

  7. Web Interface Deployment: The Open WebUI interface is deployed as a Docker container with GPU support enabled and linked to the local Ollama instance.

  8. Reverse Proxy Setup: An Nginx configuration is generated and managed via Docker Compose to handle SSL termination and traffic routing.

Docker Containers and Their Deployment

The deployment utilizes two primary containers:

Open WebUI Container

  • Image: ghcr.io/open-webui/open-webui:cuda

  • Ports: 8080:8080

  • Volumes: open-webui:/app/backend/data

  • Environment Variables:

  • ENV=dev

  • OLLAMA_BASE_URLS=http://host.docker.internal:11434

  • Restart Policy: always

  • Additional Config: Uses --gpus all for hardware acceleration and --add-host=host.docker.internal:host-gateway to communicate with the host's Ollama service.

Nginx Proxy Container

  • Image: jonasal/nginx-certbot:latest

  • Network Mode: host

  • Volumes:

  • nginx_secrets:/etc/letsencrypt

  • /data/nginx/user_conf.d:/etc/nginx/user_conf.d

  • Restart Policy: unless-stopped

Custom Scripts and Additional Setup

The following actions are performed during the setup process:

  • Service Configuration: The Ollama systemd service is updated to include specific environment variables for network accessibility and performance optimization.

  • Nginx Proxy Routing: A configuration file is dynamically generated in /data/nginx/user_conf.d/ that routes incoming traffic from port 443 to the internal application running on http://127.0.0.1:8080.

Application Update Instructions

To update the main web interface component, execute the following commands in the directory containing the container configuration:

docker compose pull && docker compose up -d

Location of configuration files and data

  • Nginx Configuration Directory: /root/nginx

  • Nginx User Configurations: /data/nginx/user_conf.d

  • Open WebUI Data Volume: open-webui (Docker managed volume)

  • Ollama Models: Managed by the Ollama service in the system path.

Available ports for connection

Port Service Access Type
443 Nginx (HTTPS) External
8080 Open WebUI Internal/Local
question_mark
Is there anything I can help you with?
question_mark
AI Assistant ×