Skip to content

Deployment Overview of gpt-oss-20b on Server

Prerequisites and Basic Requirements

To ensure the successful deployment and operation of the application, the following requirements must be met:

  • Operating System: Ubuntu

  • Privileges: Root or sudo access is required for package installation and service management.

  • Hardware Acceleration: NVIDIA GPU support (CUDA) is required for running the containerized interface with hardware acceleration.

  • Network Ports:

  • 443/TCP: External HTTPS traffic.

  • 8080/TCP: Internal application access.

FQDN of the final panel on the hostkey.in domain

The application is accessible via a specific subdomain template based on the server ID:

Parameter Value
Prefix gpt-oss
Domain hostkey.in
Full template gpt-oss{Server_ID}.hostkey.in

Application installation process

The installation follows a multi-stage process involving system-level package management and container orchestration:

  1. System Dependencies: The system is updated, and required packages such as curl are installed.

  2. Ollama Installation:

  3. The Ollama service is installed via the official installation script.

  4. A dedicated system user named ollama is created.

  5. The ollama.service configuration is modified to allow remote connections (OLLAMA_HOST=0.0.0.0) and enable Flash Attention (LLAMA_FLASH_ATTENTION=1).

  6. Model Deployment: The specific model gpt-oss:20b is pulled into the local Ollama storage.

  7. Docker Environment: Docker and NVIDIA CUDA drivers are configured to allow containerized applications to utilize the GPU.

  8. Container Orchestration:

  9. An Nginx proxy container is deployed via a compose.yml file located in /root/nginx.

  10. The Open WebUI container is launched using docker run with specific environment variables and volume mappings.

Docker Containers and Their Deployment

The deployment utilizes the following containers:

open-webui

This container provides the web interface for interacting with the LLM.

  • Image: ghcr.io/open-webui/open-webui:cuda

  • Ports: 8080:8080

  • Volumes: open-webui:/app/backend/data

  • Environment Variables:

  • ENV=dev

  • OLLAMA_BASE_URLS=http://host.docker.internal:11434

  • Restart Policy: always

  • Additional Configuration: Uses --gpus all for hardware acceleration and --add-host=host.docker.internal:host-gateway to communicate with the host's Ollama service.

nginx

This container acts as a reverse proxy to handle SSL termination and routing.

  • Image: jonasal/nginx-certbot:latest

  • Network Mode: host

  • Volumes:

  • nginx_secrets:/etc/letsencrypt

  • /data/nginx/user_conf.d:/etc/nginx/user_conf.d

  • Restart Policy: unless-stopped

Custom Scripts and Additional Setup

The following actions are performed during the setup process:

  • Ollama Service Modification: The system service file for Ollama is backed up to /etc/systemd/system/ollama.service.bak before being updated with specific environment variables (OLLAMA_ORIGINS, OLLAMA_HOST).

  • Nginx Configuration Injection: A configuration line is dynamically added to the Nginx user configuration files in /data/nginx/user_conf.d/ to proxy traffic from the external domain to the internal service at http://127.0.0.1:8080.

Application Update Instructions

To update the main application interface, execute the following command in the directory containing the compose file:

docker compose pull && docker compose up -d

Location of configuration files and data

Component Path / Resource
Nginx Compose File /root/nginx/compose.yml
Nginx User Configs /data/nginx/user_conf.d/
Open WebUI Data Volume open-webui (Docker managed)
Ollama Models /usr/share/ollama/.ollama/models/

Available ports for connection

  • External Access: 443 (HTTPS via Nginx proxy)

  • Internal API: 8080 (Open WebUI)

  • Local LLM API: 11434 (Ollama)

question_mark
Is there anything I can help you with?
question_mark
AI Assistant ×