Skip to content

Deployment Overview of Gemma-3-27B on Server

Prerequisites and Basic Requirements

To ensure the successful deployment of the application, the following requirements must be met:

  • Operating System: Ubuntu (recommended for compatibility with NVIDIA drivers).

  • Privileges: Root or sudo access is required for package installation and service management.

  • Hardware Support: An NVIDIA GPU is required to utilize CUDA acceleration for model inference.

  • Network Requirements:

  • Outbound access to ollama.com for the installation script.

  • Outbound access to ghcr.io to pull Docker images.

  • Port 443 (HTTPS) must be open for external traffic.

FQDN of the final panel on the hostkey.in domain

The application is accessible via a specific subdomain template based on your Server ID:

Parameter Value
Prefix gemma
Domain hostkey.in
Full template gemma{Server_ID}.hostkey.in

File and Directory Structure

The following directories are used for configuration, certificates, and application data:

  • /root/nginx: Contains the Nginx reverse proxy configuration files (compose.yml).

  • /data/nginx/user_conf.d: Custom Nginx configuration directory.

  • /data/nginx/nginx-certbot.env: Environment file for Certbot email settings.

  • /etc/systemd/system/ollama.service.bak: Backup of the original Ollama service configuration.

Application installation process

The application is installed using a combination of system package management, shell scripts, and Docker containers:

  1. Ollama Installation: The official Ollama installation script is executed via curl to install the model runner on the host system.

  2. System Service Configuration:

  3. A dedicated ollama user is created.

  4. The ollama.service file is modified to allow remote connections (OLLAMA_HOST=0.0.0.0) and enable Flash Attention (LLAMA_FLASH_ATTENTION=1).

  5. Model Download: The gemma3:27b model is pulled directly via the Ollama CLI.

  6. GPU Acceleration Setup:

  7. NVIDIA Container Toolkit and NVIDIA Container Runtime are installed to allow Docker containers to access host GPU hardware.

  8. The Docker daemon is configured with /etc/docker/daemon.json to set nvidia as the default runtime.

  9. Web Interface Deployment: The Open WebUI interface is deployed as a Docker container using the CUDA-enabled image.

Access Rights and Security

  • Firewall: Ensure port 443 is open for HTTPS traffic.

  • Service Isolation: Ollama runs as a system service with restricted environment variables to manage origin access (OLLAMA_ORIGINS=*).

  • Docker Runtime: The NVIDIA runtime is configured globally in Docker to ensure all containers can leverage GPU resources securely.

Databases

The application uses an internal volume for data persistence:

  • Volume Name: open-webui

  • Mount Path: /app/backend/data inside the container.

Docker Containers and Their Deployment

Nginx Proxy Container

This container handles SSL termination and reverse proxying to the web interface.

Property Value
Image jonasal/nginx-certbot:latest
Network Mode host
Restart Policy unless-stopped
Volumes nginx_secrets:/etc/letsencrypt, /data/nginx/user_conf.d:/etc/nginx/user_conf.d

Open WebUI Container

This container provides the user interface for interacting with the Gemma model.

Property Value
Image ghcr.io/open-webui/open-webui:cuda
Name open-webui
Ports 8080:8080
Restart Policy always
Volumes open-webui:/app/backend/data
Environment Variables ENV='dev', OLLAMA_BASE_URLS='http://host.docker.internal:11434'

Custom Scripts and Additional Setup

The following actions are performed during the initial setup to prepare the environment:

  • NVIDIA Runtime Configuration: The command nvidia-ctk runtime configure --runtime=docker is executed to integrate NVIDIA drivers with the Docker engine.

  • Ollama Environment Modification: Systemd service parameters are injected into /etc/systemd/system/ollama.service to enable network accessibility and performance optimizations.

Application Update Instructions

To update the web interface, execute the following command in the directory containing your compose files:

docker compose pull && docker compose up -d

For updating the underlying model runner (Ollama), use:

ollama pull gemma3:27b

Location of configuration files and data

  • Nginx Configuration: /root/nginx/compose.yml

  • Docker Data Volume: open-webui (managed by Docker)

  • Ollama Service Config: /etc/systemd/system/ollama.service

Available ports for connection

Port Protocol Description
443 HTTPS External access to the Web UI via Nginx
8080 HTTP Internal application port (mapped from 8080)

Starting and Stopping the application

Starting the Application

To start the web interface and proxy:

docker compose up -d

To ensure the Ollama service is running:

sudo systemctl start ollama

Stopping the Application

To stop the web interface and proxy:

docker compose down

To stop the Ollama service:

sudo systemctl stop ollama

question_mark
Is there anything I can help you with?
question_mark
AI Assistant ×