Deployment Overview of Gemma-3-27B on Server¶
Prerequisites and Basic Requirements¶
To ensure the successful deployment of the application, the following requirements must be met:
-
Operating System: Ubuntu (recommended for compatibility with NVIDIA drivers).
-
Privileges: Root or sudo access is required for package installation and service management.
-
Hardware Support: An NVIDIA GPU is required to utilize CUDA acceleration for model inference.
-
Network Requirements:
-
Outbound access to
ollama.comfor the installation script. -
Outbound access to
ghcr.ioto pull Docker images. -
Port
443(HTTPS) must be open for external traffic.
FQDN of the final panel on the hostkey.in domain¶
The application is accessible via a specific subdomain template based on your Server ID:
| Parameter | Value |
|---|---|
| Prefix | gemma |
| Domain | hostkey.in |
| Full template | gemma{Server_ID}.hostkey.in |
File and Directory Structure¶
The following directories are used for configuration, certificates, and application data:
-
/root/nginx: Contains the Nginx reverse proxy configuration files (compose.yml). -
/data/nginx/user_conf.d: Custom Nginx configuration directory. -
/data/nginx/nginx-certbot.env: Environment file for Certbot email settings. -
/etc/systemd/system/ollama.service.bak: Backup of the original Ollama service configuration.
Application installation process¶
The application is installed using a combination of system package management, shell scripts, and Docker containers:
-
Ollama Installation: The official Ollama installation script is executed via
curlto install the model runner on the host system. -
System Service Configuration:
-
A dedicated
ollamauser is created. -
The
ollama.servicefile is modified to allow remote connections (OLLAMA_HOST=0.0.0.0) and enable Flash Attention (LLAMA_FLASH_ATTENTION=1). -
Model Download: The
gemma3:27bmodel is pulled directly via the Ollama CLI. -
GPU Acceleration Setup:
-
NVIDIA Container Toolkit and NVIDIA Container Runtime are installed to allow Docker containers to access host GPU hardware.
-
The Docker daemon is configured with
/etc/docker/daemon.jsonto setnvidiaas the default runtime. -
Web Interface Deployment: The Open WebUI interface is deployed as a Docker container using the CUDA-enabled image.
Access Rights and Security¶
-
Firewall: Ensure port
443is open for HTTPS traffic. -
Service Isolation: Ollama runs as a system service with restricted environment variables to manage origin access (
OLLAMA_ORIGINS=*). -
Docker Runtime: The NVIDIA runtime is configured globally in Docker to ensure all containers can leverage GPU resources securely.
Databases¶
The application uses an internal volume for data persistence:
-
Volume Name:
open-webui -
Mount Path:
/app/backend/datainside the container.
Docker Containers and Their Deployment¶
Nginx Proxy Container¶
This container handles SSL termination and reverse proxying to the web interface.
| Property | Value |
|---|---|
| Image | jonasal/nginx-certbot:latest |
| Network Mode | host |
| Restart Policy | unless-stopped |
| Volumes | nginx_secrets:/etc/letsencrypt, /data/nginx/user_conf.d:/etc/nginx/user_conf.d |
Open WebUI Container¶
This container provides the user interface for interacting with the Gemma model.
| Property | Value |
|---|---|
| Image | ghcr.io/open-webui/open-webui:cuda |
| Name | open-webui |
| Ports | 8080:8080 |
| Restart Policy | always |
| Volumes | open-webui:/app/backend/data |
| Environment Variables | ENV='dev', OLLAMA_BASE_URLS='http://host.docker.internal:11434' |
Custom Scripts and Additional Setup¶
The following actions are performed during the initial setup to prepare the environment:
-
NVIDIA Runtime Configuration: The command
nvidia-ctk runtime configure --runtime=dockeris executed to integrate NVIDIA drivers with the Docker engine. -
Ollama Environment Modification: Systemd service parameters are injected into
/etc/systemd/system/ollama.serviceto enable network accessibility and performance optimizations.
Application Update Instructions¶
To update the web interface, execute the following command in the directory containing your compose files:
For updating the underlying model runner (Ollama), use:
Location of configuration files and data¶
-
Nginx Configuration:
/root/nginx/compose.yml -
Docker Data Volume:
open-webui(managed by Docker) -
Ollama Service Config:
/etc/systemd/system/ollama.service
Available ports for connection¶
| Port | Protocol | Description |
|---|---|---|
443 | HTTPS | External access to the Web UI via Nginx |
8080 | HTTP | Internal application port (mapped from 8080) |
Starting and Stopping the application¶
Starting the Application¶
To start the web interface and proxy:
To ensure the Ollama service is running:
Stopping the Application¶
To stop the web interface and proxy:
To stop the Ollama service: