Skip to content

Deployment Overview of Llama-3.3-70B on Server

Prerequisites and Basic Requirements

To ensure the successful deployment of the application, the following requirements must be met:

  • Operating System: Ubuntu (recommended).

  • Privileges: Root or sudo access is required for package installation and service management.

  • Hardware Support: NVIDIA GPU support is required for CUDA acceleration.

  • Network Requirements:

  • Access to curl and the ability to download from ollama.com.

  • Outbound connectivity to ghcr.io for Docker images.

  • Ports:

  • 443: External HTTPS access.

  • 8080: Internal application traffic (Open WebUI).

  • 11434: Local Ollama API communication.

FQDN of the final panel on the hostkey.in domain

The application is accessible via a specific subdomain template:

Parameter Value
Prefix llama
Domain hostkey.in
Full template llama{Server_ID}.hostkey.in

Application installation process

The installation follows a multi-stage process involving system package management, service configuration, and container orchestration:

  1. System Dependencies: The system is updated, and necessary packages such as curl are installed via the apt package manager.

  2. Ollama Installation:

  3. The Ollama installation script is executed via curl -fsSL https://ollama.com/install.sh | sh.

  4. A dedicated system user named ollama is created.

  5. The ollama.service configuration is modified to allow remote connections (OLLAMA_HOST=0.0.0.0) and enable Flash Attention (LLAMA_FLASH_ATTENTION=1).

  6. Model Download: The llama3.3 model is pulled into the local Ollama storage using the command ollama pull llama3.3.

  7. Docker Environment Setup:

  8. Docker and NVIDIA CUDA drivers are installed to support GPU acceleration.

  9. An Nginx proxy container is configured with SSL certificates via Certbot for secure web access.

  10. Container Deployment: The Open WebUI application is deployed as a Docker container using the ghcr.io/open-webui/open-webui:cuda image.

Access Rights and Security

  • Firewall: Only necessary ports are exposed; internal services like Ollama communicate via the local bridge or host gateway.

  • User Management: A system user ollama is created to run the backend service.

  • SSL/TLS: Secure communication is enforced through an Nginx proxy with SSL certificates managed by Certbot.

Databases

The application uses Docker volumes for persistent data storage:

  • Open WebUI Data: Stored in a Docker volume named open-webui located at /app/backend/data within the container.

Docker Containers and Their Deployment

Open WebUI Container

  • Image Name: ghcr.io/open-webui/open-webui:cuda

  • Ports: 8080:8080

  • Volumes: open-webui:/app/backend/data

  • Environment Variables:

  • ENV=dev

  • OLLAMA_BASE_URLS=http://host.docker.internal:11434

  • Restart Policy: always

  • Additional Configuration: Uses --gpus all for hardware acceleration and --add-host=host.docker.internal:host-gateway to communicate with the host's Ollama service.

Nginx Proxy Container

  • Image Name: jonasal/nginx-certbot:latest

  • Network Mode: host

  • Volumes:

  • nginx_secrets:/etc/letsencrypt

  • /data/nginx/user_conf.d:/etc/nginx/user_conf.d

  • Restart Policy: unless-stopped

Custom Scripts and Additional Setup

The following setup actions are performed to finalize the environment:

  • Service Configuration: The ollama.service file is updated with specific environment variables (OLLAMA_ORIGINS="*") to ensure cross-origin compatibility for the web interface.

  • Proxy Redirection: An Nginx configuration is dynamically generated and modified to proxy traffic from the external domain to http://127.0.0.1:8080.

Application Update Instructions

To update the main application (Open WebUI), perform the following steps:

docker compose pull && docker compose up -d

(Note: This assumes you are in the directory containing the compose.yml file for the web interface).

Location of configuration files and data

  • Nginx Configuration: /data/nginx/user_conf.d/

  • Nginx Compose File: /root/nginx/compose.yml

  • Ollama Service Config: /etc/systemd/system/ollama.service

  • Open WebUI Data Volume: open-webui (Docker managed)

Available ports for connection

  • HTTPS (External): 443

  • Web UI (Internal): 8080

  • Ollama API (Local): 11434

Starting and Stopping the application

Open WebUI

To manage the web interface container:

docker stop open-webui
docker start open-webui

Ollama Service

To manage the model backend service:

sudo systemctl restart ollama
sudo systemctl stop ollama
sudo systemctl start ollama

question_mark
Is there anything I can help you with?
question_mark
AI Assistant ×