Skip to content

Deployment Overview of PyTorch on Server

Prerequisites and Basic Requirements

The deployment requires a server running either Debian or Ubuntu. The following system requirements and tools must be present:

  • Operating System: Debian or Ubuntu.

  • Privileges: Root or sudo access is required for all installation steps.

  • Hardware Requirement: For optimal performance, an NVIDIA GPU (specifically H100 SXM5 80GB models are supported via specialized kernel packages) is recommended to utilize CUDA acceleration.

  • Required Packages: curl, wget, and sudo.

FQDN of the final panel on the hostkey.in domain

Parameter Value
Prefix pytorch
Domain hostkey.in
Full template pytorch{Server_ID_from_Invapi}.hostkey.in

File and Directory Structure

The following files and directories are created or utilized during the deployment process:

  • /root/install_script.sh: The primary installation script.

  • /root/user_credentials: Contains the generated password for the user account.

  • /home/user/venv: The Python virtual environment directory.

  • /home/user/pytorch.sh: A helper script to activate the PyTorch virtual environment.

  • /home/user/pytorch_install.sh: An automated script used to initialize the Python environment and verify installation.

Application Installation Process

The application is installed through a multi-stage process involving system configuration, driver installation, and Python environment setup:

  1. System Preparation: The package manager updates all existing packages, performs a safe upgrade, and removes unused dependencies.

  2. Driver and CUDA Installation:

  3. If an H100 GPU is detected via PCI ID 10de:2330, the linux-generic-hwe-22.04 kernel package is installed.

  4. The system installs ubuntu-drivers-common.

  5. The recommended NVIDIA driver for the hardware is automatically identified and installed.

  6. CUDA toolkit is installed via the official NVIDIA repository corresponding to the current Ubuntu release version.

  7. User Configuration: A new user named user is created with a randomly generated password. This user is granted sudo privileges.

  8. Environment Setup:

  9. Python 3.10, pip, and venv are installed.

  10. CUDA environment variables (PATH and LD_LIBRARY_PATH) are added to the .bashrc file for the user.

  11. PyTorch Installation: A virtual environment is created in /home/user/venv, and the core libraries torch, torchvision, and torchaudio are installed via pip3.

Access Rights and Security

  • User Account: A dedicated user named user is created to manage the application environment.

  • Privileges: The user account is added to the sudo group for administrative tasks when necessary.

  • Credentials: An 8-character random password is generated during installation and stored in /root/user_credentials.

Custom Scripts and Additional Setup

The deployment utilizes several scripts to configure the environment:

  • install_script.sh: Located in /root, this script automates driver detection, CUDA installation, user creation, and system dependency management.

  • pytorch_install.sh: Located in /home/user, this script manages the lifecycle of the Python virtual environment, including creation, package installation (PyTorch), and a functional test to verify GPU availability.

Application Update Instructions

To update the PyTorch environment or its dependencies:

  1. Navigate to the user directory: cd /home/user.

  2. Activate the virtual environment using the provided script: ./pytorch.sh.

  3. Use pip to upgrade the packages within the active environment:

    pip install --upgrade torch torchvision torchaudio
    

Location of Configuration Files and Data

  • Virtual Environment: /home/user/venv

  • User Credentials: /root/user_credentials

  • System Binaries: CUDA binaries are located in /usr/local/cuda.

Available Ports for Connection

The application relies on standard system communication. No specific web ports are opened by the core PyTorch installation, as it functions primarily as a computational library within the Python environment.

Starting and Stopping the Application

Since the application runs within a Python virtual environment, management is handled via the shell:

  • To start/enter the environment:

    cd /home/user
    ./pytorch.sh
    

  • To stop the environment: Simply exit the terminal session or type deactivate if inside the Python shell.

question_mark
Is there anything I can help you with?
question_mark
AI Assistant ×