Deployment Overview of TensorFlow on Server¶
Prerequisites and Basic Requirements¶
To ensure a successful deployment, the server must meet the following requirements:
-
Operating System: Ubuntu or Debian.
-
Privileges: Root or sudo access is required for all installation steps.
-
Hardware Requirement: For GPU acceleration, an NVIDIA H100 (or compatible) video card is recommended to trigger specific kernel optimizations.
-
Required Packages: The system must have
curl,wget, andsudoinstalled.
File and Directory Structure¶
The deployment process creates several files and directories to manage the environment and credentials:
| Path | Description |
|---|---|
/root/install_script.sh | Main installation script for drivers and system configuration. |
/root/user_credentials | Contains the generated password for the user account. |
/home/user/venv | Python virtual environment for TensorFlow. |
/home/user/TensorRT-8.6.1.6.Linux.x86_64-gnu.cuda-12.0.tar.gz | Extracted TensorRT binaries. |
/home/user/tensorflow.sh | Activation script for the TensorFlow environment. |
Application Installation Process¶
The installation is performed through a multi-stage process involving system configuration and specialized setup scripts.
1. System and Driver Configuration¶
An initial installation script (install_script.sh) performs the following actions:
-
Updates the system package list and upgrades existing packages.
-
Installs
ubuntu-drivers-commonand detects/installs the recommended NVIDIA drivers. -
Installs the CUDA toolkit via the official NVIDIA repository.
-
Configures environment variables for CUDA in the user's
.bashrc. -
Creates a dedicated service user named
userwith sudo privileges. -
Installs Python 3.10,
pip, andvenv.
2. TensorFlow Environment Setup¶
Once the system is configured, a secondary installation script (tensorflow_install.sh) is executed under the user account to set up the machine learning environment:
-
Creates a Python virtual environment in
~/venv. -
Installs
tensorflow[and-cuda]via pip. -
Downloads and extracts TensorRT 8.6.1 for CUDA 12.0.
-
Configures the TensorRT wheel and updates the
tensorrtpackage.
Access Rights and Security¶
Security is managed through the following measures:
-
User Isolation: A dedicated user named
useris created to run application processes, minimizing the risk associated with running as root. -
Credential Management: A random 8-character password is generated during installation and stored in
/root/user_credentials. -
Sudo Access: The
useraccount is added to thesudogroup for administrative tasks when necessary.
Docker Containers and Their Deployment¶
This deployment does not utilize Docker containers; it relies on a native Python virtual environment installed directly on the host operating system to manage dependencies and hardware acceleration.
Custom Scripts and Additional Setup¶
The deployment utilizes two primary scripts to automate the complex setup of GPU-accelerated environments:
-
install_script.sh: Located in/root/, this script handles low-level system requirements, including kernel updates for H100 hardware, NVIDIA driver installation, CUDA toolkit configuration, and user creation. -
tensorflow_install.sh: A temporary script generated during the process to handle Python-level dependencies. It manages the virtual environment, installs TensorFlow with CUDA support, and sets up TensorRT.
Additionally, a helper script tensorflow.sh is created in /home/user/. This script automates the activation of the virtual environment and correctly exports necessary paths for CUDNN_PATH and LD_LIBRARY_PATH to ensure the application can locate the installed libraries.
Application Update Instructions¶
To update the TensorFlow environment, follow these steps:
-
Navigate to the user directory:
-
Activate the environment using the provided script:
-
Update the packages via pip:
Location of Configuration Files and Data¶
-
Environment Activation:
/home/user/tensorflow.sh -
Python Virtual Environment:
/home/user/venv -
TensorRT Binaries:
/home/user/TensorRT-8.6.1.6.Linux.x86_64-gnu.cuda-12.0/
Available Ports for Connection¶
The application utilizes standard system resources and does not expose specific network ports by default unless configured via an external proxy or web framework.
Starting and Stopping the Application¶
To start using the TensorFlow environment, execute the activation script:
Once activated, you can run Python scripts directly within the environment. To stop using the environment, simply exit the shell or use the deactivate command if inside a sub-shell.