A high-performance, hardware-aware desktop Web application for training custom real-time voice conversion models using Project Beatrice V2. Featuring a modern Cyan Blue minimalist UI, dynamic Light/Dark theme toggle, Single-Choice Dataset Import Switcher (Local Drop-Zone, Hugging Face, Kaggle), automated NVIDIA GPU CUDA Hardware Tuning, and direct Google Colab & Kaggle Cloud Launchpad integration.
- ๐ Cyan Blue Minimalist UI & Vector SVGs โ Ultra-sleek Linear/Vercel-inspired desktop design system with custom vector SVG icons and sharp typography (
Inter,Outfit,JetBrains Mono). - ๐ Light & Dark Mode Theme Switcher โ Instant theme switching with persistent user preferences stored in local storage (
localStorage). - ๐ฏ Single-Choice Dataset Import Selector โ Segmented method picker allowing users to easily choose between Local Audio Upload, Hugging Face Import, or Kaggle Dataset Import.
- โก NVIDIA CUDA Hardware Tuning โ Deep profiling for NVIDIA GPUs (RTX 4090, 4080, 4070, 3090, 3080, 3060, 2060, etc.) and Apple Silicon MPS with automatic batch size and worker optimization.
- โ๏ธ Cloud Training Launchpad โ Direct integration of
Project-Beatrice-V2/Beatrice-colabwith one-click Google Colab (.ipynb) and Kaggle launch buttons and local notebook downloads. - ๐ Real-Time Monitoring & Metrics โ Live loss plotting on HTML Canvas, stdout logging console, and memory usage gauges.
Dashboard Overview โ Live hardware monitor, memory allocation gauges, and system diagnostics
Dataset Manager โ Single-choice import selector (Local Audio Drag & Drop, Hugging Face, Kaggle)
Training Control โ Hardware auto-tuning, custom training steps, and live console logger
Cloud Training Launchpad โ One-click Google Colab and Kaggle launch cards with local .ipynb downloads
- Operating System: Windows 10 or Windows 11 (64-bit) / macOS 12+ / Linux.
- GPU: NVIDIA GPU with CUDA support recommended (e.g., RTX 2060+, 3060+, 4060+). CPU fallback is supported.
- Python: Python 3.10 or newer (make sure to check "Add python.exe to PATH" during installation).
Double-click start_windows.bat (or start.bat) in File Explorer, or run in Command Prompt / PowerShell:
start_windows.batFor macOS / Linux systems:
chmod +x start.sh
./start.sh- Creates a local Python virtual environment (
venv/). - Installs CUDA-accelerated PyTorch (
torchandtorchaudiousing the CUDA 12.1 wheel index). - Installs backend dependencies (
fastapi,uvicorn,python-multipart,aiofiles,huggingface_hub,psutil). - Downloads the core Beatrice Trainer repository from Hugging Face (
fierce-cats/beatrice-trainer). - Launches the FastAPI backend server and automatically opens your browser at
http://localhost:8000.
When selecting a speaker dataset in the Training tab, the system profiles hardware to compute optimal hyperparameters:
| Hardware Tier | GPU / VRAM | Batch Size | Grad Accum | Effective Batch | Workers | Mixed Precision |
|---|---|---|---|---|---|---|
| Ultra Enthusiast | RTX 3090 / 4090 (24GB) | 16 | 1 | 16 | 4 | Enabled (FP16) |
| High Performance | RTX 4080 / 3080 Ti (16GB) | 8 | 2 | 16 | 4 | Enabled (FP16) |
| Standard Gaming | RTX 3060 / 4070 (12GB) | 4 | 4 | 16 | 2 | Enabled (FP16) |
| Entry Level GPU | RTX 3060 8GB / 4060 (8GB) | 2 | 4 | 8 | 2 | Enabled (FP16) |
| CPU Fallback | System CPU | 2 | 4 | 8 | 2 | Disabled |
โโโ assets/ # App branding and logo assets
โ โโโ logo.jpg # Official Project Beatrice V2 Logo
โโโ colab_repo/ # Cloned Beatrice-Colab repository
โ โโโ BeatriceV2_Trainer_Notebook_Colab.ipynb
โ โโโ BeatriceV2_Trainer_Notebook_Kaggle.ipynb
โโโ app.js # Client-side JavaScript application & UI handlers
โโโ index.html # Modern HTML layout & component views
โโโ index.css # Cyan Blue Design System & Light/Dark themes
โโโ server.py # FastAPI backend server & REST endpoints
โโโ start_windows.bat # Automated Windows batch startup script
โโโ start.bat # Batch launcher shortcut
โโโ start.sh # Cross-platform shell startup script
โโโ requirements.txt # Backend Python dependencies
โโโ README.md # Documentation manual
โโโ beatrice-trainer/ # Downloaded Hugging Face trainer files
| Endpoint | Method | Description |
|---|---|---|
/api/status |
GET |
System health check & backend availability |
/api/system/memory |
GET |
Live RAM, VRAM, and CPU usage metrics |
/api/dataset/list |
GET |
List available speaker datasets and file counts |
/api/dataset/upload |
POST |
Upload WAV or ZIP audio files |
/api/dataset/import/hf |
POST |
Download dataset directly from Hugging Face |
/api/dataset/import/kaggle |
POST |
Import dataset from Kaggle |
/api/train/auto-tune |
POST |
Compute optimal training hyperparameters |
/api/train/start |
POST |
Launch local model training process |
/api/train/stop |
POST |
Terminate active model training process |
/api/models/list |
GET |
List trained VST3 paraphernalia voice models |
This project is licensed under the MIT License.
