Skip to content

Latest commit

Β 

History

21 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸŽ™οΈ VoiceFlow AI β€” Universal Voice Intelligence Web Studio

Next.js React TypeScript Whisper Gemini Claude License


An ultra-fast, privacy-first Voice-to-Text & Multimodal Speech Intelligence Web Application.
Real-Time Dictation β€’ 48kHz DSP Noise Cancellation β€’ 2-Way Babel Live Duplex Translator β€’ Corporate Meeting Minutes (MoM) PDF Generator β€’ Native Urdu & Multilingual Neural TTS β€’ Studio EQ Mastering.


🌟 Star on GitHub β€’ πŸš€ Live Web Studio β€’ πŸ“– Documentation β€’ ⚑ Quick Start


πŸ“‘ Table of Contents


🌟 Key Highlights

  • ⚑ Zero-Latency Dictation: Live continuous speech streaming with interim preview bubbles.
  • πŸ‡΅πŸ‡° Native Urdu & Multilingual Speech: Speaks aloud in genuine, fluent Urdu (اردو), Hindi (ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯€), Arabic (Ψ§Ω„ΨΉΨ±Ψ¨ΩŠΨ©), Spanish (EspaΓ±ol), French, and 20+ languages.
  • πŸ”‡ DSP Noise Annihilation: Hardware high-pass, low-pass, 60Hz hum notch, and dynamics compression directly in the audio capture pipeline.
  • πŸ›‘οΈ Guarded Intelligence: Never gets stuck; alerts users if an API key is needed with 10-second instant setup links.
  • πŸš€ Modern Web Stack: Built with Next.js 16 App Router, React 19, Tailwind CSS, and Web Audio API.

πŸ› οΈ Complete Feature Guide & How Each Feature Works

1. Real-Time Live Speech Dictation & Spectrum Visualizer

  • What It Does: Transcribes your spoken words into text instantly with zero lag.
  • How It Works: Connects directly to the high-performance browser speech recognition engine with automatic silence recovery. Simultaneously runs a 48-band Web Audio AnalyserNode rendering a real-time reactive neon audio visualizer.
  • How to Use: Click the central "Start Recording" button, select your language, and speak naturally.

2. Universal SweetAlert API Key Guard & Quick Setup

  • What It Does: Protects all cloud-powered AI transformations (Fix Grammar, Summarize, Translate, Repurpose, Mindmap, Ask AI, and Babel Mode).
  • How It Works: When an AI button or tab is clicked without an active key, VoiceFlow blocks the action and pops up an interactive SweetAlert modal with direct 1-click links to get free keys for Gemini, GPT-4o, or Claude 3.7.
  • How to Use: Click any AI tool β€” if no key is entered, the modal will guide you to paste your key in Settings in under 10 seconds.

3. 2-Way "Babel Mode" Live Duplex Universal Translator

  • What It Does: Enables two people who speak completely different languages (e.g. English and Urdu, or Spanish and Arabic) to have a live, translated spoken conversation on a single device.
  • How It Works: Person A speaks or types in Language A. VoiceFlow translates the input into Person B's language native script and automatically speaks the translation aloud in Person B's native accent using our dedicated /api/tts neural voice engine.
  • How to Use: Click "Babel Mode" in the top navigation bar, select languages for Person A and Person B, and start chatting.

4. Multi-Stage 48kHz Web Audio DSP Noise Cancellation

  • What It Does: Eliminates room noise, computer fan hiss, desk vibrations, and electrical hum from recordings.
  • How It Works: Builds a multi-node Web Audio processing graph:
    1. 85Hz High-Pass Biquad Filter: Cuts desk bumps, handling thuds, and low rumble.
    2. 8500Hz Low-Pass Filter: Cuts high-frequency electronic hiss and fan buzz.
    3. 60Hz Notch Filter: Cancels AC power line hum.
    4. Dynamics Compressor Node: Evens out vocal peaks for crystal-clear clarity.
    5. The MediaRecorder records from the filtered DSP stream so saved audio is pristine.

5. Original Voice Replay & Studio Vocal Mastering Presets

  • What It Does: Allows you to listen back to your original recorded voice with professional broadcast mastering presets and 1-click audio download (.webm).
  • How It Works: Features an integrated scrub bar, speed multiplier (1.0x - 2.0x), volume control, and 4 audio presets:
    • πŸŽ™οΈ Clean DSP: Balanced noise-filtered voice.
    • πŸ“» Podcast Warmth: +3dB low-end boost (150Hz) for deep radio resonance.
    • ⚑ Broadcast Radio: +2.5dB high-mid presence (3.5kHz) for voice clarity.
    • ✨ Crisp Articulation: Enhanced consonants for technical lectures.

6. Official Corporate Meeting Minutes (MoM) & Executive PDF

  • What It Does: Generates an official, print-ready Corporate Meeting Minutes PDF with executive formatting and formal signature blocks.
  • How It Works: Parses the transcript into structured sections: Meeting Context, Objectives, Discussion Breakdown, Action Items Table, and sign-off lines for Host / Speaker 1 and Executive / Client Approver.
  • How to Use: Click the "MoM PDF" button in the workspace toolbar.

7. "Ask My Voice Vault" Semantic Audio Search

  • What It Does: Allows you to search across weeks and months of saved audio notes by topic, concept, or keyword.
  • How It Works: Indexes all saved transcript sessions with duration, tags, and timestamps. Matches keywords in real-time and loads the full session back into the workspace with 1 click.
  • How to Use: Click "Saved Notes" in the top action bar and use the search bar.

8. Smart Voice Macro Action Board (Spoken Kanban)

  • What It Does: Automatically turns your spoken voice into interactive Kanban task cards.
  • How It Works: The NLP parser detects spoken trigger phrases:
    • "Task: [something]" βž” Creates a task card in the Todo column.
    • "Idea: [something]" βž” Pins an idea card.
    • "Important: [something]" βž” Flags an urgent priority item.
  • How to Use: Switch to the "Action Board" tab in the workspace to view and check off items.

9. Multi-Speaker Diarization (Dialogue Splitter)

  • What It Does: Segregates conversation segments between two speakers with color-coded speech bubbles and custom names.
  • How It Works: Alternates turns between Speaker 1 (Violet) and Speaker 2 (Cyan) with real-time timestamping.
  • How to Use: Click the "Speakers Dialogue" tab and type custom names for both speakers.

10. Floating Dual-Language Theater Subtitles

  • What It Does: Displays a cinematic, floating subtitle overlay with original speech on top and translated subtitles below.
  • How It Works: Streams live transcription into an ultra-clean glassmorphic overlay designed for live presentations, zoom calls, or video recordings.
  • How to Use: Click "Live Subtitles" in the quick action bar.

11. Video & Lecture Audio Stream URL Importer

  • What It Does: Directly transcribes audio from online video links, podcast feeds, or lecture URLs.
  • How It Works: Accepts streaming audio URLs or audio file uploads (.mp3, .wav, .m4a, .webm, .ogg) and passes them through OpenAI Whisper for automated transcription.
  • How to Use: Click "Upload Audio" and switch to the "Paste Audio / Video URL" tab.

12. "Ask AI" Contextual Transcript Chat

  • What It Does: Allows you to chat with an AI assistant that has full context over your recorded voice note.
  • How It Works: Routes your question and transcript through our server-side /api/ai gateway to Gemini 2.0, GPT-4o, or Claude 3.7.
  • How to Use: Click "Ask AI" and ask questions like "What were the key numbers mentioned?" or "Draft a follow-up email based on this".

13. Concept Mindmap Visualizer (Mermaid.js)

  • What It Does: Automatically converts unstructured speech into a structured visual concept mindmap.
  • How It Works: Analyzes central themes, sub-topics, and branches, rendering an interactive SVG mindmap diagram.
  • How to Use: Click the "Mindmap" tab in the workspace.

14. 1-Click Multi-Format Content Repurposer

  • What It Does: Transforms spoken notes into 6 distinct content formats with 1 click:
    • βœ‰οΈ Executive Email
    • 🧡 Viral Twitter / X Thread
    • πŸ’Ό Engaging LinkedIn Post
    • πŸ“‹ Meeting Minutes & Summary
    • πŸŽ“ Study Flashcards & Q&A
    • πŸ“ Blog Post Outline
  • How to Use: Click the "Repurpose" tab and select any format.

πŸ”‘ Supported AI Models & Step-by-Step API Key Setup

VoiceFlow AI is equipped with the latest frontier models:

Provider Frontier Model Capabilities How to Get Your API Key
🌟 Google Gemini gemini-2.0-flash
gemini-2.0-flash-lite
Ultra-low latency, generous free tier, 1M+ context window 1. Go to Google AI Studio
2. Sign in and click Create API key
3. Paste into VoiceFlow Settings.
⚑ OpenAI gpt-4o
gpt-4o-mini
Flagship omni multimodal reasoning 1. Go to OpenAI API Keys
2. Create a new Secret Key
3. Paste into VoiceFlow Settings.
🧠 Anthropic Claude claude-3-7-sonnet
claude-3-5-sonnet
Hybrid reasoning & superior prose formatting 1. Go to Anthropic Console
2. Generate an API Key
3. Paste into VoiceFlow Settings.

πŸ‡΅πŸ‡° Authentic Multilingual Voice Synthesis & TTS Architecture

  • True Native Script Translation: Outputs authentic native scripts: Urdu (اردو), Arabic (Ψ§Ω„ΨΉΨ±Ψ¨ΩŠΨ©), Hindi (ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯€), Spanish (EspaΓ±ol), French (FranΓ§ais), German (Deutsch), Chinese (δΈ­ζ–‡), Japanese (ζ—₯本θͺž), and more.
  • Dedicated /api/tts Server Audio Gateway: Eliminates browser limitations and provides high-fidelity audio streams for Urdu and 20+ languages with native accents.
  • Multi-Tier Voice Fallback: Server Audio Stream βž” OpenAI TTS (if BYOK key provided) βž” Browser SpeechSynthesis with regional voice matching.

πŸ’» Web Architecture & Tech Stack

graph TD
    User([User Voice / Microphone]) --> DSP[48kHz Web Audio DSP Filters]
    DSP --> AudioPlayer[Original Voice Player & EQ Mastering]
    DSP --> Whisper[Speech-to-Text / Whisper]
    Whisper --> Workspace[Transcription Workspace]
    Workspace --> ServerAPI["/api/ai & /api/tts Gateway"]
    ServerAPI --> Gemini["Google Gemini 2.0 Flash"]
    ServerAPI --> OpenAI["OpenAI GPT-4o"]
    ServerAPI --> Claude["Anthropic Claude 3.7"]
    ServerAPI --> Babel["2-Way Babel Translator (Urdu, etc.)"]
Loading
Layer Web Application Technology
Framework Next.js 16 (App Router) + React 19
Language TypeScript 5 (Strict Mode)
Speech-to-Text Web Speech API + Groq / OpenAI Whisper
Audio Processing 48kHz High-Pass, Low-Pass, 60Hz Notch & Dynamics Compressor
TTS Engine Next.js /api/tts Neural Stream + Kokoro-82M + Edge Neural
Iconography Font Awesome SVG Icons (@fortawesome/react-fontawesome)
State Store Zustand 5 (IndexedDB / LocalStorage)
Document Export jsPDF (MoM & Transcript) + Canvas-Confetti

πŸš€ Step-by-Step Installation & Running Guide

Prerequisites

  • Node.js: v18.18.0 or higher (v20+ recommended)
  • npm or yarn / pnpm
  • Git

Running the Web Studio

# 1. Clone the repository
git clone https://github.com/Humaam-04-06/Voice_To_Text_App.git

# 2. Navigate to the web application folder
cd Voice_To_Text_App/web

# 3. Install dependencies
npm install

# 4. Start the local development server
npm run dev

Open http://localhost:3000 in your browser!

# 5. Build for production
npm run build

# 6. Run production server
npm start

πŸ“ Project Directory Structure

Voice_To_Text_App/
β”œβ”€β”€ web/                           # Next.js 16 Web Application
β”‚   β”œβ”€β”€ src/
β”‚   β”‚   β”œβ”€β”€ app/
β”‚   β”‚   β”‚   β”œβ”€β”€ api/
β”‚   β”‚   β”‚   β”‚   β”œβ”€β”€ ai/            # Server gateway for Gemini, GPT-4o, Claude
β”‚   β”‚   β”‚   β”‚   β”œβ”€β”€ transcribe/    # Whisper audio transcription route
β”‚   β”‚   β”‚   β”‚   └── tts/           # Multilingual Neural TTS streaming route
β”‚   β”‚   β”‚   β”œβ”€β”€ globals.css        # Glassmorphic dark theme styles
β”‚   β”‚   β”‚   β”œβ”€β”€ layout.tsx         # Root layout with Font Awesome CSS & SEO
β”‚   β”‚   β”‚   └── page.tsx           # Main application page
β”‚   β”‚   β”œβ”€β”€ components/            # UI Components with Font Awesome SVG icons
β”‚   β”‚   β”‚   β”œβ”€β”€ ApiKeyRequiredModal.tsx    # Universal SweetAlert API key guard
β”‚   β”‚   β”‚   β”œβ”€β”€ AudioFileUploadModal.tsx   # File & Video URL importer
β”‚   β”‚   β”‚   β”œβ”€β”€ AudioPlaybackPlayer.tsx    # Real voice replay & EQ presets
β”‚   β”‚   β”‚   β”œβ”€β”€ BabelTranslatorModal.tsx   # 2-Way live duplex translator
β”‚   β”‚   β”‚   β”œβ”€β”€ ChatDrawer.tsx             # Ask AI transcript chat
β”‚   β”‚   β”‚   β”œβ”€β”€ HeroRecordingZone.tsx      # Recording controls & waveform
β”‚   β”‚   β”‚   β”œβ”€β”€ HistoryDrawer.tsx          # Semantic Voice Vault search
β”‚   β”‚   β”‚   β”œβ”€β”€ MindmapViewer.tsx          # Concept mindmap generator
β”‚   β”‚   β”‚   β”œβ”€β”€ Navbar.tsx                 # Gemini 2.0, GPT-4o, Claude 3.7 switcher
β”‚   β”‚   β”‚   β”œβ”€β”€ SettingsModal.tsx          # Key settings with step-by-step guides
β”‚   β”‚   β”‚   β”œβ”€β”€ SpeechCoachWidget.tsx      # Clarity, WPM, and filler words
β”‚   β”‚   β”‚   β”œβ”€β”€ TheaterSubtitleModal.tsx   # Floating dual subtitles
β”‚   β”‚   β”‚   β”œβ”€β”€ ToneSentimentRadar.tsx     # Live emotional mood & tone HUD
β”‚   β”‚   β”‚   β”œβ”€β”€ TranscriptionWorkspace.tsx # Workspace, tabs, & MoM PDF exporter
β”‚   β”‚   β”‚   β”œβ”€β”€ VoiceMacroActionBoard.tsx  # Spoken triggers & Kanban tasks
β”‚   β”‚   β”‚   └── WaveformVisualizer.tsx     # Real-time audio spectrum
β”‚   β”‚   β”œβ”€β”€ hooks/
β”‚   β”‚   β”‚   β”œβ”€β”€ useAudioRecorder.ts        # 48kHz Web Audio DSP noise cancellation
β”‚   β”‚   β”‚   └── useSpeechRecognition.ts   # Continuous streaming speech recognition
β”‚   β”‚   β”œβ”€β”€ lib/
β”‚   β”‚   β”‚   β”œβ”€β”€ ai/                        # AI dispatchers & local NLP engines
β”‚   β”‚   β”‚   β”œβ”€β”€ audio/                     # TTS and audio filter engines
β”‚   β”‚   β”‚   └── constants/                 # Supported multilingual list
β”‚   β”‚   β”œβ”€β”€ store/                         # Zustand Global Voice Store
β”‚   β”‚   └── types/                         # TypeScript type definitions
β”‚   β”œβ”€β”€ package.json
β”‚   └── tsconfig.json
└── README.md                      # Master Project Documentation

πŸ“„ License & Attribution

  • Whisper: OpenAI (MIT License)
  • Application Code: Licensed under the MIT License.

VoiceFlow AI β€” Built with ❀️ for next-generation speech intelligence.
If you find this project helpful, please give it a ⭐ Star on GitHub!

About

πŸŽ™οΈ The open-source AI Voice-to-Text & Universal Speech Studio. Real-time dictation, 48kHz Web Audio DSP noise cancellation, 2-Way Babel Live Translator with authentic Urdu/multilingual Neural TTS, Executive MoM PDF generator, Studio EQ mastering, and 100% offline audio tools. Powered by Gemini 2.0, GPT-4o & Claude 3.7.

Topics

Resources

Stars

8 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages