An ultra-fast, privacy-first Voice-to-Text & Multimodal Speech Intelligence Web Application.
Real-Time Dictation β’ 48kHz DSP Noise Cancellation β’ 2-Way Babel Live Duplex Translator β’ Corporate Meeting Minutes (MoM) PDF Generator β’ Native Urdu & Multilingual Neural TTS β’ Studio EQ Mastering.
π Star on GitHub β’ π Live Web Studio β’ π Documentation β’ β‘ Quick Start
- π Key Highlights & Architecture
- π οΈ Complete Feature Guide & How Each Feature Works
- 1. Real-Time Live Speech Dictation & Spectrum Visualizer
- 2. Universal SweetAlert API Key Guard & Quick Setup
- 3. 2-Way "Babel Mode" Live Duplex Universal Translator
- 4. Multi-Stage 48kHz Web Audio DSP Noise Cancellation
- 5. Original Voice Replay & Studio Vocal Mastering Presets
- 6. Official Corporate Meeting Minutes (MoM) & Executive PDF
- 7. "Ask My Voice Vault" Semantic Audio Search
- 8. Smart Voice Macro Action Board (Spoken Kanban)
- 9. Multi-Speaker Diarization (Dialogue Splitter)
- 10. Floating Dual-Language Theater Subtitles
- 11. Video & Lecture Audio Stream URL Importer
- 12. "Ask AI" Contextual Transcript Chat
- 13. Concept Mindmap Visualizer (Mermaid.js)
- 14. 1-Click Multi-Format Content Repurposer
- π Supported AI Models & Step-by-Step API Key Setup
- π΅π° Authentic Multilingual Voice Synthesis & TTS Architecture
- π» Web Architecture & Tech Stack
- π Step-by-Step Installation & Running Guide
- π Project Directory Structure
- π License & Attribution
- β‘ Zero-Latency Dictation: Live continuous speech streaming with interim preview bubbles.
- π΅π° Native Urdu & Multilingual Speech: Speaks aloud in genuine, fluent Urdu (
Ψ§Ψ±Ψ―Ω), Hindi (ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯), Arabic (Ψ§ΩΨΉΨ±Ψ¨ΩΨ©), Spanish (EspaΓ±ol), French, and 20+ languages. - π DSP Noise Annihilation: Hardware high-pass, low-pass, 60Hz hum notch, and dynamics compression directly in the audio capture pipeline.
- π‘οΈ Guarded Intelligence: Never gets stuck; alerts users if an API key is needed with 10-second instant setup links.
- π Modern Web Stack: Built with Next.js 16 App Router, React 19, Tailwind CSS, and Web Audio API.
- What It Does: Transcribes your spoken words into text instantly with zero lag.
- How It Works: Connects directly to the high-performance browser speech recognition engine with automatic silence recovery. Simultaneously runs a 48-band Web Audio
AnalyserNoderendering a real-time reactive neon audio visualizer. - How to Use: Click the central "Start Recording" button, select your language, and speak naturally.
- What It Does: Protects all cloud-powered AI transformations (Fix Grammar, Summarize, Translate, Repurpose, Mindmap, Ask AI, and Babel Mode).
- How It Works: When an AI button or tab is clicked without an active key, VoiceFlow blocks the action and pops up an interactive SweetAlert modal with direct 1-click links to get free keys for Gemini, GPT-4o, or Claude 3.7.
- How to Use: Click any AI tool β if no key is entered, the modal will guide you to paste your key in Settings in under 10 seconds.
- What It Does: Enables two people who speak completely different languages (e.g. English and Urdu, or Spanish and Arabic) to have a live, translated spoken conversation on a single device.
- How It Works: Person A speaks or types in Language A. VoiceFlow translates the input into Person B's language native script and automatically speaks the translation aloud in Person B's native accent using our dedicated
/api/ttsneural voice engine. - How to Use: Click "Babel Mode" in the top navigation bar, select languages for Person A and Person B, and start chatting.
- What It Does: Eliminates room noise, computer fan hiss, desk vibrations, and electrical hum from recordings.
- How It Works: Builds a multi-node Web Audio processing graph:
- 85Hz High-Pass Biquad Filter: Cuts desk bumps, handling thuds, and low rumble.
- 8500Hz Low-Pass Filter: Cuts high-frequency electronic hiss and fan buzz.
- 60Hz Notch Filter: Cancels AC power line hum.
- Dynamics Compressor Node: Evens out vocal peaks for crystal-clear clarity.
- The
MediaRecorderrecords from the filtered DSP stream so saved audio is pristine.
- What It Does: Allows you to listen back to your original recorded voice with professional broadcast mastering presets and 1-click audio download (
.webm). - How It Works: Features an integrated scrub bar, speed multiplier (
1.0x-2.0x), volume control, and 4 audio presets:- ποΈ Clean DSP: Balanced noise-filtered voice.
- π» Podcast Warmth: +3dB low-end boost (150Hz) for deep radio resonance.
- β‘ Broadcast Radio: +2.5dB high-mid presence (3.5kHz) for voice clarity.
- β¨ Crisp Articulation: Enhanced consonants for technical lectures.
- What It Does: Generates an official, print-ready Corporate Meeting Minutes PDF with executive formatting and formal signature blocks.
- How It Works: Parses the transcript into structured sections: Meeting Context, Objectives, Discussion Breakdown, Action Items Table, and sign-off lines for Host / Speaker 1 and Executive / Client Approver.
- How to Use: Click the "MoM PDF" button in the workspace toolbar.
- What It Does: Allows you to search across weeks and months of saved audio notes by topic, concept, or keyword.
- How It Works: Indexes all saved transcript sessions with duration, tags, and timestamps. Matches keywords in real-time and loads the full session back into the workspace with 1 click.
- How to Use: Click "Saved Notes" in the top action bar and use the search bar.
- What It Does: Automatically turns your spoken voice into interactive Kanban task cards.
- How It Works: The NLP parser detects spoken trigger phrases:
"Task: [something]"β Creates a task card in the Todo column."Idea: [something]"β Pins an idea card."Important: [something]"β Flags an urgent priority item.
- How to Use: Switch to the "Action Board" tab in the workspace to view and check off items.
- What It Does: Segregates conversation segments between two speakers with color-coded speech bubbles and custom names.
- How It Works: Alternates turns between Speaker 1 (Violet) and Speaker 2 (Cyan) with real-time timestamping.
- How to Use: Click the "Speakers Dialogue" tab and type custom names for both speakers.
- What It Does: Displays a cinematic, floating subtitle overlay with original speech on top and translated subtitles below.
- How It Works: Streams live transcription into an ultra-clean glassmorphic overlay designed for live presentations, zoom calls, or video recordings.
- How to Use: Click "Live Subtitles" in the quick action bar.
- What It Does: Directly transcribes audio from online video links, podcast feeds, or lecture URLs.
- How It Works: Accepts streaming audio URLs or audio file uploads (
.mp3,.wav,.m4a,.webm,.ogg) and passes them through OpenAI Whisper for automated transcription. - How to Use: Click "Upload Audio" and switch to the "Paste Audio / Video URL" tab.
- What It Does: Allows you to chat with an AI assistant that has full context over your recorded voice note.
- How It Works: Routes your question and transcript through our server-side
/api/aigateway to Gemini 2.0, GPT-4o, or Claude 3.7. - How to Use: Click "Ask AI" and ask questions like "What were the key numbers mentioned?" or "Draft a follow-up email based on this".
- What It Does: Automatically converts unstructured speech into a structured visual concept mindmap.
- How It Works: Analyzes central themes, sub-topics, and branches, rendering an interactive SVG mindmap diagram.
- How to Use: Click the "Mindmap" tab in the workspace.
- What It Does: Transforms spoken notes into 6 distinct content formats with 1 click:
- βοΈ Executive Email
- π§΅ Viral Twitter / X Thread
- πΌ Engaging LinkedIn Post
- π Meeting Minutes & Summary
- π Study Flashcards & Q&A
- π Blog Post Outline
- How to Use: Click the "Repurpose" tab and select any format.
VoiceFlow AI is equipped with the latest frontier models:
| Provider | Frontier Model | Capabilities | How to Get Your API Key |
|---|---|---|---|
| π Google Gemini | gemini-2.0-flashgemini-2.0-flash-lite |
Ultra-low latency, generous free tier, 1M+ context window | 1. Go to Google AI Studio 2. Sign in and click Create API key 3. Paste into VoiceFlow Settings. |
| β‘ OpenAI | gpt-4ogpt-4o-mini |
Flagship omni multimodal reasoning | 1. Go to OpenAI API Keys 2. Create a new Secret Key 3. Paste into VoiceFlow Settings. |
| π§ Anthropic Claude | claude-3-7-sonnetclaude-3-5-sonnet |
Hybrid reasoning & superior prose formatting | 1. Go to Anthropic Console 2. Generate an API Key 3. Paste into VoiceFlow Settings. |
- True Native Script Translation: Outputs authentic native scripts: Urdu (
Ψ§Ψ±Ψ―Ω), Arabic (Ψ§ΩΨΉΨ±Ψ¨ΩΨ©), Hindi (ΰ€Ήΰ€Ώΰ€¨ΰ₯ΰ€¦ΰ₯), Spanish (EspaΓ±ol), French (FranΓ§ais), German (Deutsch), Chinese (δΈζ), Japanese (ζ₯ζ¬θͺ), and more. - Dedicated
/api/ttsServer Audio Gateway: Eliminates browser limitations and provides high-fidelity audio streams for Urdu and 20+ languages with native accents. - Multi-Tier Voice Fallback: Server Audio Stream β OpenAI TTS (if BYOK key provided) β Browser SpeechSynthesis with regional voice matching.
graph TD
User([User Voice / Microphone]) --> DSP[48kHz Web Audio DSP Filters]
DSP --> AudioPlayer[Original Voice Player & EQ Mastering]
DSP --> Whisper[Speech-to-Text / Whisper]
Whisper --> Workspace[Transcription Workspace]
Workspace --> ServerAPI["/api/ai & /api/tts Gateway"]
ServerAPI --> Gemini["Google Gemini 2.0 Flash"]
ServerAPI --> OpenAI["OpenAI GPT-4o"]
ServerAPI --> Claude["Anthropic Claude 3.7"]
ServerAPI --> Babel["2-Way Babel Translator (Urdu, etc.)"]
| Layer | Web Application Technology |
|---|---|
| Framework | Next.js 16 (App Router) + React 19 |
| Language | TypeScript 5 (Strict Mode) |
| Speech-to-Text | Web Speech API + Groq / OpenAI Whisper |
| Audio Processing | 48kHz High-Pass, Low-Pass, 60Hz Notch & Dynamics Compressor |
| TTS Engine | Next.js /api/tts Neural Stream + Kokoro-82M + Edge Neural |
| Iconography | Font Awesome SVG Icons (@fortawesome/react-fontawesome) |
| State Store | Zustand 5 (IndexedDB / LocalStorage) |
| Document Export | jsPDF (MoM & Transcript) + Canvas-Confetti |
- Node.js: v18.18.0 or higher (v20+ recommended)
- npm or yarn / pnpm
- Git
# 1. Clone the repository
git clone https://github.com/Humaam-04-06/Voice_To_Text_App.git
# 2. Navigate to the web application folder
cd Voice_To_Text_App/web
# 3. Install dependencies
npm install
# 4. Start the local development server
npm run devOpen http://localhost:3000 in your browser!
# 5. Build for production
npm run build
# 6. Run production server
npm startVoice_To_Text_App/
βββ web/ # Next.js 16 Web Application
β βββ src/
β β βββ app/
β β β βββ api/
β β β β βββ ai/ # Server gateway for Gemini, GPT-4o, Claude
β β β β βββ transcribe/ # Whisper audio transcription route
β β β β βββ tts/ # Multilingual Neural TTS streaming route
β β β βββ globals.css # Glassmorphic dark theme styles
β β β βββ layout.tsx # Root layout with Font Awesome CSS & SEO
β β β βββ page.tsx # Main application page
β β βββ components/ # UI Components with Font Awesome SVG icons
β β β βββ ApiKeyRequiredModal.tsx # Universal SweetAlert API key guard
β β β βββ AudioFileUploadModal.tsx # File & Video URL importer
β β β βββ AudioPlaybackPlayer.tsx # Real voice replay & EQ presets
β β β βββ BabelTranslatorModal.tsx # 2-Way live duplex translator
β β β βββ ChatDrawer.tsx # Ask AI transcript chat
β β β βββ HeroRecordingZone.tsx # Recording controls & waveform
β β β βββ HistoryDrawer.tsx # Semantic Voice Vault search
β β β βββ MindmapViewer.tsx # Concept mindmap generator
β β β βββ Navbar.tsx # Gemini 2.0, GPT-4o, Claude 3.7 switcher
β β β βββ SettingsModal.tsx # Key settings with step-by-step guides
β β β βββ SpeechCoachWidget.tsx # Clarity, WPM, and filler words
β β β βββ TheaterSubtitleModal.tsx # Floating dual subtitles
β β β βββ ToneSentimentRadar.tsx # Live emotional mood & tone HUD
β β β βββ TranscriptionWorkspace.tsx # Workspace, tabs, & MoM PDF exporter
β β β βββ VoiceMacroActionBoard.tsx # Spoken triggers & Kanban tasks
β β β βββ WaveformVisualizer.tsx # Real-time audio spectrum
β β βββ hooks/
β β β βββ useAudioRecorder.ts # 48kHz Web Audio DSP noise cancellation
β β β βββ useSpeechRecognition.ts # Continuous streaming speech recognition
β β βββ lib/
β β β βββ ai/ # AI dispatchers & local NLP engines
β β β βββ audio/ # TTS and audio filter engines
β β β βββ constants/ # Supported multilingual list
β β βββ store/ # Zustand Global Voice Store
β β βββ types/ # TypeScript type definitions
β βββ package.json
β βββ tsconfig.json
βββ README.md # Master Project Documentation
- Whisper: OpenAI (MIT License)
- Application Code: Licensed under the MIT License.
VoiceFlow AI β Built with β€οΈ for next-generation speech intelligence.
If you find this project helpful, please give it a β Star on GitHub!