● In Active Build Hinglish Code-Mixing Dual Script Engine Local WebAssembly

SwaLekha

A privacy-first, local-speech subtitle generator engineered specifically for Hinglish, Indian English dialects, and code-mixed vernacular video workflows that mainstream speech tools mangle.

● In Active Build Phonetic Engine ↓

Dialect Intelligence

Accurate code-mixed transcription that mainstream AI gets wrong.

Standard speech-to-text models assume speakers stick to one language. Indian creators naturally switch between English and Hindi mid-sentence. SwaLekha is purpose-built to recognize and transcribe code-mixed speech without garbling terms.

00:01:24,180 --> 00:01:27,420
“Aaj ka product demo ready hai, bas audio synchronization check karni hai.”
“आज का product demo ready है, बस audio synchronization चेक करनी है।”

Key Capabilities

Built for Indian creators, podcasters, and educators.

Linguistics

Hinglish Code-Mixing

Tuned with extensive phonetic dictionaries covering common vernacular phrases, slang, and rapid mid-sentence English-Hindi transitions that break conventional engines.

Scripting

Dual Script Sync

Switch seamlessly between Romanized Latin script ("Aap kaise hain?") and authentic Devanagari ("आप कैसे हैं?"), keeping millisecond timecodes perfectly synchronized.

Privacy

100% Local Inference

Runs quantized Whisper models via WebAssembly and WebGPU directly on your machine. Zero audio is ever transmitted to cloud servers or remote APIs.

Editor

Waveform Subtitle Editor

Interactive timeline block editor with draggable handle boundaries, quick typo corrections, and instant real-time video synchronization preview.

Export

SRT, VTT & Hardcode

Export standard .srt and .vtt files for YouTube and Premiere, or render hardcoded subtitles directly into video with custom fonts, colors, and outlines.

Diarization

Multi-Speaker Detection

Acoustic speaker change detection that automatically tags conversational turns between interviewers and guests on podcasts and talk shows.

Acoustic Architecture

Transcription specifications.

Acoustic Pipeline Technical Method Language / Dialect Focus Privacy & Latency
Speech Model Quantized Whisper model running in WebAssembly (Wasm + WebGPU) English, Hindi, and Hinglish hybrid audio Real-time factor: ~0.35x on Apple Silicon / modern x86
Phonetic Mapping Custom bilingual vocabulary table with code-mix normalization Common tech terms, conversational Hindi loanwords Zero external dictionary calls
Dual-Script Rendering Bi-directional transliteration engine with context-aware casing Romanized Hinglish & Devanagari Hindi Instant sub-millisecond script switching
Timecode Alignment Cross-attention weight inspection for word-level timestamps Universal 10ms frame resolution Exact start/end sync without manual dragging
Data Sovereignty In-memory audio processing buffer Completely offline after initial model cache Zero network footprint

Workflow

From raw video to synchronized subtitles.

01 · Drop

Load Video

Drop your MP4, MKV, or audio file directly into the local browser app.

02 · Transcribe

Local AI Run

Watch local WebAssembly Whisper transcribe speech with code-mixing awareness.

03 · Review

Script Toggle

Inspect Romanized vs Devanagari outputs and adjust any timecode boundaries.

04 · Export

Save Subtitles

Export standard SRT/VTT files or burn subtitles into your final video clip.