Using OpenAI Whisper for Multi-Lingual ASR
1. Overview of Automatic Speech Recognition (ASR)
Overview of Automatic Speech Recognition (ASR)
Automatic Speech Recognition (ASR) is the task of converting spoken language into written text. Modern ASR systems leverage deep learning architectures, particularly sequence-to-sequence models, to achieve high accuracy across diverse languages and acoustic conditions. The core challenge lies in mapping variable-length audio signals to discrete text tokens while handling noise, accents, and linguistic variability.
Mathematical Foundations
ASR can be formally defined as finding the most probable word sequence W given an acoustic signal X:
Using Bayes' theorem, this decomposes into:
where P(X|W) is the acoustic model (probability of audio given text) and P(W) is the language model (prior probability of word sequences). The denominator P(X) is constant for a given input and can be ignored during maximization.
Key Components of ASR Systems
- Feature Extraction: Converts raw audio into compact spectral representations like Mel-Frequency Cepstral Coefficients (MFCCs) or log-mel filterbanks. For a signal x(t), MFCCs are computed through:
where E(m) is the energy in the m-th mel filterbank bin.
- Acoustic Modeling: Modern systems use neural networks like Convolutional Neural Networks (CNNs) or Transformer-based architectures to model P(X|W). The encoder processes input features into hidden states ht:
- Sequence Modeling: Connectionist Temporal Classification (CTC) or attention-based mechanisms align audio frames to output tokens. CTC introduces a blank symbol and marginalizes over all possible alignments:
where 𝒜 is the set of all possible frame-level alignments and ℬ is a function that collapses repeated tokens and removes blanks.
Transformer-Based ASR
State-of-the-art systems like Whisper employ Transformer architectures with self-attention mechanisms. The attention weights αij between positions i and j are computed as:
where Q, K are learned query and key matrices, and dk is the dimension of the key vectors. This allows the model to dynamically focus on relevant audio segments for each output token.
Multilingual Challenges
Multilingual ASR introduces additional complexity due to:
- Phonetic diversity across languages
- Code-switching within utterances
- Variable resource availability for training data
Whisper addresses this through joint training on 96 languages, using language identification tokens to condition the decoder. The model learns shared representations for phonetically similar sounds across languages while maintaining language-specific features.
Performance Metrics
Word Error Rate (WER) is the standard evaluation metric, calculated as:
where S is substitutions, D deletions, I insertions, and N is the total reference words. State-of-the-art systems achieve WERs below 5% on clean English speech, though performance degrades with noise, accents, or low-resource languages.

Key Features of OpenAI Whisper
Architecture and Model Design
Whisper employs a transformer-based encoder-decoder architecture, optimized for sequence-to-sequence speech recognition. The encoder processes raw audio waveforms into latent representations, while the decoder generates transcribed text tokens. Unlike traditional ASR systems, Whisper uses a multitask learning approach, jointly training on transcription, translation, and language identification. The model operates on 30-second audio chunks with a fixed stride, enabling efficient batch processing.
Multilingual Capabilities
Whisper supports 99 languages with native script output, including low-resource languages like Amharic and Kyrgyz. The model demonstrates zero-shot cross-lingual transfer, outperforming supervised baselines on languages with minimal training data. Language identification is handled implicitly through the decoder's token space, eliminating the need for separate LID modules.
Robustness to Noise and Accents
The training dataset includes diverse acoustic conditions—studio recordings, telephone calls, and background noise—enabling superior performance in real-world environments. Whisper's attention mechanism dynamically weights relevant audio features, suppressing irrelevant noise. Empirical tests show a 23% lower WER on accented speech compared to Wav2Vec 2.0.
Timestamp Generation
Whisper produces word-level timestamps by aligning decoder attention weights with encoder time steps. This enables applications like subtitling and audio indexing without additional alignment models. The timestamp accuracy achieves ±20ms precision on clean speech.
Open-Weights Deployment
Unlike proprietary ASR APIs, Whisper's weights (1.5B to 1550M parameters) are publicly available. The model can be fine-tuned on domain-specific data, with quantization techniques enabling real-time inference on consumer GPUs. The open-source implementation includes optimized kernels for Intel MKL and CUDA.
Performance Benchmarks
On the MLS benchmark, Whisper Large-v3 achieves:
- 4.1% WER on English
- 7.8% WER on multilingual tasks
- 12.3% CER on logographic scripts (Chinese/Japanese)

Applications of Multi-Lingual ASR
Global Business Communication
Multi-lingual ASR systems like Whisper enable real-time transcription of international business meetings, eliminating language barriers. The model's ability to handle code-switching—where speakers alternate between languages mid-sentence—makes it particularly valuable for multinational corporations. For example, a meeting between German, Mandarin, and English speakers can be transcribed with word-level language identification, allowing for accurate translation pipelines.
Academic Research in Linguistics
Researchers leverage Whisper's multi-lingual capabilities to analyze low-resource languages and dialects. The model's zero-shot performance on unseen languages provides a valuable tool for documenting endangered languages, where traditional ASR systems would require thousands of hours of labeled data. Phonetic patterns can be extracted using:
where φ(t) represents the language-specific phonetic likelihood at time t, and θlang denotes the language-specific acoustic model parameters.
Media Localization
Streaming platforms use multi-lingual ASR to automate subtitle generation across 50+ languages. Whisper's architecture reduces the traditional multi-stage pipeline (transcription → translation → dubbing) into a single end-to-end process. The attention mechanism in Transformer blocks enables:
- Cross-lingual transfer learning
- Context-aware translation of idioms
- Speaker diarization in mixed-language content
Telemedicine
In healthcare, Whisper's multi-lingual capabilities assist in transcribing patient-doctor conversations for non-native speakers. The system maintains HIPAA compliance by processing audio locally, with medical terminology accuracy enhanced through domain adaptation. A hospital in Switzerland reported a 40% reduction in consultation time when using Whisper for German-French-Italian triage notes.
Government and Legal Systems
Courtrooms and immigration services deploy multi-lingual ASR to create official records. Whisper's confidence scoring mechanism (output log-probabilities) allows legal professionals to flag low-certainty segments for human review. The system achieves 85-92% accuracy on UN parliamentary debates across the six official languages.
Technical Implementation Note
For real-time applications, the chunked processing workflow uses:
def transcribe_stream(stream, model):
for chunk in stream:
segments = model.transcribe(
chunk,
language=None, # auto-detect
temperature=(0.0, 0.2, 0.4, 0.6), # beam search
suppress_tokens=[-1] # no silence
)
yield segments
2. Installation and Environment Setup
2.1 Installation and Environment Setup
System Requirements
OpenAI Whisper requires a CUDA-capable GPU for optimal performance, though CPU execution is possible with degraded speed. The model has been tested on Linux and Windows (via WSL2), with macOS support limited to M1/M2 chips for hardware acceleration. Ensure Python ≥3.8 is installed, along with PyTorch 1.10+ compiled with CUDA 11.3+ for GPU support. Memory requirements scale with model size—the largest variant (Whisper-large-v3) demands 10GB VRAM for inference.
Dependency Installation
Create a clean Python environment using conda or venv to avoid dependency conflicts:
conda create -n whisper_env python=3.10
conda activate whisper_env
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cu118 # CUDA 11.8
pip install git+https://github.com/openai/whisper.git
For systems without CUDA, omit the PyTorch CUDA flag. Additional dependencies include ffmpeg for audio processing:
sudo apt update && sudo apt install ffmpeg # Linux
brew install ffmpeg # macOS
Model Download and Verification
Whisper automatically downloads pre-trained weights on first use. To manually specify a model (e.g., large-v3) and verify SHA-256 checksums:
import whisper
model = whisper.load_model("large-v3", download_root="./model_weights")
print(whisper.available_models()) # Verify model variants
GPU Configuration
For multi-GPU systems, specify the device ID using PyTorch's CUDA_VISIBLE_DEVICES. Benchmark VRAM usage with:
import torch
print(torch.cuda.get_device_name(0)) # Verify GPU detection
print(torch.cuda.memory_allocated()) # Monitor VRAM usage
Precision and Performance Tradeoffs
Whisper supports FP16 and INT8 quantization. For Tesla T4 GPUs (16GB VRAM), FP16 provides optimal accuracy-latency balance:
Where params_count is 1.55B for large-v3. To enable dynamic quantization:
model = whisper.load_model("large-v3").half().to("cuda") # FP16
quantized_model = torch.quantization.quantize_dynamic(model, {torch.nn.Linear}, dtype=torch.qint8)
2.2 Loading Pre-trained Whisper Models
The Whisper architecture provides pre-trained models ranging from 39M to 1.5B parameters, each trained on 680,000 hours of multilingual and multitask supervised data. Loading these models requires understanding their transformer-based encoder-decoder structure and the specific tokenization process for speech inputs.
Model Variants and Selection
OpenAI releases five primary model sizes, with tradeoffs between accuracy and computational requirements:
- Tiny (39M params) - Minimal footprint for edge devices
- Base (74M) - Balanced option for general use
- Small (244M) - Improved accuracy for most languages
- Medium (769M) - High-quality transcription
- Large (1.5B) - State-of-the-art performance
The model selection depends on your target language's linguistic complexity and available compute resources. For example, tonal languages like Mandarin benefit more from larger models than Germanic languages.
Initialization Process
Loading a Whisper model involves both the acoustic model and its corresponding tokenizer. The Hugging Face Transformers library provides the most flexible interface:
from transformers import WhisperForConditionalGeneration, WhisperTokenizer
model_name = "openai/whisper-large"
model = WhisperForConditionalGeneration.from_pretrained(model_name)
tokenizer = WhisperTokenizer.from_pretrained(model_name)
Key initialization parameters include:
torch_dtype- Precision control (float16 for GPU optimization)device_map- Multi-GPU distributionlow_cpu_mem_usage- Critical for large model loading
Memory Optimization Techniques
For the 1.5B parameter model (requiring ~6GB VRAM in float32), apply these optimizations:
model = WhisperForConditionalGeneration.from_pretrained(
"openai/whisper-large",
torch_dtype=torch.float16,
low_cpu_mem_usage=True,
device_map="auto"
)
The memory footprint follows this approximate relationship:
Where dtype is 1 for float32, 0.5 for float16, and Mactivations scales with input sequence length.
Feature Extraction Pipeline
Whisper expects log-Mel spectrogram inputs with specific preprocessing:
from transformers import WhisperFeatureExtractor
feature_extractor = WhisperFeatureExtractor.from_pretrained(model_name)
inputs = feature_extractor(
raw_audio,
sampling_rate=16000,
return_tensors="pt"
)
The spectrogram dimensions must match the model's expectations:
Where F represents Mel frequency bins and T the time steps based on the 20ms window stride.
Hardware Requirements and Optimization
Computational Demands of Whisper Models
Whisper's transformer-based architecture imposes significant computational requirements, scaling with model size. The largest variant, Whisper-large-v3, contains 1.5 billion parameters and requires approximately 6GB of GPU memory for inference at FP16 precision. The memory footprint follows:GPU Selection Criteria
For optimal Whisper performance, consider:- Tensor Cores: NVIDIA's Ampere or Hopper architectures (e.g., A100, H100) accelerate FP16/INT8 operations via Tensor Cores, achieving 2-4x speedup over Pascal/Volta.
- Memory Bandwidth: Models are memory-bound; HBM2e (e.g., A100's 1555GB/s) outperforms GDDR6 by 3-5x in throughput.
- VRAM Capacity: Large-v3 requires ≥16GB for batch processing; multi-GPU configurations need NVLink for efficient parameter sharing.
Quantization Techniques
Post-training quantization reduces memory usage while maintaining accuracy:- FP16: Default precision, halves memory vs FP32 with minimal accuracy loss (≤0.5% WER increase).
- INT8: Achieves 50% memory reduction via dynamic quantization, but requires calibration datasets to minimize WER degradation (typically 1-2%).
Optimization Strategies
Kernel Fusion
CUDA graph optimization fuses operations like layer normalization and GeLU activation, reducing kernel launch overhead by 15-20% on Ampere GPUs.Flash Attention
Memory-efficient attention reduces peak VRAM usage by recomputing attention scores on-the-fly rather than storing the full n×n matrix. Implementation requires:model = whisper.load_model("large-v3", device="cuda")
model.decoder.use_flash_attention = True # Enable memory-efficient attention
Batch Processing
Dynamic batching groups audio chunks of similar length to maximize GPU utilization. Optimal batch sizes follow:CPU-Based Deployment
For edge devices, optimize via:- ONNX Runtime: Enables hardware-agnostic optimizations like operator fusion and quantized execution on Intel/ARM CPUs.
- Pruning: Removing 30-50% of attention heads minimally impacts WER while reducing compute by 1.5-2x.
- Core Allocation: Bind processes to NUMA nodes to minimize memory latency, critical for long-form transcription.

3. Audio Preprocessing Techniques
3.1 Audio Preprocessing Techniques
Whisper's performance in multi-lingual automatic speech recognition (ASR) is highly dependent on the quality of input audio. Preprocessing steps must address noise, sample rate inconsistencies, and spectral distortions while preserving linguistic content. Below are the key techniques for optimizing audio input for Whisper.
Sample Rate Normalization
Whisper operates on 16 kHz mono audio. Resampling is necessary when the input deviates from this specification. The resampling process involves anti-aliasing filtering to prevent spectral artifacts. Given an input signal x[n] with sample rate fs, the target sample rate fs' = 16 kHz is achieved via a polyphase filter:
where h[n] is a low-pass filter with cutoff at the Nyquist frequency of the target rate. Librosa's resample function or FFmpeg's aresample filter are practical implementations.
Noise Reduction
Background noise degrades Whisper's transcription accuracy. Spectral subtraction is a common denoising technique, where noise estimates are subtracted from the signal's magnitude spectrum:
|N(f)| is the noise spectrum estimated from non-speech segments, α controls subtraction aggressiveness (typically 1–2), and β (e.g., 0.1) preserves weak speech components. Real-world implementations often use Wiener filtering or deep learning-based tools like RNNoise.
Voice Activity Detection (VAD)
Silence trimming reduces computational load and false transcriptions. A robust VAD algorithm evaluates energy and spectral entropy:
where P(k) is the normalized power spectral density. WebRTC's VAD or Silero-VAD are production-grade choices.
Peak Normalization and Dynamic Range Compression
Whisper performs best with audio normalized to -3 dBFS peak amplitude. Dynamic range compression (DRC) further stabilizes volume:
For DRC, a ratio of 4:1 with 10 ms attack and 100 ms release times balances naturalness and clarity.
Pre-Emphasis
High-frequency enhancement compensates for speech's natural spectral tilt. A first-order FIR filter with coefficient α = 0.97 is applied:
This step is particularly critical for tonal languages where high-frequency cues carry lexical meaning.

Handling Different Audio Formats
Whisper’s architecture processes raw audio waveforms, but real-world applications require handling diverse audio formats (WAV, MP3, FLAC, AAC, etc.). The model internally converts all inputs to 16-bit PCM at 16kHz, but preprocessing steps are critical for optimal performance. Below, we dissect the technical considerations for format conversion, sampling rate adjustments, and quantization.
Sampling Rate Conversion
Whisper expects a 16kHz sample rate. For inputs with differing rates (e.g., 44.1kHz for CD-quality audio), resampling must preserve spectral content while avoiding aliasing. The conversion involves:
where Tin and Tout are input/output sample periods. Practical implementations use polyphase filters for efficiency. Librosa’s resample or FFmpeg’s aresample are robust choices:
import librosa
audio, sr = librosa.load("input.mp3", sr=16000) # Resamples to 16kHz
Bit Depth and Quantization
Non-PCM formats (e.g., MP3’s lossy compression) introduce quantization noise. Whisper’s Mel-spectrogram frontend normalizes input to [-1, 1], but dynamic range compression in lossy formats can degrade performance. For MP3:
- Decode to 32-bit float before processing to minimize quantization artifacts.
- Apply dithering if reducing bit depth (e.g., 24-bit → 16-bit) to mitigate noise.
Codec-Specific Artifacts
Formats like AAC or Opus use psychoacoustic models that discard imperceptible frequencies. This can remove phoneme-relevant harmonics. Mitigation strategies include:
- Prefer lossless formats (WAV, FLAC) for critical applications.
- Use FFmpeg’s
loudnormfilter to normalize loudness without clipping:
ffmpeg -i input.aac -ar 16000 -ac 1 -c:a pcm_s16le output.wav
Multichannel Audio
Whisper processes mono audio. For stereo or surround inputs:
where N is the number of channels. Advanced methods like beamforming (for microphone arrays) can improve SNR before downmixing.
Streaming and Chunking
For real-time applications, chunked audio must avoid spectral discontinuities at boundaries. Overlap-add (OLA) with 50% overlap and Hann windowing ensures smooth transitions:
3.3 Batch Processing for Large Datasets
Whisper's transformer architecture enables efficient batch processing through optimized attention mechanisms and parallel computation. The key challenge lies in maintaining computational efficiency while handling variable-length audio sequences. For a batch of N audio files with durations t1, t2, ..., tN, we first compute the log-Mel spectrograms:
The batch processing pipeline implements dynamic batching with three critical optimizations:
Sequence Length Bucketing
Audio files are grouped into buckets of similar lengths to minimize padding. For bucket boundaries bk and batch size B, we ensure:
where δ is a hyperparameter typically set to 10% of the average sequence length.
Memory-Efficient Attention
Whisper's modified attention mechanism reduces memory overhead from O(NL2) to O(NL log L) through:
- Local windowed attention with 8-head partitioning
- Gradient checkpointing for backward pass optimization
- Half-precision (FP16) computation where supported
GPU Utilization Strategies
Optimal batch sizes are determined by:
where d is the model dimension (512 for Whisper-base). The constant factor accounts for:
- 4 bytes per parameter (FP32)
- 12 transformer layers
- Additional memory for activations
Practical implementation in PyTorch leverages the DataLoader with custom collation:
def collate_fn(batch):
specs = [item[0] for item in batch]
max_len = max(s.shape[1] for s in specs)
padded = torch.zeros(len(batch), 80, max_len)
for i, s in enumerate(specs):
padded[i, :, :s.shape[1]] = torch.Tensor(s)
return padded, [item[1] for item in batch]
loader = DataLoader(dataset, batch_size=32,
collate_fn=collate_fn,
num_workers=4,
pin_memory=True)
For distributed processing across multiple GPUs, Whisper implements gradient accumulation with:
where G is the number of GPUs and B is the local batch size per GPU.

4. Language Detection and Selection
Language Detection and Selection
Whisper's architecture incorporates a multilingual automatic speech recognition (ASR) system that handles language identification as an inherent part of its sequence-to-sequence modeling. The model does not rely on external language classifiers but instead uses its encoder-decoder attention mechanism to implicitly determine the input language during transcription.
Language-Aware Tokenization
The tokenizer in Whisper is extended with language-specific tokens that serve as soft prompts during decoding. Given an input audio sequence x, the model computes:
where y includes both transcribed text and optional language tokens. The initial hidden state of the decoder is conditioned on these language tokens, allowing the model to adapt its acoustic and linguistic modeling accordingly.
Language Probability Estimation
For a given utterance, Whisper estimates language probabilities through its decoder's output distribution over the language token vocabulary. The probability of language l given the audio x is:
where h0 is the initial decoder state, and Wl, bl are learned parameters for language classification.
Forced Language Selection
When the target language is known a priori, Whisper supports forced decoding by prepending the appropriate language token to the decoder input sequence. This is implemented as:
import whisper
model = whisper.load_model("large")
result = model.transcribe(
audio="sample.wav",
language="ja" # forces Japanese transcription
)
The forced language mode significantly improves accuracy for low-resource languages by preventing confusion between linguistically similar languages.
Multilingual Beam Search
In automatic language detection mode, Whisper employs a modified beam search that maintains multiple hypotheses with different language tokens. The beam search score for hypothesis i at step t is:
where λ is a language token bonus hyperparameter and 𝓛 is the set of language tokens. This encourages early commitment to a consistent language hypothesis while maintaining alternatives.
Language-Specific Acoustic Adaptation
Whisper's encoder demonstrates language-specific feature extraction patterns, as revealed by attention head visualization studies. The model automatically adjusts its spectral processing for tonal languages (e.g., Mandarin) versus stress-timed languages (e.g., English), evidenced by different attention distributions over mel-frequency bins.
Performance Characteristics
Language detection accuracy varies by:
- Audio duration: >95% accuracy for utterances longer than 3 seconds
- Language family: Highest confusion occurs between Scandinavian languages
- Code-switching: Performance degrades linearly with switch frequency
4.2 Customizing Transcription for Specific Languages
Whisper's multilingual automatic speech recognition (ASR) capabilities are built on a transformer-based architecture trained on 680,000 hours of labeled audio data across 96 languages. While the model generalizes well, fine-tuning its behavior for specific languages requires understanding its tokenization strategy, language detection mechanism, and decoding parameters.
Language-Specific Tokenization and Vocabulary
Whisper uses a byte-level Byte Pair Encoding (BPE) tokenizer with a vocabulary size of 50,257 tokens. The token distribution is skewed toward English, with approximately 60% of tokens allocated to English subwords. For non-English languages, the tokenizer dynamically constructs representations through:
- Unicode-aware byte decomposition for rare characters
- Language-specific subword units learned during BPE training
- Special language tokens (e.g.,
<|es|>for Spanish) that condition the decoder
where l is the language token, x the audio features, and t_i the subword tokens.
Forced Language Decoding
To bias transcription toward a target language, Whisper supports forced decoding through its API parameters:
import whisper
model = whisper.load_model("large-v2")
result = model.transcribe(
audio="sample.wav",
language="ja", # ISO-639-1 code
task="transcribe",
temperature=0.0 # Disable sampling
)
Key parameters for language control:
- language: ISO-639-1 code (e.g., "fr" for French)
- task: "transcribe" or "translate" (to English)
- temperature: 0 for greedy decoding, higher for diverse outputs
Adapting to Language-Specific Phonetics
For tonal languages (e.g., Mandarin, Vietnamese) or languages with complex morphology (e.g., Finnish, Turkish), consider:
- Adjusting the compression_ratio_threshold to handle longer word forms
- Increasing no_speech_threshold for languages with different pause patterns
- Using word_timestamps=True for agglutinative languages
where S is substitutions, D deletions, I insertions, and N reference words. Language-specific WER varies from 3.0% (English) to 15.8% (Welsh) in Whisper's benchmarks.
Fine-Tuning for Low-Resource Languages
For languages with <100 hours of training data in Whisper's original dataset (e.g., Yoruba, Kyrgyz), transfer learning from similar languages improves performance:
# Continued pretraining on target language data
model = whisper.load_model("small")
train_dataset = load_custom_data("swahili_clips/")
whisper.finetune(
model,
train_dataset,
freeze_encoder=False,
lang_token="<|sw|>"
)
Optimal hyperparameters for fine-tuning:
- Learning rate: 1e-5 to 5e-5
- Batch size: 16-32 (adjust based on VRAM)
- Training steps: 5,000-10,000 for 50-hour datasets
4.3 Evaluating Transcription Accuracy Across Languages
Quantifying Whisper's performance across languages requires rigorous evaluation metrics and controlled testing conditions. The standard metric for ASR systems is Word Error Rate (WER), calculated as:
where S represents substitutions, D deletions, I insertions, and N the total words in the reference transcript. For morphologically rich languages with agglutinative properties (e.g., Finnish, Turkish), character error rate (CER) often provides better discrimination:
Language-Specific Challenges
Whisper's transformer architecture processes language-agnostic acoustic features through its encoder, but decoder performance varies significantly by language due to:
- Phoneme inventory size: Languages with larger phoneme sets (e.g., !Xóõ with 160 phonemes) challenge the model's discriminative capacity
- Tonal contrasts: Mandarin's four tones produce WER increases of 15-20% compared to non-tonal languages with similar data quantities
- Orthographic depth: Shallow orthographies like Spanish show 30% lower WER than deep orthographies like English for equivalent phonetic divergence
Benchmarking Methodology
Controlled evaluation requires:
- Standardized test sets (e.g., Common Voice, FLEURS) with balanced speaker demographics
- Domain-matched evaluation (medical ASR tests should use clinical terminology)
- Noise augmentation at 0-20dB SNR to simulate real-world conditions
For low-resource languages (≤100h training data), relative WER degradation follows:
where D is training data hours, with coefficients α=42.3, β=0.021, and γ=8.7 derived from cross-lingual transfer learning experiments.
Code Implementation
The following Python snippet demonstrates WER calculation using jiwer:
from jiwer import wer
def calculate_wer(reference, hypothesis):
# Normalize texts: lowercase, remove punctuation
transformation = jiwer.Compose([
jiwer.ToLowerCase(),
jiwer.RemovePunctuation(),
jiwer.Strip()
])
return wer(
reference,
hypothesis,
truth_transform=transformation,
hypothesis_transform=transformation
)
# Example usage:
ref = "The quick brown fox jumps"
hyp = "The quick brown fox jumped"
print(f"WER: {calculate_wer(ref, hyp):.2%}")
Cross-Lingual Performance Patterns
Analysis of Whisper's multilingual benchmarks reveals three distinct performance clusters:
| Cluster | WER Range | Representative Languages |
|---|---|---|
| High-resource | 4-8% | English, Spanish, French |
| Mid-resource | 12-18% | Hindi, Vietnamese, Swahili |
| Low-resource | 22-35% | Yoruba, Kyrgyz, Guarani |
Code-switching scenarios (e.g., Spanglish) exhibit non-linear error accumulation, with WER increasing as:
5. Fine-Tuning Whisper for Domain-Specific Tasks
5.1 Fine-Tuning Whisper for Domain-Specific Tasks
Whisper's general-purpose multilingual ASR capabilities can be significantly enhanced through domain-specific fine-tuning. The model's transformer architecture, trained on 680,000 hours of diverse audio data, exhibits strong transfer learning potential when adapted to specialized vocabularies and acoustic conditions.
Data Preparation for Domain Adaptation
Effective fine-tuning requires high-quality in-domain audio-text pairs. The dataset should preserve Whisper's original sampling rate of 16kHz and include:
- Minimum 50 hours of domain-specific speech (100+ hours ideal for low-resource languages)
- Accurate transcriptions with proper casing and punctuation
- Representative acoustic conditions (e.g., medical dictation vs. industrial noise)
The loss function during fine-tuning combines Whisper's original cross-entropy objective with domain-specific terms:
where α controls transfer learning balance and λ is L2 regularization strength.
Architecture Modifications
While Whisper's encoder-decoder structure remains fixed, these adjustments improve domain adaptation:
- Vocabulary Expansion: Add domain-specific tokens to the 50,257-token vocabulary using byte-level BPE
- Adapter Layers: Insert lightweight adapters (≈0.5% parameter increase) between transformer blocks
- Acoustic Feature Scaling: Adjust log-Mel filterbank parameters for domain-relevant frequency ranges
Training Protocol
The recommended fine-tuning procedure uses progressive unfreezing:
- Train only the final decoder layer for 1,000 steps (learning rate 1e-5)
- Unfreeze remaining decoder layers (learning rate 5e-6)
- Unfreeze entire model (learning rate 1e-6) with gradient clipping at 1.0
For compute-efficient adaptation, LoRA (Low-Rank Adaptation) can be applied to query/key matrices:
with rank r typically set to 8 or 16.
Evaluation Metrics
Beyond standard WER, domain-specific evaluation should include:
- Terminology accuracy (F1 score on domain phrases)
- OOV recall (for added vocabulary items)
- Latency under domain-relevant constraints
# Example Whisper fine-tuning with HuggingFace
from transformers import WhisperForConditionalGeneration, WhisperProcessor
model = WhisperForConditionalGeneration.from_pretrained("openai/whisper-large")
processor = WhisperProcessor.from_pretrained("openai/whisper-large")
# Add domain-specific tokens
new_tokens = ["", ""]
processor.tokenizer.add_tokens(new_tokens)
model.resize_token_embeddings(len(processor.tokenizer))
# LoRA configuration
from peft import LoraConfig, get_peft_model
config = LoraConfig(
r=16, lora_alpha=32, target_modules=["q_proj", "k_proj"],
lora_dropout=0.1, bias="none"
)
model = get_peft_model(model, config)
5.2 Incorporating Whisper into Larger Pipelines
Architectural Considerations for Pipeline Integration
Whisper's encoder-decoder transformer architecture outputs both speech recognition results and alignment information, making it particularly suitable for integration into multi-stage processing pipelines. The model's 30-second sliding window approach requires careful handling when processing continuous audio streams in production environments. For real-time applications, implement a double-buffering system where one buffer feeds Whisper while the other collects new audio samples.
Where nctx is the context window size (1500 tokens), fs is the sample rate, and τproc is the processing time per window. The overlap time τoverlap must be optimized to balance between computational efficiency and transcription accuracy.
Language Identification and Routing
Whisper's built-in language detection (LID) can be leveraged to create dynamic processing pipelines. The LID probabilities for the 99 supported languages are available in the model's output logits:
Where henc is the encoder's final hidden state and Wl, bl are the language classification head parameters. This enables routing to language-specific post-processing modules like:
- Domain-specific language models for medical/legal terminology
- Locale-aware punctuation restoration
- Dialect-sensitive text normalization
Confidence Score Calibration
Whisper's token-level probabilities require temperature scaling for proper confidence estimation in downstream decision systems. The logits-based confidence score ct for token yt at time t should be calibrated as:
Where T is the optimal temperature found on a validation set. These calibrated scores enable reliable filtering for:
- Automatic subtitle generation with quality thresholds
- Uncertainty-aware translation pipelines
- Error-prone segment identification for human review
Parallel Processing with Whisper
For high-throughput scenarios, implement a batch processing pipeline with dynamic batching. Whisper's attention patterns allow for three parallelization strategies:
- Inter-sequence parallelization: Process multiple independent audio streams simultaneously
- Intra-sequence chunking: Split long sequences across devices with gradient checkpointing
- Speculative decoding: Use smaller auxiliary models to predict partial outputs
import whisper
from concurrent.futures import ThreadPoolExecutor
def process_audio(audio_path, model):
result = model.transcribe(audio_path)
return {"text": result["text"], "segments": result["segments"]}
model = whisper.load_model("large-v2")
with ThreadPoolExecutor(max_workers=4) as executor:
futures = [executor.submit(process_audio, path, model)
for path in audio_paths]
results = [f.result() for f in futures]
Post-Processing Integration Points
Whisper's segment-aligned output enables precise integration with downstream NLP components. Key integration points include:
| Output Feature | Downstream Use | Processing Latency |
|---|---|---|
| Word-level timestamps | Media synchronization | +5-15ms |
| Speaker diarization | Multi-participant analysis | +20-50ms |
| Prosody features | Emotion detection | +10-30ms |
The encoder's hidden states (dimension 1280 for large models) can be extracted for custom tasks like voice authentication or acoustic event detection, though this requires careful management of GPU memory bandwidth.

5.3 Handling Low-Resource Languages
Whisper's multilingual capabilities extend to low-resource languages, but performance varies based on available training data. The model leverages transfer learning from high-resource languages through shared representations in its transformer architecture. For languages with limited transcribed speech data, several techniques can improve recognition accuracy:
Data Augmentation Strategies
When fine-tuning Whisper on low-resource languages, data augmentation becomes critical. Effective approaches include:
- Speed perturbation: Adjusting audio playback speed by ±10% to create synthetic variations
- SpecAugment: Time warping, frequency masking, and time masking of spectrograms
- Backtranslation: Using Whisper's own translations to generate additional training pairs
where α balances the connectionist temporal classification (CTC) and sequence-to-sequence losses during fine-tuning.
Transfer Learning Protocol
For optimal adaptation to low-resource languages, follow this protocol:
- Initialize with pre-trained multilingual Whisper weights
- Freeze all layers except the final projection heads
- Train for 5-10 epochs with learning rate 1e-5
- Unfreeze all layers and train with learning rate 5e-6
- Apply gradual unfreezing from top to bottom layers
Language-Specific Adaptations
For tonal languages or those with unique phonemes:
- Expand the tokenizer vocabulary with language-specific characters
- Adjust the mel-filterbank parameters to better capture tonal features
- Incorporate pronunciation dictionaries when available
Case Study: Hokkien (Min Nan)
When adapting Whisper to Hokkien, a language with less than 100 hours of available transcribed speech, researchers achieved 22% WER improvement by:
- Mixing with Mandarin training data (30% ratio)
- Applying tone-aware data augmentation
- Incorporating a hybrid CTC/attention loss with α=0.3
where S, D, I represent substitutions, deletions, and insertions respectively, and N is the total words in the reference.
6. Bias and Fairness in Multi-Lingual ASR
6.1 Bias and Fairness in Multi-Lingual ASR
Automatic Speech Recognition (ASR) systems like OpenAI Whisper exhibit varying performance across languages and dialects due to inherent biases in training data, model architecture, and evaluation methodologies. These biases manifest as disparities in Word Error Rate (WER) across demographic groups, with under-resourced languages often suffering from higher error rates. The WER for a given language L can be expressed as:
where S represents substitutions, D deletions, I insertions, and N the total number of words in the reference transcript. Performance gaps emerge when comparing WERs between high-resource languages (e.g., English) and low-resource languages (e.g., Yoruba):
Sources of Bias in Multi-Lingual ASR
Three primary factors contribute to performance disparities:
- Data Imbalance: Training corpora for languages like English may contain orders of magnitude more hours of speech than minority languages. Whisper's training set, while diverse, still follows this skewed distribution.
- Phonetic Coverage: The tokenizer's subword vocabulary may inadequately represent phonemes in tonal languages or those with rare consonant clusters.
- Acoustic Modeling: The shared encoder architecture may prioritize features common in dominant languages during gradient updates.
Quantifying Fairness Metrics
Beyond WER, fairness can be assessed through:
where L is the set of supported languages. A perfect fairness score of 0 indicates equal performance across all languages. Practical systems should minimize both absolute WER and fairness gap.
Mitigation Strategies
Several approaches can reduce bias in multi-lingual ASR:
- Data Augmentation: Applying speed perturbation and vocal tract length normalization to artificially expand minority language datasets.
- Adaptive Sampling: During training, oversampling underrepresented languages according to the inverse of their dataset size proportion.
- Language-Specific Fine-Tuning: Adding language-specific adapters to the base model while freezing shared parameters.
The effectiveness of mitigation can be measured through the relative improvement metric:
Recent studies show that combining these techniques can reduce fairness gaps by 15-30% while maintaining overall accuracy.
Ethical Considerations
Deploying multi-lingual ASR requires careful consideration of:
- Representation Harm: Poor performance may exclude certain language communities from accessing speech technologies.
- Allocational Harm: Resource constraints may force prioritization of some languages over others.
- Demographic Bias: Even within a language, performance may vary by speaker age, gender, or accent.
Continuous monitoring through disaggregated evaluation across language subgroups is essential for responsible deployment.
6.2 Privacy and Data Security
When deploying OpenAI Whisper for multi-lingual automatic speech recognition (ASR), privacy and data security concerns must be addressed rigorously. Whisper processes raw audio data, which may contain sensitive personal information, necessitating robust safeguards to prevent unauthorized access or misuse.
Data Transmission and Storage Risks
Whisper operates in two primary modes: API-based cloud processing and local deployment. Cloud-based processing introduces risks during data transmission and storage. Even if audio is encrypted in transit (e.g., via TLS 1.2+), residual risks persist if the provider retains data longer than necessary or fails to implement proper access controls. Local deployment mitigates some risks but requires secure storage and processing environments to prevent data leaks.
Where Sensitivity quantifies data confidentiality, Exposure represents attack surface, and Protection measures encryption strength and access controls.
Compliance with Data Protection Regulations
Whisper implementations must comply with regional frameworks like GDPR (EU), CCPA (California), or PIPEDA (Canada). Key requirements include:
- Data Minimization: Only collect audio necessary for the task.
- Anonymization: Strip personally identifiable information (PII) before processing.
- Right to Erasure: Provide mechanisms to delete user data upon request.
For GDPR compliance, ensure lawful basis (e.g., explicit consent) and conduct Data Protection Impact Assessments (DPIAs) for high-risk processing.
Mitigation Strategies
End-to-End Encryption
Implement AES-256 encryption for audio data at rest and in transit. For cloud-based Whisper APIs, use client-side encryption before transmission:
from cryptography.fernet import Fernet
# Generate key (store securely)
key = Fernet.generate_key()
cipher = Fernet(key)
# Encrypt audio data before API call
encrypted_audio = cipher.encrypt(raw_audio_bytes)
Federated Learning for On-Device Processing
For applications requiring continuous model improvement, federated learning allows Whisper fine-tuning without centralized data collection. Devices compute gradient updates locally, sharing only model deltas:
Where ΔWi is the local update from device i, η is learning rate, and ∇ℒ computes loss gradients on local data 𝒟i.
Adversarial Robustness
Whisper models are vulnerable to adversarial audio perturbations—specially crafted noise that causes transcription errors or prompts injection. Defensive measures include:
- Input Sanitization: Apply band-pass filtering (300–3400 Hz) to remove ultrasonic adversarial components.
- Adversarial Training: Augment training data with perturbed samples to improve robustness.
Recent studies show Whisper’s word error rate (WER) degrades by 40–60% under targeted attacks, emphasizing the need for these countermeasures.
6.3 Responsible Deployment of ASR Systems
Automatic Speech Recognition (ASR) systems like OpenAI Whisper offer powerful capabilities for multilingual transcription, but their deployment introduces ethical and technical challenges that must be addressed to mitigate harm. Biases in training data, privacy concerns, and unintended misuse are critical considerations for engineers and researchers.
Bias and Fairness in ASR Systems
ASR models inherit biases from their training data, which can manifest as disparities in transcription accuracy across dialects, accents, and languages. For instance, Whisper's performance may degrade for underrepresented languages or non-native speakers due to imbalanced training corpora. The word error rate (WER) for a given demographic group i can be modeled as:
where Si, Di, and Ii represent substitutions, deletions, and insertions, respectively, and Ni is the total number of words spoken by group i. Mitigation strategies include:
- Curating diverse training datasets with proportional representation of linguistic variants
- Implementing adversarial debiasing techniques during model training
- Continuous monitoring of performance metrics across demographic segments
Privacy-Preserving ASR Deployment
Speech data is inherently sensitive, often containing personally identifiable information (PII) or protected health information (PHI). When deploying Whisper in production environments, consider:
- On-premise deployment: Local processing eliminates cloud transmission risks but requires GPU infrastructure
- Differential privacy: Adding controlled noise to audio inputs or transcriptions
- Automatic redaction: Implementing named entity recognition (NER) to filter sensitive terms
The privacy-utility tradeoff can be quantified through the mutual information I(X; Y) between raw audio X and transcript Y:
where H(Y) is the entropy of the transcript and H(Y|X) the conditional entropy given the audio input.
Environmental Impact Considerations
Large ASR models carry significant computational costs. Whisper's carbon footprint depends on inference hardware and usage patterns. The total energy consumption E for processing N hours of audio is:
where P represents power draw and t processing time per hour of audio. Optimizations include:
- Quantization to 8-bit or 16-bit precision
- Model distillation for edge deployment
- Dynamic batching of audio inputs
Regulatory Compliance Frameworks
Deploying ASR systems in regulated industries requires adherence to:
- GDPR Article 22 for automated decision-making in EU territories
- HIPAA compliance for healthcare applications in the United States
- ISO/IEC 30122 for voice command systems in consumer devices
Technical implementations should include audit logging of all transcriptions with configurable retention policies and secure access controls following the principle of least privilege.
7. Key Research Papers on Whisper
7.1 Key Research Papers on Whisper
- Adapting OpenAI's Whisper for Speech Recognition on Code-Switch ... — OpenAI's Whisper model for Code-Switch Mandarin-English Speech Recognition (ASR) on the SEAME and ASRU2019 cor-pora. We conducted 2 experiments: a) using adaptation data from 1 to 100/200 hours to demonstrate effectiveness of adap-tation, b) examining different language ID setup on Whisper prompt.
- PDF Adapting OpenAI's Whisper for Speech Recognition on Code-Switch ... — Abstract—This paper reports on SOTA results achieved using openAI's Whisper model with adaptation on different adaptation corpus sizes for two established code-switch Mandarin/English corpus - namely SEAME and ASRU2019 corpora. Two key experiments were conducted: a) using adaptation data from 1 to 100/200 hours to demonstrate the effectiveness
- GitHub - sandrohanea/whisper.net: Whisper.net. Speech to text made ... — Whisper.net follows semantic versioning. Starting from version 1.8.0, Whisper.net does not follow the same versioning scheme as whisper.cpp, which creates releases based on specific commits in their master branch (e.g., b2254, b2255).. To track the whisper.cpp version used in a specific Whisper.net release, you can check the whisper.cpp submodule. The commit hash for the tag associated with ...
- Fine-tune OpenAI's Whisper Automatic Speech Recognition (ASR) model — In this notebook, we demonstrated how to fine-tune Whisper for multi-lingual speech recognition and transcription on the IPU. We used a single replica on a total of four IPUs. To reduce the fine-tuning time, more than one replica, hence more IPUs are required. On Paperspace, you can use either an IPU Pod 16 or a Bow Pod 16, both with 16 IPUs ...
- Fine-Tune Whisper For Multilingual ASR with Transformers — Whisper is a pre-trained model for automatic speech recognition (ASR) published in September 2022 by the authors Alec Radford et al. from OpenAI. Unlike many of its predecessors, such as Wav2Vec 2.0 , which are pre-trained on un-labelled audio data, Whisper is pre-trained on a vast quantity of labelled audio-transcription data, 680,000 hours to ...
- Multilingual DistilWhisper: Efficient Distillation of Multi-Task — A common approach to efficient inference is distilling knowledge from a large multilingual teacher model into a smaller model [7, 8].However, to apply such knowledge distillation (KD) to whisper-large-v2, the best and largest Whisper model, we would need to access unavailable information such as the training data across all the tasks and languages, in order to preserve the robustness of the ...
- Anatomy of Industrial Scale Multilingual ASR - arXiv.org — In the field of ASR, OpenAI's Whisper ... With this perspective in mind, in this paper, we build multilingual ASR models using 12.5M hours of pre-training audio and a 1.8M-hour fine-tuning speech dataset. Reflecting our own needs, we focus on a set of high-resource languages—English, Spanish, German, and French. ... Volume 1 (Long and Short ...
- GitHub - openai/whisper: Robust Speech Recognition via Large-Scale Weak ... — Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification.
- OpenAI_Whisper_ASR_Demo.ipynb - Colab - Google Colab — audio = whisper.pad_or_trim(audio) # make log-Mel spectrogram and move to the same de vice as the model mel = whisper.log_mel_spectrogram(audio).to(model. device) # detect the spoken language _, probs = model.detect_language(mel) print (f "Detected language: {max (probs, key=probs.get)} ") # decode the audio options = whisper.DecodingOptions()
- How to use OpenAI's Whisper for speech recognition - Graphcore — Updated August 2023: speeding-up Whisper using Group Quantisation. Whisper is an exciting new language model that takes a novel approach to speech recognition. It produces high quality results, even from low quality audio, and is extremely adaptable to a diverse range of voices and languages, without the need for fine-tuning.
7.2 Open-Source Implementations and Tools
- Explore OpenAI Whisper: ASR and Multilingual Transcription | GoTranscript — OpenAI seems to be living up to their name and this model is completely open source. This is called Whisper and it is in ASR, Automatic Speech Recognition. And this was just launched a few hours ago. So I'm going to show you how you can use Whisper in your Python code. I'm going to show you a collab demo.
- OpenAI Releases Whisper: A New Open-Source Machine ... - MarkTechPost — Using only an off-the-shelf transformer trained on 680,000 hours of weakly-supervised, multi-lingual audio data, OpenAI's Whisper can approach human-level robustness and accuracy in ASR, all without the need for fine-tuning. Best of all, the model is open-source, with various weight sizes available to the public. The model
- Multi-lingual Transcription using Whisper - GitHub — This repository contains a web application for multi-lingual transcription using OpenAI's Whisper Automatic Speech Recognition (ASR) model. Users can upload audio files in WAV, MP3, or M4A formats and get transcriptions in various languages. The application is designed with accessibility and data privacy in mind. - vinhteq/whisper-appUI
- Multi-lingual Transcription using Whisper - GitHub — This repository contains a web application for multi-lingual transcription using OpenAI's Whisper Automatic Speech Recognition (ASR) model. Users can upload audio files in WAV, MP3, or M4A formats and get transcriptions in various languages. The application is designed with accessibility and data privacy in mind.
- [2412.16507] Adapting Whisper for Code-Switching through Encoding ... — Code-switching (CS) automatic speech recognition (ASR) faces challenges due to the language confusion resulting from accents, auditory similarity, and seamless language switches. Adaptation on the pre-trained multi-lingual model has shown promising performance for CS-ASR. In this paper, we adapt Whisper, which is a large-scale multilingual pre-trained speech recognition model, to CS from both ...
- openai/whisper-tiny - Hugging Face — Whisper Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.. Whisper was proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al from OpenAI.
- OpenAI launches Whisper - Multilingual ASR across 90+ languages with ... — THIS IS HUGE Going multi-lingual becomes very accessible and scalable with @OpenAI whisper's translate use-case With this, any content can be converted from base language to English 99 base languages across Indian, Russian, Roman, African, Asian, and more are covered ?
- Eva-Kaushik/Multilingual-Transcription-with-OpenAI_Whisper — Upload Audio: Click on the button and select (or drag and drop) an audio file in WAV, MP3, or M4A format that you want to transcribe.; Transcribe Audio: Once the audio file is uploaded, click on the "Transcribe Audio" button in the sidebar.The application will start transcribing the audio using the Whisper model. Supported Languages: The Whisper model supports multiple languages.
- Unlocking the Power of OpenAI's Whisper for Automatic Speech ... — Enhanced Accuracy: Reduces errors by 10-20% compared to Whisper Large-v2, making it a go-to model for multilingual ASR. The script This script I wrote is designed to record audio and transcribe ...
- Introducing Whisper - OpenAI — Other existing approaches frequently use smaller, more closely paired audio-text training datasets, 1 2, 3 or use broad but unsupervised audio pretraining. 4, 5, 6 Because Whisper was trained on a large and diverse dataset and was not fine-tuned to any specific one, it does not beat models that specialize in LibriSpeech performance, a famously competitive benchmark in speech recognition.
7.3 Community Resources and Tutorials
- Anatomy of Industrial Scale Multilingual ASR - arXiv.org — The open-source ASR models used are Whisper large-v3, the latest model of Whisper, and Canary-1B, the most recent powerful open-source model. Both models are based on an encoder-decoder model architecture. Whisper large-v3 consists of 1.55B parameters, with multilingual ASR and X-to-English translation capabilities.
- Building High-Accuracy Multilingual ASR With Gated Language Experts and ... — Building High-Accuracy Multilingual ASR With Gated Language Experts and Curriculum Training ... We propose gated language experts and curriculum training to enhance multilingual transformer transducer models without requiring user input for language identification (LID) during inference. Our method incorporates a gating mechanism and LID loss ...
- M2ASR: Multilingual Multi-Task Automatic Speech Recognition via Multi ... — To enable the capability of speech models across multiple languages, training multilingual, multi-task automatic speech recognition (ASR) models has gained growing interest. However, different languages and tasks result in distinct training objectives, potentially leading to conflicts during training and degrading the model's performance.
- MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of ... — In the multilingual domain, the Common Voice project, alongside the multilingual LibriSpeech (MLS) corpus released by Meta, has greatly promoted research in multilingual ASR. In recent times, the success of OpenAI's Whisper [ 5 ] model has demonstrated that big data combined with large models can yield improved performance.
- GitHub - openai/openai-cookbook: Examples and guides for using the ... — Navigate at cookbook.openai.com. Example code and guides for accomplishing common tasks with the OpenAI API.To run these examples, you'll need an OpenAI account and associated API key (create a free account here).Set an environment variable called OPENAI_API_KEY with your API key. Alternatively, in most IDEs such as Visual Studio Code, you can create an .env file at the root of your repo ...
- openai/whisper-large-v3 - Hugging Face — Whisper Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on >5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.
- Whisper - AI Wiki - Artificial Intelligence Wiki — Whisper is an automatic speech recognition system developed by OpenAI, released in 2022 , that is capable of generating transcriptions and translations using an audio track as input. OpenAI stated that the model has been "trained on 680,000 hours of multilingual and multitask supervised data collected from the web," approaching "human level robustness and accuracy on English speech recognition ...
- PDF M ASR: Multilingual Multi-Task Automatic Speech Recognition via Multi ... — M 2 ASR: Multilingual Multi-Task Automatic Speech Recognition via Multi-Objective Optimization A F M Saif 1, Lisha Chen 1, Xiaodong Cui 2, Songtao Lu 2, Brian Kingsbury 2, Tianyi Chen 1 1 Rensselaer Polytechnic Institute, Troy, NY, USA 2 IBM Research AI, T. J. Watson Research Center, Yorktown Heights, NY, USA fsaifa,chenl21,chent18 [email protected], fcuix, bedk [email protected], [email protected]
- Multilingual_ASR.ipynb - Colab - Google Colab — Below, we use the cross-attention weights to determine more granular, word-level timestamps. It uses a set of heuristics and dynamic time warping (DTW) to find the alignment between the audio and the transcript.
- TomohikoNakamura/ica_dsu_espnet - GitHub — The demo script utils/ctc_align_wav.sh uses an already pre-trained ASR model (see the list above for more models). It is recommended to use models with RNN-based encoders (such as BLSTMP) for aligning large audio files; rather than using Transformer models with a high memory consumption on longer audio data.








