Neural Topic Modeling with BERTopic

#topic modeling #bert #nlp #unsupervised learning #clustering #dimensionality reduction #text analysis #python #hugging face #umap

1. What is Topic Modeling?

What is Topic Modeling?

Topic modeling is an unsupervised machine learning technique for discovering latent semantic structures in large text corpora. It operates by statistically analyzing word co-occurrence patterns to group documents into clusters, or topics, where each topic is represented as a probability distribution over words. Formally, given a corpus of N documents and a vocabulary of size V, topic modeling aims to infer:

$$ P(w|d) = \sum_{k=1}^K P(w|z=k)P(z=k|d) $$

where w denotes a word, d a document, and z a latent topic. The model learns two distributions: P(w|z) (words per topic) and P(z|d) (topics per document).

Historical Context and Evolution

Early approaches like Latent Semantic Indexing (LSI) used singular value decomposition to project documents into a lower-dimensional space. Probabilistic Latent Semantic Analysis (pLSA) introduced a generative framework but suffered from overfitting. Latent Dirichlet Allocation (LDA), proposed by Blei et al. in 2003, became the dominant method by introducing Dirichlet priors to model document-topic and topic-word distributions, enabling better generalization.

Neural Advancements

Traditional methods rely on bag-of-words representations, discarding word order and contextual semantics. Neural topic models, such as those leveraging BERT embeddings, address this by:

Applications

Topic modeling is widely used in:

For example, BERTopic combines BERT embeddings with dimensionality reduction (UMAP) and clustering (HDBSCAN) to produce interpretable topics while preserving semantic nuances lost in count-based models.

Traditional vs. Neural Topic Modeling

Probabilistic Foundations of Traditional Topic Models

Traditional topic modeling approaches like Latent Dirichlet Allocation (LDA) operate on bag-of-words representations using discrete probability distributions. The core generative process assumes:

$$ P(w|d) = \sum_{z=1}^K P(w|z)P(z|d) $$

where w represents words, d documents, and z latent topics. This formulation suffers from the curse of dimensionality in vocabulary space and cannot capture semantic relationships between words. The Dirichlet prior:

$$ P(\theta_d) = \frac{\Gamma(\sum_{i=1}^K \alpha_i)}{\prod_{i=1}^K \Gamma(\alpha_i)} \prod_{i=1}^K \theta_{di}^{\alpha_i - 1} $$

imposes strong assumptions about topic distributions that may not hold for real-world text data.

Neural Paradigm Shift in Topic Modeling

Neural topic models replace discrete distributions with continuous embeddings from deep neural networks. BERTopic leverages transformer architectures to produce dense document representations:

$$ h_d = \text{BERT}(d)[\text{CLS}] $$

where the [CLS] token embedding captures document-level semantics. This enables:

Dimensionality Reduction Tradeoffs

Traditional LDA operates directly in vocabulary space (typically 104-105 dimensions), while neural approaches first project documents into lower-dimensional spaces (typically 128-768 dimensions). BERTopic employs UMAP for non-linear dimensionality reduction:

$$ Y = \text{UMAP}(X; \text{min\_dist}=0.0, n\_neighbors=15) $$

This preserves both local and global structure better than linear methods like PCA or LSA used in traditional pipelines.

Clustering Algorithm Evolution

Where LDA uses collapsed Gibbs sampling for inference:

$$ P(z_i = k|z_{-i}, w) \propto \frac{n_{k,-i}^{(w_i)} + \beta}{n_{k,-i}^{(\cdot)} + V\beta} \frac{n_{d,-i}^{(k)} + \alpha}{n_{d,-i}^{(\cdot)} + K\alpha} $$

BERTopic employs density-based clustering (HDBSCAN) on the reduced embeddings:

$$ C = \text{HDBSCAN}(Y; \text{min\_cluster\_size}=10) $$

This eliminates the need to pre-specify topic numbers and handles outliers more effectively.

Representational Power Comparison

Traditional models capture word co-occurrence statistics, while neural models learn hierarchical semantic features. The table below contrasts key characteristics:

Feature Traditional (LDA) Neural (BERTopic)
Representation Space Discrete (Vocabulary) Continuous (Embedding)
Context Handling Bag-of-words Full sequence context
Dimensionality O(V) O(d), d ≪ V
Training Objective Likelihood maximization Representation learning

Practical Considerations

Neural approaches require GPU acceleration for transformer inference, while traditional models can run on CPUs. However, BERTopic's two-phase design (embedding then clustering) enables:

Traditional vs. Neural Topic Modeling – Neural Topic Modeling with BERTopic – Tutorial Diagram
Diagram Description: The diagram would show the comparative pipeline architectures of LDA (discrete probability distributions in vocabulary space) versus BERTopic (continuous embeddings through transformers and UMAP reduction).

Why BERTopic?

Traditional topic modeling techniques like Latent Dirichlet Allocation (LDA) and Non-Negative Matrix Factorization (NMF) rely on bag-of-words representations, which discard semantic relationships between words. While these methods are computationally efficient, they struggle with polysemy, synonymy, and contextual meaning—critical aspects of natural language understanding. BERTopic addresses these limitations by leveraging transformer-based embeddings, specifically sentence-BERT (SBERT), to capture dense, context-aware representations of text.

Semantic Richness Through Embeddings

BERTopic's core strength lies in its use of pre-trained language models to generate document embeddings. Unlike LDA, which operates on term-frequency matrices, BERTopic first maps documents to a high-dimensional semantic space using SBERT. The resulting embeddings preserve syntactic and semantic relationships, enabling the model to cluster documents based on meaning rather than lexical overlap. Mathematically, given a document d, its embedding e is computed as:

$$ \mathbf{e} = \text{SBERT}(d) \in \mathbb{R}^{768} $$

where the dimensionality depends on the underlying transformer model (e.g., 768 for all-MiniLM-L6-v2). This approach captures fine-grained semantic distinctions, such as differentiating "bank" as a financial institution versus a riverbank.

Dimensionality Reduction and Clustering

After embedding, BERTopic reduces dimensionality using UMAP (Uniform Manifold Approximation and Projection), which preserves local and global structures better than linear methods like PCA. The reduced embeddings are then clustered using HDBSCAN, a density-based algorithm that automatically detects the number of topics and handles outliers. The combined pipeline optimizes:

$$ \mathcal{L} = \sum_{i=1}^N \min_{c_j \in C} \|\text{UMAP}(\mathbf{e}_i) - \mathbf{c}_j\|^2 $$

where C is the set of cluster centroids. HDBSCAN's soft clustering further allows documents to remain unassigned if they lack topical relevance, reducing noise in the final output.

Dynamic Topic Representation

BERTopic generates interpretable topic labels by applying class-based TF-IDF (c-TF-IDF) to clustered documents. This technique reweights terms based on their importance within a topic relative to the entire corpus, avoiding the need for manual label curation. For a term w in topic k, its score is:

$$ \text{c-TF-IDF}(w, k) = f_{w,k} \cdot \log\left(1 + \frac{N}{n_w}\right) $$

where fw,k is the frequency of w in topic k, N is the total number of topics, and nw is the number of topics containing w. This results in discriminative labels that reflect each topic's unique vocabulary.

Practical Advantages

Why BERTopic? – Neural Topic Modeling with BERTopic – Tutorial Diagram
Diagram Description: The diagram would show the BERTopic pipeline from document embeddings (SBERT) to UMAP dimensionality reduction, HDBSCAN clustering, and c-TF-IDF labeling.

2. Embedding Generation with BERT

Embedding Generation with BERT

BERT (Bidirectional Encoder Representations from Transformers) generates contextualized word embeddings by leveraging transformer architecture. Unlike traditional word embeddings (e.g., Word2Vec, GloVe), BERT captures bidirectional context, making it highly effective for semantic representation in topic modeling. The embedding process involves tokenization, positional encoding, and multi-head attention mechanisms.

Tokenization and Input Representation

BERT uses WordPiece tokenization to split text into subword units, addressing out-of-vocabulary issues. Each input sequence is prepended with a [CLS] token and separated by a [SEP] token for sentence pairs. The input embedding E is the sum of:

$$ E = E_{\text{token}} + E_{\text{position}} + E_{\text{segment}} $$

Transformer Encoder Layers

BERT's transformer encoder stacks multiple identical layers, each containing:

The self-attention mechanism for a single head is defined as:

$$ \text{Attention}(Q, K, V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V $$

where Q, K, and V are learned query, key, and value matrices, and dk is the dimension of the key vectors.

Pooling Strategies for Document Embeddings

To derive a fixed-length document embedding from BERT's token-level outputs, common approaches include:

For BERTopic, mean pooling is often preferred for its balance between computational efficiency and semantic retention:

$$ \mathbf{d} = \frac{1}{N} \sum_{i=1}^{N} \mathbf{h}_i $$

where N is the number of tokens, and hi is the hidden state of the i-th token.

Dimensionality Reduction

BERT embeddings are high-dimensional (e.g., 768 for BERT-base), necessitating reduction for clustering. UMAP (Uniform Manifold Approximation and Projection) is commonly used due to its ability to preserve local and global structures:

$$ \mathbf{d}_{\text{reduced}} = \text{UMAP}(\mathbf{d}; n_{\text{components}}=5) $$

where ncomponents is typically set between 2 and 5 for topic modeling.

Practical Considerations

BERT Transformer Encoder Architecture A block diagram illustrating the BERT Transformer Encoder Architecture, showing input embeddings, multi-head attention, feed-forward networks, residual connections, and layer normalization. BERT Transformer Encoder Architecture Token Embeddings Multi-Head Attention Q/K/V Softmax Add & Layer Norm Feed Forward Add & Layer Norm ... Output Embeddings
Diagram Description: The diagram would physically show the transformer encoder architecture with multi-head attention, feed-forward networks, and residual connections, illustrating how token embeddings flow through the layers.

Dimensionality Reduction with UMAP

BERTopic leverages Uniform Manifold Approximation and Projection (UMAP) for dimensionality reduction of high-dimensional document embeddings before clustering. UMAP is preferred over traditional methods like PCA due to its ability to preserve both local and global structure in the data, making it particularly effective for topic modeling where semantic relationships between documents are crucial.

UMAP Theoretical Foundations

UMAP operates on the principle of constructing a high-dimensional weighted graph representation of the data and optimizing a low-dimensional layout to preserve the graph's topological structure. The mathematical foundation consists of three key components:

  1. Fuzzy topological representation: For each point xi, a neighborhood is defined using an adaptive exponential kernel:
$$ \rho_i = \min\{d(x_i, x_j) | 1 \leq j \leq k, d(x_i, x_j) > 0\} $$
$$ \sigma_i = \text{argmin}_\sigma \left( \sum_{j=1}^k \exp\left(\frac{-\max(0, d(x_i, x_j) - \rho_i}{\sigma}\right) - \log_2(k) \right) $$
  1. Graph construction: The high-dimensional probabilities are symmetrized to create a fuzzy simplicial set:
$$ p_{j|i} = \exp\left(\frac{-\max(0, d(x_i, x_j) - \rho_i)}{\sigma_i}\right) $$
$$ p_{ij} = p_{j|i} + p_{i|j} - p_{j|i}p_{i|j} $$
  1. Low-dimensional optimization: The cross-entropy between the high and low-dimensional representations is minimized using stochastic gradient descent.

Practical Implementation in BERTopic

In BERTopic, UMAP serves two critical functions:

The key parameters that significantly impact topic quality are:

from bertopic import BERTopic
from umap import UMAP

umap_model = UMAP(n_neighbors=15, 
                 n_components=5, 
                 min_dist=0.0, 
                 metric='cosine', 
                 random_state=42)
                 
topic_model = BERTopic(umap_model=umap_model)

Advanced Optimization Techniques

For research-grade implementations, consider these optimizations:

The choice of UMAP parameters significantly affects HDBSCAN clustering performance downstream. Empirical studies show that n_components between 5-20 works best for most text corpora, with higher values preserving more variance but increasing computational cost.

Dimensionality Reduction with UMAP – Neural Topic Modeling with BERTopic – Tutorial Diagram
Diagram Description: The diagram would show the transformation process from high-dimensional BERT embeddings to low-dimensional UMAP space, illustrating how local and global structures are preserved.

2.3 Clustering with HDBSCAN

HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise) extends DBSCAN by converting it into a hierarchical clustering algorithm and extracting stable clusters through a density-based approach. Unlike traditional centroid-based methods like k-means, HDBSCAN identifies clusters of varying densities while automatically determining the optimal number of clusters.

Core Algorithm

The algorithm operates in four key steps:

$$ d_{\text{mreach}}(a,b) = \max \left\{ \text{core}_k(a), \text{core}_k(b), d(a,b) \right\} $$

where $$\text{core}_k(x)$$ is the distance to the k-th nearest neighbor of point $$x$$.

$$ \lambda = \frac{1}{\text{distance}} $$

Key Advantages in Topic Modeling

HDBSCAN's non-parametric nature makes it particularly suitable for topic modeling because:

Practical Implementation

The BERTopic implementation uses HDBSCAN's soft clustering capabilities to:

from bertopic import BERTopic
from hdbscan import HDBSCAN

# Custom HDBSCAN configuration
hdbscan_model = HDBSCAN(min_cluster_size=15, 
                        metric='euclidean',
                        cluster_selection_method='eom',
                        prediction_data=True)

topic_model = BERTopic(hdbscan_model=hdbscan_model)

Parameter Optimization

Critical parameters for topic modeling applications include:

The optimal parameterization depends on the embedding space dimensionality and document corpus characteristics. A grid search over these parameters combined with topic coherence validation often yields the best results.

Clustering with HDBSCAN – Neural Topic Modeling with BERTopic – Tutorial Diagram
Diagram Description: The diagram would show the hierarchical cluster tree formation from the MST and how stable clusters are extracted based on density thresholds.

Topic Representation and Visualization

BERTopic generates interpretable topic representations by leveraging the semantic embeddings from transformer models. Each topic is characterized by a set of representative terms, which are extracted using a class-based TF-IDF (c-TF-IDF) approach. This method reweights term frequencies within topics relative to their frequencies across the entire corpus, ensuring discriminative and meaningful topic descriptors.

Class-based TF-IDF Formulation

The c-TF-IDF score for term t in topic k is computed as:

$$ W_{t,k} = f_{t,k} \times \log \left(1 + \frac{N}{f_t}\right) $$

where:

This formulation emphasizes terms that are frequent within a specific topic but relatively rare in others, enhancing topic distinctiveness.

Visualization Techniques

BERTopic supports multiple visualization methods to facilitate topic interpretation:

1. Intertopic Distance Map

This visualization projects topics into a 2D space using dimensionality reduction techniques like UMAP or PCA, where the distance between topics reflects their semantic similarity. Topics are represented as bubbles, with size proportional to their prevalence in the corpus.

2. Topic Word Scores

A bar chart displays the top n terms per topic with their c-TF-IDF scores, allowing quick assessment of term relevance. The chart is interactive in Jupyter environments, enabling dynamic exploration of topic compositions.

3. Hierarchical Clustering

Topics can be organized into a dendrogram to reveal their hierarchical relationships. This is particularly useful for identifying meta-topics or merging similar topics during post-processing.

Practical Implementation

The visualization tools are accessible through BERTopic's API. For example, generating an intertopic distance map requires:

from bertopic import BERTopic
topic_model = BERTopic()
topics, probs = topic_model.fit_transform(docs)
topic_model.visualize_topics()

Customizations like adjusting the UMAP parameters (n_neighbors, min_dist) or switching to PCA can refine the layout. For hierarchical visualization:

topic_model.visualize_hierarchy()
Topic Representation and Visualization – Neural Topic Modeling with BERTopic – Tutorial Diagram
Diagram Description: The Intertopic Distance Map visualization is inherently spatial, showing topics as bubbles in a 2D UMAP/PCA projection with distances representing semantic similarity.

3. Installing and Setting Up BERTopic

Installing and Setting Up BERTopic

BERTopic requires Python 3.7 or higher and leverages several key dependencies including sentence-transformers for embedding generation, UMAP for dimensionality reduction, and HDBSCAN for clustering. The package can be installed via pip:

pip install bertopic

GPU Acceleration

For optimal performance with large datasets, enable GPU acceleration by installing PyTorch with CUDA support. The following command installs PyTorch 1.12+ with CUDA 11.3:

pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu113

Custom Embedding Models

BERTopic supports custom sentence transformers from HuggingFace. To use a non-default model, specify it during initialization:

from bertopic import BERTopic
from sentence_transformers import SentenceTransformer

embedding_model = SentenceTransformer("all-mpnet-base-v2")
topic_model = BERTopic(embedding_model=embedding_model)

Dimensionality Reduction Configuration

The default UMAP parameters can be adjusted to optimize topic separation. Key hyperparameters include:

$$ n\_neighbors = 15 $$ $$ min\_dist = 0.0 $$ $$ n\_components = 5 $$ $$ metric = \text{'cosine'} $$

These can be passed as a dictionary to the UMAP model:

umap_params = {
    "n_neighbors": 15,
    "min_dist": 0.0,
    "n_components": 5,
    "metric": "cosine"
}
topic_model = BERTopic(umap_model=umap_params)

Clustering with HDBSCAN

HDBSCAN's density-based clustering requires careful parameter tuning. The minimum cluster size (min_cluster_size) and minimum samples (min_samples) significantly impact results:

hdbscan_params = {
    "min_cluster_size": 10,
    "min_samples": 5,
    "metric": "euclidean",
    "cluster_selection_method": "eom"
}
topic_model = BERTopic(hdbscan_model=hdbscan_params)

Verifying Installation

Confirm successful installation by generating a basic topic model on sample text:

from bertopic import BERTopic
docs = ["Example document one", "Another example document"]
topic_model = BERTopic()
topics, probs = topic_model.fit_transform(docs)

3.2 Preprocessing Text Data

Effective preprocessing is critical for neural topic modeling, as raw text often contains noise that degrades the quality of embeddings and clustering. BERTopic leverages transformer-based embeddings, which are robust to minor variations, but systematic preprocessing ensures optimal performance.

Text Normalization

Text normalization standardizes linguistic variations while preserving semantic meaning. Key steps include:

Noise Removal

Domain-specific noise patterns require targeted removal strategies:

$$ \text{NoiseScore}(t) = \frac{f_d(t)}{f_c(t)} \cdot \frac{1}{\log(1 + \text{df}(t))} $$

Where \( f_d(t) \) is term frequency in domain corpus, \( f_c(t) \) in common corpora, and \( \text{df}(t) \) is document frequency. Terms with high NoiseScore are candidates for removal.

Specialized Filters

Tokenization Refinement

BERT's WordPiece tokenizer benefits from preprocessing adjustments:

from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")

# Custom preprocessing pipeline
def preprocess(text):
    text = text.replace("_", " ")  # Preserve subword boundaries
    tokens = tokenizer.tokenize(text)
    return [t for t in tokens if not t.startswith("##") 
            and len(t) > 1]  # Filter subword fragments

This approach maintains syntactic coherence while reducing vocabulary sparsity. For technical domains, preserving hyphenated compounds (e.g., "state-of-the-art") through selective whitespace manipulation improves embedding quality.

Stopword Optimization

Traditional stopword lists often remove semantically rich terms in specialized domains. A dynamic approach computes term relevance:

$$ \text{Relevance}(w) = \text{IDF}(w) \cdot \text{TopicCoherence}(w) $$

Where TopicCoherence measures the term's association strength with other high-probability terms in candidate topics. Domain-specific stopwords are identified when Relevance(w) < θ (typically θ ≈ 0.1 for scientific texts).

Lemmatization Trade-offs

While lemmatization ("running" → "run") reduces morphological variation, it can:

An optimized pipeline uses POS-tagging to apply lemmatization only to verbs and adjectives, preserving noun forms. SpaCy's rule-based system provides sufficient accuracy for this filtered approach:

import spacy
nlp = spacy.load("en_core_web_sm", disable=["parser", "ner"])

def selective_lemmatize(text):
    doc = nlp(text)
    return [token.lemma_ if token.pos_ in {"VERB", "ADJ"} 
            else token.text for token in doc]

Training the BERTopic Model

BERTopic leverages transformer-based embeddings and clustering techniques to generate interpretable topics. The training process consists of three primary stages: document embedding, dimensionality reduction, and clustering. Each stage is optimized for scalability and semantic coherence.

Document Embedding with Sentence Transformers

BERTopic uses pre-trained language models from the Sentence Transformers library to generate dense vector representations of documents. The default model is all-MiniLM-L6-v2, which balances speed and performance. Given a corpus of documents D = {d₁, d₂, ..., dₙ}, each document dᵢ is mapped to an embedding vector eᵢ ∈ ℝᵈ, where d is the embedding dimension (384 for MiniLM-L6). The embedding process minimizes the cosine distance between semantically similar documents:

$$ \text{sim}(e_i, e_j) = \frac{e_i \cdot e_j}{\|e_i\| \|e_j\|} $$

Dimensionality Reduction via UMAP

High-dimensional embeddings are projected into a lower-dimensional space using Uniform Manifold Approximation and Projection (UMAP). UMAP preserves both local and global structures, making it superior to t-SNE for large datasets. The optimization objective involves minimizing the cross-entropy between high-dimensional and low-dimensional pairwise probabilities:

$$ \mathcal{L}_{\text{UMAP}} = \sum_{i \neq j} \left[ p_{ij} \log \left( \frac{p_{ij}}{q_{ij}} \right) + (1 - p_{ij}) \log \left( \frac{1 - p_{ij}}{1 - q_{ij}} \right) \right] $$

where pij and qij represent the probabilities of neighborhood preservation in the original and reduced spaces, respectively. Key hyperparameters include:

Clustering with HDBSCAN

The reduced embeddings are clustered using HDBSCAN, a density-based algorithm that identifies topics as high-density regions. Unlike k-means, HDBSCAN automatically determines the number of clusters and handles outliers. The algorithm constructs a minimum spanning tree (MST) from the mutual reachability distances:

$$ d_{\text{mreach}}(e_i, e_j) = \max \left[ \text{core}_k(e_i), \text{core}_k(e_j), d(e_i, e_j) \right] $$

where corek(eᵢ) is the distance to the k-th nearest neighbor. Clusters are extracted by condensing the MST and applying a stability-based selection criterion.

Key Hyperparameters

Topic Representation with c-TF-IDF

BERTopic employs class-based TF-IDF (c-TF-IDF) to derive interpretable topic descriptors. For each cluster C, a topic-term matrix is computed as:

$$ W_{t,c} = \text{tf}_{t,c} \times \log \left( 1 + \frac{A}{\text{df}_t} \right) $$

where tft,c is the term frequency in cluster c, dft is the document frequency across all clusters, and A is the average number of documents per cluster. The top n terms with the highest weights define the topic labels.

Implementation Example

from bertopic import BERTopic
from sklearn.datasets import fetch_20newsgroups

# Load sample data
docs = fetch_20newsgroups(subset='all')['data'][:1000]

# Initialize and train BERTopic
topic_model = BERTopic(
    embedding_model="all-MiniLM-L6-v2",
    umap_model={"n_neighbors": 15, "min_dist": 0.1, "metric": "cosine"},
    hdbscan_model={"min_cluster_size": 10, "min_samples": 5}
)
topics, probs = topic_model.fit_transform(docs)
BERTopic Training Pipeline A three-stage pipeline diagram showing BERTopic's workflow: document embeddings, UMAP dimensionality reduction, and HDBSCAN clustering with topic-term matrix generation. 1. Embedding Document Collection all-MiniLM-L6-v2 384D Space 2. Dimensionality Reduction UMAP (n_neighbors=15) 2D Projection 3. Clustering HDBSCAN (min_cluster_size=10) Topic Clusters Topic-Term Matrix c-TF-IDF
Diagram Description: The diagram would show the three-stage pipeline of BERTopic's workflow (embedding → UMAP reduction → HDBSCAN clustering) with vector space transformations and cluster formation.

Interpreting and Evaluating Topics

Topic Coherence Metrics

Quantitative evaluation of topic quality relies heavily on coherence measures, which assess the semantic relatedness of words within a topic. The normalized pointwise mutual information (NPMI) is a widely adopted metric due to its robustness against corpus size biases. Given a topic represented by its top N words W = {w1, w2, ..., wN}, NPMI is computed as:

$$ \text{NPMI}(w_i, w_j) = \frac{\log \frac{P(w_i, w_j)}{P(w_i)P(w_j)}}{-\log P(w_i, w_j)} $$

where P(wi) and P(wi, wj) are the empirical probabilities of word occurrence and co-occurrence within a sliding window (typically 10-15 words) across the corpus. The topic coherence score aggregates pairwise NPMI values:

$$ C_{\text{NPMI}} = \frac{2}{N(N-1)} \sum_{i=1}^{N-1} \sum_{j=i+1}^{N} \text{NPMI}(w_i, w_j) $$

Diversity and Saliency

While coherence measures intra-topic word relatedness, diversity metrics evaluate inter-topic distinctiveness. The most effective approach combines:

For a topic t with M documents, significance is:

$$ S_t = \frac{|\{d : P(t|d) > 0.7\}|}{M} $$

Visual Interpretation Techniques

BERTopic's visualization toolkit employs:

Human-in-the-Loop Validation

For mission-critical applications, combine automated metrics with expert evaluation:

  1. Sample 20-30 documents per topic and assess label accuracy
  2. Compute Krippendorff's alpha for inter-annotator agreement
  3. Iteratively refine the number of topics k until reaching α > 0.8

Case Study: Biomedical Literature

When applied to 50,000 PubMed abstracts, BERTopic with k=100 achieved:

$$ \text{Optimal } k = \argmax_k [C_{\text{NPMI}}(k) \times \log S(k)] $$

4. Fine-Tuning Embeddings for Domain-Specific Data

Fine-Tuning Embeddings for Domain-Specific Data

Pre-trained language models like BERT provide generalized semantic embeddings, but domain-specific data often contains specialized terminology and linguistic patterns that generic embeddings fail to capture optimally. Fine-tuning the underlying transformer model on in-domain corpora aligns the embedding space with the target domain's semantics, improving topic coherence and separation.

Why Fine-Tuning Matters

BERT's pre-training objectives (masked language modeling and next sentence prediction) optimize for general language understanding. However, domains like biomedical research, legal documents, or technical manuals exhibit:

Fine-tuning adjusts the model's parameters to better represent these domain characteristics. The process minimizes:

$$ \mathcal{L}(\theta) = -\mathbb{E}_{(x,y)\sim\mathcal{D}} \left[ \sum_{t=1}^T \log P(y_t|x, y_{

where θ represents the model parameters, x is the input sequence, and y are the target predictions over T tokens.

Implementation Strategies

1. Continued Pre-Training

Also called domain-adaptive pre-training, this approach further trains the base model on in-domain text using the original MLM objective before generating embeddings. For a corpus D with documents d₁...dₙ:

$$ \theta^* = \underset{\theta}{\text{argmin}} \sum_{d\in D} \sum_{i=1}^{|d|} \mathbb{I}_{m_i} \cdot \text{CE}(v_i, f_\theta(d_{\setminus i})) $$

where m_i indicates masked positions and CE is cross-entropy loss.

2. Contrastive Fine-Tuning

This method optimizes the embedding space directly using triplet loss:

$$ \mathcal{L}_{triplet} = \max(0, \delta + \text{sim}(a,p) - \text{sim}(a,n)) $$

where a is an anchor document, p a positive (semantically similar) document, n a negative document, and δ a margin hyperparameter.

Practical Considerations

When fine-tuning for BERTopic:

  • Corpus size: ≥10k documents recommended for stable fine-tuning
  • Batch construction: For contrastive learning, ensure meaningful positive/negative pairs
  • Learning rate: Typically 2e-5 to 5e-5 for transformer layers
  • Early stopping: Monitor validation loss to prevent overfitting

from transformers import BertForMaskedLM, BertTokenizer
import torch

model = BertForMaskedLM.from_pretrained('bert-base-uncased')
tokenizer = BertTokenizer.from_pretrained('bert-base-uncased')

# Domain corpus loading
corpus = load_domain_specific_text()  

# MLM fine-tuning setup
optimizer = torch.optim.AdamW(model.parameters(), lr=5e-5)
for batch in create_mlm_batches(corpus, tokenizer):
    inputs = tokenizer(batch, return_tensors='pt', padding=True, truncation=True)
    outputs = model(**inputs, labels=inputs['input_ids'])
    loss = outputs.loss
    loss.backward()
    optimizer.step()
    optimizer.zero_grad()
  

Evaluation Metrics

Assess fine-tuned embeddings using:

  • Topic coherence (C_v): Measures semantic consistency of top topic words
  • Pairwise distance distribution: Compare intra/inter-topic distances
  • Downstream task performance: Document classification F1 when using embeddings as features

Domain-adapted embeddings typically show 15-30% improvement in topic coherence scores compared to generic embeddings for specialized corpora.

Customizing Topic Representations

BERTopic's flexibility allows for fine-grained control over topic representations through custom embeddings, dimensionality reduction, and hierarchical clustering. The key components—embedding model, vector reduction, and clustering algorithm—can be independently modified to optimize topic coherence and interpretability.

Modifying Embedding Representations

The default Sentence-BERT embeddings can be replaced with any embedding model that produces dense vector representations. For domain-specific applications, models like BioBERT (biomedical) or LegalBERT (legal) often outperform generic embeddings. The embedding step is defined as:

$$ E(d) = \text{Transformer}_{\theta}(d) \in \mathbb{R}^{768} $$

where d is a document and θ are the pretrained transformer parameters. Custom embeddings are injected during BERTopic initialization:

from bertopic import BERTopic
from sentence_transformers import SentenceTransformer

custom_embedder = SentenceTransformer("allenai/scibert_scivocab_uncased")
topic_model = BERTopic(embedding_model=custom_embedder)

Dimensionality Reduction Control

BERTopic uses UMAP by default, but alternatives like PCA or t-SNE can be substituted. The reduction process projects embeddings E into a lower-dimensional space Z:

$$ Z = \text{UMAP}_{\alpha}(E) \in \mathbb{R}^{n \times k} $$

where α are UMAP hyperparameters (n_neighbors, min_dist). The reduction can be customized via:

from umap import UMAP

umap_model = UMAP(n_components=5, metric='cosine')
topic_model = BERTopic(umap_model=umap_model)

Hierarchical Topic Refinement

The default HDBSCAN clustering supports automatic topic hierarchy construction. The algorithm computes a density-based hierarchy where topics are nodes in a dendrogram. The merge threshold λ controls granularity:

$$ \lambda = \frac{\text{min\_cluster\_size}}{\text{min\_samples}} $$

Adjusting these parameters trades off between topic specificity and coverage. For example, increasing min_cluster_size reduces outlier topics but may merge semantically distinct clusters.

from hdbscan import HDBSCAN

hdbscan_model = HDBSCAN(min_cluster_size=50)
topic_model = BERTopic(hdbscan_model=hdbscan_model)

Topic Representation Customization

BERTopic generates topic labels using class-based TF-IDF (c-TF-IDF), which weights terms by their frequency within a topic relative to their corpus frequency:

$$ \text{c-TF-IDF}(w, t) = f_{w,t} \times \log\left(1 + \frac{N}{f_w}\right) $$

where fw,t is term frequency in topic t, and fw is corpus frequency. The n-gram range and stopwords can be adjusted to refine labels:

topic_model = BERTopic(
    n_gram_range=(1, 3),
    stop_words=["example", "stopword"]
)

4.3 Handling Large-Scale Datasets

Processing large-scale datasets in BERTopic requires specialized techniques to manage computational constraints while maintaining model performance. The primary bottlenecks arise from embedding generation, dimensionality reduction, and clustering—each scaling non-linearly with dataset size.

Embedding Optimization

BERT-based embeddings, while powerful, exhibit O(n²) memory complexity due to attention mechanisms. For a corpus of N documents, the memory requirement grows as:

$$ M = 4 \times L \times H \times N \times (N + L) $$

where L is sequence length and H is hidden dimension size. Two mitigation strategies prove effective:

Approximate Nearest Neighbors

Exact pairwise similarity calculations become infeasible beyond ~100k documents. Approximate Nearest Neighbor (ANN) algorithms like HNSW or FAISS reduce the complexity from O(n²) to O(n log n). The trade-off between recall and speed is governed by:

$$ \text{Recall} = 1 - e^{-\frac{k \cdot \text{efSearch}}{N}} $$

where k is the number of neighbors and efSearch controls the search depth. For most applications, efSearch = 200 provides >95% recall while being 40× faster than exact search.

Incremental Clustering

Traditional HDBSCAN requires recomputing the full hierarchy for new data. The online variant implements:

  1. Core sample extraction via κ-nearest neighbor density estimation
  2. Incremental cluster tree construction using mutual reachability
  3. Stability-based pruning with dynamic λ cutoff

The online version processes new documents in O(m log n) time versus O(n log n) for batch processing, where m is the batch size.

Implementation Example


from bertopic import BERTopic
from umap import UMAP
from hdbscan import HDBSCAN

# Configure for large datasets
umap_model = UMAP(n_neighbors=15, n_components=5, metric='cosine', low_memory=True)
cluster_model = HDBSCAN(min_cluster_size=50, 
                       prediction_data=True,
                       approx_min_span_tree=False)

topic_model = BERTopic(umap_model=umap_model,
                      hdbscan_model=cluster_model,
                      verbose=True)
    

Memory-Efficient Topic Reduction

Post-processing with c-TF-IDF scales linearly with the number of topics k and average document length l̄:

$$ C = k \times \bar{l} \times (v + 1) $$

where v is vocabulary size. Hierarchical topic merging via cosine similarity on reduced embeddings (e.g., 16-bit floats) cuts memory usage by 75% without significant quality loss.

5. Analyzing Customer Feedback

5.1 Analyzing Customer Feedback

BERTopic's neural approach to topic modeling enables sophisticated analysis of unstructured customer feedback data by leveraging transformer-based embeddings and density-based clustering. The key advantage lies in its ability to capture semantic relationships between words and phrases that traditional LDA-based models miss, particularly for short-text inputs like survey responses or product reviews.

Embedding Customer Feedback

The first stage involves converting raw text into dense vector representations using sentence transformers. For customer feedback analysis, the all-MiniLM-L6-v2 model is particularly effective due to its balance between performance and computational efficiency:

$$ \mathbf{e}_i = \text{Transformer}_{\text{enc}}(\text{"Product stopped working after 2 days"}) \in \mathbb{R}^{384} $$

These embeddings preserve semantic similarity - complaints about similar issues cluster together in the embedding space regardless of exact wording variations. The dimensionality reduction via UMAP is crucial for handling the high sparsity of real-world feedback data:

$$ \mathbf{z}_i = \text{UMAP}(\mathbf{e}_i, n_{\text{neighbors}}=15, n_{\text{components}}=5) $$

Cluster-Topic Formation

HDBSCAN identifies dense regions in the reduced space as topic clusters, with key benefits for feedback analysis:

The cluster-topic assignment follows:

$$ t_i = \begin{cases} \arg\max_j \text{density}(\mathbf{z}_i|\mathcal{C}_j) & \text{if } \max p(\mathbf{z}_i|\mathcal{C}_j) > 0.7 \\ -1 & \text{(outlier)} \end{cases} $$

Interpretable Topic Representation

BERTopic employs class-based TF-IDF (c-TF-IDF) to extract most representative terms per topic, weighting words by:

$$ w_{t,c} = \text{tf}_{t,c} \times \log\left(\frac{N}{\text{df}_t}\right) \times \frac{\text{avg doc length}}{\text{doc length}_c} $$

Where N is total topics and dft is document frequency across topics. This surfaces both frequent and distinctive terms - crucial for distinguishing similar complaints like "battery life" vs "charging issues".

Dynamic Topic Tracking

For longitudinal feedback analysis, BERTopic's reduce_topics_over_time method aligns topics across time periods by:

  1. Projecting all period embeddings into shared space
  2. Computing centroid drift vectors
  3. Merging topics with cosine similarity > 0.85
from bertopic import BERTopic
from sklearn.datasets import fetch_20newsgroups

# Load customer feedback data 
feedback = fetch_20newsgroups(subset='all')['data'][:1000]

# Initialize and fit model
topic_model = BERTopic(embedding_model="all-MiniLM-L6-v2",
                      umap_n_neighbors=15,
                      hdbscan_min_cluster_size=50)
topics, probs = topic_model.fit_transform(feedback)

# Analyze results
topic_model.get_topic_info()

Practical Considerations

When applying BERTopic to customer feedback:

The resulting topic model enables quantitative tracking of complaint frequencies while preserving the nuanced qualitative differences between feedback categories - a significant improvement over manual tagging or keyword-based approaches.

BERTopic's Customer Feedback Processing Pipeline A block diagram illustrating the transformation pipeline from raw text to embeddings to reduced dimensions to clustered topics in BERTopic, with dimensional annotations and mathematical symbols. BERTopic's Customer Feedback Processing Pipeline Raw Text Transformer Embeddings all-MiniLM-L6-v2 ℝ³⁸⁴ UMAP Reduction (n_neighbors=15) ℝ⁵ HDBSCAN Clusters c-TF-IDF Topics 384D → → 5D
Diagram Description: The diagram would show the transformation pipeline from raw text to embeddings to reduced dimensions to clustered topics, with mathematical mappings between stages.

5.2 Topic Modeling for Academic Research

BERTopic leverages transformer-based embeddings to capture semantic relationships in academic texts, enabling researchers to uncover latent themes across large corpora. Unlike traditional methods like Latent Dirichlet Allocation (LDA), which rely on bag-of-words representations, BERTopic employs sentence embeddings from pre-trained models like all-MiniLM-L6-v2 or paraphrase-multilingual-MiniLM-L12-v2 to preserve contextual meaning. The pipeline consists of three stages: embedding generation, dimensionality reduction via UMAP, and clustering using HDBSCAN.

Embedding Generation and Dimensionality Reduction

Given a corpus of academic papers D = {d1, d2, ..., dn}, BERTopic first computes document embeddings E ∈ ℝn×d where d is the embedding dimension (e.g., 384 for MiniLM). UMAP then projects these into a lower-dimensional space E' ∈ ℝn×k (k typically 5-50) while preserving local and global structure:

$$ \min_{E'} \sum_{i,j} w_{ij} \|E'_i - E'_j\|_2^2 + \gamma \sum_{i,j} (1 - w_{ij}) \exp(-\|E'_i - E'_j\|_2) $$

where wij are neighborhood weights derived from the high-dimensional space, and γ balances local/global trade-offs.

Clustering with HDBSCAN

HDBSCAN identifies dense regions in E' by constructing a minimum spanning tree (MST) from mutual reachability distances:

$$ d_{\text{mreach}}(E'_i, E'_j) = \max\{\text{core}_k(E'_i), \text{core}_k(E'_j), \|E'_i - E'_j\|_2\} $$

where corek(E'_i) is the distance to the k-th nearest neighbor. Clusters emerge by pruning the MST at varying density thresholds, automatically determining the optimal number of topics.

Topic Representation

Each cluster’s topic is represented by extracting the top N terms (default N=10) using a class-based TF-IDF (c-TF-IDF) formulation:

$$ W_{t,c} = f_{t,c} \cdot \log \left(1 + \frac{A}{f_t}\right) $$

where ft,c is the frequency of term t in cluster c, A is the average number of documents per cluster, and ft is the overall term frequency.

Case Study: Analyzing arXiv Papers

When applied to 50,000 arXiv abstracts in physics, BERTopic identified 120 topics with minimal overlap. For example, a cluster dominated by terms like "quantum", "entanglement", "qubit" was automatically labeled Quantum Information, while another with "dark matter", "halo", "WIMP" mapped to Cosmology. The model’s dynamic topic reduction feature merged related subfields (e.g., ML Theory and Optimization) into broader themes when requested.

from bertopic import BERTopic
from sklearn.datasets import fetch_20newsgroups

docs = fetch_20newsgroups(subset='all')['data']
model = BERTopic(embedding_model="all-MiniLM-L6-v2")
topics, probs = model.fit_transform(docs)
Topic Modeling for Academic Research – Neural Topic Modeling with BERTopic – Tutorial Diagram
Diagram Description: The diagram would show the three-stage BERTopic pipeline (embedding generation, UMAP dimensionality reduction, HDBSCAN clustering) with visual representations of high-dimensional embeddings transforming into lower-dimensional clusters.

5.3 Real-Time Topic Monitoring

Real-time topic monitoring with BERTopic involves dynamically updating topic models as new documents stream in, enabling applications like social media trend analysis, customer feedback tracking, and live news aggregation. The core challenge lies in maintaining model consistency while adapting to evolving semantic distributions without full retraining.

Incremental Topic Modeling

BERTopic's online learning capability leverages the following components:

$$ c\text{-}TF\text{-}IDF_{t} = \frac{\sum_{i \in D_t} f_{t,i} \times \log \left(1 + \frac{N}{n_t}\right)}{\|D_t\|} $$

where Dt is the set of documents in topic t, ft,i is the frequency of term i in topic t, and N/nt is the inverse document frequency adjusted for topic prevalence.

Drift Detection and Adaptation

Semantic drift is monitored through:

$$ JSD(P\|Q) = \frac{1}{2} D_{KL}(P\|M) + \frac{1}{2} D_{KL}(Q\|M) $$

where M = (P+Q)/2 and DKL is the Kullback-Leibler divergence. Thresholds above 0.3 typically trigger model revision.

Implementation Pipeline

The real-time processing workflow in BERTopic involves:


from bertopic import BERTopic
from sklearn.feature_extraction.text import CountVectorizer

# Initialize with online learning
topic_model = BERTopic(
    embedding_model="all-MiniLM-L6-v2",
    min_topic_size=15,
    vectorizer_model=CountVectorizer(stop_words="english"),
    calculate_probabilities=True
)

# Stream processing loop
for batch in document_stream:
    embeddings = topic_model.embedding_model.encode(batch)
    topics, probs = topic_model.topics_over_time(
        docs=batch,
        embeddings=embeddings,
        timestamps=[datetime.now()]*len(batch)
    )
    # Update visualization
    topic_model.visualize_topics_over_time()
  

Key parameters for optimization include min_topic_size (controls granularity), n_gram_range (for phrase detection), and nr_samples (for HDBSCAN's approximation quality).

Performance Considerations

Latency-critical deployments require:

Real-Time Topic Monitoring – Neural Topic Modeling with BERTopic – Tutorial Diagram
Diagram Description: The diagram would show the real-time processing workflow with document streaming, embedding updates, and topic visualization in a sequential pipeline.

6. Key Research Papers on BERTopic

6.1 Key Research Papers on BERTopic

6.2 Recommended Books and Articles

6.3 Online Resources and Tutorials