Text Index
Text Index
Level 7 — Indexes & Query Performance The specialized database index type built on string fields to support full-text search, featuring word tokenization, stop-words filtering, relevance priority weights, and wildcard text indexing.
1. Prerequisites
- Text Search (
$text/$search) — The parent query operation. createIndex()/dropIndex()— The index creation triggers.
2. Term Category
Index / Performance (Inverted Word Indexing): A Text Index creates an inverted word-stemmed index over string fields to support full-text search queries ($text).
3. Explanation
Environment Context
- MongoDB Core (Calculated on the database engine. String values are tokenized and stemmed using language dictionaries (like Snowball stemmers) before being written to the index files).
(1) Design Motivation — "Why did we design this?"
As learned in text_search.md, regular expressions cannot handle advanced full-text search patterns (stemming, stop words, or relevance scoring).
We designed the Text Index to enable high-performance text searches natively in MongoDB.
Instead of treating a string as a single value, a Text Index splits the string into individual words (tokenization), stems them to their root forms, discards noise words (like "the"), and catalog-indexes the remaining terms.
This allows fast lookups on large text bodies.
(2) Text Index Rules & Parameters
Rule 1: One Text Index per Collection
A collection can only have one text index.
If you try to build a second text index, MongoDB will throw an error.
However, your single text index can be composite (cover multiple fields at the same time).
Rule 2: Search Weights
When indexing multiple fields (like title and body), you can assign Weights to specify relevance priority.
For example, a keyword match in the title field is usually more important than a match in the body text.
MongoDB uses these weights to calculate the relevance score (textScore).
Rule 3: Wildcard Text Indexing
If your documents contain dynamic, unknown fields (such as user-submitted product specifications) and you want to search all string fields, you can use the wildcard specifier:
db.collection.createIndex({ "$**": "text" })
(3) Reality Metaphor (Book Index Glossaries)
- Text Index: The index glossary at the back of a library book. Every entry is a unique root word (e.g.
"gravity"), followed by a list of pages where it is mentioned. - Weights: A Gold Star system.
- If the word
"gravity"appears in the book's Title, the index gives the book 10 points (Gold Star). - If it only appears in the Bibliography, it gets 1 point (Silver Star).
- Books with higher scores are placed at the front of the display case.
- If the word
(4) Code Examples
Creating a Weighted Text Index
Let's build a text index on articles, prioritizing title matches over body text:
db.articles.createIndex(
{
title: "text", // String 'text' declares the index type
body: "text"
},
{
// Assign weights (relevance priority)
weights: {
title: 10, // Title matches are 5x more important than body!
body: 2
},
name: "article_text_search" // Custom name for management
}
);
To search all string fields dynamically:
db.products.createIndex({ "$**": "text" }); // Wildcard text index
4. Common Mistakes & Pitfalls
Mistake 1: Trying to create a second text index on a collection without dropping the old one first
The mistake: Running db.articles.createIndex({ description: "text" }) when the collection already contains an index built on title: "text".
Why it's wrong: MongoDB limits collections to exactly one text index.
The build command will fail immediately, throwing a database error.
Fix: If you need to change the fields covered by your text search, you must first drop the existing text index using its name, and then build a new, compound text index containing all the fields you want to search:
// CORRECT
db.articles.dropIndex("article_text_search"); // Drop old index
db.articles.createIndex({ title: "text", description: "text" }); // Build new composite index
Mistake 2: Creating Multiple Text Indexes on a Single Collection
The mistake: Attempting to create two separate text indexes on title and description.
Why it's wrong: MongoDB permits at most ONE Text Index per collection! Combine fields into a single compound text index { title: "text", description: "text" }.
Incorrect:
db.posts.createIndex({ title: "text" }); db.posts.createIndex({ description: "text" }); // ❌ Multiple text index error!
Fix:
db.posts.createIndex({ title: "text", description: "text" });
Mistake 3: Using $text Indexes for Real-Time Wildcard Prefix Substring Search
The mistake: Using $text for real-time "admin*" autocomplete search.
Why it's wrong: $text indexes tokenize words using language stemmers, making substring autocomplete slow. Use Wildcard Indexes or Atlas Search for autocomplete.
Incorrect:
// Using $text for real-time search box autocomplete
Fix:
Use Atlas Search edgeGram analyzer or Wildcard Indexes for autocomplete
5. Practice Exercises
Exercise 1: Multi-Field Text Index Construction
Scenario:
Create a full-text search index on title (weight: 10) and body (weight: 1) in collection articles.
Requirements:
- Execute
createIndex({ title: "text", body: "text" }, { weights: { title: 10, body: 1 } }).
Answer
Implementation
db.articles.createIndex(
{ title: "text", body: "text" },
{
weights: { title: 10, body: 1 },
name: "idx_article_text_search"
}
);
Technical Explanation
"text"creates an inverted text search index tokenizing text words.weightsassigns relative relevance importance to field matches (title matches score 10x higher than body matches).- Applies language stemming (e.g. "running" matches "run").
Exercise 2: Text Search Queries with Phrase Matching and Negation
Scenario:
Search articles for exact phrase "database design" while excluding articles containing keyword "oracle".
Requirements:
- Use
$text: { $search: ""database design" -oracle" }.
Answer
Exercise 3: Relevance Score Projection and Sorting
Scenario: Project and sort text search results by BM25 text score relevance.
Requirements:
- Project
{ score: { $meta: "textScore" } }and sort by{ score: { $meta: "textScore" } }.
Answer
Implementation
db.articles.find(
{ $text: { $search: "mongodb index" } },
{ score: { $meta: "textScore" } }
)
.sort({ score: { $meta: "textScore" } });
Technical Explanation
$meta: "textScore"calculates keyword frequency relevance scores.- Sorting by text score ranks best matching documents at the top of results.
- Native search engine capabilities.
6. Related Terms
- Text Search (
$text/$search) — The query command. createIndex()/dropIndex()— The DDL triggers.- Atlas Search — Related concept: Atlas Search.
7. Key Takeaways
- Text Indexes tokenize, stem, and filter string values for full-text search.
- Only one text index is allowed per collection.
- Text indexes can cover multiple fields at the same time (composite text index).
- Set
weightsto prioritize relevance scores for specific fields (e.g. title). - Build a wildcard text index (
{ "$**": "text" }) to search all fields. - Writing to text indexes is CPU-heavy because strings must be parsed and stemmed.
- Drop old text indexes first before attempting to build new ones.