Tech Behind ThingsHow the ordinary machinery actually works

Software

Why Searching Your Own Device Is Instant But Incomplete

Local search returns results before you finish typing because it queries an index built in advance, and everything it misses was excluded when that index was made.

Vibrant and engaging code displayed on a computer screen, showcasing programming concepts.
Photograph by Seraphfim Gallery via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Searching a device with hundreds of thousands of files returns results immediately, which is impossible if the files are being read at that moment. They are not.

The index is built long before the query

A background service walks the storage, opens files it understands, extracts their text and records which words appear where.

The result is an inverted structure: for each word, a list of the documents containing it. Answering a query means intersecting a few such lists.

That operation is fast regardless of how much data exists, because the cost depends on how many documents contain the search terms rather than on the total.

Freshness depends on change notifications

Rescanning everything constantly would be prohibitive, so the indexer subscribes to filesystem events and updates only what changed.

Where those notifications are unreliable, such as on network shares or external drives, the index drifts out of date and results become quietly wrong.

This is why a file created seconds ago sometimes cannot be found, and why forcing a rebuild resolves problems that look like corruption.

The indexer must understand the format

Extracting text from a document requires knowing its structure. Each file type needs a handler, and formats without one are indexed by filename only.

Encrypted containers, compressed archives and proprietary databases are commonly skipped, since their contents are unreadable or expensive to enumerate.

Images and scanned documents contain no text at all unless recognition has been run over them, which is why a scanned letter is invisible to search.

Exclusions are broader than most people realise

System directories, temporary folders and caches are excluded deliberately to keep the index small and to avoid churning on files that change constantly.

Removable and network locations are often excluded by default, because indexing them would depend on their being connected.

The consequence is that absence of a result proves nothing. It may mean the file does not exist, or that it sits somewhere the indexer was told to ignore.

Ranking decides what you actually see

A query commonly matches thousands of items, so ordering matters more than matching. Recency, file type, location and past selections all contribute weight.

Because ranking is tuned for the common case, an unusual query can push the correct answer below the visible results while technically finding it.

Narrowing by type or date is therefore more effective than adding words, since it removes candidates rather than asking the ranker to reconsider them.

Questions readers ask

Why does a copied folder show a different size?

Block allocation, compression and metadata differ between filesystems. The contents are identical while the space consumed is not.

Is defragmenting a solid state drive useful?

No. There is no seek penalty to remove, and rewriting every block consumes write endurance for no measurable benefit.

Softwarestoragesoftwareoperating systemsdata
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida