Tech Behind ThingsHow the ordinary machinery actually works

Software

Why Copying Many Small Files Takes Longer Than One Big One

Ten thousand tiny files move far slower than a single large file of the same size, because most of the work is bookkeeping the filesystem must repeat for every file.

Vibrant and engaging code displayed on a computer screen, showcasing programming concepts.
Photograph by Seraphfim Gallery via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Copying a single large file often runs at the full speed of the drive, while copying a folder of the same total size can take many times longer. The difference is per-file overhead, and it dwarfs the data itself.

Every file carries a fixed cost

Creating a file means allocating an entry in a directory, reserving space, writing ownership and timestamps, and updating the structures that track which blocks are free.

That work is roughly the same whether the file holds four bytes or four gigabytes. It is a fixed toll paid once per file rather than per byte.

With one large file the toll is paid once and the rest is bulk transfer. With ten thousand files it is paid ten thousand times, and it becomes the entire job.

Small transfers cannot fill the pipe

Storage devices reach their advertised speed on long sequential transfers, where a request is issued once and satisfied continuously without further decisions.

A tiny file produces a short request, and between requests the drive waits while the operating system looks up the next name and location. Flash storage loses less to this than a spinning disk, but it still loses.

Queue depth helps, which is why tools that keep many requests in flight at once finish sooner than ones that copy strictly in order.

Metadata writes have to be ordered

A filesystem cannot allow a crash to leave a directory pointing at space that was never allocated, so it enforces an order between certain writes.

Ordering means waiting. Some updates cannot begin until earlier ones are durably recorded, and those pauses are counted per file rather than per byte.

Journaling limits the damage a crash can do, and it adds another write to the sequence for every operation, which is a trade the designers made deliberately.

Networks and scanning add their own per-file tolls

A network copy negotiates each file separately, and every negotiation costs at least one round trip. With a distant server, latency alone can dominate the transfer time.

Security software typically inspects files as they are written, and inspection is per-file work involving reading the content back and evaluating it against rules.

These layers explain why the same folder copies at very different speeds on two machines with identical hardware and identical drives.

Archiving converts the problem

Packing a folder into a single archive turns thousands of small operations into one large sequential write, which is why transferring an archive and unpacking it can beat copying the folder directly.

The unpacking still pays the per-file cost, but it pays it locally on fast storage rather than across a network, and the ordering constraints are cheaper to satisfy there.

The same reasoning explains why estimates are so poor on mixed folders. The remaining bytes say little about the remaining time when file count is what actually governs it.

Questions readers ask

Why does a copied folder show a different size?

Block allocation, compression and metadata differ between filesystems. The contents are identical while the space consumed is not.

Is defragmenting a solid state drive useful?

No. There is no seek penalty to remove, and rewriting every block consumes write endurance for no measurable benefit.

Softwarestoragesoftwareoperating systemsdata
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida