Tech Behind ThingsHow the ordinary machinery actually works

Software

Adding cores stops helping at the point where the work refuses to divide

Every program contains a portion that must happen in order, and that portion sets a ceiling no amount of hardware can lift.

A female engineer works on code in a contemporary office setting, showcasing software development.
Photograph by ThisIsEngineering via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

This works through parallel processing in the order the parts actually depend on each other.

The short version

  • The part that cannot be divided sets the maximum possible speedup.
  • Coordination between workers is itself work that grows with worker count.
  • Memory bandwidth often limits scaling before the cores do.

The serial fraction sets a hard ceiling

Almost every task contains steps that must happen in a fixed order, because each one depends on the result of the previous one. Those steps cannot be spread across cores, so their duration remains the same however much hardware is applied to the problem.

As the parallel portion is divided among more workers, the serial portion becomes a larger share of the total time. Eventually the serial part dominates and adding workers changes the answer by an amount too small to notice. This is why doubling core count rarely halves the time for anything except workloads that were designed to divide cleanly.

Coordination is work that grows

Splitting a task requires deciding how to split it, distributing the pieces and combining the results afterwards. That overhead grows with the number of workers, so beyond some point each additional worker adds more cost than capability.

Under load, small tasks suffer most, because the effort of dispatching work to another core can exceed the effort of simply doing it. Runtimes therefore batch small items into larger chunks, trading some load imbalance for far less coordination overhead. Choosing the chunk size is a genuine engineering decision, and getting it wrong makes parallel code slower than serial code.

Shared data is where it goes wrong

Whenever two workers can modify the same data, something must ensure they do not interleave in a way that corrupts it. Locks do that by allowing one worker through at a time, which converts a parallel section back into a serial one. A lock held for a long time or acquired very frequently becomes the bottleneck, and adding cores makes the contention worse.

Lock-free structures avoid explicit waiting but are extremely difficult to write correctly and are easy to get subtly wrong. Designs that give each worker its own data and combine only at the end avoid the problem rather than managing it.

Memory is the wall people forget

Cores share a path to main memory, and that path has a fixed capacity regardless of how many cores are demanding data. A workload that streams large amounts of data saturates that path quickly, after which idle cores wait rather than compute.

Caches hide the problem for data that fits, which is why performance can fall off a cliff as a dataset grows past a threshold. Cores writing to nearby memory locations also interfere, because cache coherence moves whole lines between cores repeatedly.

That effect can make a parallel version slower than the original despite no visible sharing in the source code at all.

Some work divides beautifully

Tasks where each item is independent, such as processing many images, scale almost linearly until some other resource runs out. Graphics workloads are the extreme case, which is why processors dedicated to them contain thousands of very simple cores.

Those cores work best when every one of them performs the same operation on different data at the same moment. Any branching that sends different cores down different paths wastes much of that hardware while the paths are executed in turn. This is why some algorithms are rewritten in awkward branch-free forms specifically to suit that style of hardware.

What it means for your device

Everyday use consists largely of short bursts of work that cannot be divided, so many cores frequently sit idle. Devices therefore mix a few fast cores with several efficient ones, running background work on the efficient ones to save power. The scheduler decides which core runs what, and a misplaced task can make a fast device feel unexpectedly sluggish.

At the protocol level, extra cores genuinely help when several demanding programs run at once, which is a different benefit from making one program faster. A core count is therefore a poor predictor of how quickly any particular task will finish on any particular machine.

The takeaway

Parallelism multiplies the divisible part and leaves the rest exactly as it was.

The constraint is almost always physical, and marketing rarely mentions which one.

Questions readers ask

Why does my machine feel slow with cores sitting idle?

The work in front of it is serial, waiting on storage or waiting on a network. Idle cores cannot help with any of those.

Do more cores help with games?

Up to a point. Rendering and simulation divide reasonably well, but a serial main loop usually limits how much extra cores contribute.

Softwareprocessorsperformancesoftwareconcurrency
Junko Ishida
Contributing writer, Tech Behind Things

Junko covers batteries, charging and energy density, and is unimpressed by most battery claims.

Also by Junko Ishida