Fan-Out / Fan-In¶
Split one stream of work across multiple concurrent workers (fan-out), then merge their results back into one stream (fan-in). The in-process ancestor of the scatter-gather pattern and Spark's own parallel-stage execution model.
flowchart LR
Junior["Junior: splitting work across workers"] --> Middle["Middle: merging results back correctly"] --> Senior["Senior: handling partial failures across fanned-out workers"]
Senior --> Professional["Professional: fan-out/fan-in at scale - bounding concurrency and result ordering"]
flowchart LR
Input["Input stream"] --> FanOut{Fan-out}
FanOut --> W1[Worker 1]
FanOut --> W2[Worker 2]
FanOut --> W3[Worker 3]
W1 & W2 & W3 --> FanIn{Fan-in}
FanIn --> Output["Merged output"]
Choose a level¶
| Level | Guide | You are done when |
|---|---|---|
| Junior | Splitting work | You can explain why fanning out work to independent workers gives you parallelism. |
| Middle | Merging results correctly | You can implement fan-in without losing or reordering results incorrectly. |
| Senior | Partial failures | You can design a strategy for what happens when one of N fanned-out workers fails. |
| Professional | Bounding concurrency at scale | You can design fan-out with a concurrency limit and deterministic result ordering. |
Practice rule¶
Before fanning out work to N concurrent workers, ask: "if worker 3 out of 10 fails, what should happen to the other 9's results, and to the overall operation?" This decision (fail-fast vs. partial-success) should be made explicitly, not left as an accident of implementation.