File Format¶
How bytes are physically laid out inside a file — row-major vs. column-major, schema embedding, compression — determines whether reading "one column out of fifty" costs almost nothing or costs a full-file scan. Parquet, Avro, and ORC each made different trade-offs here on purpose.
flowchart LR
Junior["Junior: row-oriented vs. column-oriented file layout"] --> Middle["Middle: Parquet's structure - row groups, column chunks, footer"]
Middle --> Senior["Senior: schema evolution and compression codec trade-offs"]
Senior --> Professional["Professional: format internals at scale - predicate pushdown and dictionary encoding"]
flowchart LR
subgraph RowMajor["Avro (row-major)"]
R1["Row 1: all columns"] --> R2["Row 2: all columns"]
end
subgraph ColMajor["Parquet (column-major)"]
C1["Column A: all values"] --> C2["Column B: all values"]
end
Choose a level¶
| Level | Guide | You are done when |
|---|---|---|
| Junior | Row-major vs. column-major layout | You can explain why reading one column is cheap in Parquet but not in Avro. |
| Middle | Parquet's internal structure | You can describe row groups, column chunks, and the footer's role. |
| Senior | Schema evolution and compression | You can explain how Avro handles schema evolution differently from Parquet, and why codec choice matters. |
| Professional | Predicate pushdown and encoding at scale | You can explain how min/max statistics and dictionary encoding let engines skip data without decompressing it. |
Practice rule¶
Before picking a file format for a new pipeline, ask: "will this data mostly be read column-by-column for analytics (favor Parquet/ORC), or written and read as whole records for streaming/RPC (favor Avro)?" The answer should drive the format choice, not familiarity or default tooling.