Skip to content

File Format

How bytes are physically laid out inside a file — row-major vs. column-major, schema embedding, compression — determines whether reading "one column out of fifty" costs almost nothing or costs a full-file scan. Parquet, Avro, and ORC each made different trade-offs here on purpose.

flowchart LR Junior["Junior: row-oriented vs. column-oriented file layout"] --> Middle["Middle: Parquet's structure - row groups, column chunks, footer"] Middle --> Senior["Senior: schema evolution and compression codec trade-offs"] Senior --> Professional["Professional: format internals at scale - predicate pushdown and dictionary encoding"]
flowchart LR subgraph RowMajor["Avro (row-major)"] R1["Row 1: all columns"] --> R2["Row 2: all columns"] end subgraph ColMajor["Parquet (column-major)"] C1["Column A: all values"] --> C2["Column B: all values"] end

Choose a level

Level Guide You are done when
Junior Row-major vs. column-major layout You can explain why reading one column is cheap in Parquet but not in Avro.
Middle Parquet's internal structure You can describe row groups, column chunks, and the footer's role.
Senior Schema evolution and compression You can explain how Avro handles schema evolution differently from Parquet, and why codec choice matters.
Professional Predicate pushdown and encoding at scale You can explain how min/max statistics and dictionary encoding let engines skip data without decompressing it.

Practice rule

Before picking a file format for a new pipeline, ask: "will this data mostly be read column-by-column for analytics (favor Parquet/ORC), or written and read as whole records for streaming/RPC (favor Avro)?" The answer should drive the format choice, not familiarity or default tooling.