DP-900 · 1 Describe core data concepts
Data file formats
Exam objective: Describe common formats for data files
Different file formats trade off human readability against storage and processing efficiency. Pick the format by how the data will be read and written.
The format you choose for a data file depends on who or what has to read it. A format can favor being readable by a person, or being compact and fast for a machine to process. Few formats do both well.
Delimited text, most often comma separated values (CSV), is plain text with a fixed separator between fields and a line break between rows. It is easy to read and almost every tool can open it.
JSON represents each record as an object made of named attributes, and an attribute can itself be another object or a list. That nesting lets JSON describe structured records and semi-structured ones, where fields vary between records.
| Format | Organized by | Good for |
|---|---|---|
| CSV | plain text rows | human readability, wide support |
| Parquet | column | fast analytical reads over large files |
| Avro | row | compact storage, schema travels with the data |
Delta Lake builds on Parquet by adding a transaction log, so updates to files in a data lake can be versioned and ACID compliant.
On the exam, look for columnar versus row based, or a scenario that names CSV, JSON, Parquet, Avro or Delta Lake directly.
Key points
- Delimited text such as CSV separates fields with commas and rows with a line break. An optional first line can name the fields.
- JSON nests objects and lists inside each other, so it can hold both structured and semi-structured data.
- Parquet is columnar: it groups each row group's data by column, with metadata that lets a query jump straight to the columns and rows it needs.
- Avro is row based. Each file starts with a JSON header that describes the structure, followed by binary records.
- Delta Lake adds a transaction log on top of Parquet files, which brings versioning and reliable updates to a data lake.
Exam trap
Parquet and Avro get swapped on the exam. Parquet stores data by column, which suits analytical reads of a few columns across many rows. Avro stores data by row, which suits fast writes and carrying the schema with the file.
Check yourself
Which file format organizes a file into row groups where each row group's data is stored together by column, with metadata that lets a query jump straight to the columns and rows it needs?
Go deeper on Microsoft Learn
Checked against Microsoft Learn on October 1, 2026.