Skip to content
datastudyguides

DP-900 · 1 Describe core data concepts

Data file formats

Exam objective: Describe common formats for data files

Different file formats trade off human readability against storage and processing efficiency. Pick the format by how the data will be read and written.

The format you choose for a data file depends on who or what has to read it. A format can favor being readable by a person, or being compact and fast for a machine to process. Few formats do both well.

Delimited text, most often comma separated values (CSV), is plain text with a fixed separator between fields and a line break between rows. It is easy to read and almost every tool can open it.

JSON represents each record as an object made of named attributes, and an attribute can itself be another object or a list. That nesting lets JSON describe structured records and semi-structured ones, where fields vary between records.

Format Organized by Good for
CSV plain text rows human readability, wide support
Parquet column fast analytical reads over large files
Avro row compact storage, schema travels with the data

Delta Lake builds on Parquet by adding a transaction log, so updates to files in a data lake can be versioned and ACID compliant.

On the exam, look for columnar versus row based, or a scenario that names CSV, JSON, Parquet, Avro or Delta Lake directly.

Key points

  • Delimited text such as CSV separates fields with commas and rows with a line break. An optional first line can name the fields.
  • JSON nests objects and lists inside each other, so it can hold both structured and semi-structured data.
  • Parquet is columnar: it groups each row group's data by column, with metadata that lets a query jump straight to the columns and rows it needs.
  • Avro is row based. Each file starts with a JSON header that describes the structure, followed by binary records.
  • Delta Lake adds a transaction log on top of Parquet files, which brings versioning and reliable updates to a data lake.

Exam trap

Parquet and Avro get swapped on the exam. Parquet stores data by column, which suits analytical reads of a few columns across many rows. Avro stores data by row, which suits fast writes and carrying the schema with the file.

Check yourself

Which file format organizes a file into row groups where each row group's data is stored together by column, with metadata that lets a query jump straight to the columns and rows it needs?

Go deeper on Microsoft Learn

Checked against Microsoft Learn on October 1, 2026.

How well do you know this?