DP-900 · Study guide
Every objective, one card each
30 cards, about 30 minutes of reading. Each card explains one objective from the official outline of July 21, 2026, names the trap, and ends with a question.
1 Describe core data concepts
25-30%Describe ways to represent data
- Structured data Structured data follows a fixed schema, so every record has the same fields. It is usually shown as tables of rows and columns.
- Semi-structured data Semi-structured data has some structure, but the fields can differ from one record to the next. JSON is the standard example.
- Unstructured data Unstructured data has no specific structure. Think of documents, images, audio, video and other binary files.
Identify options for data storage
- Data file formats Different file formats trade off human readability against storage and processing efficiency. Pick the format by how the data will be read and written.
- Common data stores A database is a dedicated system for storing and querying records. Relational databases use tables and keys, while nonrelational databases skip the fixed schema.
- Azure datastore options Azure offers a datastore for almost every shape of data and workload. Match the use case to the service rather than memorizing every feature.
Describe common data workloads
- Transactional workloads Transactional processing, or OLTP, records individual business events quickly and reliably. ACID properties guarantee that each transaction is trustworthy.
- Analytical workloads Analytical processing, or OLAP, reads large volumes of historical data to support reporting. Data typically moves through a data lake, a warehouse and an aggregated model.
Identify roles and responsibilities for data workloads
- Database administrator role A database administrator keeps databases available, secure and recoverable. The role covers operations and protection, not building the data pipelines themselves.
- Data engineer role A data engineer builds and monitors the pipelines that integrate and prepare data across an organization, working across many kinds of data stores.
- Data analyst role A data analyst explores data to find trends, then turns them into reports and visualizations that help an organization decide and act.
2 Identify considerations for relational data on Azure
20-25%Describe relational concepts
- Relational data model A relational database models real world entities as tables, where each row is one instance and each column holds one attribute with its own data type.
- Normalization Normalization is a schema design process that splits data into tables by entity and attribute, and links them with primary and foreign keys, to remove duplicate data.
- SQL statement types SQL statements fall into three groups, defining database objects, controlling access to them, and reading or changing the rows inside them.
- Database objects Beyond tables, a relational database can hold views that present a query as a virtual table, stored procedures that package reusable logic, and indexes that speed up lookups.
Describe relational Azure data services
- Azure SQL family Azure SQL is a family of three SQL Server based services, SQL Server on Azure VMs, Azure SQL Managed Instance and Azure SQL Database, that trade administrative control for less management work.
- Azure open-source databases Azure Database for MySQL and Azure Database for PostgreSQL are managed PaaS versions of these open source engines, each with a Flexible Server option for more control.
3 Describe considerations for working with non-relational data on Azure
15-20%Describe the capabilities of Azure storage
- Azure Blob storage Azure Blob storage holds unstructured data as blobs inside containers, with access tiers that balance storage cost against retrieval speed.
- Azure Files Azure Files creates cloud based network file shares that multiple users or apps can mount over SMB or NFS, in place of an on-premises file server.
- Azure Table storage Azure Table storage is a NoSQL key value store where every row has a partition key and a row key, but rows in the same table can hold different columns.
Describe the capabilities and features of Azure Cosmos DB
- Cosmos DB use cases Azure Cosmos DB is a fully managed, schema agnostic NoSQL database built for globally distributed apps that need low latency reads and writes.
- Cosmos DB APIs Azure Cosmos DB offers five APIs, NoSQL, MongoDB, Table, Cassandra and Gremlin, so an application can keep a familiar query language while gaining Cosmos DB scale.
4 Describe an analytics workload
25-30%Describe common elements of large-scale analytics
- Data ingestion and processing Ingestion moves data from sources into an analytical store, reshaping it along the way. ETL transforms before loading, ELT transforms after.
- Analytical data stores A data warehouse is a relational store with a fixed, analytics-ready schema. A data lake is file storage that applies schema on read.
- Fabric and Azure Databricks Microsoft Fabric is a unified SaaS workspace built around one shared lake, OneLake. Azure Databricks is a Spark platform that runs inside your own subscription.
Describe considerations for real-time data analytics
- Batch versus streaming data Batch processing groups records and processes them together. Stream processing handles each event as it arrives, with much lower latency.
- Real-time analytics services Microsoft covers streaming with several services: Fabric Real-Time Intelligence, Spark Structured Streaming, and the standalone Azure Stream Analytics.
Describe data visualization in Microsoft Power BI
- Power BI capabilities Power BI reports are built in Power BI Desktop, published to the Power BI service, and consumed through a browser or the phone app.
- Power BI data models A semantic model defines measures, the numbers you analyze, and dimensions, the entities you group them by, linked in a star or snowflake schema.
- Data visualization types Picking a visualization means matching the chart to the question being asked, such as a trend, a proportion, or a relationship between measures.