newsletter

How to Model Relationships in NoSQL Databases

10 min read

After you move from a SQL database to NoSQL, you usually forget about:

  • Joins
  • Foreign keys
  • Normalization and Denormalization
  • Database migrations

Yet, most developers have no idea how to model relationships in NoSQL correctly.

Most developers come to MongoDB (for example) after years of SQL. So they model the data the way SQL taught them.

Customers, Orders and Order Items - all go in multiple collections by habit. Every link becomes an ID, just like a foreign key.

Then they write the joins by hand in C#, because MongoDB has no JOIN keyword.

The result is a document database that behaves like a slow relational one. A single read needs four database round trips.

I've run MongoDB in production for years. Almost every performance problem I've seen there came from the wrong data model.

Today you will learn how to model relationships in a document database.

In this post, we will explore:

  • Why Relationships Work Differently in NoSQL
  • Embed or Reference: The Core Decision
  • One-to-One: Embed by Default
  • One-to-Many: The Three Sizes
    • One-to-Few: Embed the Array
    • One-to-Many: Reference from the Child
    • One-to-Zillions: Never Grow an Array Forever
  • Many-to-Many: Orders and Products
  • Reading Related Data Without Joins
  • Keeping Duplicated Data Correct
  • 4 Mistakes That Break NoSQL Data Models
  • The Same Rules in DynamoDB, Cassandra and Cosmos DB

All the code in this article I show using MongoDB and C#, but the rules apply to any NoSQL database.

Let's dive in.

Copied

Why Relationships Work Differently in NoSQL

In SQL databases, you have tables and rows. In NoSQL databases, you have collections and documents.

A document is one record. It's a JSON-like object that can contain other objects and arrays.

A collection is a group of documents. It's the closest thing MongoDB has to a table.

MongoDB has no foreign key constraints. Nothing stops you from deleting a customer who still has orders.

There are no cascade deletes either. Delete a parent, and the children stay behind as orphans.

And a normal query can't join. MongoDB can join inside the aggregation pipeline, but that isn't how you read data most of the time.

In SQL you design the tables first and write the queries later. You normalize the data, which means you store each fact once and link to it. The query planner works out the rest.

In a document database, you do the opposite. You shape the documents around the queries you already know you need.

Here's the same shipment in both worlds. In SQL it's three tables:

sql
shipments (id, number, order_id, status) shipment_address (shipment_id, street, city, zip) shipment_items (id, shipment_id, product, quantity)

Reading one shipment means joining all three.

In MongoDB, it's one document:

json
{ "_id": "8e3b1a6c-2f45-4d17-9b0a-51c9e2f7a410", "number": "12345678", "orderId": "ORD-9931", "status": "Dispatched", "address": { "street": "Main St 15", "city": "Warsaw", "zip": "00-001" }, "items": [ { "product": "Laptop", "quantity": 1 }, { "product": "Mouse", "quantity": 2 } ] }

One read returns the whole shipment, without joins.

Side by side, the two shapes look like this:

Loading diagram…
Copied

Embed or Reference: The Core Decision

In a NoSQL database, you have just two tools.

Embed means you put the related data inside the parent document. The address above is embedded in the shipment.

Reference means you store an id and keep the data in its own collection. That's the same idea as a foreign key, except the database doesn't enforce it.

Every relationship in your model is one of those two. Picking the right one is most of the work.

Here's the rule I use:

Embed whenReference when
You read both together almost every timeThe child is often read on its own
The list has a known upper limitThe list can grow infinitely
The child changes when the parent changesThe child is updated on its own schedule
The child belongs to one parentMany parents share the same child
The data is smallThe data is large

One hard limit settles a lot of arguments. A MongoDB document can't be bigger than 16 MB. If a list can outgrow that, it can't be embedded, and there's nothing left to discuss.

Speed pushes the other way. Embedded data comes back in the same read, so it costs nothing extra to fetch.

Note: A good model embeds some things and references others. Don't pick one style and apply it everywhere.

Now I will show you the three shapes a relationship can take.

Copied

One-to-One: Embed by Default

A shipment has one delivery address. That address belongs to one shipment.

Embed it:

json
{ "number": "12345678", "status": "Dispatched", "address": { "street": "Main St 15", "city": "Warsaw", "zip": "00-001" } }

In C#, that's a plain nested class:

csharp
public class Shipment { public Guid Id { get; set; } public required string Number { get; set; } public required string OrderId { get; set; } public required Address Address { get; set; } public required ShipmentStatus Status { get; set; } public required List<ShipmentItem> Items { get; set; } = []; } public class Address { public required string Street { get; set; } public required string City { get; set; } public required string Zip { get; set; } }

The address is part of the shipment, so it's saved and loaded with it.

If you split it into a second collection and every read costs a second round trip. You gain nothing, because no other query wants an address on its own.

Two cases still deserve a split.

The first is a large field you rarely read. A scanned delivery slip can be hundreds of kilobytes in size. Every read of the shipment would drag it along, even when nobody needs it.

Keep it in its own collection, or in blob storage with only the URL in the document.

The second is a field with different access rules. If tax numbers or personal data need separate permissions or their own retention policy, a separate collection makes that much easier to enforce.

One-to-one is the easy case. One-to-many is where many devs get models wrong.

Copied

One-to-Many: The Three Sizes

"One-to-many" hides three very different problems. The right answer depends on how big the "many" gets.

MongoDB's own guidance splits them into one-to-one, one-to-few, one-to-many, and one-to-zillions.

Copied

One-to-Few: Embed the Array

A shipment holds a handful of items. Ten, maybe twenty. It never holds 1,000 or 10,000.

You always show the items together with the shipment, so embed them:

csharp
public class ShipmentItem { public required string Product { get; set; } public required int Quantity { get; set; } }

Adding an item is one update, and MongoDB appends to the array in place:

csharp
var update = Builders<Shipment>.Update.Push(x => x.Items, newItem); await mongoDbContext.Shipments.UpdateOneAsync( x => x.Id == shipmentId, update, cancellationToken: cancellationToken);

Push maps to MongoDB's $push operator. The server appends to the array without first loading the document into your application. Two users adding items at the same time won't overwrite each other.

The items have no IDs of their own here, and they don't need any. Nothing else in the system points at a shipment item.

That's the test for one-to-few: the children are only interesting inside their parent.

Copied

One-to-Many: Reference from the Child

A customer has many orders. That number has grown for years and has no natural limit.

Orders also matter on their own. You can search them by date, status and number, for example.

So reference them. The child holds the parent's ID:

csharp
public class Order { public Guid Id { get; set; } public required string Number { get; set; } public required Guid CustomerId { get; set; } public required DateTime CreatedAtUtc { get; set; } public required List<OrderLine> Lines { get; set; } = []; }

Notice the direction. The id sits on the many side (Order). The customer document stays the same size no matter how many orders that customer places.

Reading a customer's orders is a normal query:

csharp
var orders = await mongoDbContext.Orders .Find(x => x.CustomerId == customerId) .SortByDescending(x => x.CreatedAtUtc) .Limit(50) .ToListAsync(cancellationToken);

This query needs an index, and this is the step teams skip most often. Without one, MongoDB reads every order in the collection to find the ones that match:

csharp
await mongoDbContext.Orders.Indexes.CreateOneAsync( new CreateIndexModel<Order>( Builders<Order>.IndexKeys .Ascending(x => x.CustomerId) .Descending(x => x.CreatedAtUtc), new CreateIndexOptions { Name = "IX_CustomerId_CreatedAtUtc" }), cancellationToken: cancellationToken);

The index covers the filter and the sort in one shot. A reference without an index is the NoSQL version of a missing foreign key index, and it hurts just as much.

Copied

One-to-Zillions: Never Grow an Array Forever

Now take tracking events: every scan, every status change, every carrier update. Each shipment collects a lot of them.

An embedded array, of course, is not an option (remember the 16 MB ceiling).

Put the events in their own collection and reference the shipment:

csharp
public class TrackingEvent { public Guid Id { get; set; } public required Guid ShipmentId { get; set; } public required string Location { get; set; } public required ShipmentStatus Status { get; set; } public required DateTime OccurredAtUtc { get; set; } }

The shipment stays small and fast. The events grow on their own, and you page through them when someone opens the tracking screen.

If the shipment screen already needs the last few events, you can keep a short copy on the shipment itself. Store the most recent 5 in an embedded array and drop the rest using the $slice operator. That's the subset pattern: the hot part lives with the parent, and the full history lives in its own collection.

There's also the Bucket Pattern, where you group many events into one document per day or per hour. It's worth knowing for time-series data, and it's overkill for most applications.

Also, if your collection grows too big (to zillions), you can use partitioning and sharding to split the data. Or you can use the Archive Data Pattern, and move old records into separate collections, to keep your main collection a bit smaller.

Copied

Many-to-Many: Orders and Products

An order contains many products. A product appears in many orders. In SQL you'd add a join table. In MongoDB you usually don't need one.

Store the link on the side you query from, and copy the few fields you display:

csharp
public class OrderLine { public required Guid ProductId { get; set; } public required string Name { get; set; } public required decimal Price { get; set; } public required int Quantity { get; set; } }

The order document now carries the product ID plus the name and the price:

json
{ "number": "ORD-9931", "customerId": "b1f0e4a2-77c3-4c81-9f2d-0d3a5c6b81e7", "lines": [ { "productId": "7a21c9d4-...", "name": "Laptop", "price": 1299.00, "quantity": 1 }, { "productId": "9c88f31b-...", "name": "Mouse", "price": 25.50, "quantity": 2 } ] }

This is the Extended reference pattern. You keep the ID so you can always load the full product, and you copy the few fields you need so the common read costs one query.

An order has to remember what the customer paid, not what the product costs today.

That's the general rule for copies. Ask whether the field is a snapshot or a live value. Prices, names on an invoice and the address at delivery time are snapshots. A product's current stock level should be read in real time.

You still need a real linking collection in two cases.

First: when you query the relationship from both directions and both sides are large. "Which orders contain this product?" Over millions of orders, each wants its own indexed collection. Second: when the link carries data that belongs to neither side. A discount that applied to this product in this order is stored on the link, not on the order or the product.

Now let's look at what happens when a copy isn't enough, and you really do need the other document.

Copied

You have three ways to read across documents. They're not equally good.

Two queries from your application. Load the orders, collect the ids you need, then load those in a second batched query:

csharp
var orders = await mongoDbContext.Orders .Find(x => x.CustomerId == customerId) .ToListAsync(cancellationToken); var productIds = orders.SelectMany(o => o.Lines) .Select(l => l.ProductId) .Distinct() .ToList(); var products = await mongoDbContext.Products .Find(x => productIds.Contains(x.Id)) .ToListAsync(cancellationToken);

Two round trips, plain C#, easy to read and easy to debug. This is the default, and it's the right answer far more often than people expect.

The trap is doing it in a loop. One query per order is the N+1 problem that leads to slow queries. Always batch the second query.

$lookup in the aggregation pipeline. The aggregation pipeline is a list of steps the server runs over a collection, and $lookup is the step that joins:

csharp
var ordersWithCustomers = await mongoDbContext.Orders .Aggregate() .Match(x => x.CustomerId == customerId) .Lookup<Order, Customer, OrderWithCustomer>( mongoDbContext.Customers, // the foreign collection order => order.CustomerId, // local field customer => customer.Id, // foreign field result => result.Customers) // where the matches land .ToListAsync(cancellationToken);

The driver also takes field names as strings, but I prefer the strongly typed lambdas.

The result type holds the joined documents in an array, because a lookup can match more than one:

csharp
public class OrderWithCustomer : Order { public List<Customer> Customers { get; set; } = []; }

$lookup is a left outer join done on the server, so a document with no match still comes back, just with an empty array.

It isn't free. $lookup runs once per input document, so it requires an index on the joined field, and it can become expensive as the input grows. Use it for reports and admin screens, not for your hottest read path.

If you're joining three collections on every page load, that tells you the model is wrong (not your query).

No second read at all. With the extended reference from the last section, the name and the price are already in the order—zero extra queries.

Copied

Keeping Duplicated Data Correct

Copying data is called denormalization: you store the same fact in multiple places to make reads faster.

It works. It also means a product rename now has to touch two collections.

You have three ways to handle that, and you pick per field.

Accept the staleness. Ask what a wrong answer costs. A slightly old product name on an old order costs nothing, and most copies are like this. Snapshots such as the price at purchase time aren't stale at all, because they were never meant to change.

Update both places in one transaction. MongoDB has supported multi-document ACID transactions since version 4.0 on replica sets, and since 4.2 on sharded clusters.

ACID here means both writes land or neither does (like in an SQL database):

csharp
var orderFilter = Builders<Order>.Filter.ElemMatch(x => x.Lines, l => l.ProductId == productId); using var session = await mongoClient.StartSessionAsync(cancellationToken: cancellationToken); await session.WithTransactionAsync(async (s, ct) => { await mongoDbContext.Products.UpdateOneAsync(s, x => x.Id == productId, Builders<Product>.Update.Set(x => x.Name, newName), cancellationToken: ct); await mongoDbContext.Orders.UpdateManyAsync(s, orderFilter, Builders<Order>.Update.Set("lines.$.name", newName), cancellationToken: ct); return true; }, cancellationToken: cancellationToken);

Use WithTransactionAsync rather than calling StartTransaction and CommitTransaction yourself. It replays the callback on a transient error and retries the commit on an unknown result, both of which are easy to get wrong by hand.

Two rules matter here. Every call inside the transaction has to get the session s, or that write simply doesn't take part in it. And side effects such as cache eviction, logging and outgoing API calls belong outside the transaction, because the driver can replay the callback.

Note: lines.$.name is the positional operator. It updates the first array element that the filter matched, which is what you want when a product appears once per order. If the same product can sit on several lines, use lines.$[line].name with UpdateOptions.ArrayFilters instead.

Transactions need a replica set. A standalone mongod throws an error, so local setups usually run a single-node replica set.

Propagate the change in the background. For a big fan-out, updating thousands of orders inside one transaction is the wrong shape. Write to the product first, then push the change out with a change stream or an outbox. A change stream is a live feed of writes to a collection, and your handler reacts to each one.

Two rules keep this sane, and they're the same ones I use for any system with copies.

Keep one system of record per fact. Every piece of data has exactly one authoritative home, and every copy is a projection you can rebuild from it.

And never dual-write. If a single request writes to two collections without a transaction or outbox, one of those writes will eventually fail, and the two copies will drift apart in silence.

Copied

4 Mistakes That Break NoSQL Data Models

These are the 4 most common mistakes I've seen developers make.

1. Embedded unbounded arrays. An array that grows with traffic will hit the 16 MB limit sooner or later, and it slows writes long before that. Comments, events, logs and audit trails all belong in their own collection. Before you embed a list, ask what its largest possible size is. If you don't know the number, don't embed it.

2. Modeling by entity instead of by query. Drawing the entities first and then wondering how to read them is a classic SQL habit that hurts most here. Write down your top five queries first, then design documents that answer them in one read.

3. Referencing everything, then joining in C#. This is the SQL model over a document database, and it gives you the worst of both. You get no joins from the engine, no constraints, and no single-read advantage either. If two things are always read together and the child is small and bounded, embed it.

4. Updating copies with two independent writes. Once you denormalize, a rename becomes two writes. Without a transaction or an outbox, a crash between them leaves the data wrong, and nothing will tell you. Decide up front which copies are snapshots, which can be stale, and which have to stay in sync.

Here is the whole choice as a flowchart:

Loading diagram…
Copied

The Same Rules in DynamoDB, Cassandra and Cosmos DB

MongoDB isn't the only NoSQL database, and the same thinking carries over.

DynamoDB pushes it furthest with single-table design, where several entity types share one table and the keys carry the relationships. There's no join at all, so every access pattern has to be planned before you create the table.

Cassandra treats query-first modeling as a rule. You create one table per query and write the same data into several of them. Duplication is the normal design there.

Cosmos DB gives you a document model close to MongoDB's, plus the partition key. A partition key is the field that decides which physical partition a document lands on. Related data that shares a partition key is cheap to read and write together.

For a wider view of which database engine fits which workload, see How Architects Choose a Database in 2026.

Copied

Summary

Let's recap the key takeaways:

  • You have two tools: embed and reference. Embed when the data is read together, stays small and has a known limit. Reference when the child is read on its own, is shared, or can grow infinitely.
  • Model by query, not by entity. Write down the queries your screens need, then design documents that answer them in one read.
  • Size decides one-to-many. A few children get embedded. Many get referenced from the child side. An unbounded list always gets its own collection.
  • Copy the fields you display, and keep the id. The extended reference removes a need for the second query, and a copied price is a snapshot you want.
  • Index every reference. A reference with no index on the child's parent id makes MongoDB scan the whole collection.
  • Pay for denormalization on purpose. Accept staleness where it's cheap, use a transaction where it isn't, and use a change stream or an outbox for large fan-outs. Never dual-write.

There is no single correct way to model relationships; you need to pick whatever works best in each particular project and case. What you can't do is skip the decision, because the AI won't make it for you.

Hope you find this newsletter useful. See you next time.

Whenever you're ready, here's how I can help you:

The .NET Senior Playbook is built to:

  • Fast-track you from junior or mid-level to senior
  • Keep you growing as a senior
  • Help you beat any .NET interview

Covers everything: C#, ASP.NET Core, EF Core, system design — answer each question first, reveal the solution, and a test after every chapter proves it stuck. Finish, and you earn a verifiable certificate for your LinkedIn.

The .NET Senior Playbook
Join 500+ developers

Not sure where you stand? Take the free .NET interview test:

  • Find out your real level — Junior to Senior+
  • A realistic mock .NET interview — across 13 areas of C#, .NET, ASP.NET Core and System Design

No credit card required. When you finish, you get a personalized report: your level, your strongest and weakest areas, and where to focus next — the perfect way to benchmark yourself before diving into the Playbook.

Start the free test

Enjoyed this article? Share it with your network

Improve Your .NET and Architecture Skills

Join my community of 28,000+ developers and architects.

Each week you will get 1 practical tip with best practices and real-world examples.

Learn how to craft better software with source code available for my newsletter.

Join 28,000+ developers already reading
No spam. Unsubscribe any time.