Working with Data Modeling

1. Understanding Document Model

PrincipleDetail
Schema-flexibleNo fixed schema; validators optional
Aggregate-orientedDesign around access patterns
Joins minimizedEmbed related data when read together
16MB capPer document; design with this in mind

2. Modeling One-to-One Relationships

ApproachWhen
EmbedAlways accessed together
ReferenceLarge, optional, or rarely needed
// Embedded
{ _id: 1, name: "Alice", profile: { bio: "...", avatar: "..." } }

3. Modeling One-to-Many Relationships (embedded)

ApproachWhen
Embed arrayBounded (< few hundred), fetched with parent
Subset patternEmbed top N; full set referenced

4. Modeling One-to-Many Relationships

ApproachWhen
Child stores parent refUnbounded / large many-side
Parent stores ref arrayBounded many; common query lists ids

5. Modeling Many-to-Many Relationships

ApproachTrade-off
Two-way refsRead flexibility; write to two places
Join collectionLike SQL; use $lookup
Embed on dominant sideRead pattern dictates

6. Using Embedded Documents

ProsCons
Single readDoc size grows
Atomic updatesCannot independently shard
No joinsDuplication if shared

7. Using References with $lookup

AspectDetail
Form{$lookup: {from, localField, foreignField, as}}
Index foreignFieldCritical for performance
ShardedSupported on sharded foreign in 5.1+

8. Choosing Embedding vs Referencing

QuestionEmbedReference
Always fetched together?YesNo
Bounded size?YesUnbounded
Atomic updates needed?YesNo
Shared across parents?NoYes
High write contention?NoYes

9. Handling Large Documents

StrategyDetail
Bucket patternGroup time-series readings
Subset patternEmbed hot subset; reference rest
GridFSFor files > 16MB
External storageS3 / blob storage + reference URL

10. Understanding Schema Flexibility

AspectDetail
Add fieldNo migration; backfill optional
Remove field$unset over time
Schema versionAdd schemaVersion field
ValidationJSON Schema validator for guardrails