The Cosmos DB default indexing policy indexes every path of every document. For an event store
that is the wrong default: the data property holds a serialised domain event or snapshot, it is
the largest property in the document, and no query Memoria issues can filter on it — CONTAINS
never uses the index. Every write pays to index it anyway.
CosmosSetup.CreateDatabaseAndContainerIfNotExist already creates containers with the better
policy, so if Memoria provisions your container there is nothing to do here. This guide is for
containers provisioned elsewhere — infrastructure as code, a portal, a DBA — and for containers that
already exist, which keep whatever policy they were created with.
All Memoria containers use /streamId as the partition key, and every read passes the partition key
in the request options, so each query is scoped to one logical partition. Within that partition the
store filters and sorts on a small, fixed set of paths:
| Path | Used by |
|---|---|
documentType |
every query — the container mixes events, aggregates, and projections |
sequence |
event range reads, ORDER BY, and SELECT VALUE MAX(c.sequence) on every save |
createdDate |
the date-bounded event reads (GetEventsUpToDate, FromDate, BetweenDates) |
eventType |
the EventTypeFilter on aggregates and projections |
streamId |
the partition predicate in each WHERE clause, and the MAX(sequence) aggregate |
Do not remove
/streamId. Every query already scopes itself with a partition key, so indexing the partition key path looks redundant. Measured, excluding it takes 4.9% off writes but costs 119% more onSELECT VALUE MAX(c.sequence)— the concurrency check that runs on everySaveAggregateandSaveEvents. That is +4.43 RU per save against a saving of about 0.38 RU per event written, so it loses for anything but enormous batches.
Nothing else is queried. data, version, latestEventSequence, aggregateType,
projectionType, createdBy, updatedBy, and updatedDate are read from the returned documents,
never filtered or sorted on.
The policy lives at
scripts/install/1.7.0-cosmos-indexing-policy.json.
It excludes /* and includes only the paths in the table above. It defines no composite indexes —
see why there are no composite indexes below, which is a measurement, not
an oversight.
id is not listed: Cosmos DB always indexes it and rejects a policy that tries to override it.
Match the policy to your Memoria version. Until 1.7.0 the container also held aggregate-event link documents, and the policy indexed
/aggregateId/?and/appliedDate/?to serve the one query that read them. 1.7.0 writes no link documents, so this policy drops both paths — applying it to an older deployment would leave that query scanning. On 1.5.0 or 1.6.0, use1.6.0-cosmos-indexing-policy.jsoninstead; it remains correct for those versions.Applying the 1.7.0 policy after upgrading is safe and worthwhile: the two dropped paths cost write RU on every document and now index nothing.
For an Azure account, run either script — they do the same thing:
./scripts/install/1.7.0-cosmos-apply-indexing-policy.ps1 `
-ResourceGroup rg-shop -Account cosmos-shop -Wait
./scripts/install/1.7.0-cosmos-apply-indexing-policy.sh \
--resource-group rg-shop --account cosmos-shop --wait
Both default to database Memoria and container Domain, matching CosmosOptions. Pass
-Database/--database and -Container/--container if you changed them. Both require the
Azure CLI and an az login that can
write to the account.
Applying the same policy twice is a no-op, so the scripts are safe in a deployment pipeline.
If you provision infrastructure declaratively, take the JSON straight into your template instead —
Microsoft.DocumentDB/databaseAccounts/sqlDatabases/containers accepts it verbatim under
properties.resource.indexingPolicy in Bicep and ARM, as does indexing_policy in Terraform’s
azurerm_cosmosdb_sql_container.
The Azure CLI cannot reach the Cosmos DB emulator, so use the API instead:
await cosmosSetup.ReplaceIndexingPolicy(CosmosIndexingPolicy.CreateRecommended());
That works against any account, not just the emulator. You can also paste the JSON into Data Explorer → your container → Settings → Indexing Policy → Save, or delete and recreate the container.
Cosmos DB reindexes in the background. The container stays online and writes keep succeeding, but
queries can return incomplete results until the transformation finishes, so apply it during a
quiet period on a container that already holds data. Both scripts accept --wait / -Wait to poll
indexTransformationProgress until it reaches 100%.
There is no rollback script. To go back, apply the Cosmos DB default policy:
{
"indexingMode": "consistent",
"automatic": true,
"includedPaths": [{ "path": "/*" }],
"excludedPaths": [{ "path": "/\"_etag\"/?" }]
}
Because the policy excludes /*, a query of your own that filters on an unlisted path — say
c.aggregateType or c.updatedBy — will fall back to a scan of the partition rather than fail.
That is correct but expensive. Add the path to includedPaths before relying on it:
{ "path": "/aggregateType/?" }
Keep /data excluded regardless. It is a serialised string, and CONTAINS(c.data, ...) — which is
how eventPropertyFilter is translated — cannot use an index in any case, so indexing it buys
nothing and costs write RU on every event.
An earlier draft of this policy defined three, on (documentType, sequence),
(documentType, createdDate) and — while aggregate-event links still existed —
(documentType, aggregateId, appliedDate), reasoning that each served a filter-and-order-by read the
store issues. Measured, they cost more than they returned. The third became moot in 1.7.0 when the
link documents went; the measurement below stands as taken, on the schema of the day.
Against the Cosmos DB emulator: 200 event documents of roughly 600 bytes written in two batches, then each read issued once. Request charge, against the default index-everything policy:
| Write 200 events | Read whole stream | MAX(sequence) |
Sequence range | Date range | Type filter | |
|---|---|---|---|---|---|---|
| Default policy | 1600.00 | 10.54 | 3.55 | 6.89 | 11.11 | 11.37 |
| Exclusions only (shipped) | 1561.90 | 9.94 | 3.71 | 6.69 | 10.55 | 10.77 |
| Composites only | 1714.28 | 10.54 | 3.63 | 6.69 | 10.75 | 10.87 |
| Exclusions + composites | 1676.20 | 9.94 | 3.55 | 6.59 | 10.55 | 10.77 |
The composite indexes add about 7% to every write and return essentially nothing: the reads are
as cheap with exclusions alone. Every query here is single-partition with an equality filter on
documentType and an ORDER BY c.sequence, and within one partition the range index on
/sequence already serves that ordering — so the composite is maintained on every write and then
not used.
Two caveats on those numbers. They come from the emulator, not a real account, and from one payload
shape on one partition of 200 events. A workload with much larger partitions or more selective
filters could tip the other way. If you think yours might, add a composite index back and measure
before keeping it — turn on
index metrics with
QueryRequestOptions.PopulateIndexMetrics and compare RequestCharge.
The exclusions are the part that pays, and they pay on both sides: writes drop 2.4% and reads 3–6%.
Memoria already reports the RU charge of every operation. Each Cosmos Read Item,
Cosmos Feed Iterator, and Cosmos Transactional Batch activity event carries a
cosmos.requestCharge tag — see Cosmos DB configuration.
Capture those before and after applying the policy; the write path (Save Aggregate,
Save Events) is where the largest change should show.