Back to all writing

Designing idempotent bulk ingestion in Microsoft Fabric Eventhouse

Learn how deterministic Eventhouse batch identities and ingest-by extent tags make retries safer, and why extent idempotency is not row deduplication.

BY TIAGO BALABUCH
Three ingestion attempts with the same batch key converge on one Eventhouse extent, while a retry is marked as already ingested.
On this page

Retries are a normal part of data ingestion.

A network call times out. An orchestrator loses the completion response. A job restarts after a transient failure. The safest action often looks like running the same ingestion again.

That is also how duplicate data enters a system.

Reliable ingestion needs more than retry logic. It needs a stable way to recognize that two attempts represent the same logical batch.

The central design rule is:

Give each logical batch a deterministic identity, and reuse that identity for every retry.

Microsoft Fabric Eventhouse ingestion properties support this pattern through ingest-by: extent tags and ingestIfNotExists.

Retry safety starts before the command

An ingestion request and a logical batch are not the same thing.

The request is one execution attempt. The batch is the data unit you intend to load exactly once from the application’s perspective.

For example:

Logical batch: orders / 2026-10-06 / partition-0001
Attempt 1: timed out
Attempt 2: retried by the orchestrator
Attempt 3: manually replayed

All three attempts must use the same batch identity.

If every attempt generates a new identifier, the destination cannot distinguish a retry from new data.

flowchart LR
    A[Logical batch] --> B[Deterministic batch key]
    B --> C[Attempt 1]
    B --> D[Attempt 2]
    B --> E[Manual replay]
    C --> F[Same ingestion identity]
    D --> F
    E --> F

Choose a deterministic batch key

A useful batch key comes from stable source attributes.

Common inputs include:

  • Source system
  • Dataset or entity
  • Business date or extraction window
  • Source partition
  • Immutable file name
  • Source file checksum
  • Upstream transaction or manifest ID

A synthetic key might look like:

orders-2026-10-06-part-0001

The same source batch must always produce the same key. A genuinely new batch must produce a different key.

Avoid using:

  • Current ingestion time
  • Random GUID generated on each attempt
  • Orchestrator run ID when retries create new runs
  • A file path that changes when the same content is copied

Those values describe the attempt, not the logical data.

How Eventhouse ingestion tags prevent reingestion

Eventhouse extent tags attach metadata to ingested extents.

Tags beginning with ingest-by: have a specific purpose. They work with the ingestIfNotExists ingestion property to prevent a logical batch from being ingested again.

The pair performs two jobs:

  1. ingestIfNotExists checks whether an extent already has the expected ingest-by: tag.
  2. tags assigns that tag when the new ingestion succeeds.

The values should match:

with (
  ingestIfNotExists='["orders-2026-10-06-part-0001"]',
  tags='["ingest-by:orders-2026-10-06-part-0001"]'
)

If the destination already contains an extent tagged with that batch key, the ingestion does not complete again.

Test with queued ingestion commands

Fabric supports queued ingestion commands for ingesting individual blobs, lists of blobs, folders, or containers.

These commands are useful for exploring and validating an ingestion design. Microsoft documents them as tools for prototyping and testing, not for production or high-volume ingestion.

Use them to prove:

  • The source path resolves correctly
  • The target schema and mapping work
  • The deterministic batch key is stable
  • The first attempt succeeds
  • A retry with the same key is blocked
  • A new batch with a new key succeeds
  • Operation status can be tracked

A synthetic validation command can look like this:

.ingest-from-storage-queued into table database('Operations').Orders
EnableTracking=true
with (
  format='parquet',
  ingestionMappingReference='OrdersMapping',
  ingestIfNotExists='["orders-2026-10-06-part-0001"]',
  tags='["ingest-by:orders-2026-10-06-part-0001"]'
)
<| 'https://storage.example/orders/2026/10/06/part-0001.parquet'

The URL is illustrative. Use a supported storage connection string, SAS, or managed identity according to the command documentation.

The command returns an ingestion operation ID. Use it to inspect progress:

.show queued ingestion operations "<operation-id>"

Operation tracking helps distinguish these outcomes:

  • The command was accepted and remains in progress.
  • The ingestion completed.
  • The ingestion failed.
  • The retry was skipped because the batch identity already exists.

Do not infer successful ingestion only from the client receiving a command response.

Validate the retry path explicitly

A basic test requires at least three operations:

Test Batch key Expected purpose
Initial ingestion orders-2026-10-06-part-0001 Load the logical batch
Retry orders-2026-10-06-part-0001 Confirm duplicate ingestion is prevented
New batch orders-2026-10-06-part-0002 Confirm new data still loads

Verify both the operation status and the data:

Orders
| summarize Records=count() by SourceBatchId
| order by SourceBatchId asc

SourceBatchId is an application-level column in this example. Keeping the logical batch key in the data makes validation and operational investigation easier.

Extent idempotency is not record deduplication

ingest-by: tags apply to extents. They prevent a tagged logical ingestion from being accepted again.

They do not inspect every row and decide whether its business key already exists.

This distinction matters:

Scenario Ingestion tag helps? Additional design needed?
Same batch retried with the same key Yes Monitor operation result
Same file copied to a new batch key No Stable content or manifest identity
Two files contain the same business records No Upstream or query-layer deduplication
Streaming event delivered more than once Not before extent creation Idempotent event-consumer design
Corrected batch replaces earlier data Not by itself Explicit replacement or reconciliation pattern

If the source is known to contain duplicates, Microsoft recommends addressing them before ingestion where possible.

Do not generate a unique tag for every call

It can be tempting to use the request ID as an ingest-by: tag.

That defeats duplicate prevention when a retry receives a new request ID. It can also create excessive tag cardinality.

Microsoft documentation warns that assigning unique ingest-by: tags for each ingestion call can affect performance.

Use tags at the logical batch level:

Good: orders-2026-10-06-part-0001
Weak: request-8f4c2f06-1c62-4ea8-a907-8a55827bd608

The first key remains stable across attempts. The second identifies only one execution.

Control ingestion-property fragmentation

Queued ingestion batches data according to compatible ingestion properties.

Using many distinct mapping properties or constant values can fragment ingestion and reduce performance. Idempotency design should therefore avoid unnecessary per-request variation.

Keep these properties stable when the logical data shape is stable:

  • Format
  • Ingestion mapping reference
  • Schema behavior
  • Creation-time strategy
  • Validation policy

Vary the ingestion key only when the logical batch changes.

Model the full ingestion state

An orchestrator should not reduce ingestion to Succeeded or Failed.

Use states that represent the actual workflow:

stateDiagram-v2
    [*] --> Prepared
    Prepared --> Submitted
    Submitted --> InProgress
    InProgress --> Completed
    InProgress --> Failed
    Submitted --> AlreadyIngested
    Failed --> Submitted: Retry same batch key
    Completed --> Verified
    AlreadyIngested --> Verified

This model makes a skipped duplicate a valid terminal path rather than an unexplained failure.

Store:

  • Logical batch key
  • Source manifest or file set
  • Target database and table
  • Ingestion operation ID
  • Attempt number
  • Submission time
  • Final operation state
  • Record or extent validation result

Do not store credentials or SAS tokens in the operational log.

Use a manifest for multi-file batches

When one logical batch contains many files, define the batch before submitting it.

A manifest can include:

{
  "batchId": "orders-2026-10-06-hour-09",
  "files": [
    "part-0001.parquet",
    "part-0002.parquet",
    "part-0003.parquet"
  ],
  "mapping": "OrdersMapping",
  "target": "Operations.Orders"
}

Generate the batch key from stable manifest content or an upstream immutable identifier.

Do not silently add files to an existing completed batch. Create a new batch or a documented correction operation so that retries remain distinguishable from new data.

Plan for partial and ambiguous outcomes

A client timeout does not prove that ingestion failed.

Before resubmitting:

  1. Check the queued operation status when an operation ID exists.
  2. Check whether the logical batch tag already exists.
  3. Validate the destination using the batch column or manifest.
  4. Retry with the same deterministic key only when the state remains unresolved.

This pattern protects against the classic ambiguity:

The client did not receive success.
The service might still have completed the ingestion.

Changing the batch key during that retry removes the protection.

Know the streaming limitation

Extent tags cannot be assigned to streaming records before those records are stored in extents.

For streaming scenarios, use event-level identity and idempotent consumer logic. If the event format follows a delivery model that can retry, persist the event identity and make repeated processing produce the same result.

Do not assume that the queued file-ingestion pattern applies unchanged to streaming ingestion.

The key takeaway

Retries do not create reliable ingestion by themselves.

Reliable ingestion needs a stable identity for the data being loaded. In Eventhouse, ingestIfNotExists and ingest-by: tags provide a documented extent-level mechanism for preventing the same logical batch from being ingested again.

The design works only when the batch key survives retries.

Define the batch first. Derive a deterministic key. Use the same key on every attempt. Track the operation. Validate the destination. Handle record-level duplicates as a separate concern.

That turns retries from a duplication risk into a controlled recovery path.

References

Share this post