Deadlines and attempt timeouts
A retried call needs two different time bounds - mixing them up causes common timeout bugs. A 30-second per-attempt timeout with three retries could run for 90 seconds in total.
The deadline is the ceiling for the entire operation, including every attempt and backoff delay. The attempt timeout is the ceiling for a single attempt.
Both are enabled by default:
- Deadline: 30 seconds for the whole call.
- Attempt timeout: 10 seconds for any single attempt, and usually far less: the ceiling is measured from the dependency's own latency, and 10 seconds is where the lowering stops.
Use Timeout.InfiniteTimeSpan to disable either bound.
CAUTION
A timeout cannot terminate a callback that ignores its cancellation token. If a callback never observes the token, the policy must wait for the task to complete because the executor is awaiting that task.
To prevent this, every execution overload requires a callback that accepts a CancellationToken. The analyzers (NRES001 and NRES002) report cases where a callback is handed the wrong token at build time. If an attempt overruns its ceiling by more than one second, an OrphanedWork event fires retrospectively when the work finally returns.
The two bounds
The effective ceiling for any attempt is the minimum of the AttemptTimeout and the time remaining on the Deadline.
var api = Resilience.Default with
{
Deadline = TimeSpan.FromSeconds(value: 10), // the whole call
AttemptTimeout = TimeSpan.FromSeconds(value: 3), // one attempt
};
// Attempt 1 gets 3 s. An attempt starting with 2 s left on the deadline gets 2 s, not 3 -
// the effective ceiling is min(AttemptTimeout, time left), so there is no
// "is that per attempt or total?" question to get wrong.Deadline is wall-clock time from the moment you call RunAsync. It covers every attempt, every backoff delay, and every BeforeAttempt hook.
AttemptTimeout covers one attempt. If no time remains on the deadline, a retry never starts; the call fails immediately with a deadline exception rather than sleeping through a backoff delay.
Which of the two actually binds, and what the last attempt is left with, is arithmetic over the attempt count, the backoff curve and both bounds. Ask the policy: Explain() prints the worst case attempt by attempt and names the bound that ends it.
This also applies when too little time remains. If the circuit breaker measures how long a healthy call to the dependency takes (which it does by default; see Breaker.NormalLatency), a retry with less time remaining than that measurement is not started. You get the same DeadlineExceededException a few milliseconds sooner, with one fewer attempt in result.Attempts, and the dependency gets one fewer request that you would not have waited for. The first attempt of a call always runs regardless of the measurement, and a policy with no breaker or a cold baseline behaves as usual.
Measure the attempt ceiling instead of guessing it
On by default. AttemptTimeout alone is a number you pick per dependency before it runs, and update whenever it changes. AttemptCeiling measures the ceiling from the dependency's own latency instead, by default at AttemptCeiling.Above(3) - three times the recent p95. AttemptCeiling = null leaves AttemptTimeout as the only per-attempt ceiling.
var api = Resilience.Http with
{
AttemptTimeout = TimeSpan.FromSeconds(value: 5), // the ceiling. Never exceeded.
AttemptCeiling = AttemptCeiling.Above(multiple: 3), // and usually far below it: 3x the recent p95.
};
// The measured term can only lower the ceiling, so AttemptTimeout stops being a guess about how
// long this dependency takes and becomes what it reads as - the point beyond which you stop
// caring. A dependency whose p95 is 40 ms gets a 120 ms ceiling; one whose p95 is 2 s gets the
// configured 5 s, because 3x its p95 is above that and the clamp is what wins.The effective ceiling is the minimum of AttemptTimeout, the time remaining on the deadline, and the measured quantile multiplied by Multiple. Because the measured term only lowers the ceiling, the feature is safe to leave on: AttemptTimeout remains the ultimate ceiling, and a dependency slow enough that the measurement exceeds it simply gets the default behavior.
AttemptCeiling.Above(3) is a complete configuration. The properties you can change:
| Property | Default | Description |
|---|---|---|
Multiple | none - you supply it | How many times the measured quantile an attempt may take. Must be greater than 1. |
Quantile | 0.95 | The quantile of recent successful latency the ceiling is measured from. Between 0.5 and 0.99. |
Window | 5 min | How much history the estimate covers. |
MinimumSamples | 20 | How many recent successful calls the estimate needs before it bounds anything. |
Floor | 50 ms | A floor under the measured ceiling, so a dependency whose p95 is microseconds does not cancel itself on one scheduling hiccup. |
Four behaviors are worth knowing:
- A cold process does not guess. Below
MinimumSamplesthere is no measured term and the attempt getsAttemptTimeoutunchanged. - It only tightens a ceiling you set. A policy whose
AttemptTimeoutisTimeout.InfiniteTimeSpangets no default measured ceiling: you said the deadline was the only per-attempt bound, and there is nothing there to tighten. WritingAttemptCeilingyourself there is a different instruction - "bound me by the dependency's latency and nothing else" - and it is honored. - Only successful attempts are sampled. A ceiling tight enough to cancel calls that would have succeeded starves its own estimator, so the policy reverts to
AttemptTimeoutrather than tightening further. - The estimate is per policy instance. The HTTP handler derives one policy per host, so each host's ceiling is measured from that host's own latency.
- It measures wall clock, and attributes all of it to the dependency. A local incident - a thread pool deep enough that a work item waits 400 ms for a thread - reads as a dependency that got 400 ms slower, and raises the ceiling accordingly.
Saturationis the opt-in switch that tells the two apart.
Read the current value from policy.Measured.AttemptCeiling, or watch the nresilience.attempt.ceiling histogram, which is recorded when the number moves. Both report the measured ceiling before AttemptTimeout clamps it, so a value above your AttemptTimeout is the reading that says the clamp is now what bounds the attempt.
NOTE
When hedging is configured too, the ceiling is measured from at least the hedge's own quantile. A ceiling below the hedge threshold would cancel the first leg at the moment the second was due to start, and you would have bought a feature that never fires.
The third thing the attempt timeout bounds
Enabled by default. AttemptTimeout also bounds the gap between two reads of a response body, and the gap between two elements of a stream. Set BoundProgress = false to remove that bound.
This is not a new number - it is the same one, applied to the part of a call that used to escape it. An attempt ends when the response headers arrive: the body is a live stream over the connection, read after the executor has classified the attempt and returned. Before this bound existed, nothing covered it. Not the deadline, not the attempt timeout, not the breaker's slow-call detection, and not the retry - and because the handler takes ownership of HttpClient.Timeout, not that either. A dependency that sends headers and then stops writing produces a call that never completes.
var api = Resilience.Http with
{
// The attempt, and the gap between two reads of the body that attempt returned.
AttemptTimeout = TimeSpan.FromSeconds(value: 10),
};
// Nothing else to configure. A body that stops arriving for longer than AttemptTimeout fails
// the read with AttemptStalledException rather than hanging, and a body that keeps arriving is
// never cut off however long it takes - the bound is on the gap, not on the total.
//
// Two ways to change that. BoundProgress = false removes the bound entirely. BufferResponses
// reads the body inside the attempt, so a stall becomes one more transient failure and is
// retried - at the cost of holding the whole body in memory.
using var client = HttpResilience.CreateClient(api, new HttpResilienceOptions { BufferResponses = true });The bound is on the gap, not on the total. A 4 GB download that keeps arriving is never cut off, however long it takes: a deadline is a budget for reaching an answer, and the body is what the answer was. Only a read that has been outstanding longer than AttemptTimeout is a stall - and a read that is outstanding is, by definition, waiting on the far side. A consumer that spends a minute processing each buffer is never blamed for it.
Where the stall surfaces decides whether it can be retried
| Where the read happens | What a stall raises | Retried? |
|---|---|---|
| A response body you read yourself | AttemptStalledException, at your read | No - the call already succeeded |
| A stream, between two elements | AttemptStalledException, at your MoveNextAsync | No - the handover is over |
A body read under BufferResponses | AttemptTimeoutException, inside the attempt | Yes, like any transient failure |
The first two are the honest floor: the retry loop was over before the stall existed, so the hang becomes finite rather than retryable, and AttemptStalledException.Attempts is empty because there is no attempt to report. BufferResponses is how to take the other side of that trade - it reads the body inside the attempt, where the deadline and the attempt timeout both already apply, at the cost of holding the whole body in memory.
A stall raises CallEventKind.Stalled, which is the one event that can arrive after a terminal event. It is not itself terminal, so a listener counting terminal events per call still counts exactly one.
NOTE
A policy whose AttemptTimeout is Timeout.InfiniteTimeSpan gets no progress bound, because the number this feature applies is the one that is not there. That is the same step-aside rule AttemptCeiling follows, and it is why Resilience.None states BoundProgress = false rather than leaving it to be inferred.
If you have an exact SLA
Two different requirements hide behind "we need an exact timeout", and they have different answers.
A hard upper bound - "this call must never take longer than 3 seconds" - is Deadline, and a measured ceiling never touches it. The effective ceiling is min(AttemptTimeout, time left, measured), so the time left always clamps and the measured term only ever operates inside the deadline. Your bound is exact to the tick whether or not AttemptCeiling is set, and the measured term can only ever cancel an attempt earlier than you configured, never later.
Counter-intuitively, a tight SLA is the strongest case for measuring the ceiling. Take a 3-second deadline, three attempts, and a dependency whose p95 is 40 ms:
With AttemptTimeout = 10 s alone | With AttemptCeiling = AttemptCeiling.Above(3) | |
|---|---|---|
| First attempt hangs | Capped at min(10 s, 3 s left) = 3 s | Cancelled at ~120 ms |
| Attempts you actually get | One. The deadline is gone. | Three, all inside ~660 ms |
An attempt timeout far above the dependency's real latency is not a safety margin under a tight deadline; it is a guarantee that one hung attempt spends the whole budget. NRES004 warns about the extreme form of this - an AttemptTimeout longer than the Deadline - and a measured ceiling handles the cases an analyzer cannot see, because they depend on what the dependency actually does.
A guaranteed allowance - "every attempt must be allowed a full 2 seconds before we give up on it" - is the requirement a measured ceiling would genuinely fight, and Floor is the answer:
// An exact SLA: this call has 10 seconds, full stop. Deadline is that bound, and nothing here
// lowers or raises it.
var api = Resilience.Http with
{
Deadline = TimeSpan.FromSeconds(value: 10),
AttemptTimeout = TimeSpan.FromSeconds(value: 5),
// And this endpoint legitimately takes up to 2 s sometimes, so no attempt may be
// cancelled before then. Adaptation is confined to [2 s, 5 s]: it can trim the dead time
// above 2 s and can never cut into the allowance below it.
AttemptCeiling = AttemptCeiling.Above(multiple: 3) with { Floor = TimeSpan.FromSeconds(value: 2) },
};Note that a Floor at or above AttemptTimeout is refused at validation. That combination pins the ceiling to exactly AttemptTimeout, which makes AttemptCeiling do nothing at all, and the library refuses configurations that silently have no effect - so the honest way to say "an exact attempt timeout, always" is AttemptCeiling = null. An AttemptTimeout at or below the default 50 ms Floor works the same way: the default steps aside rather than turning your policy into an error.
Bounding one request, not one policy
AttemptCeiling measures across calls, so the estimate lives on the policy instance. If you need a bound that differs per request, publish it rather than deriving a policy per request:
AmbientDeadline.Begin(remaining)withUseAmbientDeadlinegives that request an exact deadline, resolved once asmin(Deadline, remaining). See propagating the deadline.- Deriving
policy with { Deadline = ... }per request also works, but the latency estimate is keyed by the policy instance - so a policy built per request is permanently cold andAttemptCeilingsilently does nothing. It fails safe, back toAttemptTimeout, but it fails quietly.NRES008reports the cases the compiler can see.
Propagate the deadline across a hop
A deadline stops at the process edge unless something carries it across. A service with 200 ms left that sends a request the peer works on for 10 seconds has already produced garbage, and neither side can tell. Two halves fix that, and each is useful without the other.
Send the deadline
Set PropagateDeadline on the HTTP options, and every attempt carries how long this side is going to wait for it:
// The outbound half. Every attempt carries the time this side is prepared to wait:
// min(AttemptTimeout, time left on the deadline). This allows peers to stop
// work that is no longer needed. Off by default.
var api = Resilience.Http with
{
Deadline = TimeSpan.FromSeconds(value: 10),
AttemptTimeout = TimeSpan.FromSeconds(value: 3),
};
var options = new HttpResilienceOptions { PropagateDeadline = true };
using var client = new HttpClient(handler: new HttpResilienceHandler(innerHandler: transport, policy: api, options: options));
using var response = await client.GetAsync(requestUri: uri, cancellationToken: cancellationToken);
// X-Deadline-Ms: 3000 on the first attempt, and less on every attempt after it.The value is the attempt's own ceiling - min(AttemptTimeout, time left on the deadline) - in whole milliseconds, recomputed for every attempt and every hedged leg. DeadlineHeader changes the header name, which defaults to X-Deadline-Ms.
NOTE
grpc-timeout is not a drop-in name for it. gRPC's value carries a unit suffix rather than a bare count of milliseconds, and the gRPC client stack already propagates its own deadlines from CallOptions.Deadline.
The gRPC integration carries the same switch under the same name, and the one difference is deliberate: it is on there and off here. grpc-timeout is a protocol field every gRPC peer already honors; X-Deadline-Ms is a convention this library invented, and a header the other side does not read is not worth sending by default.
Inherit the deadline
Set UseAmbientDeadline on the policy, and the effective deadline becomes min(Deadline, the time the caller is still waiting):
// The inbound half. The policy is bounded by the inherited deadline, so its
// effective deadline is min(Deadline, time the caller is still waiting), resolved once
// at the start of the call.
var api = Resilience.Http with { UseAmbientDeadline = true };
// In an ASP.NET Core app, UseResilienceDeadline() publishes what the caller sent. Anywhere else -
// a queue consumer reading a deadline off a message, or a test - publish it yourself.
using var inbound = AmbientDeadline.Begin(remaining: TimeSpan.FromMilliseconds(value: 200));Nothing else in the model changes. AttemptTimeout is already min(configured, time left), so a shorter deadline shortens the attempts with it, and a call whose inherited deadline has already expired fails immediately with DeadlineExceededException without contacting the dependency.
In an ASP.NET Core app, install NResilience.AspNetCore and read the header with one line:
app.UseResilienceDeadline();Register it before anything that makes an outbound call. One pass reads both propagated values - the deadline, and how much the work matters; see Criticality for the second half.
UseResilienceDeadline also takes a callback: Header changes the header it reads, Maximum caps what it believes from a caller, Reserve keeps part of the deadline back for this service's own work, and RejectExpired refuses a request that arrives with nothing left.
Refuse a request that arrives too late
By default the middleware runs a request whose deadline has already expired, because it may well be answerable from cache. RejectExpired takes the other side of that trade: a request arriving with less time than answering costs is refused with 504 and an RFC 9457 problem document instead of running.
app.UseResilienceDeadline(o =>
{
o.Reserve = TimeSpan.FromMilliseconds(200); // what answering actually costs
o.RejectExpired = true; // refuse anything that arrives with less
});Pair it with Reserve, which is what states that cost. A deadline header carries a positive number or nothing at all, so with no reserve set there is no request this can refuse.
UseAmbientDeadline is off by default and stays off in every preset, because reading the ambient value costs an AsyncLocal<T> read on calls that mostly have no inbound deadline to read. For what that costs and why the read happens once per call rather than once per attempt, see the cancellation contract.
Handle timeout exceptions
Both DeadlineExceededException and AttemptTimeoutException derive from TimeoutException, so you can catch them together or separately. Both include the attempt log.
// DeadlineExceededException and AttemptTimeoutException are both TimeoutException, so one
// catch covers "it did not answer in time" and the two are still distinguishable.
try
{
result.ValueOrThrow();
}
catch (DeadlineExceededException deadline)
{
Console.WriteLine(value: $"gave up after {deadline.Deadline.TotalSeconds}s and {deadline.Attempts.Count} attempt(s)");
}
catch (TimeoutException attempt)
{
Console.WriteLine(value: $"one attempt overran: {attempt.Message}");
}The executor classifies attempt timeouts as Transient internally rather than through your classifier, so it can tell its own timeout apart from caller cancellation.
Caller cancellation
Cancelling the token you passed to the call aborts it immediately. Caller cancellation is not a failure:
- It is never retried.
- It is not counted against a breaker or a budget.
- It is never converted into a timeout.
- Classifiers cannot override it.
The call returns an OperationCanceledException, even when using TryRunAsync.
If a token is cancelled while an attempt is already succeeding, NResilience does not throw away the completed work. The post-attempt check only prevents the loop from starting another attempt.
Work that ignores the token
Both bounds work through the cancellation token, so callbacks must observe it to be terminated. The required CancellationToken parameter on every execution overload, the analyzers, and the OrphanedWork event are the safeguards against callbacks that ignore cancellation.
For the full picture, see The cancellation contract.
