Race Conditions Between Kalshi Order Cancel and Fill Confirmation
Only fills can tell you whether your canceled order actually closed.

Kalshi's API has a race condition between order cancellation and fill confirmation built into it by design, and no cancel logic that treats a cancel response as final can avoid it. This piece lays out why the window exists, what it breaks when ignored, and what a cancel-handling system has to do to survive it, with fills, not cancel acknowledgements, as the only record that settles what actually happened to an order.
The Cancel-Fill Window as a Structural Property of the Kalshi API
Every order cancel sent to Kalshi is a request, not a command. Between the moment a client sends the REST DELETE and the moment a fill confirmation arrives over WebSocket, the exchange is free to match the order against the book. If that match happens inside the window, the fill wins, and the cancel arrives too late to undo it. Nothing about this is a bug in Kalshi's implementation. It's the ordinary consequence of running a matching engine and a client over a network with nonzero latency: the client's cancel request and the exchange's matching process are two independent streams of events, and the exchange has no obligation to pause matching while a cancel is in flight.
The clearest evidence that this is structural rather than incidental comes from the FIX protocol itself, which long predates Kalshi and governs order entry across much of the trading industry. FIX defines a specific rejection reason, "Too late to cancel," coded as 0, whose entire purpose is to tell a client that the order it tried to cancel had already filled. Kalshi's own FIX Order Entry documentation follows the same pattern: it defines the full lifecycle of an OrderCancelRequest and spells out the conditions under which an OrderCancelReject (35=9) carries that rejection. The protocol assumes the race exists and puts the burden of handling it on the client, not on the exchange.
Two distinct failure surfaces fall out of this single structural fact, and they matter because they look identical from the outside but demand different recovery logic. The first is a cancel that the exchange acknowledges, but where a fill was already in flight at the moment the cancel was processed. Treating them the same is the mistake almost every broken cancel flow makes, and it's the mistake the rest of this piece works through.
What happens to a trading bot's state when cancel logic treats the cancel response as final
A cancel acknowledgement from Kalshi is not a terminal state. And depending on what the strategy does next, it may place a brand new order directly on top of a position it doesn't know it's carrying.
This is exactly the sequence that broke Hummingbot's XEMM V2 executor strategy. The executor's state machine didn't account for that outcome. It broke, and rather than booking the fill and sending the taker-side hedge order the strategy depends on, it proceeded straight to creating the next maker order. Code that read "cancel attempted" as "position closed" left a naked position sitting on the book with no hedge behind it.
Partial fills compound the problem by introducing a second quantity to track. If code doesn't separate these two quantities, it ends up either double-counting contracts that were never actually bought, or silently losing contracts that were. This isn't a hypothetical failure mode confined to one project. Two codebases, two different exchanges, the same underlying cause: cancel logic that trusted a response as confirmation when only the ledger could confirm it.
Fills as the Authoritative Source of Truth
The only record that can settle what actually happened to an order is the fill record. A cancel response tells a client what the exchange intends to do, or has started doing, with a request. It does not tell the client what happened on the book. Treating a cancel acknowledgement as settlement confuses an intent signal with an outcome, and the fix isn't cleverer cancel handling, it's routing every cancel decision through the one data source that actually reflects what the matching engine did.
Kalshi's own API surface is built around that separation. As of October 8, 2026, GET /portfolio/fills accepts up to 100 comma-separated market tickers in a single call, making it practical to reconcile fills across an entire portfolio of resting orders in one pass. These two endpoints, taken together, form the reconciliation layer the system is meant to run against. A cancel response was never meant to double as that ledger.
This matters as much when a cancel request times out as when it succeeds. A timeout is not evidence that the cancel failed. Reading fills is the only way to find out which happened before deciding what to do next. Retrying blindly on a timeout carries a specific risk here: sending the same cancel request twice is harmless, since a cancel on an already-canceled or already-filled order simply fails or no-ops, but retrying the original order as though it were lost can open a second position directly on top of one that already matched. Put simply, a cancel moves an order into an unresolved state, and only a fill read can resolve it.
Building a reconciliation loop that handles both cancel outcomes and in-flight fills
A cancel flow that holds up under this race has three distinct layers, and almost every bug described so far comes from collapsing them into one step.
The first layer is issuing the cancel itself, carrying a stable idempotency key rather than relying on the cancel call alone to be safe. That key is the order's client_order_id, or the exchange-assigned order_id you pull from GET /portfolio/orders. On the FIX side, an OrderCancelRequest is required to carry OrigClOrdID in tag 41, the ClOrdID of the original order, as the binding reference between the cancel and the order it targets. For FCM-sponsored access, the customer-account identifier has to match exactly between the original order and the cancel request, or the exchange has no reliable way to tie the two together.
The second layer holds state in place while waiting for resolution. Once the cancel request goes out, the order should move to a CANCEL_PENDING marker and sit there. An unresolved marker is a more honest description of reality than any guess about what the terminal state will turn out to be, and every decision the strategy makes afterward has to treat that order's quantity as still at risk until the fill data says otherwise.
The third layer is resolution against fills, and it's where the actual numbers come in. The system reads GET /portfolio/fills, filtered down to the relevant ticker, and compares fill_count_fp against initial_count_fp from GET /portfolio/orders. Three outcomes fall out of that comparison. If fill_count_fp equals initial_count_fp, the fill won the race completely: the order resolves as FILLED, and the full quantity gets booked at its fill cost. If fill_count_fp is greater than zero and remaining_count_fp reads zero once the cancel is confirmed, a partial fill got booked and the remainder was genuinely canceled, so only the filled quantity and its actual cost get recorded. If fill_count_fp comes back zero and the cancel is confirmed, the cancel won cleanly and the capacity it was holding can be released.
The cost side of that resolution has to come from the fills data itself, not from assuming the limit price is what the order actually paid. The order object returned by GET /portfolio/orders carries the authoritative cost fields: taker_fill_cost_dollars, maker_fill_cost_dollars, taker_fees_dollars, and maker_fees_dollars. The fills endpoint uses a different schema entirely, with fee_cost alongside yes_price_dollars and no_price_dollars, so a reconciliation loop built against the wrong endpoint's field names will quietly pull incomplete or mismatched cost data.
WebSocket data has a role in this loop, but a narrow one. The user_orders channel, which as of October 1, 2026 carries a sending_ts_ms field for establishing message order, is useful for detecting that a fill event happened quickly. No order should get placed, and no validated state should get updated, from inside a WebSocket message handler directly. The sequence should run state update first, strategy evaluation second, and any follow-on order sent through the REST client with its own risk checks and its own idempotency key, kept separate from whatever triggered it over the socket. After any disconnect, the right recovery is to reconnect and pull a fresh order snapshot over REST before resuming cancel or fill logic at all, rather than trying to replay whatever WebSocket deltas were missed during the gap. The REST snapshot is ground truth; a reconstructed stream of missed messages is a guess.
Five fault scenarios worth injecting to prove the reconciliation loop is correct
A reconciliation loop that only passes happy-path tests has proven almost nothing, because the race this entire piece is about appears only under specific timing conditions that ordinary test runs rarely produce on their own. Proving the loop correct means deliberately forcing the conditions that expose it, and a documented taxonomy (Vigil, GitHub Issue #51, September 2026) lays out five fault scenarios that any serious test suite needs to cover as a baseline, not as a stretch goal.
The first scenario is a cancellation that actually took effect on the exchange, but whose acknowledgement never reached the client. The fourth is the same race but with a full fill winning outright, so the order resolves as FILLED rather than CANCELED, and there is no capacity to release afterward since none remains. The fifth is a partially filled resting order that simply expires, with neither a subsequent fill nor an actual cancel ever resolving the remainder, a case that has to be recognized as its own distinct terminal event rather than quietly logged as a cancel that never technically happened.
A documented acceptance standard (TradeX Issue #102, October 2026) turns this taxonomy into something testable: exercise the cancellation, partial-fill, and full-fill races in both possible orderings, fill arriving before the cancel acknowledgement and cancel acknowledgement arriving before the fill, and assert four things hold across every run. And reservation settlement happens exactly once, only after the terminal evidence for that order is complete. Those four assertions, run against all five fault scenarios in both orderings, are the actual bar a reconciliation loop has to clear before anyone should trust it in production.
Why a stateless mock or live sandbox cannot exercise these fault scenarios reliably
Testing this race condition properly requires a system that can produce a fill concurrently with a pending cancel, on command, and neither a stateless mock nor a live sandbox is built to do that.
A stateless mock validates that a request is shaped correctly and hands back a canned response. A cancel call against a stateless mock that returns a clean success response leaves the entire fill-race assertion untested: the mock confirmed the cancel, but it never had the fill to send in the first place, so the scenario that actually matters never ran.
A live sandbox solves part of that problem by exercising the real API mechanics, but it introduces a different limitation: nobody controls the timing. Rate limits and the live credentials a sandbox typically requires also mean these tests, if they run at all, run occasionally rather than on every commit, and nothing about them is deterministic from one run to the next.
A simulator that holds order state across calls, rather than responding to each request in isolation, is what closes this gap. An order placed against it is a real resting order inside the simulator's own model of the book. A cancel request gets evaluated against that live state, and a fill can be injected at a precise, chosen moment relative to it. Each of the five fault scenarios can be reproduced on demand, deterministically, inside a CI pipeline, without touching live credentials or running into rate limits. Kalshi's API itself changes over time, so a simulator is only worth trusting if you check it against the real API before every release; if it's drifted out of sync with live behavior, it gives you the same false confidence a stateless mock does, just with better production values.
A practical checklist for cancel logic that is safe against the race window
Cancel logic that survives this race treats every cancel acknowledgement as the start of a reconciliation process, never its conclusion. Once a cancel request goes out, the order belongs in an explicit CANCEL_PENDING state, and nothing downstream, capacity release, position updates, or the placement of a follow-on order, should happen while it sits there. Resolution has to come from reading GET /portfolio/fills and GET /portfolio/orders and comparing fill_count_fp, remaining_count_fp, and initial_count_fp directly, rather than from inferring an outcome out of whatever the cancel call itself returned. WebSocket fill notifications are useful only as a trigger to go re-read the fills endpoint, never as a stand-in for that read, and no state should update and no order should be placed from inside a socket handler. A timeout on a cancel request calls for a reconciliation read before any retry, never an assumption that the cancel failed. After any disconnect, the system should pull a fresh REST snapshot of orders before resuming cancel or fill logic, rather than trying to reconstruct what happened from missed WebSocket messages. And before any of this logic ships, it needs to run against all five fault scenarios, cancellation effective with acknowledgement lost, cancellation ineffective with the order still live, partial fill winning the race, full fill winning the race, and partial fill followed by expiry, in both possible message orderings, with no lost fills, no fabricated cancellations, and no capacity released before the order's terminal state is actually known.