Skip to content

Road to State Machines Part V - But How Does the Machine Wait for Something That Happens Later?

Model asynchronous work with waiting states, commands, events, correlation identifiers, timeouts, and a queue that serializes local reactions.

On this page

Our shipping transition now considers more than the order’s status.

It requires a valid paid order and a supported delivery address. The operation that changes the order and the function that lists available actions consult the same transition definition.

That works because the address guard can answer from information already available in the program.

But Part IV ended with a different requirement:

The carrier must confirm whether it can accept a shipment to this destination.

That answer may arrive later. The request may fail. Another command may arrive while we wait.

How do we represent the order during that interval?

Start by making the external question explicit

We should first distinguish three operations that are easy to collapse into one:

OperationWhat it establishes
Check availabilityThe carrier reports whether it currently supports the proposed shipment.
Book a shipmentThe carrier accepts a request to arrange a shipment.
Confirm dispatchThere is evidence that the package has actually been dispatched.

For this part, we will implement only a read-only availability inquiry. It reserves nothing and dispatches nothing. Our existing ship_order() operation will remain a local simulation.

This boundary matters when handling failures. Repeating a read-only inquiry is different from potentially booking the same package twice.

We will also keep the example’s availability policy fixed for an unchanged destination. Real bookings may require fresh validation, expiring quotations, or capacity reservations. Those are additional requirements, not guarantees we get from recording an availability response.

Why not just wait inside the shipping function?

We could call the carrier, wait for an answer, and then continue: Validate the order, await carrier availability, and then continue with the local shipping simulation.

For a small request that has no other work to handle while waiting, this may be sufficient. A state machine does not require a message queue merely because a network call exists.

But our current order would remain PAID throughout the wait.

Another caller could not distinguish:

Paid, with no availability inquiry started

from:

Paid, with an inquiry already in progress

Both could attempt to start the operation.

Making the function asynchronous does not resolve that modeling problem. In Python, an await can suspend a task while the event loop runs other tasks. It does not, by itself, establish which task owns the order or prevent another task from acting on the same record.[1]

We need to represent the unfinished work.

Give waiting a place in the lifecycle

Let’s separate asking the question from using its answer.

We add a CHECK_SHIPPING command. An accepted command starts an availability inquiry and moves the order into CHECKING_SHIPPING.

A positive response moves it to READY_TO_SHIP. Only then does our simulated SHIP_ORDER command become eligible.

PAID
  │ CHECK_SHIPPING
  ▼
CHECKING_SHIPPING
  │ availability confirmed
  ▼
READY_TO_SHIP
  │ SHIP_ORDER
  ▼
SHIPPED

This deliberately changes the rule from Part IV. PAID + SHIP_ORDER is no longer sufficient.

The new states represent different behavior:

StateWhat the machine should accept next
PAIDA request to check shipping availability
CHECKING_SHIPPINGA result for the outstanding inquiry
READY_TO_SHIPThe simulated shipping command
SHIPPEDNo further command in our current lifecycle

We are not adding states merely because an operation takes time. We are adding them because the set of eligible inputs changes while that operation is outstanding.

Replace the state and command definitions with:

class OrderState(Enum):
    CREATED = "created"
    PAID = "paid"
    CHECKING_SHIPPING = "checking_shipping"
    READY_TO_SHIP = "ready_to_ship"
    SHIPPED = "shipped"


class OrderCommand(Enum):
    CAPTURE_PAYMENT = "capture_payment"
    CHECK_SHIPPING = "check_shipping"
    SHIP_ORDER = "ship_order"

READY_TO_SHIP does not claim that a package has been booked or dispatched. It records that our prerequisite inquiry succeeded.

An answer to a question is not the completion of the entire workflow.

Remember what we are waiting for

A waiting state alone is insufficient.

Suppose the carrier returns an answer. Which order does it concern? Which inquiry? Which destination did we submit?

Let’s retain the request:

@dataclass(frozen=True)
class ShippingCheck:
    order_id: str
    request_id: str
    delivery_address: str

The request contains a copy of the destination, not a reference to the mutable order.

We then extend Order:

@dataclass
class Order:
    id: str
    customer_name: str
    delivery_address: str
    status: OrderState = OrderState.CREATED
    payment: Optional[Payment] = None
    shipping_check: Optional[ShippingCheck] = None

Payment remains unchanged.

While checking availability, shipping_check identifies the outstanding request. After a successful response, we retain it to identify the destination that was checked.

This is an example of a correlation identifier: a value carried from a request into its reply so the receiver can connect the two. Enterprise Integration Patterns describes this request–reply relationship explicitly.[2]

The order identifier answers:

Which order?

The request identifier answers:

Which particular inquiry for that order?

Those are not interchangeable.

Extend the invariants with the new states

The payment rule still applies: every state after CREATED requires the captured payment record.

We also need rules for the shipping inquiry:

Lifecycle stateRequired shipping-check data
CREATED, PAIDNo active or accepted check
CHECKING_SHIPPINGThe request currently awaiting a result
READY_TO_SHIP, SHIPPEDThe request whose positive result was accepted

For this model, the recorded destination must match the order’s destination. We have not introduced an address-editing operation that invalidates and restarts a check.

Replace the validator with:

def validate_order(order: Order) -> None:
    if not isinstance(order.status, OrderState):
        raise InvalidOrderState(
            "status must be an OrderState member."
        )

    if order.status is OrderState.CREATED:
        if order.payment is not None:
            raise InvalidOrderState(
                "A created order cannot contain a captured payment."
            )
    elif not isinstance(order.payment, Payment):
        raise InvalidOrderState(
            "This state requires a Payment record."
        )

    needs_check = order.status in (
        OrderState.CHECKING_SHIPPING,
        OrderState.READY_TO_SHIP,
        OrderState.SHIPPED,
    )

    check = order.shipping_check

    if needs_check:
        if not isinstance(check, ShippingCheck):
            raise InvalidOrderState(
                "This state requires a ShippingCheck."
            )

        if (
            not check.request_id
            or check.order_id != order.id
            or check.delivery_address != order.delivery_address
        ):
            raise InvalidOrderState(
                "Shipping check does not match the order."
            )

    elif check is not None:
        raise InvalidOrderState(
            "No shipping check belongs in this state."
        )

This prevents a result for one destination from being silently used after someone changes the order to another destination.

It does not prove that the carrier sent a positive response. That depends on how we control the transition into READY_TO_SHIP.

As before, validating a snapshot and validating the path into that snapshot are separate responsibilities.

A reply is an event, not another shipping command

The user requests a check. The carrier reports an outcome.

Let’s give those outcomes names:

class OrderEvent(Enum):
    SHIPPING_AVAILABLE = "shipping_available"
    SHIPPING_UNAVAILABLE = "shipping_unavailable"
    SHIPPING_CHECK_FAILED = "shipping_check_failed"
    SHIPPING_CHECK_TIMED_OUT = "shipping_check_timed_out"

They have different meanings.

SHIPPING_UNAVAILABLE means the inquiry completed and returned a negative answer.

SHIPPING_CHECK_FAILED means an expected lookup failure prevented us from obtaining an answer.

SHIPPING_CHECK_TIMED_OUT means our application stopped waiting for that inquiry. It does not mean the carrier answered negatively.

The last event comes from our timer, not the carrier.

All four can use the same envelope:

@dataclass(frozen=True)
class ShippingCheckResult:
    order_id: str
    request_id: str
    event: OrderEvent

For this part, a negative answer, a lookup failure, or a timeout returns the order to PAID. That means no usable availability confirmation is retained, and another check may be requested.

This policy is safe for our read-only inquiry. We must not generalize it to a shipment-booking request: a lost booking response could leave an actual booking in the external system.

Keep the new behavior in the transition table

We retain the Transition class and the local has_supported_address() guard from Part IV.

The local guard remains useful as an initial filter. It can reject a destination before we contact the carrier. Passing it no longer claims that all shipping conditions have been satisfied.

Replace the transition table with:

TRANSITIONS: Mapping[
    tuple[OrderState, OrderCommand | OrderEvent],
    Transition,
] = MappingProxyType({
    (
        OrderState.CREATED,
        OrderCommand.CAPTURE_PAYMENT,
    ): Transition(
        target=OrderState.PAID,
    ),

    (
        OrderState.PAID,
        OrderCommand.CHECK_SHIPPING,
    ): Transition(
        target=OrderState.CHECKING_SHIPPING,
        guard=has_supported_address,
        rejection_reason="The delivery address is not supported.",
    ),

    (
        OrderState.READY_TO_SHIP,
        OrderCommand.SHIP_ORDER,
    ): Transition(
        target=OrderState.SHIPPED,
    ),

    (
        OrderState.CHECKING_SHIPPING,
        OrderEvent.SHIPPING_AVAILABLE,
    ): Transition(
        target=OrderState.READY_TO_SHIP,
    ),

    (
        OrderState.CHECKING_SHIPPING,
        OrderEvent.SHIPPING_UNAVAILABLE,
    ): Transition(
        target=OrderState.PAID,
    ),

    (
        OrderState.CHECKING_SHIPPING,
        OrderEvent.SHIPPING_CHECK_FAILED,
    ): Transition(
        target=OrderState.PAID,
    ),

    (
        OrderState.CHECKING_SHIPPING,
        OrderEvent.SHIPPING_CHECK_TIMED_OUT,
    ): Transition(
        target=OrderState.PAID,
    ),
})

Commands and events are now both inputs to the machine. We keep their types separate because requesting an operation and reporting a result are different responsibilities.

Update the selector accordingly:

def next_state(
    order: Order,
    trigger: OrderCommand | OrderEvent,
) -> OrderState:
    validate_order(order)

    if not isinstance(trigger, (OrderCommand, OrderEvent)):
        raise TypeError(
            "trigger must be an OrderCommand or OrderEvent."
        )

    transition = TRANSITIONS.get((order.status, trigger))

    if transition is None:
        raise InvalidOrderOperation(
            f"Cannot handle {trigger.value} "
            f"while the order is {order.status.value!r}."
        )

    if not transition.is_enabled(order):
        raise InvalidOrderOperation(
            transition.rejection_reason
        )

    return transition.target

capture_payment() and ship_order() from Part IV can remain unchanged. Their destinations and eligibility now come from this table.

The previous available_actions() implementation also continues to work: it iterates over OrderCommand, not OrderEvent. Carrier outcomes therefore do not appear as buttons a user may press.

Start the check before starting the external work

The new operation prepares a request and records that the machine is waiting:

from uuid import uuid4


def check_shipping(order: Order) -> ShippingCheck:
    target = next_state(
        order,
        OrderCommand.CHECK_SHIPPING,
    )

    request = ShippingCheck(
        order_id=order.id,
        request_id=uuid4().hex,
        delivery_address=order.delivery_address,
    )

    order.shipping_check = request
    order.status = target

    return request

The returned value describes work to perform. The function does not contact the carrier.

After it returns:

status = CHECKING_SHIPPING
shipping_check = the request we prepared

Only then should another component submit the inquiry.

Why this order?

If the external component can produce a result before we record the pending request, the result handler could receive an answer the order does not yet recognize.

We are establishing local state before allowing external work to respond.

This is still an in-memory sequence of assignments. It does not survive a crash between recording the request and submitting it. We will need a durable boundary later.

For now, it makes the normal execution order explicit.

Accept a result only for the current inquiry

Receiving SHIPPING_AVAILABLE is not enough.

A late answer for an older request must not complete a newer request.

Our handler therefore checks both the lifecycle state and request identifier:

def apply_shipping_result(
    order: Order,
    result: ShippingCheckResult,
) -> bool:
    validate_order(order)

    if result.order_id != order.id:
        raise ValueError(
            "Result was routed to the wrong order."
        )

    if not isinstance(result.event, OrderEvent):
        raise TypeError(
            "Unknown shipping result event."
        )

    if order.status is not OrderState.CHECKING_SHIPPING:
        return False

    request = order.shipping_check

    if request is None or result.request_id != request.request_id:
        return False

    target = next_state(order, result.event)

    if target is OrderState.PAID:
        order.shipping_check = None

    order.status = target
    return True

True means that the result was applied.

False means that this well-formed result no longer belongs to the active inquiry. It does not mean shipping failed.

We treat a wrong order identifier differently: that is a routing error, not an obsolete reply.

On success, the request remains attached to the order and the state becomes READY_TO_SHIP. On an unsuccessful outcome, the request is cleared and the order returns to PAID.

The handler does not perform shipment as a side effect of receiving availability. That remains a separate command.

A late answer can be valid and still be irrelevant

Consider this sequence:

StepInput or actionCurrent condition afterward
1Start inquiry AWaiting for A
2A times out locallyPAID, no accepted check
3Start inquiry BWaiting for B
4Positive response for A arrivesStill waiting for B
5Positive response for B arrivesREADY_TO_SHIP

The answer for A may accurately describe the carrier’s response to A. It is nevertheless irrelevant to the inquiry we are now waiting for.

Checking only the order identifier would accept it.

Checking only CHECKING_SHIPPING would also accept it.

We need the request identifier because the machine may revisit the same waiting state many times.

Being in the right state is necessary. Receiving the right result is necessary too.

Who calls the result handler?

So far, our domain functions are synchronous. They inspect the order, decide, update it, and return.

The network interaction can run separately and produce a message:

import asyncio
from typing import Awaitable, Callable


class CarrierLookupError(Exception):
    pass


async def execute_check(
    request: ShippingCheck,
    carrier: Callable[[str], Awaitable[bool]],
    inbox: asyncio.Queue,
) -> None:
    try:
        available = await carrier(request.delivery_address)
    except CarrierLookupError:
        event = OrderEvent.SHIPPING_CHECK_FAILED
    else:
        if type(available) is not bool:
            raise TypeError(
                "Carrier lookup must return a bool."
            )

        event = (
            OrderEvent.SHIPPING_AVAILABLE
            if available
            else OrderEvent.SHIPPING_UNAVAILABLE
        )

    await inbox.put(
        ShippingCheckResult(
            order_id=request.order_id,
            request_id=request.request_id,
            event=event,
        )
    )

The supplied carrier adapter performs the actual lookup. Its contract is to return a Boolean or translate an expected carrier or transport failure into CarrierLookupError.

We do not catch every exception. A programming error is not automatically a negative availability answer.

More importantly, this function receives a frozen request not the mutable order. It can wait, fail, or finish without directly changing lifecycle state.

Its output goes into a queue.

The component processing that queue decides whether the result still applies.

Process one local reaction at a time

We now need an execution rule:

For this order, finish handling one input before beginning the next local state-machine reaction.

This is run-to-completion behavior.

SCXML formalizes this boundary: processing an external event finishes its resulting internal steps before the next external event is handled. Its interpreter separately describes waiting for events and invoking external processes.[3]

The useful lesson is not that we need XML. It is that a transition’s local work and an external operation’s lifetime are different intervals.

For our order:

Handle CHECK_SHIPPING:
    validate
    create request
    enter CHECKING_SHIPPING
    schedule external work
    return

Later, handle ShippingCheckResult:
    validate
    check request identity
    apply the permitted result
    return

We do not hold the first reaction open until the carrier answers.

That allows the machine to inspect another incoming command while the inquiry is outstanding. A duplicate check request is rejected because CHECKING_SHIPPING has no CHECK_SHIPPING transition.

A small event-processing loop

Python’s asyncio.Queue provides FIFO ordering and allows a consumer to wait when the queue is empty. It is intended for asynchronous code, not unsynchronized access from arbitrary threads.[4]

We will use one consumer for one order. Commands and results enter the same queue.

Payment capture needs its argument, so it gets a small message wrapper:

@dataclass(frozen=True)
class CapturePaymentRequested:
    payment: Payment


Message = (
    CapturePaymentRequested
    | OrderCommand
    | ShippingCheckResult
)

The consumer delegates to the domain operations:

async def run_order(
    order: Order,
    inbox: asyncio.Queue[Message | None],
    start_check: Callable[[ShippingCheck], None],
    report_rejection: Callable[
        [Message, InvalidOrderOperation],
        None,
    ],
) -> None:
    while True:
        message = await inbox.get()

        try:
            if message is None:
                return

            if isinstance(message, CapturePaymentRequested):
                capture_payment(order, message.payment)

            elif message is OrderCommand.CHECK_SHIPPING:
                request = check_shipping(order)
                start_check(request)

            elif message is OrderCommand.SHIP_ORDER:
                ship_order(order)

            elif isinstance(message, ShippingCheckResult):
                apply_shipping_result(order, message)

            else:
                raise TypeError(
                    "Unsupported message or missing command arguments."
                )

        except InvalidOrderOperation as exc:
            report_rejection(message, exc)

        finally:
            inbox.task_done()

start_check is an injected scheduling function. It schedules execute_check() and a timeout; it must return without waiting for the carrier. The surrounding application owns those tasks and their shutdown. Python task groups are one way to manage related asynchronous tasks.[1]

report_rejection reports an expected command rejection without changing the order. Invalid records and unexpected implementation errors are not swallowed.

There is deliberately no await inside the local handlers. The consumer waits for the next message only after the current handler returns.

Under our single-event-loop assumption, and provided all order mutations go through this consumer, another coroutine cannot interleave its own mutation into those synchronous handlers.[1]

That is a local execution guarantee not database atomicity, thread safety, or crash recovery.

task_done() is queue bookkeeping. It does not mean that an external inquiry has finished or that anything was durably committed.[4]

The None message is an orderly stop signal, intended to be sent after message producers have been stopped or drained. This loop is not a complete service-lifecycle implementation.

Waiting for a result is not the same as blocking the consumer

A tempting implementation would put this inside run_order():

available = await carrier(order.delivery_address)

Python could run other tasks during that wait. But this particular consumer would remain inside the handler and would not inspect its next queued message until the lookup finished.

Alternatively, we could create a separate task for every message handler. That would allow handlers for the same mutable order to interleave.

Our design chooses a different boundary:

External work may overlap. Decisions that mutate this order do not.

The event loop runs asynchronous work. The order consumer serializes local reactions. The state machine defines what each reaction may do.

Those are three related responsibilities, not three names for the same mechanism.

A timeout is another input

Suppose the carrier never answers.

We need a local policy for ending the wait.

A timer can enqueue the same kind of result envelope, using the request identifier it was created for. For example, a scheduler using our unbounded, same-event-loop queue could arrange:

loop = asyncio.get_running_loop()

loop.call_later(
    5.0,
    inbox.put_nowait,
    ShippingCheckResult(
        order_id=request.order_id,
        request_id=request.request_id,
        event=OrderEvent.SHIPPING_CHECK_TIMED_OUT,
    ),
)

call_later() schedules a callback using the event loop’s clock. It does not guarantee an exact wall-clock execution instant.[5]

The timer does not mutate the order. It produces an input that the owner will process.

If the carrier response is processed first, the order leaves CHECKING_SHIPPING. The later timeout becomes obsolete.

If the timeout is processed first, that inquiry is abandoned. A later response for it becomes obsolete.

This is an explicit policy: the first applicable outcome processed by the owner settles that inquiry. It is not a rule based on the carrier’s remote timestamp.

A hard deadline based on timestamps would require additional recorded timing information and a different acceptance rule.

We could cancel the timer after accepting a response to avoid unnecessary work. But cancellation cannot be our only protection against stale timeout events. SCXML’s cancellation semantics acknowledge the same issue: a delayed event may already have been delivered when cancellation is attempted.[3]

The request check remains the correctness boundary.

Test event sequences, not network timing

We can test these rules without making a network call or waiting five seconds.

Begin with a paid order:

payment = Payment(
    method="credit_card",
    provider="stripe",
    payment_id="payment-81",
    captured_amount=Decimal("120.00"),
)

order = Order(
    id="order-66",
    customer_name="Alice",
    delivery_address="Istanbul",
)

capture_payment(order, payment)

Start inquiry A, then deliver its timeout:

first = check_shipping(order)

assert order.status is OrderState.CHECKING_SHIPPING
assert available_actions(order) == ()

applied = apply_shipping_result(
    order,
    ShippingCheckResult(
        order.id,
        first.request_id,
        OrderEvent.SHIPPING_CHECK_TIMED_OUT,
    ),
)

assert applied
assert order.status is OrderState.PAID
assert order.shipping_check is None

Start inquiry B. The old positive response must not complete it:

second = check_shipping(order)

applied = apply_shipping_result(
    order,
    ShippingCheckResult(
        order.id,
        first.request_id,
        OrderEvent.SHIPPING_AVAILABLE,
    ),
)

assert not applied
assert order.status is OrderState.CHECKING_SHIPPING
assert order.shipping_check is second

Now deliver the result that actually belongs to B:

assert apply_shipping_result(
    order,
    ShippingCheckResult(
        order.id,
        second.request_id,
        OrderEvent.SHIPPING_AVAILABLE,
    ),
)

assert order.status is OrderState.READY_TO_SHIP
assert available_actions(order) == (
    OrderCommand.SHIP_ORDER,
)

Even B’s own timeout is now too late:

assert not apply_shipping_result(
    order,
    ShippingCheckResult(
        order.id,
        second.request_id,
        OrderEvent.SHIPPING_CHECK_TIMED_OUT,
    ),
)

assert order.status is OrderState.READY_TO_SHIP

ship_order(order)

assert order.status is OrderState.SHIPPED
assert order.payment is payment
validate_order(order)

The tests control the input order directly. They verify the behavior we chose rather than depending on a convenient scheduler race.

Other important sequences include duplicate check commands while waiting, duplicate replies, negative answers, lookup failures, results routed to the wrong order, and attempted shipping before confirmation.

What ordering does the queue actually provide?

FIFO ordering means the consumer receives messages in the order they enter this queue.[4]

It does not establish the order in which remote events occurred.

A carrier may generate a response before our deadline, yet network delay may cause its message to arrive after the timeout message. Our current policy accepts the timeout in that case.

We must distinguish:

When something happened elsewhere

from:

When this machine received and processed information about it

The queue gives us a local sequence. The state and request checks give that sequence a defined interpretation.

Neither mechanism makes the outside world globally ordered.

What have we gained?

The order can now represent an unfinished interaction without pretending it has already succeeded.

Its lifecycle records whether a check may begin, whether a result is expected, and whether shipping is eligible. Its supporting data identifies the exact request and destination. External work produces messages instead of mutating the order. One consumer applies those messages in a local sequence.

This is a reactive execution model: the machine progresses by handling inputs as they arrive, rather than requiring one uninterrupted call to complete the entire process.

The central guarantee is now:

Only a result for the currently outstanding inquiry may advance the waiting order.

That does not provide general event deduplication or exactly-once execution. It prevents obsolete results from completing the wrong inquiry under our current, controlled execution path.

Nor is the waiting durable. A process crash still loses this queue, its timers, and the in-memory order. Failed scheduling after a local update still needs recovery we have not implemented.

The read-only inquiry also matters. Returning to PAID after a timeout means we discarded an unanswered check. It would not prove that a timed-out booking or payment request had no external effect.

We have made waiting explicit. We have not yet made external effects recoverable.

But what if the order is waiting for more than one thing?

Our model allows one shipping inquiry at a time.

Now suppose preparation requires both carrier availability and an inventory reservation. Either result may arrive first. Both must succeed before the order can proceed.

We could add:

WAITING_FOR_CARRIER_AND_INVENTORY
CARRIER_READY_WAITING_FOR_INVENTORY
INVENTORY_READY_WAITING_FOR_CARRIER

Then introduce one more independent check.

The lifecycle begins accumulating combinations again.

How do we represent independent work that can run concurrently, finish in either order, and then join into one permitted next step?

References

[1] Python Software Foundation, Coroutines and Tasks. Cooperative scheduling, task suspension, and task groups.

[2] Gregor Hohpe and Bobby Woolf, “Correlation Identifier”, Enterprise Integration Patterns. Associating a reply with the request that produced it.

[3] World Wide Web Consortium, State Chart XML: State Machine Notation for Control Abstraction, W3C Recommendation, 2015. See Appendix D for run-to-completion execution and §6.3 for cancellation of delayed events.

[4] Python Software Foundation, Queues. FIFO behavior, asynchronous queue operations, task bookkeeping, and thread-safety limitations.

[5] Python Software Foundation, Event Loop - Scheduling Delayed Callbacks. Timer scheduling and the event loop’s monotonic clock.