How Road Runs Long-Running Processes with Durable Workflows
A charge session waits, reacts to events and fails partway, so writing it as a request never really held. Here is how treating it as one durable workflow changed the way we see the problem.


- Written as a request, a charge session's real state ends up smeared across a status column, cron sweepers, retry loops and six log streams.
- A session starts, waits, reacts to events and fails partway. That is a process, not a request, and it needs one owner that drives it end to end.
- On Temporal the session is one Go function, and its event history, not a status row, becomes the source of truth you can replay and read back.
- The real change was perspective: once a session is a process, the same shape turns up in firmware rollouts and roaming too.
A charge session is not a request
A driver taps start. We reserve money on a card we don't issue, send a command over a mobile network to a charging station we don't own, wait for it to answer, follow the session for the next three hours, stop it, wait for the roaming partner to send the final charge detail record, and only then capture the right amount. The partner's uptime is not ours to control, and their CDR may reference a tariff we haven't seen yet.
Most backend code is written as request, work, response. A charge session is none of those. It starts, waits, reacts to events, fails partway, and has to be explainable weeks later. For a long time we wrote it as if it were a request anyway, and the shape of the problem kept leaking through the code.
What the old way cost us
Written as request/response, the state of a session had to live somewhere between the requests. So it lived in a status column, nudged along by a cron sweeper. Retries were hand-rolled per call. A dead-letter queue caught what fell over. A reconciliation job ran nightly and fixed what the day broke. Each piece was reasonable on its own. Together they were a workflow engine we never meant to build: no history, no UI, and a different failure mode in every service.
And when support asked the only question that mattered, "what happened to this session?", the answer was an afternoon of reading six log streams and guessing. The session's real state was smeared across all of them.
We could have kept building that engine properly ourselves. We decided not to: a durable, replayable process engine is a product, not a feature, and we would rather spend our engineers on charging than on infrastructure that already exists.
No one owned the session
We already had events. A session ends, we publish it, and analytics, notifications and reporting each react. That is choreography, and it is the right tool for spreading the word.
It is the wrong tool for owning the session itself. In a choreographed system the sequence lives nowhere: "we reserved funds, started the station, and are waiting for a CDR that may be ten days away" is split across four consumers and a database. When a step fails there is no one place to decide what happens next, and no one place to look afterwards.
What we needed was an owner: one piece of code that calls each step in turn, waits for the answer, and decides what comes next. We kept the event bus for fan-out, and gave the session a single owner.
Writing the session as a reentrant process
Once the session has an owner, the question is how to write it so that a crash, a deploy or a three-hour wait is a non-event. We wanted it to read like the happy path and still be true when the world is not. Three properties get you there.
Resumable. The code suspends on whatever it is waiting for (a station acknowledgement, a timer, a payment result) and continues from exactly that point when it arrives, whether that is in 200 milliseconds or next Tuesday. No polling loop, no "check every minute" cron job.
Recoverable. If the machine dies, or we ship a new version, the session continues on another machine from the last step it completed. Not from the beginning, and not from a checkpoint someone remembered to write.
Reactive. A meter reading, a stop from the app, a webhook from the payment provider: each is delivered into the running session and changes its course. The code does not go looking for them.
Written this way, the session is one function that reads top to bottom.
- 1Reserve funds
- 2Start charging station
- 3Follow the session
- 4Stop
- 5Capture the right amount
The workflow suspends on each wait (a station acknowledgement, a timer, the final CDR) and resumes from the last completed step, whether that is 200 milliseconds or next Tuesday.
A charge session is like a food order: a clear start, a clear end, and all the difficulty in the middle, where the kitchen catches fire, the courier can't find the door, or the customer cancels halfway. The happy path is five steps. The real thing is those five steps plus every detour, written once, in one place.
Durable execution with Temporal
The platform that makes this true for us is Temporal, an open-source engine for what it calls durable execution. You write the session as an ordinary Go function, and the steps that touch the outside world (a payment call, an OCPP command) as separate activities. Temporal records every step in an event history. When a worker crashes, a deploy restarts it, or the session just waits three hours for a car, it replays that history and picks up where it left off. Retries, timeouts, timers and "is this the only workflow for this session?" are the platform's job, not ours.
The shift was not that we found a better library. It was that the history became the source of truth. The session is no longer a row we mutate and hope; it is a recorded sequence we can replay, read and reason about.
Temporal never runs our code. It decides what should happen next and hands the task to a worker we own, which is why scaling a domain means adding workers, not touching the server.
We started deliberately small: one namespace, one Postgres database behind it, one team, one use case, the session. The point was to prove the model on something real before asking anyone else to trust it.
The same shape shows up elsewhere
Once you see a charge session as a process, you start noticing the shape wherever a job waits, reacts and can fail partway.
Firmware updates are the clearest example. A rollout across thousands of charging stations is not one big job; it is hundreds of independent workflows, one per station, each sending the update and tracking it on its own. A station that is asleep, mid-reboot or offline for a week simply gets its update when it reconnects.
A sleeping workflow costs nothing: no worker is occupied, no cron job is polling. Roaming is another: each exchange of sessions, CDRs and tariffs with a partner over OCPI is a workflow, with retries, ordering and a record of exactly what was sent and acknowledged, against a counterparty whose availability we do not control.
Where we ended up
The first question in any design review is now "is this a workflow?" When the answer is yes, the design gets simpler: one function that reads top to bottom, with the waiting, retrying and recovering left to the platform.
"What happened to this session?" went from an afternoon of log archaeology to a two-minute answer. Every run carries its own history, every step, every retry, every signal, with timestamps. Support escalates a case, an engineer opens the history and reads it.
New work reuses what is already there. An account-free charging flow, a reservation, a fleet-wide command: each is a new workflow built from activities we already trust, not a new state machine with its own edge cases.
And a few things we would tell our earlier selves. Versioning in-flight workflows needs discipline from day one. Keep event histories small; large payloads belong in a store, not in the history. Test with the replayer against production histories before every deploy.
Not everything is a workflow. A synchronous lookup is still a function call, and forcing it into a workflow only adds latency.