Resonate
Distributed Async Await for cloud apps
How do you keep a program running when the machine under it dies?
Most cloud software is a Software whose work is split across several processes or machines that talk over a network., whether or not anyone calls it that. One operation becomes a controlling process calling a string of smaller services. Launching a single virtual machine at Amazon means calls to services that pick where it goes, create its disk volumes and network interfaces, and provision the machine. Now and then one of those calls fails.
The network sits in the middle of all of it, and people new to distributed systems tend to assume it behaves. The classic list of false assumptions starts with "The network is reliable" and "Latency is zero". Software written on those assumptions can stall or wait forever during an outage, and then fail to retry when the network comes back.
Inside one program, the tool for waiting is A way of writing code that pauses at a slow operation and resumes when its result arrives.. A A placeholder for a result that isn't ready yet., or future, stands in for a result that isn't known yet, usually because the work isn't finished. The code awaits it and carries on when the answer arrives. The catch is that all of this lives in the memory of one process.
Further reading Making retries safe with idempotent APIs (Amazon Builders' Library)Fallacies of distributed computing (Wikipedia)Futures and promises (Wikipedia)What is Durable Execution? (Resonate)
- i.
Did it happen?
Say a request times out. Did the other side do the work? There is no telling, and simply retrying could launch two machines where you wanted one. Finding out means a reconciliation step, a lot of heavy lifting for an edge case. The general version, agreeing on anything over a link that can drop messages, was the first computer communication problem proven to be unsolvable.
- ii.
Retries need idempotence
Retrying is still the best medicine, since a surprisingly large number of transient faults go away on a second try. But a retry is only safe if the operation can be repeated without extra side effects, a property called Safe to repeat: doing it twice has the same effect as doing it once.. Making operations idempotent pushes the problem to the service, which now has to recognise that a request is a repeat of one it has already seen.
- iii.
Memory that dies
A function's call stack lives in the memory of one process. When the process dies, every local variable and pending call dies with it. Deploys roll, machines get recycled and containers move, while the work itself can take hours or months. A plain retry starts the function over from the top, and repeats whatever the first attempt already did.
- iv.
Scaffolding everywhere
So teams write retries, timeouts, A unique ID sent with a request so that a repeat can be recognised and ignored. and state machines that remember where a job was when its process died. Resonate puts that defensive scaffolding at "more than two-thirds of what you ship", and whatever the true share, almost none of it is the product.
Further reading Making retries safe with idempotent APIs (Amazon Builders' Library)Two Generals' Problem (Wikipedia)What is Durable Execution? (Resonate)Resonate - Fully Open Durable Execution (Resonate)
Resonate is going after that scaffolding. The idea is Running a function so that its progress is saved and it can resume after a crash.: write the code as if failure doesn't exist, and the execution survives it anyway. The process can crash, restart or move to another machine, and the running function carries on from where it stopped.
The developer writes ordinary async functions, and the shape of the code is the shape of the execution. Resonate calls this Distributed Async Await, and it defines it as an open protocol that anyone can implement.
Further reading What is Durable Execution? (Resonate)Resonate - Fully Open Durable Execution (Resonate)
- Step 1: Promises that outlive processes
The basic unit is a durable promise: a promise with a home outside any process, written down the moment it's made and settled when the work is done. When a machine crashes, a new process on any machine reads the last promise that was kept and continues from there. The process can die. The promise is still kept.
- Step 2: Replay what's done
On recovery the function starts again from the top. Steps whose promises already resolved hand back their recorded results instantly, and the first unfinished step runs for real. A step that was mid-flight when the process died runs again, so a charge against a card network still needs to be idempotent. Each promise's ID doubles as the idempotency key for that step.
- Step 3: Three verbs
Most code rests on three calls: run a step durably, call another function, or sleep, for a second or a month. The server is a single binary whose state lives in a Postgres database the team already runs, readable with plain SQL.
- Step 4: A spec you can check
The protocol is written down as an executable, machine-checkable model in Lean 4. The server is tested by pushing the same random operations through several storage backends and a reference model at once, and every one must answer the same way.
Further reading What is Durable Execution? (Resonate)How Resonate works (Resonate)Resonate - Fully Open Durable Execution (Resonate)How Resonate is tested (Resonate)
What happens when the code changes mid-flight?
Replay only works if a function does the same thing the second time around, which is why durable functions must be Always doing the same thing, in the same order, given the same inputs.. A workflow that sleeps for a month may wake up to a new version of its own code. Temporal, another durable execution system, handles this with patches that work like feature flags, applying a change to new executions without disrupting ones already running.
Can anything run exactly once?
Over a link that can lose messages, perfect agreement is provably impossible. What these systems offer instead is a function that runs effectively once and to completion, built from recorded results, retries and idempotent steps.
Further reading Constraints and gotchas (Resonate)Temporal Workflow Execution overview (Temporal)Patching (Temporal)Two Generals' Problem (Wikipedia)How Resonate works (Resonate)
Resonate is building a developer-friendly programming model for distributed execution that replaces brittle integrations in distributed systems. The platform provides Distributed Async Await, a procedural async/await programming model that works across distributed processes, enabling developers to build reliable and scalable cloud applications.
Resonate offers distributed, durable, and composable functions that are "dead simple, formally verified, and deterministically tested." Use cases include service orchestration, transactional applications, and autonomous systems.
- distributed program
- Software whose work is split across several processes or machines that talk over a network.
- async/await
- A way of writing code that pauses at a slow operation and resumes when its result arrives.
- promise
- A placeholder for a result that isn't ready yet.
- idempotent
- Safe to repeat: doing it twice has the same effect as doing it once.
- durable execution
- Running a function so that its progress is saved and it can resume after a crash.
- idempotency key
- A unique ID sent with a request so that a repeat can be recognised and ignored.
- deterministic
- Always doing the same thing, in the same order, given the same inputs.
- 1Making retries safe with idempotent APIs · Amazon Builders' Library
- 2Fallacies of distributed computing · Wikipedia
- 3Futures and promises · Wikipedia
- 4Two Generals' Problem · Wikipedia
- 5What is Durable Execution? · Resonate
- 6Resonate - Fully Open Durable Execution · Resonate
- 7How Resonate works · Resonate
- 8Constraints and gotchas · Resonate
- 9How Resonate is tested · Resonate
- 10Temporal Workflow Execution overview · Temporal
- 11Patching · Temporal