Durable timers & child workflows
Ian — two structural tools that make the 7-stage pipeline realistic. Durable timers let a run pause for days (a review SLA, a scheduled re-crawl) with nothing running. Child workflows let you break the monolith into reusable, independently-retryable pieces. Both are decisions with real tradeoffs — this lesson is as much about when not to as how.
Durable timers: workflow.Sleep
Inside workflow code you never call time.Sleep — it blocks a worker thread and breaks determinism. You call workflow.Sleep. From the docs: the workflow is "paused and resumed by the Cadence service… durable, meaning the workflow can survive worker restarts or failures during the sleep period," and "not consuming worker resources while sleeping."
func ResearchWorkflow(ctx workflow.Context, topic string) (string, error) {
// ... produce a draft, notify a reviewer ...
// Give the reviewer 2 days — the run parks, costs nothing, survives restarts:
if err := workflow.Sleep(ctx, 48*time.Hour); err != nil {
return "", err // interrupted (e.g. workflow cancelled)
}
// ... 48h later, on some worker, execution resumes exactly here ...
return finalize(ctx, topic)
}
workflow.Selector: whichever fires first wins. That is a complete review gate — "human approves within 48h, else auto-escalate." The timer is the deadline; the signal is the human. Neither needs a process babysitting it.
Child workflows: split the pipeline
workflow.ExecuteChildWorkflow starts another workflow from inside a running one. The parent shapes the child's lifecycle; otherwise the child is independent — its own event history, its own timeouts and retry policy, and it can run on a different task list and worker pool (docs). For your pipeline, an "execution" stage that itself orchestrates several LLM sub-steps is a natural child.
cwo := workflow.ChildWorkflowOptions{
WorkflowID: "research-exec-" + topicSlug,
ExecutionStartToCloseTimeout: 30 * time.Minute, // required
}
ctx = workflow.WithChildOptions(ctx, cwo)
var deliverable string
future := workflow.ExecuteChildWorkflow(ctx, ExecutionChildWorkflow, plan)
if err := future.Get(ctx, &deliverable); err != nil {
return "", err
}
The decision that actually matters: activity vs child vs standalone
This is the judgment the docs put front and center. Reach for the lightest tool that fits:
| Use | When |
|---|---|
| Activity | A single non-deterministic operation — one LLM call, one DB write. Lowest overhead, retried independently. Your default. |
| Child workflow | A reusable, self-contained orchestration with its own history/timeouts/retries, but whose lifecycle stays tied to the parent. |
| Standalone workflow | A fully independent process that shouldn't share the parent's lifecycle. Start it top-level with a client instead. |
For most of your 7 stages, an activity is right — one stage, one main operation. Promote a stage to a child workflow only when it becomes its own multi-step orchestration you'd want to reuse or run on separate workers. Resist the urge to make all seven children; that is over-engineering the coordination.
When the parent closes: ParentClosePolicy
A live footgun. What happens to a still-running child when the parent finishes? Set it per child:
| Policy | Behavior |
|---|---|
Terminate (default) | Child is killed immediately when the parent closes. |
RequestCancel | Child gets a cancel request — a chance to clean up first. |
Abandon | Child keeps running independently after the parent closes. |
Default is Terminate. If you spawn a long detached job and let the parent return, the child dies unless you set Abandon. And the docs stress: a child is only scheduled once the parent yields to Cadence, so for an abandoned child, block on future.GetChildWorkflowExecution().Get(ctx, nil) before the parent returns — otherwise it may never start.
workflow.SignalExternalWorkflow, but run queries from an activity or an external client. Communication between parent and child is asynchronous — signals, not shared memory.
Retrieval check
A research run must wait 48 hours for a review deadline. What goes in the workflow?
Correct. workflow.Sleep parks the run in the server — durable, survives restarts, consumes no worker while waiting. time.Sleep and polling loops break determinism and waste a worker.
No. time.Sleep blocks a worker and breaks determinism; a polling loop wastes resources. workflow.Sleep parks the run durably for the full duration.
Your "coordination" stage is a single LLM call that produces a plan. What should it be?
Correct. A single non-deterministic operation is the textbook activity case. Reserve child workflows for reusable multi-step orchestrations; a lone LLM call doesn't need one.
No. One LLM call is a single operation — the lightest tool, an activity, fits. Child and standalone workflows are for self-contained orchestrations.
A parent spawns a long child, then returns quickly. You did not set ParentClosePolicy. What happens to the child?
Correct. The default is Terminate — the child dies when the parent closes. To let it outlive the parent, set Abandon and block until the child has actually started.
No. The default ParentClosePolicy is Terminate, so the child is killed when the parent closes. Use Abandon to let it keep running.
New terms — workflow.Sleep, ChildWorkflowFuture, ParentClosePolicy, standalone workflow — are in the glossary.