Durable timers & child workflows

Time: ~9 minutes. Tangible win: you can express a research run that waits days without a running process, and you can decide when to split a stage into its own child workflow versus keeping it a plain activity.

Ian — two structural tools that make the 7-stage pipeline realistic. Durable timers let a run pause for days (a review SLA, a scheduled re-crawl) with nothing running. Child workflows let you break the monolith into reusable, independently-retryable pieces. Both are decisions with real tradeoffs — this lesson is as much about when not to as how.

Primary source Go Client — Sleep and Go Client — Child workflows. The child-workflows page has the decision table and the ParentClosePolicy behaviors used below.

Durable timers: workflow.Sleep

Inside workflow code you never call time.Sleep — it blocks a worker thread and breaks determinism. You call workflow.Sleep. From the docs: the workflow is "paused and resumed by the Cadence service… durable, meaning the workflow can survive worker restarts or failures during the sleep period," and "not consuming worker resources while sleeping."

func ResearchWorkflow(ctx workflow.Context, topic string) (string, error) {
    // ... produce a draft, notify a reviewer ...

    // Give the reviewer 2 days — the run parks, costs nothing, survives restarts:
    if err := workflow.Sleep(ctx, 48*time.Hour); err != nil {
        return "", err // interrupted (e.g. workflow cancelled)
    }
    // ... 48h later, on some worker, execution resumes exactly here ...
    return finalize(ctx, topic)
}
Sleep + signal = the real review gate Combine Lesson 4's signal with a timer using workflow.Selector: whichever fires first wins. That is a complete review gate — "human approves within 48h, else auto-escalate." The timer is the deadline; the signal is the human. Neither needs a process babysitting it.
Don't fire a million timers at once The docs warn: very large numbers of simultaneous timers can strain the cluster. If you'd start thousands of long sleeps at the same instant (e.g. a batch of scheduled re-crawls), jitter or batch them. For a handful of research runs this never bites — but know the ceiling exists.

Child workflows: split the pipeline

workflow.ExecuteChildWorkflow starts another workflow from inside a running one. The parent shapes the child's lifecycle; otherwise the child is independent — its own event history, its own timeouts and retry policy, and it can run on a different task list and worker pool (docs). For your pipeline, an "execution" stage that itself orchestrates several LLM sub-steps is a natural child.

cwo := workflow.ChildWorkflowOptions{
    WorkflowID:                   "research-exec-" + topicSlug,
    ExecutionStartToCloseTimeout: 30 * time.Minute, // required
}
ctx = workflow.WithChildOptions(ctx, cwo)

var deliverable string
future := workflow.ExecuteChildWorkflow(ctx, ExecutionChildWorkflow, plan)
if err := future.Get(ctx, &deliverable); err != nil {
    return "", err
}

The decision that actually matters: activity vs child vs standalone

This is the judgment the docs put front and center. Reach for the lightest tool that fits:

UseWhen
ActivityA single non-deterministic operation — one LLM call, one DB write. Lowest overhead, retried independently. Your default.
Child workflowA reusable, self-contained orchestration with its own history/timeouts/retries, but whose lifecycle stays tied to the parent.
Standalone workflowA fully independent process that shouldn't share the parent's lifecycle. Start it top-level with a client instead.

For most of your 7 stages, an activity is right — one stage, one main operation. Promote a stage to a child workflow only when it becomes its own multi-step orchestration you'd want to reuse or run on separate workers. Resist the urge to make all seven children; that is over-engineering the coordination.

When the parent closes: ParentClosePolicy

A live footgun. What happens to a still-running child when the parent finishes? Set it per child:

PolicyBehavior
Terminate (default)Child is killed immediately when the parent closes.
RequestCancelChild gets a cancel request — a chance to clean up first.
AbandonChild keeps running independently after the parent closes.

Default is Terminate. If you spawn a long detached job and let the parent return, the child dies unless you set Abandon. And the docs stress: a child is only scheduled once the parent yields to Cadence, so for an abandoned child, block on future.GetChildWorkflowExecution().Get(ctx, nil) before the parent returns — otherwise it may never start.

One thing you can't do from workflow code A parent cannot query a child from inside workflow code (queries would be non-deterministic). Signal a child with workflow.SignalExternalWorkflow, but run queries from an activity or an external client. Communication between parent and child is asynchronous — signals, not shared memory.

Retrieval check

A research run must wait 48 hours for a review deadline. What goes in the workflow?

Correct. workflow.Sleep parks the run in the server — durable, survives restarts, consumes no worker while waiting. time.Sleep and polling loops break determinism and waste a worker.

No. time.Sleep blocks a worker and breaks determinism; a polling loop wastes resources. workflow.Sleep parks the run durably for the full duration.

Your "coordination" stage is a single LLM call that produces a plan. What should it be?

Correct. A single non-deterministic operation is the textbook activity case. Reserve child workflows for reusable multi-step orchestrations; a lone LLM call doesn't need one.

No. One LLM call is a single operation — the lightest tool, an activity, fits. Child and standalone workflows are for self-contained orchestrations.

A parent spawns a long child, then returns quickly. You did not set ParentClosePolicy. What happens to the child?

Correct. The default is Terminate — the child dies when the parent closes. To let it outlive the parent, set Abandon and block until the child has actually started.

No. The default ParentClosePolicy is Terminate, so the child is killed when the parent closes. Use Abandon to let it keep running.

New terms — workflow.Sleep, ChildWorkflowFuture, ParentClosePolicy, standalone workflow — are in the glossary.

Ask me anything. Good follow-ups: "Sketch the full 7 stages — which become children?" or "How does Sleep interact with the ExecutionStartToCloseTimeout?" Next lesson is the hands-on: stand up a local server (no Docker needed) and run a workflow you can crash and resume.