Activities in depth
Ian — Lesson 1 gave you the split: the workflow coordinates, activities do the messy real work. This lesson is about making activities survive reality. Your research pipeline calls LLMs and search APIs. Those are slow, flaky, and occasionally down. Activities are the layer built to absorb that, and Cadence gives you three controls: timeouts, retries, and heartbeats.
The core promise: at-least-once
Cadence does not recover activity state. If a worker dies mid-activity, or the activity errors, Cadence just runs it again. That means an activity may execute more than once. This is the single most important consequence for your pipeline:
stages/03_execution/output/draft.md twice is fine (same file). Charging a card twice, or appending to a log twice, is not. For your research workflow: prefer overwrite outputs keyed by a stable ID over append. Idempotency is your responsibility, not Cadence's.
The four timeouts (and which one you actually set)
Straight from the docs — every activity has these, in second resolution:
| Timeout | Covers | Fires when |
|---|---|---|
ScheduleToStart | Queued → picked up by a worker | All workers down / overloaded |
StartToClose | Worker starts → activity returns | The work itself takes too long |
ScheduleToClose | Queued → fully complete | End-to-end cap across retries |
Heartbeat | Max gap between heartbeats | A long activity goes silent |
The docs require either ScheduleToClose or both ScheduleToStart and StartToClose. In practice, for an LLM call, set StartToClose to the longest a single call should take (say 2 minutes) and let the retry policy handle repeated failures.
Retry policy — the fields that matter
Attach an exponential retry policy to the activity. Parameters from the docs:
ao := workflow.ActivityOptions{
StartToCloseTimeout: 2 * time.Minute, // one LLM call's ceiling
RetryPolicy: &cadence.RetryPolicy{
InitialInterval: time.Second, // wait before first retry
BackoffCoefficient: 2.0, // 1s, 2s, 4s, 8s...
MaximumInterval: time.Minute, // cap the growth
MaximumAttempts: 5, // give up after 5 tries
NonRetryableErrorReasons: []string{ // don't retry these
"BadRequestError", // e.g. a malformed prompt
},
},
}
ctx = workflow.WithActivityOptions(ctx, ao)
var draft string
err := workflow.ExecuteActivity(ctx, ExecutionActivity, plan).Get(ctx, &draft)
When MaximumAttempts is exceeded, the docs state the error is returned to the workflow that invoked it. So your workflow code decides what happens next — skip the stage, fail the run, or branch to a fallback. That decision lives in the deterministic workflow; the retrying lives in Cadence.
Long-running activities: heartbeats
If an activity runs for minutes (a big multi-source research crawl, a long generation), set a short Heartbeat timeout and call the heartbeat API periodically from inside the activity. Two payoffs, per the docs:
- Fast failure detection: if the worker dies, Cadence notices within the heartbeat interval and retries elsewhere — instead of waiting for a long
StartToClose. - Progress checkpointing: a heartbeat can carry a payload (e.g. "processed 8 of 20 sources"). On retry, the next attempt reads that progress and resumes rather than restarting the crawl.
func crawlSources(ctx context.Context, sources []string) error {
for i, s := range sources {
fetch(s)
activity.RecordHeartbeat(ctx, i) // checkpoint: last index done
}
return nil
}
Local activities — the optimization to know but rarely reach for
A local activity runs in the same worker process as the workflow, skipping the task-list round trip. The docs recommend them only for functions that are idempotent, run a few seconds, need no rate limiting or routing, and are non-critical (logging, loading config). They trade away debuggability and have a higher chance of duplicate execution. For your LLM calls: use regular activities. For a quick "load the prompt template from disk": a local activity is fine.
Retrieval check
An activity appends the LLM's answer to a shared results file. A worker crashes and Cadence retries it. What is the risk?
Correct. Activities are at-least-once; a retry re-runs the whole function. Append is not idempotent, so you get a duplicate. Overwrite by stable key instead.
No. Activities run at-least-once, so a retry re-runs the append and duplicates the entry. Idempotency is on you — overwrite rather than append.
Your prompt is malformed and the LLM returns a 400 every single time. Which setting stops five pointless retries?
Correct. A deterministic failure like a 400 will never succeed on retry. Marking it non-retryable fails it fast back to the workflow instead of burning attempts.
No. A 400 fails identically every time; timeouts and backoff still let it retry. Mark it in NonRetryableErrorReasons to fail fast.
Why heartbeat a 15-minute multi-source crawl activity?
Correct. Heartbeats let Cadence spot a dead worker within the heartbeat interval and let the retry resume from the last checkpointed index rather than restarting.
No. Heartbeats give fast failure detection and progress checkpointing. Activities never need to be deterministic — only workflows do.
New terms — at-least-once, idempotency, StartToClose, heartbeat, local activity, NonRetryableErrorReasons — are in the glossary.