Activities in depth

Time: ~8 minutes. Tangible win: you can configure an activity that calls an LLM so that when it fails — and it will — Cadence retries it safely without corrupting your research run.

Ian — Lesson 1 gave you the split: the workflow coordinates, activities do the messy real work. This lesson is about making activities survive reality. Your research pipeline calls LLMs and search APIs. Those are slow, flaky, and occasionally down. Activities are the layer built to absorb that, and Cadence gives you three controls: timeouts, retries, and heartbeats.

Primary source Cadence Concepts — Activities. Read the Timeouts and Retries sections. The docs are explicit: "Cadence does not recover an activity's state… failures are expected. Therefore, Cadence supports automatic activity retries."

The core promise: at-least-once

Cadence does not recover activity state. If a worker dies mid-activity, or the activity errors, Cadence just runs it again. That means an activity may execute more than once. This is the single most important consequence for your pipeline:

Make activities idempotent Because retries can re-run an activity, running it twice must not double the effect. Writing stages/03_execution/output/draft.md twice is fine (same file). Charging a card twice, or appending to a log twice, is not. For your research workflow: prefer overwrite outputs keyed by a stable ID over append. Idempotency is your responsibility, not Cadence's.

The four timeouts (and which one you actually set)

Straight from the docs — every activity has these, in second resolution:

TimeoutCoversFires when
ScheduleToStartQueued → picked up by a workerAll workers down / overloaded
StartToCloseWorker starts → activity returnsThe work itself takes too long
ScheduleToCloseQueued → fully completeEnd-to-end cap across retries
HeartbeatMax gap between heartbeatsA long activity goes silent

The docs require either ScheduleToClose or both ScheduleToStart and StartToClose. In practice, for an LLM call, set StartToClose to the longest a single call should take (say 2 minutes) and let the retry policy handle repeated failures.

Retry policy — the fields that matter

Attach an exponential retry policy to the activity. Parameters from the docs:

ao := workflow.ActivityOptions{
    StartToCloseTimeout: 2 * time.Minute,   // one LLM call's ceiling
    RetryPolicy: &cadence.RetryPolicy{
        InitialInterval:    time.Second,     // wait before first retry
        BackoffCoefficient: 2.0,             // 1s, 2s, 4s, 8s...
        MaximumInterval:    time.Minute,     // cap the growth
        MaximumAttempts:    5,               // give up after 5 tries
        NonRetryableErrorReasons: []string{  // don't retry these
            "BadRequestError",               // e.g. a malformed prompt
        },
    },
}
ctx = workflow.WithActivityOptions(ctx, ao)
var draft string
err := workflow.ExecuteActivity(ctx, ExecutionActivity, plan).Get(ctx, &draft)

When MaximumAttempts is exceeded, the docs state the error is returned to the workflow that invoked it. So your workflow code decides what happens next — skip the stage, fail the run, or branch to a fallback. That decision lives in the deterministic workflow; the retrying lives in Cadence.

NonRetryableErrorReasons is the sharp tool A rate-limit or timeout should retry. A malformed-prompt or auth error will fail identically forever — retrying just wastes 5 attempts and money. Mark those non-retryable so they fail fast back to the workflow.

Long-running activities: heartbeats

If an activity runs for minutes (a big multi-source research crawl, a long generation), set a short Heartbeat timeout and call the heartbeat API periodically from inside the activity. Two payoffs, per the docs:

func crawlSources(ctx context.Context, sources []string) error {
    for i, s := range sources {
        fetch(s)
        activity.RecordHeartbeat(ctx, i) // checkpoint: last index done
    }
    return nil
}

Local activities — the optimization to know but rarely reach for

A local activity runs in the same worker process as the workflow, skipping the task-list round trip. The docs recommend them only for functions that are idempotent, run a few seconds, need no rate limiting or routing, and are non-critical (logging, loading config). They trade away debuggability and have a higher chance of duplicate execution. For your LLM calls: use regular activities. For a quick "load the prompt template from disk": a local activity is fine.

Retrieval check

An activity appends the LLM's answer to a shared results file. A worker crashes and Cadence retries it. What is the risk?

Correct. Activities are at-least-once; a retry re-runs the whole function. Append is not idempotent, so you get a duplicate. Overwrite by stable key instead.

No. Activities run at-least-once, so a retry re-runs the append and duplicates the entry. Idempotency is on you — overwrite rather than append.

Your prompt is malformed and the LLM returns a 400 every single time. Which setting stops five pointless retries?

Correct. A deterministic failure like a 400 will never succeed on retry. Marking it non-retryable fails it fast back to the workflow instead of burning attempts.

No. A 400 fails identically every time; timeouts and backoff still let it retry. Mark it in NonRetryableErrorReasons to fail fast.

Why heartbeat a 15-minute multi-source crawl activity?

Correct. Heartbeats let Cadence spot a dead worker within the heartbeat interval and let the retry resume from the last checkpointed index rather than restarting.

No. Heartbeats give fast failure detection and progress checkpointing. Activities never need to be deterministic — only workflows do.

New terms — at-least-once, idempotency, StartToClose, heartbeat, local activity, NonRetryableErrorReasons — are in the glossary.

Ask me anything. Good follow-ups: "How do I make each of my 7 stages idempotent?" or "What's a sane timeout for a Claude call in practice?" Next lesson: we actually run a workflow locally in Go — start it, kill the worker mid-run, and watch it resume.