Workflow vs Activity
Ian — this is the one distinction the entire framework rests on. Cadence has exactly two kinds of code: workflows and activities. Almost every newcomer mistake is putting code in the wrong one. Get this split right and the rest of Cadence is detail. Get it wrong and your workflow silently corrupts on the first crash.
The split in one sentence
A workflow decides what happens and in what order. An activity does the actual work that touches the outside world.
| Workflow | Activity | |
|---|---|---|
| Job | Coordinate: order, branching, waiting, retries | Do one side-effecting task |
| Touches outside world? | Never | Always |
| Can call an LLM / API / DB? | No | Yes |
| Survives a crash? | Yes — state is rebuilt by replay | No state of its own; just retried |
| Must be deterministic? | Yes | No — free to be messy |
Why the workflow can't touch the outside world
This is the counterintuitive part, and it is the whole ballgame. Cadence recovers a crashed workflow by replay: it re-runs your workflow function from the start against a recorded event history, rebuilding every local variable exactly as it was. (Source: docs, "State Recovery and Determinism".)
Replay only works if the code produces the same result every time it runs. So workflow code must be deterministic. That rules out, inside the workflow function:
All of that goes into activities. The docs put it bluntly: "all communication with the external world should happen through activities." The workflow orchestrates; the activities reach out.
Mapped onto your research pipeline
Your 7-stage research-assistant-workflow splits cleanly. The orchestration — which stage runs next, whether evaluation says PASS or REVISE, looping back to revision — is workflow. The work inside each stage — calling an LLM, searching sources, writing an output file — is activities.
Notice what this buys you over the ICM folder model: the loop back to revision, the retries when an LLM call fails, and the "resume after a crash" are all handled by the workflow function instead of by you re-running a stage folder by hand.
What it looks like in Go
Rough shape only — we run this for real in a later lesson. The point is to see the split in code: the workflow calls activities through workflow.ExecuteActivity and never does I/O itself.
// ACTIVITY — ordinary Go. Free to call an LLM, hit the network, be non-deterministic.
func ExecutionActivity(ctx context.Context, plan string) (string, error) {
// e.g. call your LLM / research tools here — messy real work lives here
return callLLM(ctx, plan)
}
// WORKFLOW — pure coordination. No time.Now, no direct LLM call, no goroutines.
func ResearchWorkflow(ctx workflow.Context, topic string) (string, error) {
ao := workflow.ActivityOptions{
StartToCloseTimeout: time.Minute * 10,
// a retry policy goes here — Cadence retries the activity automatically
}
ctx = workflow.WithActivityOptions(ctx, ao)
var plan string
// ExecuteActivity is how a workflow reaches the outside world — indirectly.
if err := workflow.ExecuteActivity(ctx, CoordinationActivity, topic).Get(ctx, &plan); err != nil {
return "", err
}
var deliverable string
if err := workflow.ExecuteActivity(ctx, ExecutionActivity, plan).Get(ctx, &deliverable); err != nil {
return "", err
}
return deliverable, nil
}
Two tells that mark the boundary: the workflow uses workflow.Context (not the plain context.Context an activity gets), and it only reaches outward through workflow.ExecuteActivity. Anything a workflow needs from the world — even the current time — comes through a Cadence API or an activity, never a raw Go call.
time.Now() or an LLM call directly in the workflow function and it will often appear to work. Then a worker restarts mid-run, replay produces a different value than the recorded history, and Cadence throws a non-determinism error — or worse, silently diverges. This is the #1 Cadence bug. The split is the defense.
Retrieval check
In your research workflow, where does the actual call to the LLM that drafts the deliverable belong?
Correct. LLM calls are non-deterministic external I/O, so they live in an activity. The workflow only decides when to invoke it.
No. An LLM call is non-deterministic external I/O. It must live in an activity; the workflow only orchestrates the call.
Why is workflow code forbidden from calling time.Now() directly?
Correct. Recovery replays the code against recorded history; a live clock returns a new value each replay, breaking determinism. Use Cadence's workflow time API instead.
Not quite. The reason is replay: recovery re-runs the code against recorded history, and a live clock diverges. Cadence provides a deterministic workflow time API.
A crash kills the worker while stage 3 is running. What happens to the workflow?
Correct. The workflow is fault-oblivious: a worker resurrects it by replaying its event history, restoring local state, and continues from the last recorded point.
No. Cadence rebuilds the workflow from its recorded event history via replay and resumes from where it stopped — no manual re-run needed.
New terms — durable execution, workflow, activity, determinism, event sourcing, replay — are in the glossary.