Workflow vs Activity

Time: ~8 minutes. Tangible win: you can look at any line of your research pipeline and say correctly whether it belongs in a workflow or an activity — and know why getting it wrong breaks everything.

Ian — this is the one distinction the entire framework rests on. Cadence has exactly two kinds of code: workflows and activities. Almost every newcomer mistake is putting code in the wrong one. Get this split right and the rest of Cadence is detail. Get it wrong and your workflow silently corrupts on the first crash.

Primary source Read this once, slowly: Cadence Concepts — Workflows. Focus on the section "State Recovery and Determinism." Everything below is that section, made concrete for your research pipeline. (The docs' code sample is Java; ignore the syntax — the model is identical in Go.)

The split in one sentence

A workflow decides what happens and in what order. An activity does the actual work that touches the outside world.

WorkflowActivity
JobCoordinate: order, branching, waiting, retriesDo one side-effecting task
Touches outside world?NeverAlways
Can call an LLM / API / DB?NoYes
Survives a crash?Yes — state is rebuilt by replayNo state of its own; just retried
Must be deterministic?YesNo — free to be messy

Why the workflow can't touch the outside world

This is the counterintuitive part, and it is the whole ballgame. Cadence recovers a crashed workflow by replay: it re-runs your workflow function from the start against a recorded event history, rebuilding every local variable exactly as it was. (Source: docs, "State Recovery and Determinism".)

Replay only works if the code produces the same result every time it runs. So workflow code must be deterministic. That rules out, inside the workflow function:

calling an external API or LLM → result changes each run ✗ time.Now() → different every replay ✗ rand / uuid → different every replay ✗ raw `go` goroutines → scheduling not reproducible ✗ reading a file / DB directly → content can change ✗

All of that goes into activities. The docs put it bluntly: "all communication with the external world should happen through activities." The workflow orchestrates; the activities reach out.

The mental test Ask of any line: "If I ran this twice, could it give a different answer?" If yes, it is an activity. If it is pure coordination logic — if/else, loops, calling activities in order — it is workflow code.

Mapped onto your research pipeline

Your 7-stage research-assistant-workflow splits cleanly. The orchestration — which stage runs next, whether evaluation says PASS or REVISE, looping back to revision — is workflow. The work inside each stage — calling an LLM, searching sources, writing an output file — is activities.

WORKFLOW (the durable coordinator — pure decision logic) │ run intentActivity → get intent doc │ run coordinationActivity → get plan │ loop: │ run executionActivity → deliverable │ result = run evaluationActivity → PASS or REVISE │ if PASS: break │ run revisionActivity │ run knowledgeActivity │ run followThroughActivity │ ACTIVITIES (the messy real work — each can call an LLM, hit the web, write files) intentActivity · coordinationActivity · executionActivity evaluationActivity · revisionActivity · knowledgeActivity · followThroughActivity

Notice what this buys you over the ICM folder model: the loop back to revision, the retries when an LLM call fails, and the "resume after a crash" are all handled by the workflow function instead of by you re-running a stage folder by hand.

What it looks like in Go

Rough shape only — we run this for real in a later lesson. The point is to see the split in code: the workflow calls activities through workflow.ExecuteActivity and never does I/O itself.

// ACTIVITY — ordinary Go. Free to call an LLM, hit the network, be non-deterministic.
func ExecutionActivity(ctx context.Context, plan string) (string, error) {
    // e.g. call your LLM / research tools here — messy real work lives here
    return callLLM(ctx, plan)
}

// WORKFLOW — pure coordination. No time.Now, no direct LLM call, no goroutines.
func ResearchWorkflow(ctx workflow.Context, topic string) (string, error) {
    ao := workflow.ActivityOptions{
        StartToCloseTimeout: time.Minute * 10,
        // a retry policy goes here — Cadence retries the activity automatically
    }
    ctx = workflow.WithActivityOptions(ctx, ao)

    var plan string
    // ExecuteActivity is how a workflow reaches the outside world — indirectly.
    if err := workflow.ExecuteActivity(ctx, CoordinationActivity, topic).Get(ctx, &plan); err != nil {
        return "", err
    }

    var deliverable string
    if err := workflow.ExecuteActivity(ctx, ExecutionActivity, plan).Get(ctx, &deliverable); err != nil {
        return "", err
    }
    return deliverable, nil
}

Two tells that mark the boundary: the workflow uses workflow.Context (not the plain context.Context an activity gets), and it only reaches outward through workflow.ExecuteActivity. Anything a workflow needs from the world — even the current time — comes through a Cadence API or an activity, never a raw Go call.

The failure you are avoiding Put a raw time.Now() or an LLM call directly in the workflow function and it will often appear to work. Then a worker restarts mid-run, replay produces a different value than the recorded history, and Cadence throws a non-determinism error — or worse, silently diverges. This is the #1 Cadence bug. The split is the defense.

Retrieval check

In your research workflow, where does the actual call to the LLM that drafts the deliverable belong?

Correct. LLM calls are non-deterministic external I/O, so they live in an activity. The workflow only decides when to invoke it.

No. An LLM call is non-deterministic external I/O. It must live in an activity; the workflow only orchestrates the call.

Why is workflow code forbidden from calling time.Now() directly?

Correct. Recovery replays the code against recorded history; a live clock returns a new value each replay, breaking determinism. Use Cadence's workflow time API instead.

Not quite. The reason is replay: recovery re-runs the code against recorded history, and a live clock diverges. Cadence provides a deterministic workflow time API.

A crash kills the worker while stage 3 is running. What happens to the workflow?

Correct. The workflow is fault-oblivious: a worker resurrects it by replaying its event history, restoring local state, and continues from the last recorded point.

No. Cadence rebuilds the workflow from its recorded event history via replay and resumes from where it stopped — no manual re-run needed.

New terms — durable execution, workflow, activity, determinism, event sourcing, replay — are in the glossary.

I am your teacher — ask me anything. Good next questions: "What exactly does a retry policy look like on an activity?" or "Show me how the workflow gets the current time the legal way." Next lesson digs into activities: retries, timeouts, and how they fail safely.