Hands-on — run the real thing
Ian — every prior lesson was a mental model. This one you run. The payoff is the crash-and-resume demo you can't get from the ICM folder approach. Do this in a scratch checkout first; we port to the research pipeline afterward. Follow the steps in order and watch the feedback at each one.
Step 1 — Start a local server
Clone and build the server, install the SQLite schema, and start it:
# Clone and build (needs Go + make)
git clone https://github.com/cadence-workflow/cadence.git
cd cadence
make bins
# Install the schema and start the server on the embedded SQLite backend
make install-schema-sqlite
./cadence-server --zone sqlite start
Step 2 — Open the Web UI
In a browser, go to http://localhost:8088. This is Cadence Web — you'll watch your workflow's event history here, which is the single best way to see durable execution working.
Step 3 — Register a domain
Workflows live in a domain (Lesson 1's glossary). Register one with the CLI before running anything:
# From the cadence repo; registers a domain named "research"
./cadence --domain research domain register
# Verify it exists:
./cadence --domain research domain describe
Step 4 — Run the Hello World workflow
In a second checkout, get the samples, point them at your local server, and run the worker + starter:
git clone https://github.com/cadence-workflow/cadence-samples.git
cd cadence-samples
make # builds the sample binaries
# Start the worker (hosts the workflow + activity code):
./bin/helloworld -m worker &
# In another terminal, trigger a run (the starter):
./bin/helloworld -m trigger
Step 5 — The moment of truth: crash and resume
This is the whole point. Use a workflow with a durable sleep (the sleep sample, or add a workflow.Sleep(ctx, 60*time.Second) between two activities), then:
127.0.0.1:7933). The Workflow Troubleshooting page and the CNCF Slack are the fastest unblocks. And ask me — that's what I'm here for.
Step 6 — Port a slice to your pipeline
Once the sample crash-resumes, you've proven the engine. Now build the smallest real slice: a ResearchWorkflow that runs two of your seven stages as activities (say coordination → execution), with one workflow.Sleep or one signal gate between them. Run it, crash it, resume it. That slice — not the whole pipeline — is enough evidence to make the mission's call: does durable execution earn its operational cost for your research workflow, or does the simpler ICM/folder model still win?
Retrieval check
What does the SQLite quickstart let you skip compared to a production Cadence setup?
Correct. The embedded SQLite backend replaces Cassandra/MySQL/Postgres for local dev. You still register a domain and write the workflow/activity code.
No. You still register a domain and write the code. What SQLite removes is needing a Cassandra/database cluster behind the server.
After you Ctrl-C the worker mid-sleep and restart it, what does the event history show for the first activity?
Correct. Replay rebuilds workflow state from recorded history without re-executing completed activities. The first activity shows a single execution.
No. Replay reconstructs state from history without re-running completed activities, so the first activity appears once — not twice, and not removed.
What is the smallest slice that proves the mission's point for your pipeline?
Correct. A two-stage slice with a sleep or signal gate, crashed and resumed, is enough evidence to judge Cadence vs ICM. Rebuilding all seven or going to production is premature.
No. You don't need all seven stages or a cluster to decide. A crash-tested two-stage slice with one gate is the minimum that answers the question.
Terms used here — domain, worker, starter, event sourcing, workflow.Sleep — are in the glossary.