When a product changes in the fashion store, a background Vercel Workflows run (code whose progress is saved so it can resume after a crash) refreshes search records for linked merchandising and navigation items. Each step is one language: load CMS copy, read the commerce API catalog slice, push documents to the search provider. If a language still fails after retries, the workflow emits a degradation signal tagged with which upstream failed and an error code, so on-call can tell a CMS 400 from a commerce timeout.
Who gets stuck: the engineer staring at the degradation dashboard where every failed language showed source: unknown and reason: fetch_failed. A CMS validation error and a commerce API timeout looked identical. The workflow body never saw the rich error object that the step threw.
Why the workflow body only sees recorded results
On replay, the workflow function runs again from the top, but it does not re-execute finished steps. The event log stores each step outcome. Deterministic replay means the body reads those recorded results instead of calling the dependency again. When a step throws, the runtime stores a serialized failure and rebuilds an Error when the workflow catches it. In the workflow package version I ran (^4.5.0), that rebuilt error carried the message only: no cause, no custom fields, and instanceof checks against my app error class failed. I then passed that error into a second step, which serialized it again.
The public docs describe hydration of cause on WorkflowRunFailedError when you await run.returnValue outside the workflow. That is a different surface from the error object inside the workflow catch after await someStep(). Check what your runtime preserves on both paths before you trust typed fields.
What looked correct inside the step
Inside the step I did the textbook wrap: map typed app errors to FatalError when they are not retryable, keep retryable ones as ordinary errors, and attach the original on cause for local debugging.
const wrapped = e.retryable ? new Error(msg) : new FatalError(msg)
wrapped.cause = eUnit tests called wrap and unwrap helpers in one Node process. cause was intact. Every test passed. Production attribution still said unknown.
I had shipped unwrapStepError that read error.cause and a round-trip test asserting equality after toStepError. Green in CI. Still wrong on the dashboard until I replaced that helper about fifteen minutes later.
A quieter bug lived in the same patch: retryable typed errors were rethrown without encoding, so a failure that exhausted retries also arrived without its code. The fix encodes attribution for retryable failures too.
Treat the step boundary like JSON over the wire
A Vercel Workflows step boundary is a serialization boundary. Anything not in the surviving fields is gone. Tests that never cross the boundary cannot catch this class of bug.
What survived for me was the message string. I encode code, source, and the human text in a parseable prefix, parse it back in the workflow body, and keep cause as a bonus for in-process callers only.
const ATTRIBUTED = /^(\S+) \[([^\]]*)\]: /
function stepFailureAttribution(e: unknown, fallback: string) {
const m = e instanceof Error ? ATTRIBUTED.exec(e.message) : null
return m ? { source: m[2] || 'unknown', reason: m[1] } : { source: 'unknown', reason: fallback }
}
// Test helper: what the workflow body actually sees
const acrossStepBoundary = (e: unknown) => new FatalError((e as Error).message)Tests should rebuild from the message (simulate acrossStepBoundary) instead of round-tripping in memory. Engines with documented typed error serialization (Temporal ApplicationError with type and details, or custom serialize hooks) can skip the string format when you trust the pipeline.
Try it in the demo
Pick CMS 400 or commerce API timeout, choose Read cause, press Cross step boundary, and watch attribution stay unknown / fetch_failed even though the message still contains the encoded prefix. Switch to Encode in message and cross again: the side panel’s “After boundary” row should show the real source and code. The in-memory unit test row stays green for both strategies, which is the trap.
Durable step error identity
cause vs encoding fields in the message.Inside the step (before boundary)
- class
- i
- code
- cms_bad_request
- source
- cms
- retryable
- false
- cause
- cms_bad_request (cms)
- message
- cms_bad_request [cms]: Invalid locale payload
Workflow body (after boundary)
Press Cross step boundary to rebuild the error from the recorded message only.
Degradation attribution
Cross the boundary to see what the dashboard would log.
The same rule applies anywhere you persist and replay: queue payloads, structuredClone across workers, server action errors on the client. When the engine documents error shapes you can rely on, use that API instead of inventing a message format.
Happy coding!
Sander