Skip to main content

October 2, 2026

Force toolChoice until measured memory write

Flash skipped record_memory on multi-step agent replies. Forcing toolChoice on step 0 only still left pending measures and empty read-back cards.

Sander Korf4 min read
llmconvextypescript

A guided movement-discovery web session runs with an on-screen AI facilitator: voice plus scripted surfaces, not therapy. After readiness, the guide asks for memories of resisting something the participant meant to do. For each memory the participant rates intensity on a 0–10 scale. A read-back card is the on-screen confirmation built from the model's own record_memory tool call in that reply. Later review lists only memories that have a real rating, not a placeholder.

I was on the third memory in a long Convex-backed chat when Flash-class models started skipping the write. Prompt wording alone could not hold the line: about one session in eight reached the rating step without a measured write (a record_memory call whose measure is a real number, not { kind: 'pending' }). The participant heard the guide acknowledge the story, then the UI asked them to confirm a memory the screen never showed. The review step was one memory short.

Why step-0 toolChoice was not enough

The AI SDK multi-step agent path already had prepareStep / toolChoice hooks so you can force the next model step to call a named tool. I forced record_memory on step 0 only and left steps 1 and 2 for spoken wrap-up (stopWhen allows three steps).

That still failed in two ways I did not expect. Step 0 sometimes returned spoken text only, so the forced step never produced a tool call the UI could read. Step 0 sometimes called record_memory with measure: { kind: 'pending' } while the spoken line already asked for a 0–10 rating. Helpers like memoryCallsOf only treat non-pending measures as measured, so the read-back card stayed blank even though a tool ran.

Forcing the tool name once is not the same as forcing the data shape the product needs.

Use the demo below. Toggle Step-0 only vs Until measured, then pick a Flash outcome: Text-only, Pending write, or Measured write. Watch the step timeline, whether the read-back card would render, and whether the review count stays at three. Under Until measured, a pending or text-only first step should still get a second forced step when the cap allows; the last step should stay free for speech once a measured write exists.

Multi-step record_memory on the third memory

AI SDK agent reply with up to three steps. The backend can force toolChoice on step 0 only, or keep forcing record_memory until a measured write lands (first two steps max).
toolChoice forcedStep 1text

Flash returns spoken text (no tool call)

Step-0 onlyText-only step 0Read-back card: noReview count: 2/3

No read-back card: UI asks for confirmation on a memory the screen never showed. Review stays at two.

prepareStep until measured, cap two

The fix lives in prepareStep({ steps }) on the generate path. I read prior steps, flatten tool calls, and keep returning { toolChoice: { toolName: 'record_memory', type: 'tool' } } until memoryCallsOf(...).some(c => c.measured) is true. I stop forcing after two steps so the third can stay speech-only.

Shape only, not a production paste:

const forcedSteps = 2 // stopWhen allows 3; last stays free for speech
 
export function measureStepOptions(
  measureFirst: boolean,
  steps: readonly { toolCalls: readonly { input: unknown; toolName: string }[] }[],
) {
  const written = memoryCallsOf(steps.flatMap((s) => s.toolCalls)).some((c) => c.measured)
  if (!measureFirst || written || steps.length >= forcedSteps) return {}
  return { toolChoice: { toolName: 'record_memory', type: 'tool' as const } }
}

When the participant is answering the intensity question, measureFirst is true. When they are not in that beat, the helper returns {} and the model can talk normally.

Same session, different bug class

An earlier post on this site covers withholding ended LLM steps after move-on: stale review copy in the system prompt after the participant moved on. That is prompt composition. This post is about multi-step toolChoice until the backend sees a measured tool write before it paints the read-back card.

If you ship agent replies with tools, decide what "done" means in data, not just in the tool name. Cap forced steps so you still leave room for natural speech on the final hop.


Happy coding! Sander