generative-a11y

Accessible streaming AI for screen readers

Make streaming AI responses understandable to screen-reader users by announcing meaningful text segments instead of tokens or repeated transcripts.

A visual stream and an audible stream need different pacing

Token-by-token streaming is difficult for screen readers because each visual update can become a separate accessibility-tree change. Partial words, repeated accumulated text, and a long queue of low-value changes make the audible stream hard to follow. A better pattern keeps the visible transcript current while announcing only confirmed, meaningful phrases and lifecycle boundaries.

The visible response should still update normally. The accessible announcement stream is a separate representation of the same confirmed lifecycle, optimized for useful listening rather than visual immediacy.

WAI-ARIA defines live-region semantics, but those semantics are hints to user agents; they do not segment generative text or model response, retry, and tool lifecycles. That policy belongs in application logic or an accessibility runtime.

Avoid announcing the growing transcript

This pattern makes the live region contain the entire accumulated response after every token. Depending on the browser and assistive technology, users may hear fragments, repeated content, or inconsistent results.

Avoid: live region around accumulated text
function StreamingMessage({ text }: { text: string }) {
  return <div aria-live="polite">{text}</div>;
}

Send append-only deltas and an explicit terminal event

Dispatch response.started once, send only newly arrived text in response.text.delta, and finish with the terminal event the application actually observed. Core buffers partial text and emits meaningful segments according to the selected policy.

Dispatch confirmed response events
runtime.dispatch({ type: "response.started", responseId: "r1" });
runtime.dispatch({
  type: "response.text.delta",
  responseId: "r1",
  delta: "The migration completed successfully. ",
});
runtime.dispatch({ type: "response.completed", responseId: "r1" });

Expected behavior and test boundaries

With the balanced policy, complete phrases can be announced without waiting for every token or repeating earlier text. Higher-priority failures and approval requests can take precedence over routine progress. Ordinary streaming never moves focus.

Use deterministic tests to verify events, announcement intents, queue bounds, and DOM delivery. Then test the integrated application with representative browser and screen-reader combinations because spoken output is controlled by software outside the library.

Sources and evidence