Skip to main content

How Flow TTS Generates Speech Locally in Chrome

9 min read

How Flow TTS Generates Speech Locally in Chrome

Flow TTS turns a web page into speech on this computer. Readable text is collected in Chrome, cleaned, chunked, and run through a local ONNX model in the extension. The article is not sent to a cloud text-to-speech API to generate audio.

That is the privacy difference versus many Chrome read-aloud extensions. Those tools still live in the browser, but the speech step happens on someone else's server: your paragraph goes out, audio comes back. Flow TTS keeps that step here. Accounts, usage checks, and first-time model downloads still use the network. The words you wanted to hear do not leave the device to become speech.

How is Flow TTS different from extensions that upload your text?

A Chrome badge does not tell you where audio is made. Three designs all get sold as “browser TTS.”

1. Chrome's built-in voices. An extension can call the Web Speech API and let Chrome pick a system or vendor voice. Some of those voices are local. Some are remote and send text to a vendor. You do not control the model.

2. A cloud TTS extension. Play uploads the article — or a large chunk of it — to a speech API. A server runs a large model and returns audio. Voice catalogs can be huge. The cost is the data flow: the text you wanted to hear left this computer.

3. Flow TTS. ONNX model files live on this machine, bundled or installed once. Inference runs inside the extension with WebGPU or WebAssembly. The in-page player plays audio that was written here. For supported installed or bundled models, page text does not need a cloud TTS service.

The private part is specific. Flow TTS is more private for the article because synthesis does not require uploading that article. It is not a claim that the product never uses the network. Local vs cloud text-to-speech is the category comparison. This page is how Flow actually does the local path.

How does Flow TTS turn a page into speech?

Once you press play, the article stays inside the extension.

1. Collect readable text. Auto mode takes main readable content from the current page. It is not a promise that menus, cookie banners, and footers stay out of the queue. A selection uses the context-menu read action. Flow Templates reuse saved Read and Skip targets on sites you visit again.

2. Apply pronunciation rules. Substitutions run on the text before the model sees it — abbreviations, names, jargon — so the engine speaks what you configured, not a guess.

3. Clean and chunk. Emojis, noisy punctuation, and awkward breaks get normalized. The remaining text is split at sentence or phrase boundaries. That split is what lets audio start before a long article is finished.

4. Hand the chunk to an offscreen worker. Manifest V3 Chrome extensions cannot keep a hidden page alive in the background tab the way older extensions did. Flow TTS uses an offscreen document plus a worker so synthesis and playback continue without taking over the article you are reading.

5. Run the local ONNX model. The worker loads the selected model and writes audio for that chunk on this computer. There is no cloud TTS endpoint in that step. A cloud extension's equivalent step is an HTTPS upload.

6. Start playback, keep synthesizing. The first ready chunk plays in the in-page player. Later chunks continue in the worker. That is why a long page can start speaking before the whole thing has been rendered to audio.

If you are still deciding *whether* to listen at all, start with how to listen to web pages in Chrome. The rest of this page stays on *how Flow makes the audio*.

Where does the Flow TTS model actually run?

The model is a file on this machine. The runtime is how Chrome executes it. Neither one is a speech vendor's API.

Flow TTS routes Supertonic-family models through a WebGPU-oriented Transformers path. WebGPU defaults on for that family where the browser exposes it. Kitten, Kokoro, Piper, and Tiny TTS use ONNX Runtime in WebAssembly. Kitten and Kokoro can enable WebGPU where the model supports it. Piper and Tiny TTS do not.

WebGPU is a speed path, not a privacy path. It runs inference on the GPU in this browser. It does not send the article to a speech vendor. WASM is the same privacy boundary on the CPU. Pick the path your hardware can sustain. The privacy policy is about data flow. The Voice tab is about whether the machine keeps up.

Each catalog entry also carries a light / balanced / demanding label. That is a hardware hint, not a quality score. A demanding model can sound better and still stall. A light model starts sooner and may sound thinner. Voice quality follows the model, not whether the runtime used a GPU.

Is downloading a Flow TTS model the same as uploading an article?

No. That confusion is how a “private” Chrome extension can still be a cloud TTS client.

A model download in Flow TTS fetches engine files once — often from a public host such as Hugging Face — checks them with SHA-256, and stores them in the browser (IndexedDB) for the next session. Later plays read those files. They do not re-upload the article to generate speech.

A per-play upload is what many other read-aloud extensions do. This page's text goes to a server every time you press play. That is cloud TTS, even if the toolbar icon sits in Chrome.

Flow TTS ships Supertonic v1 bundled, so the first listen does not wait on Models. Every other family — Kitten, Kokoro, Piper, Supertonic v2, Tiny TTS — installs from Settings → Models and must reach Complete before Voice selection will play it. The download uses the network. The synthesis after that does not send page text to a TTS API.

Why choose Flow TTS for private listening

Choose Flow TTS when the requirement is “do not give this article to a speech vendor.” Auto mode, a selection, and Flow Templates only change *which* text is collected. They do not change *where* it is synthesized.

What that looks like in the current product:

  • Install from the Chrome Web Store listing, then pin the toolbar icon. Getting started covers the first session.
  • Bundled Supertonic v1 is ready without a Models-tab download. The registry has 27 model IDs across five families. Finish an install before you play that model.
  • Streaming chunks start audio at natural boundaries so a long post does not wait for a full render.

Current limits — these are real, and they are not the same as uploading the article:

  • Play requires an account signed in with an email one-time code. The onboarding tour can continue unsigned-in. Playback will not start until a token is present.
  • Free usage is 10,000 characters per day. Premium unlocks unlimited characters and Glass Mode when entitlement is active.
  • Usage sync sends counts, not the words on the page. Accounts, license checks, and model downloads use the network. That is why Flow TTS is not fully offline and not zero-cloud.

If the requirement is “do not upload the article to generate speech,” try Auto mode on one clean page with bundled Supertonic. Then read the privacy policy against what you just heard.

What local synthesis in Flow TTS does not solve

Local inference does not make every machine fast. WebGPU is not available in every Chrome profile. A demanding model on a small laptop will hitch. Switching families is the control. Hoping a cloud API will “just work” is a different product — and a different privacy bargain.

Local inference also does not mean Auto mode reads only the article. Extraction can still pick up chrome. Flow Templates and selection exist because the pipeline will faithfully speak whatever text it was given, still on this computer.

And local inference does not remove accounts or usage. Playback is gated on a signed-in token and a usage snapshot. After models are cached and that snapshot is valid, synthesis can continue without a TTS API. Downloading a new family, signing in again, or refreshing usage can still need the network.

FAQ

Does Flow TTS upload the article to generate speech?

No. For supported installed or bundled models, page text is turned into audio on this computer. That is the difference from a cloud TTS extension, which sends the article to a server on every play. Accounts, usage, licenses, and model downloads may still use the network. They do not need the article text to generate speech.

Why is Flow TTS more private than a typical read-aloud extension?

Because the speech step does not require a third-party TTS API. Many extensions keep a player in Chrome and still upload the page to make audio. Flow TTS runs the model here. More private for the article is not the same as never using the internet.

Is this the same as Chrome's built-in voices?

No. Built-in read-aloud uses the Web Speech API and a voice list Chrome already knows. Some of those voices are remote. Flow TTS loads its own ONNX model into the extension and writes the audio there.

Does WebGPU send page text off the device?

No. WebGPU accelerates Flow TTS inference inside this browser. It is a hardware path. WASM is the fallback on the CPU. Neither one is a cloud TTS API.

Can Flow TTS work after the models are installed?

Synthesis can, for a supported installed or bundled model, once a valid account token and usage snapshot are present. Signing in, checking usage, activating a license, and downloading another model can still need a connection.

Why isn't Flow TTS fully offline?

Because “local” describes the speech step, not the whole product. Flow TTS generates audio on-device for supported models and still uses the network for accounts, usage, licenses, and model files. Treat those as separate pipes from “did this article go to a speech server?”

Listen to articles with Flow TTS

Turn long web pages into natural speech in Chrome, with on-device generation for supported local voices.

How to Make Text-to-Speech Read Only the Article

Why read-aloud tools speak menus, ads, and comments, and five escalating ways to make text-to-speech read only the article body in Chrome.

Read post

How to Reduce Eye Strain From Reading Online

Practical ways to make online reading easier on tired eyes, including screen breaks, workspace changes, and listening to articles with text-to-speech.

Read post

How to Listen to Articles While Working in Chrome

Which work pairs with background listening, which does not, and how to set up a single playback tab in Chrome so audio does not fight the rest of your day.

Read post