本地语音合成

浏览器本地运行 Hojo-TTS 系列神经网络模型(Hojo-TTS-Light 与 Hojo-TTS-Light-40M,可下拉切换),把中英文文字念出来:不联网、音频不出设备,首次按所选型号下载约 250MB 权重后可离线使用,内置多音色(中文 + 英文)。

本地语音合成tts语音合成文字转语音离线wasmhojo神经网络中文语音text to speech
关于「本地语音合成」

它属于「AI 本地推理」分类:用 WASM 在浏览器本地跑模型:语音转文字等能力不联网也能用,音视频不上传服务器。

完全免费、无需注册;所有处理都在浏览器本地完成,文件不会上传到任何服务器。

相关关键词:本地语音合成、AI 本地推理、tts、语音合成、文字转语音、离线、wasm、hojo、神经网络、中文语音、text to speech、在线工具、免费

如何使用「本地语音合成」
  1. 打开工具:打开「本地语音合成」页面,无需注册、无需安装任何软件。
  2. 准备输入:准备好要处理的内容或文件:文本可直接粘贴,文件可拖入页面(本工具属于「AI 本地推理」)。
  3. 填入内容:把内容粘贴进输入框,或把文件拖入页面;也可以点击「示例」先填入样例数据试一下。
  4. 设置选项并处理:按需调整页面上的选项,然后点击处理 / 转换 / 生成按钮。
  5. 获取结果:处理完全在浏览器本地完成,结果可直接复制或下载,文件不会上传到任何服务器。
常见问题
下载很慢 / 连不上 huggingface.co?
下载会自动改用镜像 hf-mirror.com 重试;「选择模型目录」后模型直接写进你选的文件夹(断点续传),下次从目录读取不再下载。也可用别的工具提前把 HF 仓库下好,放进去装载。
合成没声音 / 一直转圈?
首次合成要装载 ~250MB 权重并在单线程 WASM 上自回归推理,长文本可能要几十秒到几分钟,请耐心等待;若长时间无响应,打开控制台看是否报 BF16 / WASM 相关错误。
Local Text-to-Speech (English)

Synthesize Chinese and English text entirely in the browser with the Hojo-TTS family neural models (Hojo-TTS-Light and Hojo-TTS-Light-40M, switchable from a dropdown): no network, audio never leaves the device. Weights (~250MB, depending on the chosen model) download once, then it works offline; multiple built-in voices (Chinese + English).

Belongs to “Local AI”: Run models locally in your browser with WASM: speech-to-text and more work offline, and audio/video is never uploaded.

Free, no sign-up; everything runs locally in your browser and files are never uploaded.

Keywords: Local Text-to-Speech, Local AI, local tts, text to speech, speech synthesis, offline, wasm, hojo, neural, chinese voice, online tool, free

How to use it

  1. Open the tool:Open the “Local Text-to-Speech” page — no sign-up and nothing to install.
  2. Prepare the input:Prepare the content or file to process: paste text or drop a file onto the page (this tool belongs to “Local AI”).
  3. Provide the input:Paste your content into the input box or drop a file onto the page; you can also click “Sample” to try example data first.
  4. Adjust options and run:Adjust the options if needed, then click the process / convert / generate button.
  5. Get the result:Everything runs locally in your browser; copy or download the result — your files are never uploaded.

FAQ

Download is slow / huggingface.co is unreachable?
Downloads auto-retry through the hf-mirror.com mirror; "Pick model folder" writes models straight to your chosen folder (resumable), so next time they load from disk. You can also pre-fetch the HF repo with another tool and load it from there.
No audio / stuck spinning?
The first synthesis loads ~250MB of weights and runs autoregressive inference on a single WASM thread; long text can take tens of seconds to minutes. If it hangs, open the console for BF16 / WASM errors.

More tools in Local AI

  • Speech to Text — Live streaming and file transcription: live mode turns microphone audio into text as you speak, file mode transcribes local audio/video with timestamps. Local WASM inference with VibeVoice-ASR-BitNet can be downloaded (auto-retried via the hf-mirror.com mirror), saved into a folder you choose, or loaded straight from local files — with SHA-256 integrity checks and no network needed afterwards; capture supports WebRTC 3A (AEC / noise suppression / AGC); export to TXT/SRT.

返回崽崽工具箱首页