NeuTTS-2E

NeuTTS2E_Intro

πŸš€ Spaces Demo, πŸ”§ Github

Torch version, Q4 GGUF version

Created by Neuphonic - building faster, smaller, on-device voice AI

NeuTTS-2E is a super-fast, highly realistic, on-device emotional TTS speech language model. It is an early alpha release, English-only model, supporting six emotions plus neutral (angry, disgusted, fearful, happy, sad, surprised and neutral) across four fixed speakers (emily, paul, sophie, steven). With a compact backbone and an efficient LM + codec design, NeuTTS-2E delivers strong naturalness and expressive control at a fraction of the compute, making it ideal for embedded voice agents, assistants, toys, and privacy-sensitive applications.

This model is English only with fixed speakers: for other languages and instant voice cloning, see the NeuTTS Nano Multilingual Collection.

Key Features

  • ⚑️ Ultra-fast for on-device β€” built for real-time or better-than-real-time generation on laptop-class CPUs
  • 😠😁😭 Emotional control β€” six emotions plus neutral, selected with a single argument
  • πŸ—£ High realism for its size β€” natural, expressive speech in a compact footprint
  • πŸ“¦ GGUF/GGML-friendly deployment β€” easy to run locally via CPU-first tooling
  • πŸ”’ Local-first + compliance-friendly β€” keep audio and text on-device

Websites like neutts.com are popping up and they're not affliated with Neuphonic, our github or this repo.

We are on neuphonic.com only. Please be careful out there! πŸ™

Model Details

NeuTTS-2E is designed for maximum speed per parameter while retaining strong naturalness and expressive control:

  • Backbone: compact LM backbone tuned for emotional TTS token generation
  • Input Format: text β€” no phonemizer or system dependencies required
  • Speakers: four fixed speakers (emily, paul, sophie, steven)
  • Emotions: angry, disgusted, fearful, happy, sad, surprised and neutral
  • Audio Codec: NeuCodec - our open-source neural audio codec that achieves exceptional audio quality at low bitrates using a single codebook
  • Format: quantisations available in GGUF format for efficient on-device inference
  • Responsibility: Watermarked outputs
  • Inference Speed: Optimised for real-time generation on CPUs
  • Power Consumption: Designed for mobile and embedded devices

Parameter Count

  • Active params (backbone only): ~125M
  • Total params (backbone + tied embeddings/head): ~236M

Get Started with NeuTTS

  1. Install NeuTTS

    pip install neutts
    

    Or for a local editable install, clone the neutts repository and run in the base folder:

    pip install -e .
    

    Alternatively to install all dependencies, including onnxruntime and llama-cpp-python (equivalent to steps 2 and 3 below):

    pip install neutts[all]
    

    or for an editable install:

    pip install -e .[all]
    
  2. (Required for this model) Install llama-cpp-python to use .gguf models.

    pip install "neutts[llama]"
    

    Note that this installs llama-cpp-python without GPU support. To install with GPU support (e.g., CUDA, MPS) please refer to: https://pypi.org/project/llama-cpp-python/

  3. (Optional) Install onnxruntime to use the .onnx decoder.

    pip install "neutts[onnx]"
    

Examples

To get started with the example scripts, clone the neutts repository and navigate into the project directory:

git clone https://github.com/neuphonic/neutts.git
cd neutts

Basic Example

Run the emotional example script to synthesize speech:

python -m examples.basic_example_emotions \
  --input_text "I can't believe it's finally here!" \
  --speaker emily \
  --emotion happy \
  --backbone neuphonic/neutts-2e-q8-gguf

Simple One-Code Block Usage

from neutts import NeuTTS2E
import soundfile as sf

tts = NeuTTS2E(backbone_repo="neuphonic/neutts-2e-q8-gguf")

wav = tts.infer(
    "I can't believe it's finally here!",
    speaker="emily",
    emotion="happy",
)
sf.write("test.wav", wav, 24000)

Streaming

Speech can also be synthesised in streaming mode, where audio is generated in chunks and plays as generated. Ensure you have llama-cpp-python, onnxruntime and pyaudio installed to run this example. pyaudio needs the PortAudio library on macOS and Linux:

# macOS
brew install portaudio && pip install pyaudio

# Ubuntu/Debian
sudo apt install portaudio19-dev && pip install pyaudio

# Windows (prebuilt wheels)
pip install pyaudio
python -m examples.basic_streaming_example_emotions \
  --input_text "I can't believe it's finally here!" \
  --speaker emily \
  --emotion happy \
  --backbone neuphonic/neutts-2e-q8-gguf

Responsibility

Every audio file generated by NeuTTS-2E includes by default a Perth (Perceptual Threshold) Watermark.

Disclaimer

Don't use this model to do bad things… please.

Downloads last month
308
GGUF
Model size
0.2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including neuphonic/neutts-2e-q8-gguf