FlowSpeech
The tagging system is the reason to choose it: emotion, accent and pause markers give you control that most one-click TTS tools hide. Multi-speaker detection makes dialogue and audiobook work far less tedious. The catalogue of 30 voices is small, and the vendor is anonymous.
Pros
- Inline tags for emotion, accent and delivery, plus precise pause control
- Automatic speaker detection with a distinct voice per speaker
- PDF and Word upload rather than paste-only input
- Free tier of 10,000 credits a month when signed in
- 70+ languages covered
Cons
- Only 30 voices, which is small next to the big TTS catalogues
- No API, so it cannot be wired into a pipeline
- No company name, country or ownership published
- Per-request cap of 200,000 characters forces long books to be split